NorthCore Labs

What hardware do you need to run an LLM on your own servers?

Memory is the limit. A model needs roughly its parameter count times the bytes per weight in GPU memory, plus working space for the conversation. At 4-bit quantization that is about half a gigabyte per billion parameters, so a 70-billion-parameter model needs roughly 40 GB before overhead, which is why 48 GB-class GPUs, or two smaller ones, are the usual starting point for a business-grade model.

Updated 2026-10-09

Rough memory by model size

These are approximations for the weights alone. Real use adds memory for the conversation (the context window) and for each simultaneous user, so size above these figures.

Model sizeWeights at 16-bitWeights at 4-bitTypical hardware class
7 to 8 billion parametersabout 15 GBabout 4 to 5 GBOne mid-range GPU
13 to 14 billionabout 27 GBabout 8 GBOne 16 to 24 GB GPU
30 to 34 billionabout 65 GBabout 18 to 20 GBOne 24 GB GPU (tight) or a 48 GB GPU
70 billionabout 140 GBabout 38 to 42 GBA 48 GB-class GPU, or two 24 GB GPUs

What else decides the hardware

  • Concurrent users: each active conversation needs working memory, so ten users is not one user times ten in speed or in memory.
  • Context length: long documents in the prompt, which RAG produces, use much more memory than short chats.
  • Speed: tokens per second depends on memory bandwidth as well as capacity.
  • The rest of the machine: system memory, fast storage for model files, power and cooling.
  • Apple Silicon is an option for small teams because memory is shared, at lower speed than a datacenter GPU.
  • Private cloud: a dedicated GPU instance in your own cloud account keeps the data under your control without buying hardware.

When running locally is not worth it

At low volume a hosted model under a contract that forbids training on your data and sets retention limits can be cheaper and better than hardware you maintain. Local deployment earns its cost when the data cannot leave your control, when volume is steady and high, or when you need to tune the model.

How we size it

We test the actual workload, with your documents and your expected number of users, on candidate models before recommending hardware. The number that matters is not the largest model that fits but the smallest one that passes your test set.

Questions people ask.

How much GPU memory do I need for a 70B model?

Roughly 40 GB for the weights at 4-bit quantization, plus working memory for the conversation, so a 48 GB-class GPU or two 24 GB GPUs is a common starting point.

Can I run a business LLM without a GPU?

Small models can run on a CPU, slowly. For several users or long documents a GPU is the practical choice.

Is a local LLM cheaper than an API?

It depends on volume. At steady high volume it can be; at low volume a hosted model under suitable contract terms is often cheaper.

Want this done for your business?

Walk us through how things run today. We show you what we would build first, what it costs and how fast. If nothing here fits, we say so on the call.

No card, no contract
A calendar invite and a Zoom link the moment you book.
Prefer email?
admin@northcorelabs.io
NorthCore Labs LLC
7901 4th St N, Ste 300, St. Petersburg, FL 33702