Local AI Server for Businesses: Multi-user LLM, Cost, and Sizing

Twenty employees on a cloud AI subscription already amounts to several thousand euros per year, not to mention that every prompt, every document pasted, goes to a third-party server. A local AI server changes the equation: a single investment, the whole team connected to it, and no data leaving the company.

This guide is for IT decision-makers and executives considering deploying a shared LLM internally: the real cost compared to the cloud, technical criteria for multi-user operation, how to size it according to headcount, and the deployment checklist.


The problem of cloud at team scale

An individual AI subscription seems harmless. Multiplied by the number of employees, it is no longer so.

Cloud subscription, per user

€20 to €30
per month and per user, standard offers

Billed continuously, as long as the account exists. The cost grows with each new hire. Data transits through the provider's servers.

Local AI server

One-time investment
generally amortized in 12 to 18 months

The whole team connected to the same machine via the network. No added cost per user, no data leaving the company.

Figures for informational purposes only. Cloud offer prices vary depending on the provider, level (standard or enterprise) and negotiated volume discounts; advanced offers with extended compliance features often cost more. The amounts above serve as a reference for reasoning, not a guaranteed quote. For an accurate comparison, start with your actual cloud quotes.


What a multi-user local AI server is

It's a single machine, sized for computational power, installed on your network. Each employee accesses it from their workstation via a web interface, just as they would with a cloud tool, except that the request never leaves your premises.

The principle changes everything at scale: instead of paying a license per person, you pay for infrastructure once, sized for your team's actual usage.


Technical criteria specific to multi-user operation

A shared server does not have the same constraints as an individual workstation. Here's what really matters.

  • Concurrent requests. Several people query the model at the same time. The server must process these requests in parallel without anyone having to wait their turn, which requires available VRAM beyond the model's weight alone.
  • Throughput, not just capacity. A model that responds quickly for a single user can slow down significantly with five or ten simultaneous requests. The GPU's memory bandwidth is the factor that determines this throughput.
  • Account management. Each collaborator must have their own space, history, and access. An interface like Open WebUI natively manages separate accounts on the same server.
  • Continuous availability. A team server must remain on and reachable during business hours, with stability designed for prolonged operation, not for occasional sessions.
  • Growth margin. Headcount or usage rarely decreases. Planning for a VRAM and RAM margin avoids premature replacement.
The most common sizing error: choosing a server for the model size, without accounting for users. A model that fits in 16 GB for a single user may require two to three times more VRAM to absorb five to ten simultaneous sessions with a comfortable context. The number of users weighs as much as the model in sizing.


Which server for which headcount

The table below provides a rough estimate. Actual usage (conversation length, RAG for documents, code generation) will vary these benchmarks.

Connected users Recommended VRAM Reference machine Equivalent cloud cost (indicative, per year)
Up to 10 users 16 to 32 GB CoreAI 64 (RTX 5090) around €3,000
10 to 25 users 64 to 96 GB Mini Server GB10 or Rack 2 × RTX 5090 around €6,000 to €9,000
25 to 50 users 96 GB ECC Pro AI Ultra Threadripper around €9,000 to €18,000
50 to 100 users and more 192 GB ECC Rack 2 × RTX 6000 Blackwell around €18,000 to €36,000
Typical profitability. At constant headcount, a local server generally pays for itself within 12 to 18 months compared to an equivalent cloud subscription, even before considering that the hardware investment remains usable for several years. Beyond this period, the marginal cost per user becomes zero.


Deployment checklist

  • Authentication enabled from commissioning, individual accounts for each employee.
  • Network access restricted to the internal network or via VPN, never open to the internet without control.
  • HTTPS to encrypt exchanges between workstations and the server.
  • Regular backups of conversations and indexed documents if RAG is used.
  • Load monitoring to anticipate an increase in headcount before it degrades response times.
  • Scheduled updates of the inference engine and interface, without interrupting activity.
Turnkey installation. Upon request, we deliver the server with the complete pre-configured environment: inference engine, multi-user interface, created accounts, and possible on-site installation anywhere in France and Europe for initial network setup.


Our servers for multi-user LLMs

All these machines are hand-assembled in Auriol (13390) and fully configurable, via the online configurator or on quote for a specific need.

Small team, up to 10 workstations
Radiance CoreAI 64 RTX 5090

CoreAI 64 — RTX 5090 32 GB

32 GB VRAM · Ryzen 9 9950X3D · 64 GB DDR5

Restricted team server, models up to 70B.

€6,042starting from
Medium team, 10 to 25 workstations
ASUS Ascent GX10 NVIDIA GB10

NVIDIA GB10 — ASUS Ascent GX10

128 GB unified memory · Grace Blackwell · DGX OS

Compact and silent server, ideal for shared offices.

€3,999starting from
Radiance Rack 2x RTX 5090

CoreAI Rack — 2 × RTX 5090 (64 GB)

64 GB VRAM · Ryzen 9 9950X3D · 128 GB DDR5 · 4U Rack

Higher throughput, more simultaneous requests.

€11,221starting from
Large team, 25 to 50 workstations
Radiance Pro AI Ultra Threadripper

Pro AI Ultra — Threadripper PRO

96 GB ECC VRAM · Threadripper PRO 7955WX · 128 GB ECC expandable

ECC reliability for continuous large-scale operation.

€20,213starting from
Enterprise, 50 to 100 workstations and more
Radiance Rack 2x RTX 6000 Blackwell ECC

CoreAI 128 Rack — 2 × RTX 6000 Blackwell (192 GB ECC)

192 GB ECC VRAM · Ryzen 9 9950X3D · 128 GB DDR5 · 4U Rack

Top of the line: large models, 24/7 availability, several tens of workstations.

€27,980starting from
Everything is configurable. Each server can be precisely adjusted to your headcount and usage: VRAM, RAM, storage, power supply. Configure it directly from the online configurator, or describe your needs for a custom quote, including for headcounts over 100 workstations. Write to contact@radiancesystems.eu or via the quote form on the website.


Frequently Asked Questions


How many users can a local AI server support?

This mainly depends on the VRAM and GPU memory bandwidth, not a software limit. Our configurations cover small teams of a few workstations to deployments of several tens of simultaneous users on rack and Threadripper configurations.


Does the server really replace a cloud subscription like ChatGPT or Copilot?

For writing, summarization, analysis, and code use cases, yes: current open models rival the best proprietary tools. A local server adds total confidentiality and eliminates recurring per-user costs. Some very specific integrations into a proprietary ecosystem may require a supplement, to be evaluated based on your tools.


Do I need an IT team to manage the server?

Not on a daily basis. Upon request, we deliver the pre-configured environment, with user accounts created. On-site installation is possible for initial network setup. Routine maintenance (updates, backups) remains comparable to that of a classic file server.


What happens if the headcount increases?

Machines are designed to evolve: adding RAM, storage, or even a second GPU on configurations that allow it. Planning for a margin from the outset avoids complete replacement when the team grows.


Does data really stay within the company?

Yes. The server operates on your local network or via VPN. No request, no document, no conversation is sent to a third-party service. This is the fundamental difference from a cloud subscription.


Can I add document RAG to the server?

Yes, as an option, just like on all our workstations. The server can index a document base common to the entire team, searchable by each user, with the same guarantees of confidentiality.

Back to the blog

Your quote for a custom AI solution within 24–48 hours

Every Radiance project starts with a conversation. Fill out this form, and an expert will quickly respond with a solution tailored to your business and budget.

Response within 24–48 business hours
Delivery throughout Europe (EU)
2-year warranty included
On-site installation available
No commitment on demand
Dedicated support before and after purchase
contact@radiancesystems.eu
+33 4 65 84 48 21
Mon – Fri, 9 AM – 5 PM
01 What is your primary use of AI?
Multiple choice.
02 In what context will the system be used?
Single choice.
03 What type of system are you looking for?
Single choice.
04 Which operating system do you prefer?
Single choice.
05 What are your expectations for the software?
Multiple choice.
06 What is your approximate budget?
Single choice.
07 When would you like to receive your system?
Single choice.
08 Would you like assistance with the setup?
Single choice. A Radiance technician can assist you at your location or remotely.
09 Delivery country (EU only) *
We only deliver within the European Union (EU).
10 Additional information (optional but very useful)
Briefly describe your project, your specific constraints, or any useful information.
11 Would you like to be contacted to discuss your project?
If you choose "Quote only," you will be able to reply to our email to ask your questions and refine the quote.
12 Email *
We will send the quote to this address.

Any more questions?

Send us an email at contact@radiancesystems.eu or contact us via the contact form; we respond to all inquiries within 3 hours during business hours (Monday to Friday from 9 AM to 5 PM).

📞 +33 4 65 84 48 21