Private AI deployment
Want to self-host your company's AI? Vu can help.
We check whether a private model can handle the work you have in mind. If it can, we install it on your servers or private cloud and connect the tools your team needs.
It could answer questions from company files or work through documents. A strong enough model can also power an agent such as Codex. We test the job and expected number of users before recommending hardware.
WRITTEN AND FACT-CHECKED 9 AUGUST 2026
TECHNICAL REFERENCES MAY CHANGE AFTER THIS DATE
Agent harness
CODEX + COMPATIBLE TOOLS
Private endpoint
AUTH + API SUPPORT
INSIDE YOUR NETWORK
Local model server
WORKING CONTEXT
Code and knowledge
OPERATIONS
Logs, limits and updates
Run compatible agent harnesses on a local model
Codex can handle software and knowledge work. It can read files and use tools to complete multi-step jobs inside the permissions you give it.
A self-hosted model can sit behind Codex if its serving stack implements the Responses API and the model handles the context and tool calls the work requires. We test that complete path.
Model service
A pinned model and serving engine, selected through task tests and configured on your machines or private cloud.
Harness connection
Codex Responses API configuration, or equivalent provider setup for another compatible harness.
Work access
Controlled access to the files, sources and internal tools each job needs.
Operations
A written process for access, monitoring, software updates, regression tests and rollback.
Hardware depends on active use
An agent session alternates between model calls and tool work, so total users and active model calls are different numbers. Long contexts and parallel subagents can move the requirement quickly.
We test representative software or knowledge work under the reasoning settings and load you expect before recommending hardware. The table below shows why that step matters.
ILLUSTRATIVE CAPACITY METHOD
DeepSeek V4 Flash 0731 deployment
SGLang verifies Flash 0731 on 4 x GB300 and 4 x H200. Its current public page does not publish measured multi-user results for those setups. This illustration assumes 30% of sessions are generating, adds 25% headroom and uses a load-test target of 16 active calls per 4-GPU replica.
| CONCURRENT SESSIONS | ACTIVE CALLS TO TEST | GB300 GPUS IF 16 CALLS PASS |
|---|---|---|
| 10 | 4 | 4 |
| 25 | 10 | 4 |
| 50 | 19 | 8 |
| 100 | 38 | 12 |
These GPU totals are planning arithmetic, not benchmark results or a purchase estimate. Production sizing comes from a load test using your tasks and context lengths.
SGLANG SOURCEA working private model service
The first engagement is fixed-scope. We establish whether local deployment makes sense, test a suitable model and price the production setup before you commit to hardware.
BOOK A PRIVATE AI REVIEWReview and benchmark
Use case, current machines, candidate models, measured quality and load results.
Install and connect
Model server, private endpoint, harness configuration and any agreed tools.
Control access
Authentication, network policy, logs, usage limits and monitoring.
Hand over
Pinned versions, configuration, update and rollback procedure, plus named ownership.
Local deployment tends to suit steady workloads, restricted data or environments that must work offline. An approved hosted API may be cheaper when use is light or access to the strongest available model matters more.
Questions
Which agent harnesses can you connect?
Codex officially supports custom providers through the Responses API. Claude Code can use a private LLM gateway, but Anthropic does not support non-Claude models behind it. Other harnesses vary, so we confirm protocol and tool support first.
Do we need GB300 GPUs?
No. SGLang currently verifies Flash 0731 on 4 x GB300, 4 x H200 and 8 x B200. That shows where the checkpoint runs, not how many users each setup will support. We load-test before sizing.
Does local mean secure?
It removes the public inference route when configured that way. Access, model files, updates, logs, backups and support connections still need controls.
Can it run without internet access?
Yes. The deployment then needs an offline process for software patches and new model versions, including verification and rollback.
Tell us what you want to run
Send the model you have in mind and the number of users. Include any hardware you already own. If you want to run an agent harness, tell us which one and what it needs access to.