Skip to main content
Success
[PRO SERVICES / SECURITY & GOVERNANCE]

Private AI deployment

Want to self-host your company's AI? Vu can help.

We check whether a private model can handle the work you have in mind. If it can, we install it on your servers or private cloud and connect the tools your team needs.

It could answer questions from company files or work through documents. A strong enough model can also power an agent such as Codex. We test the job and expected number of users before recommending hardware.

WRITTEN AND FACT-CHECKED 9 AUGUST 2026
TECHNICAL REFERENCES MAY CHANGE AFTER THIS DATE

LOCAL AGENT SETUP
MODEL: INTERNAL

Agent harness

CODEX + COMPATIBLE TOOLS

Private endpoint

AUTH + API SUPPORT

INSIDE YOUR NETWORK

Local model server

WORKING CONTEXT

Code and knowledge

OPERATIONS

Logs, limits and updates

[AGENT HARNESSES]

Run compatible agent harnesses on a local model

Codex can handle software and knowledge work. It can read files and use tools to complete multi-step jobs inside the permissions you give it.

A self-hosted model can sit behind Codex if its serving stack implements the Responses API and the model handles the context and tool calls the work requires. We test that complete path.

Model service

A pinned model and serving engine, selected through task tests and configured on your machines or private cloud.

Harness connection

Codex Responses API configuration, or equivalent provider setup for another compatible harness.

Work access

Controlled access to the files, sources and internal tools each job needs.

Operations

A written process for access, monitoring, software updates, regression tests and rollback.

[HARDWARE]

Hardware depends on active use

An agent session alternates between model calls and tool work, so total users and active model calls are different numbers. Long contexts and parallel subagents can move the requirement quickly.

We test representative software or knowledge work under the reasoning settings and load you expect before recommending hardware. The table below shows why that step matters.

ILLUSTRATIVE CAPACITY METHOD

DeepSeek V4 Flash 0731 deployment

CHECKED 9 AUGUST 2026

SGLang verifies Flash 0731 on 4 x GB300 and 4 x H200. Its current public page does not publish measured multi-user results for those setups. This illustration assumes 30% of sessions are generating, adds 25% headroom and uses a load-test target of 16 active calls per 4-GPU replica.

CONCURRENT SESSIONS ACTIVE CALLS TO TEST GB300 GPUS IF 16 CALLS PASS
10 4 4
25 10 4
50 19 8
100 38 12

These GPU totals are planning arithmetic, not benchmark results or a purchase estimate. Production sizing comes from a load test using your tasks and context lengths.

SGLANG SOURCE
[WHAT VU INSTALLS]

A working private model service

The first engagement is fixed-scope. We establish whether local deployment makes sense, test a suitable model and price the production setup before you commit to hardware.

BOOK A PRIVATE AI REVIEW
01

Review and benchmark

Use case, current machines, candidate models, measured quality and load results.

02

Install and connect

Model server, private endpoint, harness configuration and any agreed tools.

03

Control access

Authentication, network policy, logs, usage limits and monitoring.

04

Hand over

Pinned versions, configuration, update and rollback procedure, plus named ownership.

Local deployment tends to suit steady workloads, restricted data or environments that must work offline. An approved hosted API may be cheaper when use is light or access to the strongest available model matters more.

[QUESTIONS]

Questions

Q.01

Which agent harnesses can you connect?

Codex officially supports custom providers through the Responses API. Claude Code can use a private LLM gateway, but Anthropic does not support non-Claude models behind it. Other harnesses vary, so we confirm protocol and tool support first.

Q.02

Do we need GB300 GPUs?

No. SGLang currently verifies Flash 0731 on 4 x GB300, 4 x H200 and 8 x B200. That shows where the checkpoint runs, not how many users each setup will support. We load-test before sizing.

Q.03

Does local mean secure?

It removes the public inference route when configured that way. Access, model files, updates, logs, backups and support connections still need controls.

Q.04

Can it run without internet access?

Yes. The deployment then needs an offline process for software patches and new model versions, including verification and rollback.

[PRIVATE AI REVIEW]

Tell us what you want to run

Send the model you have in mind and the number of users. Include any hardware you already own. If you want to run an agent harness, tell us which one and what it needs access to.

[MORE PRO SERVICES]

More from Security & Governance

Every Pro Service page covers what it is, who it fits and how to start. The full list is in the footer below.

Message us on WhatsApp