Skip to main content
Success
[PRO SERVICES / SECURITY & GOVERNANCE]

Private AI
deployment

Want to know if it makes sense to run an AI model in-house? We’ll test a model on your infrastructure and report back with an assessment of what’s required.

What we work out

As a company, you may be interested in running models privately, whether on-premise or on your own private cloud, and using agentic harnesses like Codex or Claude Code. This may be for security concerns, performance, or just plain preference. We can help you figure out:

What models and harnesses you can run on your existing hardware, and what kind of hardware is required to run the models you want

How much it would cost to get the required infrastructure

How to configure and deploy everything

How to keep everything secure

How it works

We will try to connect Codex to your preferred model. Codex natively supports other models as long as they follow the Responses API. Claude Code can connect to a private gateway, but Anthropic does not allow connecting to other models. Whatever the case, we will make sure the agentic harness you want works with your preferred model. We can help you figure out a good agent/model combo if necessary.

Codex, Claude, and other agentic models work with files and tools to accomplish tasks. We’ll test with some of your own tasks to make sure the model/harness combo works with your own stack and permissions.

Agent harnesses have a different infrastructure load than chatbots: agent sessions flip between model calls and tool work, which means that the number of users and the number of active model calls are different. This, combined with the choice of model, means we can’t just plug values into an equation to determine hardware requirements. Instead, we’ll perform a load test to see how many users your hardware can support. We’ll let you know whether your hardware is sufficient or not. It’s very possible that your current servers can handle an agentic model just fine, because it’s not always necessary to have the biggest GPUs on the market. We’ll give you a list of hardware and rough ballpark of cost if necessary.

Once we know your infrastructure can handle your target number of users, we’ll set everything up, test, and give you a bill of materials for what’s required. This may include installing and configuring servers, setting up connections between models and agent harnesses, security policies to ensure that your model cannot be accessed from the internet, logging, backups, monitoring, etc. We’ll document everything. If your setup is fully airgapped, we’ll document a process to manage model updates. We’ll document a rollback plan and leave it in the hands of a designated owner.

When it makes sense

Self-hosted models make sense if you expect steady use of AI, even if that means only a handful of users, need to keep prompts and responses within your infrastructure, or if your setup is airgapped. Hosted solutions can be much less expensive for smaller numbers of users and if you need the absolute latest models. We’ll give you an honest answer as to whether self-hosting makes sense.

Vu Agency working session

First steps

Give us a shout. Tell us what model(s) you’d like to use, approximately how many users will be using them, what hardware you have, and what agentic harnesses you’d like to connect them to, such as Codex or Claude Code, and we’ll send you more information. This initial assessment would be fixed-scope with a flat fee.

[MORE PRO SERVICES]

More from Security & Governance

Every Pro Service page covers what it is, who it fits and how to start. The full list is in the footer below.

Message us on WhatsApp