“Can we run this on our own servers?”
AI that runs inside your own network.
Where the reasoning happens is a decision you make rather than a default we set — private cloud, hybrid, or an appliance in your own rack.
Engagements typically start at $35,000, and scale with what you connect and where it runs.
Why firms ask for hardware
Almost nobody asks for on-premise hardware because they want hardware. They ask because they want a defensible answer to a different question: where does our material go, who can reach it, and what happens when a client or a regulator asks us to account for it.
On-premise is one answer to that question. It is a good one, it is not the only one, and it is not free. Worth knowing what it actually buys before it becomes a line in a procurement document that nobody can trace back to a rule.
What on-premise means here
A managed Cogneros appliance is installed in your office or data centre. Approved models run locally, the index is built and held inside your network, and egress is restricted to what your policy allows — which can be nothing at all for the categories you define.
The appliance is managed. It is kept ingesting, patched, monitored and supported the same way a cloud environment is. On-premise means the hardware is yours and the work stays inside it; it does not mean you inherit an operations burden along with the rack space.
What running locally costs you
Models that run on hardware you own are smaller than the frontier models that run in a data centre. For retrieval and citation over your own material that gap matters far less than people expect, because the hard part is finding the right passage rather than composing the sentence around it. For open-ended reasoning, the gap is real.
Capacity becomes a decision rather than an elastic default. You size for the work you expect, and adding capacity later means adding hardware rather than changing a setting.
Both of these are properties of local inference rather than of any one vendor. Anyone who tells you on-premise costs nothing is selling you the hardware.
Hybrid is what most organizations actually want
It is where the conversation usually lands once the question gets specific. Defined categories of work stay on hardware inside your network; everything else runs in a private cloud environment, where capability is highest and there is nothing to rack.
You set the boundary — by matter, by client, by role — and you can move it later. Many organizations start entirely in the cloud and pull categories local as the genuinely sensitive work becomes clear, which is a much better order than guessing at the start.
Model choice is yours either way
Cogneros is model-flexible. Work can route to local models, to approved cloud models such as Claude or GPT, or to a combination. Which models are available, to whom, and for which kind of work is configured by you.
That is usually the part that survives a security review: not a promise that nothing ever reaches a cloud model, but a written, enforced rule about which work may reach which model — and a log showing the rule held.
How to decide
Four situations and what each one points at. If none of them describes yours, the answer is a conversation rather than a page.
- Contract or policy says certain material cannot leave the building
- An on-premise appliance, for that material at least. The rest can run wherever is convenient.
- A defined subset is genuinely sensitive and the rest is ordinary work
- Hybrid. The boundary is yours to set, and yours to move once you have watched it in use.
- You want capability soon and have no policy blocker
- Private cloud. Start there — categories can move local later without starting over.
- The driver is a checkbox nobody can trace to an actual rule
- Find the rule first. On-premise chosen for its own sake buys cost rather than safety.
Questions this raises
The ones that come next in nearly every evaluation, answered plainly.
Do we have to buy the hardware?
The appliance is specified and supplied as part of the engagement, and it lives in your rack. Sizing follows the size of the corpus and the number of people using it at once, and it is settled during the architecture phase — before anything is indexed, so the number is not a surprise arriving after the decision.
Can it work with no outbound connection at all?
Egress is restricted to what your policy allows, and for the categories you define that can be nothing. Updates and support then follow a scheduled process you control rather than a continuous connection. That is a deliberate trade: a fully isolated environment moves more slowly, and for some material that is exactly the right trade.
Is on-premise more secure than a private cloud environment?
Different, not automatically better. On-premise moves the boundary to a perimeter you already defend, which is worth a great deal if you defend it well — and it also makes you responsible for that perimeter. The honest comparison is against how your own environment is actually run, not against a brochure.
Can we start in the cloud and move on-premise later?
Yes, and it is a common path. The deployment model is a configuration of the same system rather than a different product, so moving categories of work local does not mean rebuilding. What changes is where inference happens and where the index sits.
The other questions
These three come up together, in roughly this order, in nearly every evaluation.
Over your document system
Nothing moves. Nothing gets re-filed.
Cogneros is an intelligence layer over the repositories you already run — originals stay where they are, under the permissions they already carry.
Read moreConfidentiality
Three questions wearing one word.
“Is it confidential?” usually means three things at once: what trains on our material, who inside the firm can reach it, and what we can show afterwards.
Read moreCompare
Four ways to put AI into a firm.
Public assistants, point tools, private deployment, or building it yourself — what each is good at, and how to tell which one your situation calls for.
See the comparison
Every one of these is settled in writing during evaluation, before anything is indexed — so what your counsel reviews is an architecture rather than a summary of a sales conversation.
Your business has already built the knowledge.
Cogneros makes it available to the people who need it — under your permissions, traced to your own documents.
What an engagement costs
Engagements typically start at $35,000.
Scope follows the repositories you connect and where the system runs — cloud, hybrid, and on-premise are priced differently.
- Prefer email?
- hello@theravengroup.com
- Prefer to talk?
- +1 303-351-1691
- Based in
- Denver, Colorado


