Why people ask for this

  • “We cannot paste client data into a public AI tool.”

    The same workflow running on a model you control, with every connector, log and backup checked so the data stays inside your boundary.

  • “Legal and IT have both said no.”

    A written data-flow diagram that shows what goes where, so the decision rests on facts.

  • “We tried a local model and it was too slow or too weak.”

    We size the model to the hardware and the job, then measure it on your own examples before you buy anything.

  • “We do not want to depend on one provider.”

    Model routing that lets a provider be swapped, with the choice logged.

  • “The cloud bill is unpredictable.”

    We measure the cost per task in each mode, so the choice between local and cloud rests on numbers.

What we build

Deployment is the product here. The model is a part we choose to fit it.

  1. On-device deployments

    Models running on laptops and workstations, for sensitive drafting and offline work.

  2. On-premises servers

    Open-weight models sized to your GPUs or CPUs, served behind your firewall.

  3. Private cloud

    Deployments in your cloud account, with your keys, network rules and logging.

  4. Hybrid routing

    Sensitive steps stay local. Heavy or public-data steps use cloud models. The rule is written down and every routed call is logged.

  5. Model selection and sizing

    An open-weight model chosen to fit your hardware and your task, tested on your examples.

  6. Production hardening

    Serving, queuing, monitoring, alerts, evaluation and rollback, so it keeps running after the demo.

What you receive

  • The deployed system, with the infrastructure described as code where your environment allows.
  • A data-flow diagram: what stays where, and what leaves.
  • Benchmark results on your own examples for each model we considered.
  • Dashboards and alerts for latency, errors and cost.
  • A runbook covering updates, model swaps, rollback and failure.
  • A handover session for whoever operates it.

Does it connect to what we already use?

Yes. A private model is only useful when it reaches your documents, your tools and your users.

  • Document stores
  • Databases
  • Identity and single sign-on
  • Internal APIs
  • Existing cloud accounts
  • Monitoring tools
  • Ticketing
  • MCP servers

We work inside your network and identity setup instead of asking for exceptions to it.

Where it can run

This is the page where the answer differs most by mode. Here is the honest version of each.

  1. On your devices

    Fits, with limits

    Small open-weight models on a laptop or workstation. Good for drafting and private lookups. Quality and speed depend on the machine, and long or complex tasks may be out of reach.

  2. On your servers

    Fits

    Larger open-weight models on servers you control. You carry the hardware, the patching and the capacity planning.

  3. Private cloud

    Fits

    Rented GPUs or managed model services inside your own cloud account. Data stays in your tenancy. Check what a managed service retains under its terms.

  4. Hybrid

    Fits

    The usual answer for mixed data. It needs a clear rule for which step goes where, and we build the rule and its log.

  5. Cloud

    Fits

    The largest models and scale on demand. Data goes to the provider under its terms, so it is not for your most sensitive material.

Not every model or feature is available in every mode. Some capabilities, such as the very largest models, are cloud only. We say which before you commit.

Running it on your own hardware

Hardware is a scoped part of the work, separate from the equipment itself. We do not hold stock, and equipment is bought separately.

  • Assess what you have

    We check whether your servers, workstations or GPUs can run the model you need at the speed you need, before anything is bought.

  • Size it from the workload

    The model’s size and how far it is compressed set the memory it needs; how many people use it at once and how long their documents are add to that. We size from your model, your users and your response-time goal, then test.

  • Count the running cost

    Local hardware swaps per-use fees for equipment, power, cooling, storage and upkeep. You get an itemized estimate for your case, next to the cloud option for the same work.

  • Source, install and configure

    If new equipment is needed, we help you specify it and buy it from your supplier, then install and configure the serving stack and test it under your load.

Small teams or low-volume jobs can often run a smaller compressed model on a workstation GPU, or even a CPU. We test your workload before recommending anything.

Offline and isolated operation

A local model alone does not keep data local. For systems that must run offline or isolated, we verify the whole system, not just the model.

  • Models, embeddings and tokenizers

    Every model file is staged in advance and the libraries are set to offline mode, so nothing is downloaded while the system runs.

  • Connectors and outside calls

    Every connector, tool and integration is listed with where its traffic goes. In an isolated deployment, only local or on-network ones are enabled.

  • Telemetry and tracing

    Some inference servers and libraries send anonymous usage data by default. We switch it off, and hosted tracing services stay off.

  • Logs, caches and backups

    We decide with you what is logged, how long it is kept and where backups live, and keep those stores inside your boundary.

  • Proven by test

    Settings are not proof. We run the deployment with outside network access blocked and show you the results, including any attempted outside connection.

Privacy statements are made per system, from the architecture we built and the controls we verified, and written into the handover.

How it is built

For technical readers: what we choose from. The choice is made per system, from your data, your constraints and your workload.

Open-weight models
Families such as gpt-oss, Qwen, Mistral, Gemma, DeepSeek, Phi and Llama, each checked against its licence. Open-weight does not mean unrestricted: some licences depend on company size or on reselling access.
Serving
vLLM, llama.cpp or Ollama on your hardware: NVIDIA, AMD or Apple GPUs, or a CPU for small workloads.
Private cloud
Amazon Bedrock with private VPC endpoints, Microsoft Azure AI Foundry with private endpoints, or Google Cloud with VPC Service Controls.
Cloud models
OpenAI, Anthropic, Google Gemini or xAI Grok models on paid tiers, with each provider’s training, retention and processing-location terms confirmed in writing.
Fast inference
Groq, the inference provider that runs open-weight models on its own chips. Not to be confused with Grok, the model family from xAI.
Routing and caching
Rules handle what can be written down, caches handle repeats, and larger models handle only what needs them.

Model names and versions change quickly, so proposals name the exact model, version and licence.

Systems we run ourselves

We do not publish client work. Two systems we operate show the pattern.

System we operate

Daena

Routes across model providers with local and cloud runtimes as options, and keeps the routing on the record.

Beta, with access on request at daena.mas-ai.co. It is the reference architecture we draw on for governed agent builds.

Read the case study
Daena Brain view with demo data: the governance core, its six capabilities, ten departments such as Engineering, Finance and Legal and Compliance, their agents and demo tool servers, drawn as a network.
  • ragX

    Runs on our own hardware, with a local model for the critic step.

    System we operateRead the case study

Questions we get

Is a local model as good as the best cloud models?

For some tasks it is close enough. For others it is not. The very largest models are cloud only. We test open-weight models on your examples and tell you where the gap matters.

What hardware do we need?

It depends on the model, how many people use it at once and the response time you need. We can assess what you have, or size, source and set up new equipment as a scoped piece of work; the equipment is bought separately. An existing workstation is often enough to test.

Does private cloud mean the data never leaves us?

It means the data stays in your cloud account, under your keys and network rules. If a managed model service is involved, that provider retention terms still apply, so we read them with you.

Can we start in the cloud and move later?

Yes, and it is a common path. We design the system so the model call is a swappable part, which makes a later move to local or private a much smaller job than a rebuild.

How do you keep a local model reliable?

The same way as any production system: serving, queuing, monitoring and evaluation. We track quality against your examples so a model update cannot quietly make it worse.

Do you train or fine-tune models?

Not by default. Most business problems are solved by retrieval, prompting and evaluation on an existing model. If a fine-tune would clearly help, the Build Blueprint says so.

Bring us the data that cannot leave.

We start with the problem, then tell you whether it is worth building.

Bring us a bottleneck