Private AI and on-premise LLMs
Language models that run on hardware you control, for data that has no business on someone else’s cloud. Estela designs, installs and operates private AI infrastructure for companies in Uruguay and abroad.
When it makes sense
- Sensitive documentsLaw firms, finance, health: client files, contracts and records that cannot be sent to a third-party API.
- Regulation and contractsData that must stay in the country, in the building, or under a specific agreement.
- Steady, heavy useEnough daily volume that owning the compute costs less than paying per token.
What we set up
- GPU computeNVIDIA GPUs for inference and fine-tuning, sized to the models and the load.
- Self-hosted modelsOpen-weight language models served inside your network, with an API your tools already understand.
- VirtualizationVMware vSphere or Proxmox underneath, bare metal where latency matters.
- Storage and backupsFast storage for datasets and model weights, snapshots, and restores that actually restore.
- Access and securityEncryption, access control and audit logs, so you know who asked what.
- Hybrid, when it paysA cloud model only for the tasks where the numbers favor it, and never with data that must stay home.
How it works
- 1. AssessmentWe look at the tasks, the data and the rules around it, and tell you whether private AI is worth it.
- 2. PilotWe prototype the pipeline on our own GPU server, without sending your data to any third party.
- 3. DeploymentThe same setup goes onto your hardware or a dedicated machine, with monitoring and redundancy.
- 4. OperationWe keep it running, updated and backed up. Nearly a decade of infrastructure work is the part people don’t see.
We run our own
- Estela’s GPU serverNVIDIA hardware on Proxmox runs the local models we use for AI-assisted engineering and for prototyping client pipelines. We recommend what we operate.
Questions
- Is a private model as good as ChatGPT?
- For reading, classifying, extracting and drafting from your own documents, current open-weight models do the job well. For tasks where they fall short, a hybrid setup sends only non-sensitive work to a cloud model.
- Does our data leave the company?
- No. Models, weights, prompts and documents stay on your hardware or on a dedicated machine you control.
- What does it cost?
- It depends on the models, the volume and whether you already own hardware. The assessment gives you a number before you commit to anything.
- Do we need an IT team?
- No. We install, monitor and maintain it. Your team uses it.
Have data that can’t leave the building? Tell us the task. We’ll tell you what it takes to run it privately.
Send a note to: hola@este.la