Your artificial intelligence. On your infrastructure.
Under your control.
No tokens · No limits · All yours.
Value proposition
From renting answers to building your own AI node.
Modern AI is often sold as a cloud token meter. Xeretron proposes another relationship: your private AI running in your own environment — for you, your team, or your organization. A practical operating system that brings models, hardware, security, and services together in one place, from a personal workstation to an enterprise server.
What it is
A practical operating system for private AI.
Models under your control
Local LLM inference, own models from 4B to 552B, and selection by use case. No usage quotas.
Multimodal AI
Speech-to-text, text-to-speech, vision, documents, and OCR on the same node.
APIs to build
Compatible endpoints to integrate your apps, internal agents, and own products.
Users and security
MFA auth, scoped API keys, full audit, TLS, and managed firewall.
Unified panel
Which model is loaded, where it runs, how much it consumes, and whether services are healthy — without being an infra expert.
Real operations
Guided deployment: diagnosis, install, training, and ongoing support.
Public cloud tokens vs. installed capacity
The cost conversation changes.
Public cloud
- Linear cost with usage: it grows and doesn’t stop.
- Your data travels to third-party infrastructure.
- Shared usage limits and external availability.
- Every extra query adds to the monthly bill.
Xeretron in your environment
- Clear investment + predictable annual cost, for a few users or dozens.
- Data and models inside your institution, lab, or home.
- Dedicated GPU: resilience and power without shared limits.
- Unlimited local tokens: usage doesn’t create extra cost.
Who it’s for
For people, teams, and institutions accountable for their data.
Typical first-month cases: institutional assistant, AI lab for teachers, research testbed, policy lookup, committee and class transcription, clinical document support, API for own products, voice accessibility, prototyping with no token bill, reserved-document analysis, telephony voice, and character simulation.
Plans & investment
Four ways to start.
Personal
For you and up to 3 users
List price: $1,299 · limited time
+ $299 / year maintenance
- 1 GPU on your machine or server
- Up to 3 models of your choice
- Unified panel and personal API
- No tokens or usage quotas
Starter
Small teams · up to 10 users
Installation
+ $799 / year maintenance
- 1 GPU with compact or mid-size models
- Initial catalog and internal API
- Team training
Institution
Schools · NGOs · up to 20 users
Installation
+ $999 / year maintenance
- GPU sized after evaluation
- Vision-capable models and multi-user
- Sync, audit, and training
Enterprise
Organizations · up to 50 users
Installation
+ $1,399 / year maintenance
- Scalable architecture with frontier mode
- Scoped API keys by area
- Audit, service agreement, and integrations
Prices in USD per installation. Extra user beyond the plan limit: $29 USD / year. Extra GPU license: $99 / year. Final amounts are confirmed after evaluating your case.
Process
Four steps, with real support.
Diagnosis
Data, models, users, existing hardware, and expected availability. Week 1.
Installation
Node prep, CUDA, services, initial catalog, and HTTPS. Week 2.
Training & integration
Users, keys, app connections, and team enablement. Week 3.
Support
Performance tuning, support, and pilot results report. Week 4+.
Whose AI will it be?
Request a free evaluation of your case and we’ll tell you which plan makes sense — or if it isn’t the right moment yet.