Xeretron AI Server · Your private AI, in your own environment

Your own artificial intelligence, on your machine or server.

NO tokens · NO limits · ALL yours

Xeretron installs models, voice, documents, APIs and AI operations wherever you decide: your computer, your server, or your company's. Your data stays under your control — whether for personal use, a small team, or an entire organization.

  • No tokens or usage fees
  • For people, teams and companies
  • Xeretron's own models
  • On your machine or server
  • Designed to grow with you
  • Text and vision models
  • Voice to text
  • OCR and documents
  • Text to speech
  • Application API
  • Multi-GPU operation

The paradigm shift

From renting answers to building your own AI node.

Modern AI is sold as a cloud token counter. Xeretron proposes another relationship: your private AI running in your own environment, far from being a distant service that is rented on demand, becomes a capability that is yours — to you, your team or your organization.

Comparison between AI in the cloud for tokens and your private AI with Xeretron
Dimension AI in the cloud for tokens Your private AI with Xeretron
Data pathEvery request travels outside your controlProcessing in your own environment
Cost structureVariable cost per consumptionInstalled capacity, more predictable cost
ModelsSupplier dependencyModel selection and control
Information governanceData subject to external policiesLocal information governance
ConnectivityInternet as a central requirementWorks on your computer or server, even without internet

Xeretron does not intend to eliminate the public cloud. It allows you to decide which loads should remain private in your environment and where it makes sense to combine both worlds by consuming giant models with external inference providers.

15 s
To understand the product
Days
At the first inference
0
Mandatory interactions with third parties
ES/EN/PT
Interface and documentation

What is Xeretron

A practical operating system for private AI.

It's not a chat with another logo: it's a system that brings together models, hardware, security and services in one place — from a person's computer to a company's server.

Models under your control

Text and vision GGUF with upload, download and profiles managed from the node itself.

Visible hardware

GPU, CPU, RAM, temperature and power in one clear HUD - you know what's happening in real time.

Multimodal AI

Chat, images, speech to text, documents and audio generated in the same local ecosystem.

APIs to build

Endpoints compatible with the OpenAI format plus own services to integrate your applications.

Users and security

Authentication, MFA, API keys with limited scope and auditing of accesses and inferences.

Real operations

Multi-GPU, network, firewall, services and power managed from the same interface.

The product, inside

See what is happening.

What model is loaded, where it runs, how much it consumes and if the services are healthy. Each dashboard screen exists so that any person or team can operate their own AI without being engineering experts.

  • Dashboard and hardware HUD with node telemetry.
  • Catalog of models with loading, unloading and published profiles.
  • Chat with streaming, visual attachments and inference metrics.
  • Speech to text, document OCR and text to speech in the same flow.
  • Users, roles, keys and auditing in one place.
Explore the product
Actual Xeretron dashboard screenshot: model catalog and inference stack
Click to enlarge

Capabilities

Much more than a chat.

A set of specialized engines that share the same panel, the same security policy and the same infrastructure.

LLM inference

Text and vision with streaming, profiles, upload/download, lanes, stacks, frontier mode and metrics.

voice to text

Local transcription with Whisper and idle flushing to free up GPU when not in use.

Documents and OCR

PDF, Office, images and spreadsheets converted to text in Spanish and English.

text to speech

On-demand audio generation within the node, without sending the content to external services.

Platform

APIs, MCP/Unreal, SIP telephony, Sync and explicit hybrid.

Government

System and application users, PAM, MFA, scoped keys, auditing and network.

Infrastructure

Multi-GPU, firewall, services and energy consumption managed from the panel.

Multilingual

Interface and documentation in Spanish, English and Portuguese, ready for diverse teams.

Solutions

For people, teams and institutions that are responsible for their data.

From your own team to an entire organization: each context has its own language and its own limits. Xeretron adapts to everyone.

Education

An AI that strengthens the institution, not extracts its knowledge: academic assistants, regulation consultations, class transcription, accessibility and teaching support.

See cases

Universities and research

Experiment with your own Xeretron models and quantizations, build applications with no variable cost per test, take care of datasets on campus and train talent in real infrastructure.

See cases

Healthcare

AI close to the teams that care for people, with administrative and documentary cases. It does not replace medical judgment nor is it, by itself, a medical device.

See cases

Government

The AI ​​capacity of a public institution can also be public assets. Sovereignty does not mean isolation, and we do not claim automatic legal compliance.

See cases

People

Your private AI on your own computer: an assistant for writing, studying and creating, with your data under your control. Ideal to start with the Personal plan.

See cases

Teams and companies

Internal assistants, document processing, automation, voice and operational continuity in the event of supplier or price changes. From a small team to the entire organization.

See cases

Entrepreneurs and labs

Maximum flexibility of models and profiles to prototype products with your own AI, without each experiment being charged per token.

See cases

Privacy

Privacy by architecture, not just by promise.

Confidentiality does not depend on a declaration of intent: it depends on where the bytes travel, who can enter and what is recorded. Xeretron is designed around those controls.

We do not promise “inviolable” systems. We describe specific controls and the current limits of each deployment, which are defined in the pre-assessment.

Security and technology

Authentication and MFA

Second factor access and session management from the node itself.

scoped API keys

Each integration receives permission-limited and consumption-limited keys.

Audit

Security and inference logs to reconstruct what happened.

TLS and firewall

Controlled external exposure, encryption in transit and managed ports.

Economy

When AI becomes infrastructure, the cost conversation changes.

In the cloud, the cost grows linearly with your usage — and it doesn't stop. With Xeretron in your own environment, the conversation changes: a clear investment and a predictable annual cost, for a few users or for dozens.

  • Estimated breakeven point: CAPEX ÷ (monthly cloud cost − monthly operation).
  • Blended SME indicative profile: entry ~$1.50 and exit ~$6.00 per million tokens.

How it works

Four steps, with real guidance.

We don't deliver an installer and we disappear. Every deployment starts with an assessment and ends with you — or your team — operating your AI autonomously.

We assess

Diagnosis

What data will be processed, what models you need, how many users, what hardware exists and what availability you expect.

We install

Deployment

Preparation of the node, CUDA and services, initial catalog of models and HTTPS over your network.

We configure

Integration

Users, roles, API keys, connections with your applications and training of the responsible team.

We accompany

Operation

Boot support, performance adjustments, agreed updates and results reporting.

First inference in days, not months

Plans

Four ways to get started.

From personal equipment to organization: all plans include installation, training and start-up support. Additional user on any plan: $29 USD / year.

Staff

For you and up to 3 users

$599 USD for a limited time

Regular price: $1,299 USD

+ $299 / year · license maintenance and updates

  • 1 GPU on your computer or server
  • Up to 3 models of your choice
  • Up to 3 users per installation
  • Unified Dashboard and Personal API
  • No tokens or usage fees
Start with Personal

Institution

Schools · NGOs · up to 20 users

$4.999 USD Installation

+ $999 / year · license maintenance and updates

  • GPU sized based on evaluation
  • Up to 20 users per installation
  • Vision and multi-user models
  • Synchronization and audit
  • Training program
Evaluate this plan

Company

Organizations · up to 50 users

$7.999 USD Installation

+ $1,399 / year · license maintenance and updates

  • Scalable architecture with frontier mode
  • Up to 50 users per installation
  • API keys by area with limited scope
  • Security, audit and service agreement
  • Integrations with existing systems
Evaluate this plan

Prices in USD, per installation. Each additional user over the plan limit is added to $29 USD/year. Additional GPU license $99/year. Final amounts are confirmed after evaluation of your case.

Voices of the program

What people who already run their own node say.

Developers, operational teams and business owners who already run their models on their own infrastructure.

I used to pay for tokens every time the model got stuck in a large refactor. Now I run everything local with Xeretron: I iterate without fear, I try three approaches at once and the monthly cost is the same as the first month. For vibe coding, the freedom of unlimited use completely changes the flow.

Martín VegaIndependent developer · vibe coding · Mexico

We built a document agent that reads contracts, extracts key clauses and puts together summaries in minutes. Previously, it took two analysts an afternoon to review a file. With our own node, sensitive documents never leave our networks and the volume of paperwork grew without hiring more people.

Carolina RuizOperations Director · Legal SME · Colombia

I use XMail to respond to routine emails and check my schedule without touching my phone in the morning. The node responds with my tone, prioritizes what is urgent and makes my day. I regained hours every week and, above all, my business conversations do not go through a third-party cloud.

Jorge RenteriaEntrepreneur · manufacturing · Guatemala

We migrate technical product documentation to the node: engineers consult manuals, sheets and protocols with answers in seconds. Before, half a day was lost searching for the correct information. Today each team responds alone and knowledge of the area no longer depends on a single person.

Laura PinedaEngineering manager · industrial sector · Peru

As a graduate student I need to experiment without each test costing me a subscription. With my own node I test and compare models 24 hours a day, and everything remains on my server: my academic data does not go to an external provider nor does it appear in some other log.

Diego SalasPhD student · data science · Argentina

As a small law firm we could not afford an analyst to review each contract. Today our document agent reads, classifies and resolves files at night, and in the morning the team only reviews what is important. The investment paid for itself in two months, and nothing we signed leaves our network.

Andrea MoralesFounding partner · legal firm · Chile

Each case operates on its own infrastructure with Xeretron models. Results vary depending on plan, hardware, and volume of usage.

Allies Program

Limited spaces for the launch partner program.

Accompanied deployment, close access to the team and direct participation in the evolution of the product. At this stage it is not self-service: it is your private AI, installed and accompanied.

History

We spent years building Xeretron, doing the difficult so that you can imagine and create the “impossible”.

Xeretron is born from a simple conviction: The world should not limit itself to consuming artificial intelligence governed from other places. It is not a story of founders; It's the quiet work of making a node work, every day, under real-world conditions.

Meet Xeretron

Strategic question

The question is no longer “do we use AI?” Is: Whose AI will it be?


  • Sovereignty without isolation: you decide what is local and what is hybrid.
  • Power with clear limits: we say what your node can and cannot do.
  • Knowledge that remains within your institution, your lab or your home.

Frequently asked questions

Direct answers, without fine print.

Does it work without Internet?

Local components work if the models are already installed. Model downloads and external integrations may require connectivity. We define what remains offline in the evaluation of your architecture.

Does Xeretron include the models?

Yes, Xeretron has its own models from 4B to 552B and they are available according to your plan or you can purchase them separately, you can also manage and operate the models within the node. The specific selection of models and their licenses depends on your use case and is defined in the project.

Does it replace the public cloud?

Not necessarily. It can operate in private or hybrid mode: you decide which loads remain inside and where it is convenient to combine both worlds.

What hardware do I need?

It depends on the models and the expected attendance. That is why prior evaluation is mandatory: we size GPU, RAM and disk before proposing a configuration.

Does my data leave the server?

Local flows can stay entirely within your network. Any exit to an external service occurs only if you explicitly enable it.

Does Xeretron decide for doctors, teachers or civil servants?

No. It is infrastructure plus human supervision: it amplifies capacity, it does not replace the professional judgment of those who teach, research, care or serve.

How did Xerebria train the models for Xeretron?

We use the best open weight models and fine tune them so that they work correctly with our entire ecosystem.

How is it installed?

With deployments accompanied in this initial stage: diagnosis, installation, configuration, training and start-up support.

Can I integrate my applications?

Yes, through API compatible with the OpenAI format and own services, with access keys limited by scope, according to the project agreement.

Next step

Your private AI can run on your own computer, today.