Product

An AI server that is governed from a single dashboard.

Xeretron brings together models, hardware, voice, documents, APIs, security and operation into a system that lives inside your organization. This page explains the problem it solves, what you'll see on the panel, and what's needed to get it up and running.

Request a private evaluation See technical architecture

The operational problem

Having AI today means, in practice, depending on someone else.

Every conversation leaves the organization. Each bill changes with use. Each change in supplier or price forces us to redesign what was already built. And the institution's most sensitive information ends up governed by outside policies.

Data in transit

What is processed in an external cloud is no longer under your daily governance.

Variable cost

The budget depends on how much is used, and usage can rarely be predicted.

Supplier dependency

Models, limits and prices change without the organization being able to decide.

Scattered tools

Chat, transcription, OCR and APIs in different services, without unified control.

Unified view

The entire node, on a single screen.

The Xeretron dashboard displays hardware status, loaded models, active services, and inference activity. The idea is simple: if you can see it, you can operate it.

Hardware HUD

GPU and VRAM per card, power, CPU, RAM, temperatures, disks and service status.

Catalog and loading

Available models, health, upload and download, published profiles and download queues.

Chat and tests

Streaming, visual attachments, audio transcription and per-response metrics.

APIs and keys

Published endpoints, scoped keys, and documentation for your developers.

Inferences and audit

Inference history and security logs to review what happened and when.

Network, firewall and power

Ports, system services and electrical consumption managed from the panel.

Panel Tour

From loaded model to response, with metrics.

Each operation leaves useful information: which model responded, which lane it ran in, how long the first token took, and at what speed it generated. That visibility is what turns a demo into a sustained operation.

  • Real-time streaming with history on the client.
  • Visual attachments for models with supported mmproj component.
  • Audio transcription and document extraction in the same work thread.
  • Behavior indicators by model family.
Actual Xeretron dashboard screenshot 1: hardware dashboard with dual GPU
Click to enlarge

Capabilities in detail

One platform, multiple specialized engines.

Each engine solves a specific task and shares the same panel, the same security and the same infrastructure with the rest.

Local inference

Inference server with GGUF models and streaming compatible with the OpenAI format: context, threads, fast attention, KV cache and card division. Honest limit: not all GGUF runs on any binary or hardware; Compatibility is validated in the evaluation.

Text and vision

Models with a visual component and multimodal projection, accepting images by URL or embedded data, for reading documents, captures and photographs.

Lanes and concurrency

Base, GPU0, GPU1, dual and cpu execution lanes. Actual concurrency is limited by available video memory: it is sized, not blindly promised.

Profiles and multi-model stack

Draft, test and publish flow. Saved and restorable model combinations, with pre-estimated memory for each stack.

frontier mode

Suspends the regular stack, loads a larger model, and restores the configuration after inactivity without cutting active inferences. It may take several minutes.

Local voice to text

Whisper transcription on the node itself and idle offloading to return GPU to main tasks.

Documents and OCR

PDF, Office, images and spreadsheets converted to text in Spanish and English, ready to consult or summarize with the loaded models.

text to speech

On-demand audio generation within the node, for notices, accessibility and products that need to speak.

Benchmarks and downloads

Initial latency and tokens per second measurements on real hardware, contextualized; Managed weight storage and download queues.

Inference API

Endpoints over HTTPS with scope-limited access keys, to connect internal applications or existing services.

Optional modules

Integration with game development environments and Unreal, SIP telephony and synchronization between nodes.

Offline operation

With the models installed, local operation is possible. Updates and downloads require connectivity, and we specify this in each deployment.

Architecture

Decoupled services on private infrastructure.

Users and applications
Xeretron AI Server
LLM/VL Models
voice to text
OCR / documents
text to speech
Security / users
Observability
Private infrastructure · NVIDIA GPU · Linux

Simplified public diagram. We do not publish internal routes, credentials, or sensitive deployment details.

Requirements and deployment

What do you need to get started?

The requirements depend on the models and the expected attendance, which is why prior evaluation is mandatory. As a general guideline:

  • Server with Linux and NVIDIA GPU with CUDA.
  • From one to two or more GPUs depending on the package chosen.
  • RAM and disk sized to the planned model catalog.
  • Internal network with the possibility of exposing HTTPS in a controlled manner.
  • A technical manager designated by the organization.

What the client receives

  • Optimized hardware and software as a single system.
  • Xeretron panel and inference engine configured.
  • Speech to text, text to speech and file processing.
  • HTTPS, initial model catalog and on-site installation.
  • Team training and startup support from 30 to 90 days.
  • Documentation in Spanish.

Can your current infrastructure become an AI node?

Tell us what data, models and users you have in mind. We respond with an honest evaluation, including cases where Xeretron is not the best option.