Hardware HUD
GPU and VRAM per card, power, CPU, RAM, temperatures, disks and service status.
Product
Xeretron brings together models, hardware, voice, documents, APIs, security and operation into a system that lives inside your organization. This page explains the problem it solves, what you'll see on the panel, and what's needed to get it up and running.
The operational problem
Every conversation leaves the organization. Each bill changes with use. Each change in supplier or price forces us to redesign what was already built. And the institution's most sensitive information ends up governed by outside policies.
What is processed in an external cloud is no longer under your daily governance.
The budget depends on how much is used, and usage can rarely be predicted.
Models, limits and prices change without the organization being able to decide.
Chat, transcription, OCR and APIs in different services, without unified control.
Unified view
The Xeretron dashboard displays hardware status, loaded models, active services, and inference activity. The idea is simple: if you can see it, you can operate it.
GPU and VRAM per card, power, CPU, RAM, temperatures, disks and service status.
Available models, health, upload and download, published profiles and download queues.
Streaming, visual attachments, audio transcription and per-response metrics.
Published endpoints, scoped keys, and documentation for your developers.
Inference history and security logs to review what happened and when.
Ports, system services and electrical consumption managed from the panel.
Panel Tour
Each operation leaves useful information: which model responded, which lane it ran in, how long the first token took, and at what speed it generated. That visibility is what turns a demo into a sustained operation.
Capabilities in detail
Each engine solves a specific task and shares the same panel, the same security and the same infrastructure with the rest.
Inference server with GGUF models and streaming compatible with the OpenAI format: context, threads, fast attention, KV cache and card division. Honest limit: not all GGUF runs on any binary or hardware; Compatibility is validated in the evaluation.
Models with a visual component and multimodal projection, accepting images by URL or embedded data, for reading documents, captures and photographs.
Base, GPU0, GPU1, dual and cpu execution lanes. Actual concurrency is limited by available video memory: it is sized, not blindly promised.
Draft, test and publish flow. Saved and restorable model combinations, with pre-estimated memory for each stack.
Suspends the regular stack, loads a larger model, and restores the configuration after inactivity without cutting active inferences. It may take several minutes.
Whisper transcription on the node itself and idle offloading to return GPU to main tasks.
PDF, Office, images and spreadsheets converted to text in Spanish and English, ready to consult or summarize with the loaded models.
On-demand audio generation within the node, for notices, accessibility and products that need to speak.
Initial latency and tokens per second measurements on real hardware, contextualized; Managed weight storage and download queues.
Endpoints over HTTPS with scope-limited access keys, to connect internal applications or existing services.
Integration with game development environments and Unreal, SIP telephony and synchronization between nodes.
With the models installed, local operation is possible. Updates and downloads require connectivity, and we specify this in each deployment.
Architecture
Simplified public diagram. We do not publish internal routes, credentials, or sensitive deployment details.
Requirements and deployment
The requirements depend on the models and the expected attendance, which is why prior evaluation is mandatory. As a general guideline:
Tell us what data, models and users you have in mind. We respond with an honest evaluation, including cases where Xeretron is not the best option.