System tools

Each tool solves one job. They all share the same panel.

These are the tools that run inside the Xeretron node. For each we publish what it does and its real limit—we prefer commercial precision to a wish list.

Inference Voice Documents Platform Governance Operations

01 · Inference

The engine that answers.

Models, memory, and concurrency: where your node's real capacity is defined.

Core

Local inference server

Runs GGUF models inside your infrastructure with OpenAI-format-compatible streaming. Configure context, threads, fast attention, and KV cache from the panel.

Core

Text and vision

Models with a vision component and multimodal projection. Accept images via URL or embedded data to read documents, screenshots, and photos.

Performance

Execution lanes

base, gpu0, gpu1, dual, and cpu lanes to distribute loads and isolate critical tasks from ordinary traffic.

Performance

Model profiles

Full lifecycle: draft, test, and publish. What one engineer experiments with becomes stable configuration for everyone else.

Performance

Multi-model stack

Saved, restorable combinations of several models, with memory estimation before loading.

Advanced

Frontier mode

Suspends the usual stack, loads a larger model, and restores configuration after idle time without cutting active inference.

02 · Voice

Listen and speak without leaving the building.

Transcription and synthesis that don't send audio to any external service.

Core

Local speech to text

In-node audio transcription, with unload on idle to return the GPU to primary tasks when unused.

Core

Text to speech

On-demand audio generation for announcements, accessibility, training content, and products that need to speak.

03 · Documents

From file to queryable text.

The foundation of almost every institutional use case: turning paper into usable knowledge.

Core

File service and OCR

Converts PDF, Office documents, images, and spreadsheets into plain text ready to query, summarize, or extract data.

Performance

Download queues and weights

Model storage management: download queues, weight verification, and control of which versions stay available.

04 · Platform

So your applications can consume the node.

Interfaces and optional modules. Everything optional is labeled as such.

Core

Inference API

HTTPS endpoints with scope-limited access keys to connect internal systems, automations, or your own products.

Optional per project

Unreal / GameDev integration

Connect the node to game and simulation dev environments for characters, assisted dialogue, and dynamic content.

Optional per project

SIP telephony

PBX integration for phone voice scenarios on the same transcription and synthesis capacity.

Optional Phase 3

Sync between nodes

Configuration and catalog coordination across several nodes in the same organization.

Hybrid

Explicit external providers

Ability to route specific workloads to external services when the organization decides and documents it.

Performance

Multilingual interface

Panel and documentation in Spanish, English, and Portuguese for diverse teams and regions.

05 · Governance and security

Who can do what—and a record of it.

Concrete controls with clear responsibilities: neither absolute promises nor a black box.

Core

System and application users

Human and service identities with roles and least privilege, managed from the node itself.

Core

Authentication and MFA

Time-based second factor and session management to reduce risk from compromised credentials.

Core

Scoped API keys

Per-integration keys, permission-limited, revocable, with usage logging.

Core

Security audit

Logs of access, configuration changes, and relevant events to reconstruct what happened.

Performance

Inference logging

History of system queries and responses for operational traceability and usage analysis.

Core

Network and firewall

Port, rule, and external exposure management from the panel, with TLS termination.

06 · Operations

See the node—don't guess.

Telemetry, services, and power: what you need to sustain the system over time.

Core

Hardware HUD

GPU and VRAM per card, power draw, CPU, memory, temperatures, disks, and service status in one view.

Core

Service management

Status and control of system processes, with health checks for each engine.

Performance

Power

Node power consumption and its correlation with workload—direct input for total cost calculation.

Advanced

Internal benchmarks

First-token latency and generation speed measured on the client's actual hardware.

Core

Offline operation

With models installed, local flows keep working without internet egress.

Core

Phased installation

Staged deployment on a clean server, with verification and training at each phase.

Want to see these tools running live?