Technology and security

For the technical team: what's inside and what we don't promise.

This page is written for sysadmins, engineers, security architects, and developers. No unnecessary adjectives: engines, interfaces, controls, and known limits.

Public diagram

One platform, multiple specialized engines.

Users and applications · browser · integrations
Xeretron AI Server · FastAPI panel + frontend
LLM / VL models
llama-proxy + llama.cpp
Speech to text
FastWhisper
OCR / documents
File Server
Text to speech
Qwen TTS
Security / users
PAM · TOTP · keys
Observability
health · logs
Private infrastructure · NVIDIA CUDA GPU · Linux · systemd

Public view. We don't publish internal paths, credentials, sensitive network details, or unmitigated vulnerabilities.

Engines and foundation

What runs, what supports it, and where the limit is.

ComponentWhat it doesStated limit
Unified panelDashboard, chat, models, downloads, APIs, inference, security, users, network, firewall, services, and power.We don't claim full enterprise observability.
Hardware telemetryGPU and VRAM, power, CPU/RAM, temperature, disks, and service status.Sensors depend on the physical host.
Local inferenceGGUF models with OpenAI-format-compatible streaming; context, threads, fast attention, KV cache, and tensor split.Not every GGUF runs on every binary or hardware.
Catalog and proxyCatalog, health, load/unload, chat and admin endpoints, published profiles.Depends on weights downloaded beforehand.
Lanesbase, gpu0, gpu1, dual, and cpu to distribute workloads.Concurrency limited by VRAM.
Text and visionVision component with multimodal projection and images via URL or data.Requires mmproj compatible with the model.
ChatSSE streaming, client-side history, transcription, visual attachments, and metrics.History is managed on the client side.
Profiles and VRAMDraft → test → publish cycle and memory estimation before building stacks.The estimate is not an exact guarantee.
Frontier modeSuspends the usual stack, loads a large model, and restores after idle time without cutting active inference.Switching can take several minutes.
Voice a textoLocal transcription with unload on idle.Quality depends on input audio and language.
OCR / documentsPDF, Office, images, and CSV to text, ES/EN.Depends on original scan quality.
Inference APIHTTPS endpoints with scope-limited keys.Compatible only on supported endpoints and structures.
Network / HTTPSTLS termination and controlled external exposure.Network policy is defined by the organization.
Optional modulesGameDev-MCP / Unreal, SIP telephony, sync between nodes, and explicit external providers.Activated per project and agreement.

Security

Concrete controls, not slogans.

We don't say a system is "100% secure" or "impossible to hack." We describe which controls exist and who is responsible for each layer.

Authentication and MFA

Access management with time-based second factor and session administration from the node.

API keys with scopes

Per-integration keys with limited scope, revocable, with usage logging.

Audit

Security and inference logs to reconstruct access, changes, and queries.

TLS and firewall

Encryption in transit and port and rule management from the panel itself.

System and app users

Human and application identities with differentiated roles and least privilege.

Backups and deployment

Agreed backups and phased installation on a clean server, with verification.

Shared responsibility. Xeretron provides technical controls and supported operations; the organization maintains access policy, network management, and compliance with its regulatory framework. We don't claim automatic regulatory compliance.

Indicative requirements

Realistic sizing.

  • Linux OS with NVIDIA drivers and CUDA.
  • 1 to 2 or more GPUs depending on package and target models.
  • RAM and storage based on the planned weight catalog.
  • Administrative access during deployment and a maintenance window.
  • Mandatory evaluation before any commitment.

Technical presales questions

  • What data will be processed? Defines privacy, retention, network, and access.
  • Which models do you need? Defines GPU, RAM, disk, and compatibility.
  • How many users? Defines concurrency and capacity.
  • Context and documents? Defines KV cache, RAM, and OCR.
  • Will you use voice? Defines shared GPU, audio, and storage.
  • Offline operation? Defines packages, support, and updates.
  • Will you integrate applications? Defines API, authentication, TLS, and scopes.
  • What availability do you expect? Defines redundancy, monitoring, and SLA.
Schedule a technical conversation

Want to see the node running before you decide?

30-minute demo, in person or remote, with a technical lead on your side if you want.