Top-tier architectures
System design, domain boundaries, service contracts, and trade-offs that hold in production. The blueprint before code — and code that respects the blueprint.
Omega series · Brain of Brains FRONTIER
The brain that directs the other brains. Built on GLM-5.3-Flash (zai-org/GLM-5.3-Flash · Z.ai) — native multimodal ~320B with ~18B active — and finetuned by Xeretron. Competes with Frontier models: top-tier architectures, impossible bugs, device control, and agent orchestration.
18 t/s: Xeretron measurement on a node with 2× RTX 3090. Varies with quantization, context, and reasoning_effort. Base is close to Claude Opus 4.8 on public coding/agentic benches (Z.ai) — not guaranteed equivalence on your node.
Omega strengths
Not a bigger chat. The model that designs systems, directs agents, and touches the real world.
System design, domain boundaries, service contracts, and trade-offs that hold in production. The blueprint before code — and code that respects the blueprint.
Multi-layer failures, heisenbugs, states that “shouldn't exist.” Omega holds long context (up to 1M) and deep reasoning to a verifiable root cause.
Desktop UI, terminals, and operating environments: the base excels on OSWorld and Terminal-Bench. Ideal for automating real machines, not just APIs.
Orchestrate agent fleets, tools, and MCPs with discipline: Toolathlon, AutomationBench, and Agents’ Last Exam in the base's public reference.
Text, image, and video in the same loop — not a “vision bolt-on.” Documents, charts, screens, and visual evidence in the reasoning.
Inherits reasoning_effort: low, high, max (default max). Scale thinking to the mission.
Native window of 1,048,576 tokens: whole repos, long traces, multi-hour agent sessions without “losing the thread” to improvised scaling.
Tools, APIs, and MCPs as first-class citizens. Fewer broken loops; more correct calls to the right server.
Base under MIT from Z.AI: commercial use, modification, and redistribution with NOTICE. Frontier intelligence without the Community License maze.
Ready for research, complex math, and demanding computation. Fits universities, schools, and labs: deep reasoning locally, without sending data to the cloud.
Titan 125B · king of perfection
Omega 320B · brain of brains
Base · GLM-5.3-Flash
Open weights via zai-org/GLM-5.3-Flash (Z.ai). Architecture and recipe redesigned around capacity and efficiency: more intelligence per compute, affordable long context, agents that endure.
First time in the GLM series: cuts the cost of serving long contexts without giving up precision. Critical for agent traces and huge repos.
mHC to scale inter-layer connections more efficiently — more useful capacity without inflating inference cost.
Frontier intelligence with selective activation: thinking density without paying 320B dense on every token.
Pretrained on a massive multimodal corpus: vision and text aren't “add-ons,” they're the same brain.
Base reference
Figures from the Z.ai model card / blog for GLM-5.3-Flash. Omega inherits this base and specializes it for the Xeretron ecosystem.
| Benchmark | GLM-5.2 | GLM-5.3-Flash (base Omega) | Opus 4.8* |
|---|---|---|---|
| Terminal-Bench 2.1 | 81.0 | 84.3 | 85.0 |
| DeepSWE v1.1 | 46.2 | 63.4 | 58.0 |
| NL2Repo | 48.9 | 56.3 | 69.7 |
| Toolathlon Verified | 59.9 | 78.4 | 76.2 |
| AutomationBench | 26.2 | 48.8 | 41.0 |
| Agents’ Last Exam | 20.4 | 26.3 | 27.0 |
| HLE w/ tools | 54.7 | 55.3 | 57.9 |
| GDPval-AA v2 | 1504 | 1773 | 1582 |
| OSWorld 2.0 | — | 59.1 | 54.8 |
| OfficeQA Pro | — | 62.4 | 48.9 |
| CharXiv Reasoning | — | 89.4 | 89.9 |
| Chartography | — | 78.0 | 75.0 |
| Vision2Web | — | 77.8 | 76.1 |
| MMVU | — | 80.5 | 67.4 |
Source: Z.ai · GLM-5.3-Flash / model card. *Claude Opus 4.8 per the author's table. Harnesses are as declared there. Omega doesn't replace evaluation on your hardware.
Specifications
| Attribute | Detail |
|---|---|
| Nombre | Xeretron Omega 320B |
| Base | GLM-5.3-Flash · zai-org/GLM-5.3-Flash |
| Parameters | ~320B / 321B total · ~18B active per token (MoE) |
| Architecture | Hybrid sparse + linear attention · mHC · native multimodal |
| Context | 1,048,576 tokens native · recommended output up to ~131K |
| Modalities | Text · image · video · tool / MCP / agents |
| Intelligence | reasoning_effort: low · high · max (default max) |
| Xeretron specialty | Brain of brains · architecture · bugs · devices · agents |
| Local throughput | Up to ~18 t/s on 2× RTX 3090 (Xeretron measurement; varies with quantization/context) |
| Reference node | 2× RTX 3090 (24 GB each) · CPU 24 threads / 12 physical · ~92 GB system RAM |
| Base license | MIT (Z.AI) — LICENSE / NOTICE |
License · MIT
On your node
We run it on real infrastructure: up to 18 tokens/s with 2× RTX 3090 (24 GB VRAM each), a 24-thread CPU, and ~92 GB system RAM. The Xeretron node sizes quantization, context, and lanes so Omega is usable — not just “loadable”.
Next step
Omega for Frontier and orchestration. Titan for product perfection. Turbo for speed. Developer for daily craft in CodiiX.