Omega series · Brain of Brains FRONTIER

Xeretron Omega 320B

The brain that directs the other brains. Built on GLM-5.3-Flash (zai-org/GLM-5.3-Flash · Z.ai) — native multimodal ~320B with ~18B active — and finetuned by Xeretron. Competes with Frontier models: top-tier architectures, impossible bugs, device control, and agent orchestration.

FRONTIER Architecture ×10 Impossible bugs Devices Agent control
I want Omega on my node See all strengths
18 t/s up to · 2× RTX 3090
320BMoE total
18Bactive / token
1Mnative context
MITbase license
3reasoning levels

18 t/s: Xeretron measurement on a node with 2× RTX 3090. Varies with quantization, context, and reasoning_effort. Base is close to Claude Opus 4.8 on public coding/agentic benches (Z.ai) — not guaranteed equivalence on your node.

Omega strengths

When the problem is frontier-grade, Omega thinks.

Not a bigger chat. The model that designs systems, directs agents, and touches the real world.

Top-tier architectures

System design, domain boundaries, service contracts, and trade-offs that hold in production. The blueprint before code — and code that respects the blueprint.

Impossible bugs

Multi-layer failures, heisenbugs, states that “shouldn't exist.” Omega holds long context (up to 1M) and deep reasoning to a verifiable root cause.

Device control

Desktop UI, terminals, and operating environments: the base excels on OSWorld and Terminal-Bench. Ideal for automating real machines, not just APIs.

Agent control

Orchestrate agent fleets, tools, and MCPs with discipline: Toolathlon, AutomationBench, and Agents’ Last Exam in the base's public reference.

Native multimodal

Text, image, and video in the same loop — not a “vision bolt-on.” Documents, charts, screens, and visual evidence in the reasoning.

3 intelligence levels

Inherits reasoning_effort: low, high, max (default max). Scale thinking to the mission.

Native 1M context

Native window of 1,048,576 tokens: whole repos, long traces, multi-hour agent sessions without “losing the thread” to improvised scaling.

Tool calling & MCP

Tools, APIs, and MCPs as first-class citizens. Fewer broken loops; more correct calls to the right server.

MIT License

Base under MIT from Z.AI: commercial use, modification, and redistribution with NOTICE. Frontier intelligence without the Community License maze.

Research, math & academia

Ready for research, complex math, and demanding computation. Fits universities, schools, and labs: deep reasoning locally, without sending data to the cloud.

Titan 125B · king of perfection

Product polish

  • 125B · 6B active · Qwen Flash-Next
  • Canvas, landings, fine-tuned MCP
  • ~40 t/s on 2× RTX 3090
  • Ideal for “no patchwork” deliverables

Omega 320B · brain of brains

When the bar is Frontier

  • 320B · 18B active · GLM-5.3-Flash
  • Architecture · agents · devices
  • ~18 t/s on 2× RTX 3090 · MIT
  • Ideal for directing fleets and systems

Base · GLM-5.3-Flash

First native multimodal in the GLM-5 series.

Open weights via zai-org/GLM-5.3-Flash (Z.ai). Architecture and recipe redesigned around capacity and efficiency: more intelligence per compute, affordable long context, agents that endure.

Hybrid sparse + linear attention

First time in the GLM series: cuts the cost of serving long contexts without giving up precision. Critical for agent traces and huge repos.

Manifold-Constrained Hyper-Connections

mHC to scale inter-layer connections more efficiently — more useful capacity without inflating inference cost.

MoE ~320B · 18B active

Frontier intelligence with selective activation: thinking density without paying 320B dense on every token.

~30T multimodal corpus

Pretrained on a massive multimodal corpus: vision and text aren't “add-ons,” they're the same brain.

Base reference

Public numbers: close to Frontier.

Figures from the Z.ai model card / blog for GLM-5.3-Flash. Omega inherits this base and specializes it for the Xeretron ecosystem.

Benchmark GLM-5.2 GLM-5.3-Flash (base Omega) Opus 4.8*
Terminal-Bench 2.181.084.385.0
DeepSWE v1.146.263.458.0
NL2Repo48.956.369.7
Toolathlon Verified59.978.476.2
AutomationBench26.248.841.0
Agents’ Last Exam20.426.327.0
HLE w/ tools54.755.357.9
GDPval-AA v2150417731582
OSWorld 2.0—59.154.8
OfficeQA Pro—62.448.9
CharXiv Reasoning—89.489.9
Chartography—78.075.0
Vision2Web—77.876.1
MMVU—80.567.4

Source: Z.ai · GLM-5.3-Flash / model card. *Claude Opus 4.8 per the author's table. Harnesses are as declared there. Omega doesn't replace evaluation on your hardware.

Specifications

Omega technical sheet.

AttributeDetail
NombreXeretron Omega 320B
BaseGLM-5.3-Flash · zai-org/GLM-5.3-Flash
Parameters~320B / 321B total · ~18B active per token (MoE)
ArchitectureHybrid sparse + linear attention · mHC · native multimodal
Context1,048,576 tokens native · recommended output up to ~131K
ModalitiesText · image · video · tool / MCP / agents
Intelligencereasoning_effort: low · high · max (default max)
Xeretron specialtyBrain of brains · architecture · bugs · devices · agents
Local throughputUp to ~18 t/s on 2× RTX 3090 (Xeretron measurement; varies with quantization/context)
Reference node2× RTX 3090 (24 GB each) · CPU 24 threads / 12 physical · ~92 GB system RAM
Base licenseMIT (Z.AI) — LICENSE / NOTICE

License · MIT

LICENSE MIT NOTICE Omega

On your node

Runs locally. Tested by Xeretron on two RTX 3090 GPUs.

We run it on real infrastructure: up to 18 tokens/s with 2× RTX 3090 (24 GB VRAM each), a 24-thread CPU, and ~92 GB system RAM. The Xeretron node sizes quantization, context, and lanes so Omega is usable — not just “loadable”.

  • Dual RTX 3090 as a practical Frontier reference.
  • ~48 GB total VRAM + system RAM for offload.
  • OpenAI-compatible API for CodiiX, agents, and MCPs.
  • Pre-evaluation: if your profile can't reach it, we'll tell you with numbers.

Xeretron reference node

  • GPU 0 · RTX 3090 · 24 GB · up to 330 W
  • GPU 1 · RTX 3090 · 24 GB · up to 240 W
  • CPU 24 threads (12 physical) · ~92 GB RAM
  • Up to ~18 t/s · Omega profile in the panel

Next step

Install the brain of brains.

Omega for Frontier and orchestration. Titan for product perfection. Turbo for speed. Developer for daily craft in CodiiX.