Hellish bugs
Heisenbugs, race conditions, impossible states, refactors that break at the edge. Titan holds the thread to root cause and a verifiable fix.
Titan series · Perfection flagship FRONTIER
The king of perfection. Built on Qwen3.8-Flash-Next — a preview of the architecture on the path to Qwen4 — and finetuned by Xeretron. Intelligence multiplied vs 27B: solves hellish bugs, controls MCPs with precision, creates and manipulates canvas and 3D objects, and delivers builds and landings with fine polish.
*Xeretron product positioning: perceived depth and finish quality vs Developer 27B on complex tasks (bugs, agents, UI). Not a single-benchmark score.
Titan strengths
Not “another bigger chat.” The model that closes the job with product quality.
Heisenbugs, race conditions, impossible states, refactors that break at the edge. Titan holds the thread to root cause and a verifiable fix.
Inherits reasoning_effort: low, medium, xhigh. Scale thinking to the ticket — don't burn depth when you don't need it.
Create and manipulate canvas, interactive layouts, and 3D objects: visual ideation that becomes an artifact, not a paragraph.
Tool calling and MCP orchestration with discipline: fewer broken loops, more correct calls to the right server, in the right order.
Builds that go to the edge: empty states, errors, typography, responsive layout, basic accessibility, and “product detail,” not just the happy path.
Landing pages with atmosphere, hierarchy, and micro-detail. UI that feels designed, not half-generated.
Developer 27B · dense precision
Titan 125B · king of perfection
Base · Qwen3.8-Flash-Next
Published as open weights on the Qwen/Qwen3.8-Flash-Next line (Qwen family). Not “more of the same”: introduces Qwen Sparse Attention, Gated Residual, and N-gram Embedding to scale intelligence efficiently.
Attention at micro-block level (not token-by-token). Lower latency on long contexts — critical for agents and deep debug sessions.
Residuals with read/write gates: more expressivity between layers without inflating inference cost.
Offloadable parameter axis: scales capacity without the compute cost of an equivalent MoE. Ideal for memory-tight nodes.
10 routed + 1 shared active · ~6B per token. Multi-step MTP to push throughput when the runtime enables it.
Base reference
Figures from the Qwen3.8-Flash-Next model card. Titan inherits this base and specializes it for the Xeretron ecosystem; local t/s are our own measurements.
| Benchmark | Qwen3.8-27B | Flash-Next (base Titan) |
|---|---|---|
| DeepSWE 1.1 | 42.2 | 58.7 |
| SWE-bench Pro | 61.7 | 62.5 |
| SWE-bench Multilingual | 73.8 | 81.0 |
| JobBench | 33.4 | 55.7 |
| Toolathlon Verified | 67.1 | 73.5 |
| GPQA Diamond | 89.2 | 91.7 |
| LiveCodeBench v6 | 90.3 | 91.9 |
| Vision2Web | 62.9 | 64.0 |
| ClawEval-MM Pass@3 | 57.4 | 64.4 |
Source: public model card for Qwen/Qwen3.8-Flash-Next. Harnesses are as declared by the author. Titan doesn't replace evaluation on your hardware: we measure quality and tok/s on the real node.
Specifications
| Attribute | Detail |
|---|---|
| Nombre | Xeretron Titan 125B |
| Base | Qwen3.8-Flash-Next · Qwen/Qwen3.8-Flash-Next |
| Parameters | 125B backbone · ~6B active · +51B n-gram · +MTP (~4B) |
| Layers / MoE | 48 layers · 512 experts · 10 routed + 1 shared |
| Architecture | Gated DeltaNet + Qwen Sparse Attention · Gated Residual |
| Context | 262,144 native · up to ~1,000,000 (YaRN) |
| Modalities | Text · image · video · tool / MCP |
| Intelligence | 3 levels: low · medium · xhigh (reasoning_effort) |
| Xeretron specialty | Product perfection · bugs · canvas/3D · MCP · landings |
| Local throughput | Up to ~40 t/s on 2× RTX 3090 (Xeretron measurement; varies with quantization/context) |
| Base license | Qwen Community License 1.0 — see distribution notice |
License · Qwen Community 1.0
In compliance with the Qwen Community License 1.0 of the base model (Qwen3.8-Flash-Next), Xeretron Titan 125B is shared for whatever use you want on your Xeretron node. We don't charge for the model: what you evaluate or buy is infrastructure, the panel, and support — not a Titan weights license.
On your node
We run it on real infrastructure: up to 40 tokens/s with 2× RTX 3090. The Xeretron node sizes quantization, context, and lanes so Titan is usable — not just “loadable”.
Next step
Titan when the deliverable can't look “almost done.” Turbo when you need speed. Developer when you iterate all day in CodiiX.
The model is shared under Qwen Community License 1.0 for whatever use you want on your node; we don't charge for Titan. License notice.