Titan series · Perfection flagship FRONTIER

Xeretron Titan 125B

The king of perfection. Built on Qwen3.8-Flash-Next — a preview of the architecture on the path to Qwen4 — and finetuned by Xeretron. Intelligence multiplied vs 27B: solves hellish bugs, controls MCPs with precision, creates and manipulates canvas and 3D objects, and delivers builds and landings with fine polish.

FRONTIER Hellish bugs 3 intel. levels Canvas · 3D MCP mastery Total polish
I want Titan on my node See all strengths
40 t/s up to · 2× RTX 3090
125Bbackbone MoE
6Bactive / token
51Bn-gram emb.
×10vs Developer 27B*
3intel. levels

*Xeretron product positioning: perceived depth and finish quality vs Developer 27B on complex tasks (bugs, agents, UI). Not a single-benchmark score.

Titan strengths

Where others stall, Titan finishes.

Not “another bigger chat.” The model that closes the job with product quality.

Hellish bugs

Heisenbugs, race conditions, impossible states, refactors that break at the edge. Titan holds the thread to root cause and a verifiable fix.

3 intelligence levels

Inherits reasoning_effort: low, medium, xhigh. Scale thinking to the ticket — don't burn depth when you don't need it.

Canvas & 3D objects

Create and manipulate canvas, interactive layouts, and 3D objects: visual ideation that becomes an artifact, not a paragraph.

MCPs under control

Tool calling and MCP orchestration with discipline: fewer broken loops, more correct calls to the right server, in the right order.

Complete polish

Builds that go to the edge: empty states, errors, typography, responsive layout, basic accessibility, and “product detail,” not just the happy path.

Refined landings

Landing pages with atmosphere, hierarchy, and micro-detail. UI that feels designed, not half-generated.

Developer 27B · dense precision

The CodiiX craft

  • 27B dense · daily vibe coding
  • Excellent quality / VRAM ratio
  • ~70 t/s depending on GPU
  • Ideal for continuous iteration

Titan 125B · king of perfection

When the bar goes ×10

  • 125B MoE · 6B active + n-gram
  • Hard bugs · MCP · canvas · 3D
  • ~40 t/s on 2× RTX 3090
  • Ideal for “no patchwork” deliverables

Base · Qwen3.8-Flash-Next

A preview of the architecture on the path to Qwen4.

Published as open weights on the Qwen/Qwen3.8-Flash-Next line (Qwen family). Not “more of the same”: introduces Qwen Sparse Attention, Gated Residual, and N-gram Embedding to scale intelligence efficiently.

Qwen Sparse Attention

Attention at micro-block level (not token-by-token). Lower latency on long contexts — critical for agents and deep debug sessions.

Gated Residual

Residuals with read/write gates: more expressivity between layers without inflating inference cost.

N-gram Embedding (51B)

Offloadable parameter axis: scales capacity without the compute cost of an equivalent MoE. Ideal for memory-tight nodes.

MoE 512 experts · MTP

10 routed + 1 shared active · ~6B per token. Multi-step MTP to push throughput when the runtime enables it.

Base reference

Public numbers vs Qwen3.8-27B.

Figures from the Qwen3.8-Flash-Next model card. Titan inherits this base and specializes it for the Xeretron ecosystem; local t/s are our own measurements.

Benchmark Qwen3.8-27B Flash-Next (base Titan)
DeepSWE 1.142.258.7
SWE-bench Pro61.762.5
SWE-bench Multilingual73.881.0
JobBench33.455.7
Toolathlon Verified67.173.5
GPQA Diamond89.291.7
LiveCodeBench v690.391.9
Vision2Web62.964.0
ClawEval-MM Pass@357.464.4

Source: public model card for Qwen/Qwen3.8-Flash-Next. Harnesses are as declared by the author. Titan doesn't replace evaluation on your hardware: we measure quality and tok/s on the real node.

Specifications

Titan technical sheet.

AttributeDetail
NombreXeretron Titan 125B
BaseQwen3.8-Flash-Next · Qwen/Qwen3.8-Flash-Next
Parameters125B backbone · ~6B active · +51B n-gram · +MTP (~4B)
Layers / MoE48 layers · 512 experts · 10 routed + 1 shared
ArchitectureGated DeltaNet + Qwen Sparse Attention · Gated Residual
Context262,144 native · up to ~1,000,000 (YaRN)
ModalitiesText · image · video · tool / MCP
Intelligence3 levels: low · medium · xhigh (reasoning_effort)
Xeretron specialtyProduct perfection · bugs · canvas/3D · MCP · landings
Local throughputUp to ~40 t/s on 2× RTX 3090 (Xeretron measurement; varies with quantization/context)
Base licenseQwen Community License 1.0 — see distribution notice

License · Qwen Community 1.0

How we share it at Xeretron

In compliance with the Qwen Community License 1.0 of the base model (Qwen3.8-Flash-Next), Xeretron Titan 125B is shared for whatever use you want on your Xeretron node. We don't charge for the model: what you evaluate or buy is infrastructure, the panel, and support — not a Titan weights license.

  • Finetune / derivative delivered with NOTICE and Qwen license notice.
  • Use in your environment: you decide what the model is for on your node.
  • We don't sell Titan as a separate weights product or as public Qwen MaaS.
  • You must comply with Qwen Community License 1.0 in your own use; if your case is MaaS or commercial AI Work Assistant at scale, review terms with Qwen.

LICENSE Qwen Community NOTICE Titan

On your node

Runs locally. Tested by Xeretron on two RTX 3090 GPUs.

We run it on real infrastructure: up to 40 tokens/s with 2× RTX 3090. The Xeretron node sizes quantization, context, and lanes so Titan is usable — not just “loadable”.

  • Dual high-end consumer GPUs as a practical reference.
  • N-gram offload + active MoE: more intelligence per watt of attention.
  • OpenAI-compatible API for CodiiX, agents, and MCPs.
  • Pre-evaluation: if your VRAM can't hit the desired profile, we'll tell you with numbers.

Recommended stack

  • 2× RTX 3090 (o superior) · CUDA
  • Titan profile in the Xeretron panel
  • reasoning_effort by task
  • MCPs + CodiiX / agents on the same node

Next step

Install the king. Keep the perfection.

Titan when the deliverable can't look “almost done.” Turbo when you need speed. Developer when you iterate all day in CodiiX.

The model is shared under Qwen Community License 1.0 for whatever use you want on your node; we don't charge for Titan. License notice.