Hybrid architecture
40 layers · layout 10 × (3 × DeltaNet→MoE + 1 × Attention→MoE). Hidden 2048. 35B capacity at the cost of ~3B active.
Models · Turbo series
Built on Qwen3.6-35B-A3B-MTP and finetuned by Xeretron: expert in architecture and DevOps, runs all kinds of agents and excels at document vision. Our fastest finetune yet — up to 200 t/s depending on GPU.
Why it exists
Xeretron Turbo 35B is our fastest finetune: MoE with ~3B active parameters per token and MTP to push throughput. Design architectures, run DevOps, orchestrate agents, and read documents with vision — on your infrastructure, with no token meter.
Services, boundaries, data, security, and deployment: systems you can operate, not just diagram.
Pipelines, infrastructure as code, incidents, runbooks, and automation with local agents.
Runs all kinds of agents: coding, ops, documents, and multi-tool orchestration.
Reads PDFs, screenshots, forms, and diagrams; feeds agents real visual context.
Why Turbo
The Qwen3.6-35B-A3B base activates ~3B parameters per token (of 35B total) and includes Multi-Token Prediction. On top of that we tuned for architecture, DevOps, and agents: lower perceived latency, more work finished per second on the node.
Open base · Qwen3.6-35B-A3B-MTP
Qwen3.6-35B-A3B (Alibaba / Qwen) is a multimodal MoE: 35B total and ~3B active per token, 256 experts (8 routed + 1 shared), hybrid Gated DeltaNet + attention architecture, native vision, and Multi-Token Prediction. Ideal for a Turbo finetune aimed at ops and fast agents.
40 layers · layout 10 × (3 × DeltaNet→MoE + 1 × Attention→MoE). Hidden 2048. 35B capacity at the cost of ~3B active.
Thinking on by default in the Qwen3.6 family; adjustable depth so you don't burn thinking tokens when the job is ops.
Integrated visual encoder: PDFs, screenshots, forms, diagrams, and tickets. Turbo's strength in document vision.
Tool calling and harnesses (vLLM / SGLang). Built to run agent fleets with environment feedback.
Multi-step MTP: speculative decoding and more tokens/s when the runtime enables it — the foundation of the Turbo series.
Compatible with Transformers, vLLM, SGLang, and the Xeretron node stack. Apache 2.0 on the author's open line.
Xeretron finetune
On Qwen3.6-35B-A3B-MTP we tuned for architecture, DevOps, agent orchestration, and document reading — with the throughput aggression that defines Turbo.
Specialization
Turbo 35B is biased toward speed and operations: where the agent needs to finish the job now.
System design, trade-offs, contracts, and boundaries. Turbo pushes architecture decisions you can deploy and operate.
From requirements to pipelines: infra, CI/CD, incidents, and automation. Fewer one-off scripts, more reliable operations.
Runs all kinds of agents: planning, tool calling, environment feedback, and long-horizon tasks — at Turbo speed.
Documents, screenshots, and forms as agent context. The same GPU running DevOps also reads the scanned paper.
Base reference · Qwen3.6-35B-A3B
Data from the public model card for Qwen3.6-35B-A3B. Xeretron Turbo 35B inherits this MoE+MTP base and specializes it; finetune metrics (including t/s on your GPU) are validated during onboarding.
| Attribute (base) | Detail | Qwen3.6-35B-A3B |
|---|---|---|
| Parameters | Capacity / cost per token | 35B total · ~3B active |
| Experts | MoE routing | 256 · 8 routed + 1 shared |
| Layers | Hybrid DeltaNet + Attention | 40 |
| Context | Native / extended | 262K · up to ~1M (YaRN) |
| MTP | Speculative decoding | Multi-step trained |
| Modalities | Text + vision | Image / video / docs |
| License | Author's open line | Apache 2.0 |
Source: public model card for Qwen/Qwen3.6-35B-A3B. The “-MTP” in our naming highlights Multi-Token Prediction in the Turbo stack. 200 t/s are Xeretron measurements on high-performance hardware and depend on GPU, quantization, and context.
Specifications
| Attribute | Detail |
|---|---|
| Name | Xeretron Turbo 35B |
| Base | Qwen3.6-35B-A3B-MTP · Qwen/Qwen3.6-35B-A3B (MoE VLM · Apache 2.0) |
| Parameters | 35B total · ~3B active per token |
| Layers | 40 · layout 10 × (3 × DeltaNet→MoE + 1 × Attention→MoE) |
| Hidden / experts | 2048 · 256 experts · intermediate 512 |
| Vocabulary | ~248K tokens (padded) |
| Context | 262,144 native · up to ~1,000,000 with YaRN / rope scaling |
| Modalities | Text + document vision (image / video / PDF) · tool calling · MTP |
| Specialty | Architecture · DevOps · Agents · Document vision |
| Runtime Xeretron | Panel + GPU lanes · OpenAI-compatible API · agents |
| Deployment | Personal and team nodes (NVIDIA GPU / CUDA) |
| Base license | Apache 2.0 — LICENSE / NOTICE |
License · Apache 2.0
Personal nodes
The 35B-A3B MoE is built for throughput: few active parameters per token + MTP. With the right quantization it runs strong on workstations and servers. The Xeretron node sizes VRAM, context, and concurrency with you — 200 t/s is the observed ceiling, not a guarantee on every GPU.
Load the Turbo profile, publish the internal endpoint, and point your agents / CI at the node. First-token, tok/s, and model health metrics on the same screen.
Ecosystem
Agents and pipelines that don't wait: more t/s, more correction cycles in the same day.
Hooks to CI, runbooks, and local orchestrators via OpenAI-compatible API.
Code, prompts, and documents stay on your infrastructure. No endless intelligence rental.
Next step
We assess your GPU, context, and whether Turbo or Developer fits better. If your hardware can't reach 200 t/s, we'll tell you with real numbers.