Models · Turbo series

Xeretron Turbo 35B

Built on Qwen3.6-35B-A3B-MTP and finetuned by Xeretron: expert in architecture and DevOps, runs all kinds of agents and excels at document vision. Our fastest finetune yet — up to 200 t/s depending on GPU.

Architecture DevOps Agents Document vision MTP · Turbo
Try it on my node View specifications
200 t/s up to · GPU-dependent
35Btotal · MoE
3Bactive / token
262Knative context
MTPmulti-token pred.
VLdocument vision

Why it exists

Node speed. Operations brain.

Xeretron Turbo 35B is our fastest finetune: MoE with ~3B active parameters per token and MTP to push throughput. Design architectures, run DevOps, orchestrate agents, and read documents with vision — on your infrastructure, with no token meter.

  • Specialized in architecture and DevOps end to end.
  • Runs all kinds of agents: ops, docs, coding, and orchestration.
  • Expert in document vision: PDFs, screenshots, blueprints, and visual tickets.
  • The fastest finetune we've shipped: up to 200 t/s depending on GPU.

Architecture

Services, boundaries, data, security, and deployment: systems you can operate, not just diagram.

DevOps

Pipelines, infrastructure as code, incidents, runbooks, and automation with local agents.

Agents

Runs all kinds of agents: coding, ops, documents, and multi-tool orchestration.

Document vision

Reads PDFs, screenshots, forms, and diagrams; feeds agents real visual context.

Why Turbo

MoE + MTP = our fastest finetune.

The Qwen3.6-35B-A3B base activates ~3B parameters per token (of 35B total) and includes Multi-Token Prediction. On top of that we tuned for architecture, DevOps, and agents: lower perceived latency, more work finished per second on the node.

  • Up to 200 t/s measured on high-performance GPUs (varies with quantization and context).
  • Ideal when the bottleneck is agent / CI / ops speed.
  • Complements Developer 27B: Turbo pushes throughput; Developer pushes dense coding depth.

Turbo vs Developer series

  • Turbo 35B — MoE + MTP · architecture · DevOps · agents · docs.
  • Developer 27B — dense · CodiiX · deep programming.
  • Both run on Xeretron nodes with an OpenAI-compatible API.

Open base · Qwen3.6-35B-A3B-MTP

We start from an open MoE built for real-world speed.

Qwen3.6-35B-A3B (Alibaba / Qwen) is a multimodal MoE: 35B total and ~3B active per token, 256 experts (8 routed + 1 shared), hybrid Gated DeltaNet + attention architecture, native vision, and Multi-Token Prediction. Ideal for a Turbo finetune aimed at ops and fast agents.

Hybrid architecture

40 layers · layout 10 × (3 × DeltaNet→MoE + 1 × Attention→MoE). Hidden 2048. 35B capacity at the cost of ~3B active.

Thinking control

Thinking on by default in the Qwen3.6 family; adjustable depth so you don't burn thinking tokens when the job is ops.

Native vision

Integrated visual encoder: PDFs, screenshots, forms, diagrams, and tickets. Turbo's strength in document vision.

Tool calling

Tool calling and harnesses (vLLM / SGLang). Built to run agent fleets with environment feedback.

MTP

Multi-step MTP: speculative decoding and more tokens/s when the runtime enables it — the foundation of the Turbo series.

Local deployment

Compatible with Transformers, vLLM, SGLang, and the Xeretron node stack. Apache 2.0 on the author's open line.

Xeretron finetune

From open MoE to the ecosystem's turbo.

On Qwen3.6-35B-A3B-MTP we tuned for architecture, DevOps, agent orchestration, and document reading — with the throughput aggression that defines Turbo.

What the finetune adds

  • Preference for runbooks, pipelines, and actionable architecture decisions.
  • Better alignment with all kinds of agents, tool use, and operations.
  • Clear technical Spanish + frictionless infra/code English.
  • Architecture + DevOps focus: from ticket to deploy, from incident to postmortem.
  • Optimized for personal and team nodes where speed matters.

Specialization

Four fronts. One Turbo.

Turbo 35B is biased toward speed and operations: where the agent needs to finish the job now.

Architecture

System design, trade-offs, contracts, and boundaries. Turbo pushes architecture decisions you can deploy and operate.

DevOps

From requirements to pipelines: infra, CI/CD, incidents, and automation. Fewer one-off scripts, more reliable operations.

Agents

Runs all kinds of agents: planning, tool calling, environment feedback, and long-horizon tasks — at Turbo speed.

Document vision

Documents, screenshots, and forms as agent context. The same GPU running DevOps also reads the scanned paper.

Base reference · Qwen3.6-35B-A3B

What the original work contributes.

Data from the public model card for Qwen3.6-35B-A3B. Xeretron Turbo 35B inherits this MoE+MTP base and specializes it; finetune metrics (including t/s on your GPU) are validated during onboarding.

Attribute (base) Detail Qwen3.6-35B-A3B
ParametersCapacity / cost per token35B total · ~3B active
ExpertsMoE routing256 · 8 routed + 1 shared
LayersHybrid DeltaNet + Attention40
ContextNative / extended262K · up to ~1M (YaRN)
MTPSpeculative decodingMulti-step trained
ModalitiesText + visionImage / video / docs
LicenseAuthor's open lineApache 2.0

Source: public model card for Qwen/Qwen3.6-35B-A3B. The “-MTP” in our naming highlights Multi-Token Prediction in the Turbo stack. 200 t/s are Xeretron measurements on high-performance hardware and depend on GPU, quantization, and context.

Specifications

Technical sheet.

AttributeDetail
NameXeretron Turbo 35B
BaseQwen3.6-35B-A3B-MTP · Qwen/Qwen3.6-35B-A3B (MoE VLM · Apache 2.0)
Parameters35B total · ~3B active per token
Layers40 · layout 10 × (3 × DeltaNet→MoE + 1 × Attention→MoE)
Hidden / experts2048 · 256 experts · intermediate 512
Vocabulary~248K tokens (padded)
Context262,144 native · up to ~1,000,000 with YaRN / rope scaling
ModalitiesText + document vision (image / video / PDF) · tool calling · MTP
SpecialtyArchitecture · DevOps · Agents · Document vision
Runtime XeretronPanel + GPU lanes · OpenAI-compatible API · agents
DeploymentPersonal and team nodes (NVIDIA GPU / CUDA)
Base licenseApache 2.0 — LICENSE / NOTICE

License · Apache 2.0

LICENSE Apache 2.0 NOTICE Turbo

Personal nodes

Fast on real hardware, not just on benches.

The 35B-A3B MoE is built for throughput: few active parameters per token + MTP. With the right quantization it runs strong on workstations and servers. The Xeretron node sizes VRAM, context, and concurrency with you — 200 t/s is the observed ceiling, not a guarantee on every GPU.

  • Guidance: high-end consumer GPU / workstation with enough VRAM for your chosen quantization.
  • Long context (128K+) recommended to preserve thinking on complex tasks.
  • Panel lanes to isolate DevOps agents, document vision, and background loads.
  • Offline with weights installed; updates when you decide.

In the Xeretron panel

Load the Turbo profile, publish the internal endpoint, and point your agents / CI at the node. First-token, tok/s, and model health metrics on the same screen.

Ecosystem

Not just weights. Operational speed.

Throughput

Agents and pipelines that don't wait: more t/s, more correction cycles in the same day.

DevOps & agents

Hooks to CI, runbooks, and local orchestrators via OpenAI-compatible API.

Privacy

Code, prompts, and documents stay on your infrastructure. No endless intelligence rental.

Next step

Put Turbo 35B on your node.

We assess your GPU, context, and whether Turbo or Developer fits better. If your hardware can't reach 200 t/s, we'll tell you with real numbers.