Hybrid architecture
64 layers in a 3:1 pattern — Gated DeltaNet (linear attention) + full Gated Attention. Hidden size 5120, FFN 17408. Real density on every token.
Models · Developer series
Model built on Qwen3.8-27B and fine-tuned for the Xeretron ecosystem: CodiiX (our IDE for vibe coders), architectures, apps, and agents. The same quality you expect from cloud providers — with local power on your personal node.
Why it exists
Xeretron Developer 27B is not a generic chat in a box. It is the ecosystem's work brain: understands repos, designs systems, orchestrates agents, and answers with the depth you associate with paid APIs — without sending your code to third parties.
Refactor, feature generation, tests, and repo navigation with the model on your machine.
Mental diagrams turned into modules, API contracts, and traceable technical decisions.
Multi-step planning, tool calling, and reliable execution loops in a local environment.
Documents, analysis, research, and long-horizon office tasks on the same node.
In action · CodiiX
Real capture: CodiiX agent with xeretron-developer-27b loaded, exploring the repo and applying changes — no third-party tokens.
Open base · Qwen3.8-27B
Qwen3.8-27B (Alibaba / Qwen) is a 27B dense multimodal model with hybrid Gated DeltaNet + attention architecture, Multi-Token Prediction, and flexible reasoning control. Strong in agentic coding, professional work, and long-horizon tasks — the ideal starting point for a Developer fine-tune.
64 layers in a 3:1 pattern — Gated DeltaNet (linear attention) + full Gated Attention. Hidden size 5120, FFN 17408. Real density on every token.
Thinking mode by default, with reasoning_effort (xhigh / medium / low) and preserve_thinking for multi-turn agents.
Integrated visual encoder: diagrams, UI captures, documents, and video. Ideal for CodiiX when context is a screen, not just a file.
Official parsers for tools and coding harnesses (vLLM / SGLang). Ready for agent loops with environment feedback.
Multi-Token Prediction trained over multiple steps: more throughput in local inference when the runtime uses it.
Compatible with Transformers, vLLM, SGLang, and Xeretron node GGUF runtimes. Open weights; operation under your control.
Xeretron fine-tune
On top of Qwen3.8 we tune behavior to speak Xeretron: panels, lanes, model profiles, OpenAI-compatible APIs, CodiiX, and real vibe-coding flows — not lab demos.
Specialization
Developer 27B is deliberately biased toward the work that hurts most when you pay per token.
Repo-level generation, refactors, debugging, tests, and agentic terminal. The Qwen3.8 base already scores strong on SWE-bench Pro, Terminal Bench, and competitive coding; the fine-tune pushes that muscle toward CodiiX and your stack.
From requirements to modules: service boundaries, contracts, data, security, and deployment. Less “loose code”, more systems you can operate.
Planning, tool calling, environment feedback, and long-horizon tasks. Built for loops that finish the job, not just the next paragraph.
Office and knowledge work: documents, research, analysis, and multi-step flows. The same GPU that codes also works the rest of the day.
Base reference · Qwen3.8-27B
Figures reported by Qwen on the Qwen3.8-27B model card. Xeretron Developer 27B inherits this base and specializes it; fine-tune scores are published when each internal evaluation closes.
| Benchmark | Area | Qwen3.8-27B |
|---|---|---|
| SWE-bench Pro | Agentic coding | 61.7 |
| Terminal Bench 2.1 | Agentic terminal | 73.0 |
| QwenSWEBench | Software engineering | 79.0 |
| LiveCodeBench v6 | Coding competitivo | 90.3 |
| NL2Repo-Bench | Repo-level codegen | 42.3 |
| CoWorkBench | Long-horizon work | 70.7 |
| GPQA Diamond | Scientific reasoning | 89.2 |
| IFBench | Instruction following | 79.5 |
| OSWorld-Verified | Computer use | 84.3 |
| OmniDocBench 1.5 | Documents / vision | 91.1 |
Source: public model card for Qwen/Qwen3.8-27B. Harnesses and protocols are as declared by the model author. Developer 27B does not replace evaluation on your hardware: we measure latency and quality on the real node during onboarding.
Specifications
| Attribute | Detail |
|---|---|
| Name | Xeretron Developer 27B |
| Base | Qwen3.8-27B · Qwen/Qwen3.8-27B (dense VLM · Apache 2.0) |
| Parameters | 27B dense active |
| Layers | 64 · layout 16 × (3 × DeltaNet→FFN + 1 × Attention→FFN) |
| Hidden / FFN | 5120 / 17408 |
| Vocabulary | ~248K tokens (padded) |
| Context | 262,144 native · up to ~1,000,000 with YaRN / rope scaling |
| Modalities | Text + vision (images / video) · tool calling |
| Specialty | Work · Programming · Agents · Architecture |
| Xeretron runtime | Panel + GPU lanes · OpenAI-compatible API · CodiiX |
| Deployment | Personal and team nodes (NVIDIA GPU / CUDA) |
| Base license | Apache 2.0 — LICENSE / NOTICE |
License · Apache 2.0
Personal nodes
The Qwen3.8 dense 27B family is built to be deployment-friendly: with quantization (e.g. FP8 / GGUF) it runs on workstations and one- or few-GPU servers. The Xeretron node sizes VRAM, context, and concurrency with you — no magic promises.
Load the Developer profile, publish the internal endpoint, and point CodiiX at the node. First-token, tok/s, and model health metrics on the same screen as the rest of the stack.
Ecosystem
IDE for vibe coders connected to the local model: iterate without fear of the token meter.
Try in the panel or consume from your apps with scoped keys.
Code, prompts, and documents stay on your infrastructure. No eternal intelligence rental.
Next step
We evaluate your GPU, the context you need, and your CodiiX flow. If hardware is not enough, we tell you before install.