Models · Developer series

Xeretron Developer 27B

Model built on Qwen3.8-27B and fine-tuned for the Xeretron ecosystem: CodiiX (our IDE for vibe coders), architectures, apps, and agents. The same quality you expect from cloud providers — with local power on your personal node.

Work Programming Agents Architecture Personal nodes
Try on my node View specifications
70 t/s up to · GPU dependent
27Bdense parameters
262Knative context
1Mextended context
64hybrid layers
VLvision + text

Why it exists

A development engine that lives with you.

Xeretron Developer 27B is not a generic chat in a box. It is the ecosystem's work brain: understands repos, designs systems, orchestrates agents, and answers with the depth you associate with paid APIs — without sending your code to third parties.

  • Integrated with CodiiX, the Xeretron IDE for vibe coding.
  • Built for architectures, full apps, and agentic flows.
  • Runs on personal or team nodes: your GPU, your network, your policy.
  • Open base weights + Xeretron fine-tune for the full stack.

CodiiX

Refactor, feature generation, tests, and repo navigation with the model on your machine.

Architecture

Mental diagrams turned into modules, API contracts, and traceable technical decisions.

Agents

Multi-step planning, tool calling, and reliable execution loops in a local environment.

Work

Documents, analysis, research, and long-horizon office tasks on the same node.

In action · CodiiX

Developer 27B inside the IDE.

Real capture: CodiiX agent with xeretron-developer-27b loaded, exploring the repo and applying changes — no third-party tokens.

CodiiX with Xeretron Developer 27B: explorer, editor, and agent panel on a real project
Click to enlarge

Open base · Qwen3.8-27B

We start from one of the most capable dense open models.

Qwen3.8-27B (Alibaba / Qwen) is a 27B dense multimodal model with hybrid Gated DeltaNet + attention architecture, Multi-Token Prediction, and flexible reasoning control. Strong in agentic coding, professional work, and long-horizon tasks — the ideal starting point for a Developer fine-tune.

Hybrid architecture

64 layers in a 3:1 pattern — Gated DeltaNet (linear attention) + full Gated Attention. Hidden size 5120, FFN 17408. Real density on every token.

Thinking control

Thinking mode by default, with reasoning_effort (xhigh / medium / low) and preserve_thinking for multi-turn agents.

Native vision

Integrated visual encoder: diagrams, UI captures, documents, and video. Ideal for CodiiX when context is a screen, not just a file.

Tool calling

Official parsers for tools and coding harnesses (vLLM / SGLang). Ready for agent loops with environment feedback.

MTP

Multi-Token Prediction trained over multiple steps: more throughput in local inference when the runtime uses it.

Local deployment

Compatible with Transformers, vLLM, SGLang, and Xeretron node GGUF runtimes. Open weights; operation under your control.

Xeretron fine-tune

From open base to ecosystem brain.

On top of Qwen3.8 we tune behavior to speak Xeretron: panels, lanes, model profiles, OpenAI-compatible APIs, CodiiX, and real vibe-coding flows — not lab demos.

What the fine-tune adds

  • Preference for actionable solutions in CodiiX and the panel.
  • Better alignment with agents, tool use, and software architecture.
  • Work style for LatAm: clear technical Spanish without losing code English.
  • Work + Programming focus: from ticket to PR, from brief to app.
  • Optimized to run on personal nodes, not just clusters.

Specialization

Four fronts. One model.

Developer 27B is deliberately biased toward the work that hurts most when you pay per token.

Programming

Repo-level generation, refactors, debugging, tests, and agentic terminal. The Qwen3.8 base already scores strong on SWE-bench Pro, Terminal Bench, and competitive coding; the fine-tune pushes that muscle toward CodiiX and your stack.

Architecture

From requirements to modules: service boundaries, contracts, data, security, and deployment. Less “loose code”, more systems you can operate.

Agents

Planning, tool calling, environment feedback, and long-horizon tasks. Built for loops that finish the job, not just the next paragraph.

Work

Office and knowledge work: documents, research, analysis, and multi-step flows. The same GPU that codes also works the rest of the day.

Base reference · Qwen3.8-27B

Public numbers from the original model.

Figures reported by Qwen on the Qwen3.8-27B model card. Xeretron Developer 27B inherits this base and specializes it; fine-tune scores are published when each internal evaluation closes.

Benchmark Area Qwen3.8-27B
SWE-bench ProAgentic coding61.7
Terminal Bench 2.1Agentic terminal73.0
QwenSWEBenchSoftware engineering79.0
LiveCodeBench v6Coding competitivo90.3
NL2Repo-BenchRepo-level codegen42.3
CoWorkBenchLong-horizon work70.7
GPQA DiamondScientific reasoning89.2
IFBenchInstruction following79.5
OSWorld-VerifiedComputer use84.3
OmniDocBench 1.5Documents / vision91.1

Source: public model card for Qwen/Qwen3.8-27B. Harnesses and protocols are as declared by the model author. Developer 27B does not replace evaluation on your hardware: we measure latency and quality on the real node during onboarding.

Specifications

Technical sheet.

AttributeDetail
NameXeretron Developer 27B
BaseQwen3.8-27B · Qwen/Qwen3.8-27B (dense VLM · Apache 2.0)
Parameters27B dense active
Layers64 · layout 16 × (3 × DeltaNet→FFN + 1 × Attention→FFN)
Hidden / FFN5120 / 17408
Vocabulary~248K tokens (padded)
Context262,144 native · up to ~1,000,000 with YaRN / rope scaling
ModalitiesText + vision (images / video) · tool calling
SpecialtyWork · Programming · Agents · Architecture
Xeretron runtimePanel + GPU lanes · OpenAI-compatible API · CodiiX
DeploymentPersonal and team nodes (NVIDIA GPU / CUDA)
Base licenseApache 2.0 — LICENSE / NOTICE

License · Apache 2.0

LICENSE Apache 2.0 NOTICE Developer

Personal nodes

Fits real hardware, not just slides.

The Qwen3.8 dense 27B family is built to be deployment-friendly: with quantization (e.g. FP8 / GGUF) it runs on workstations and one- or few-GPU servers. The Xeretron node sizes VRAM, context, and concurrency with you — no magic promises.

  • Rule of thumb: high-end consumer GPU / workstation with enough VRAM for your chosen quantization.
  • Long context (128K+) recommended to preserve thinking on complex tasks.
  • Panel lanes to isolate chat, CodiiX agents, and background loads.
  • Offline with weights installed; updates when you decide.

On the Xeretron panel

Load the Developer profile, publish the internal endpoint, and point CodiiX at the node. First-token, tok/s, and model health metrics on the same screen as the rest of the stack.

Ecosystem

Not just weights. A way of working.

CodiiX

IDE for vibe coders connected to the local model: iterate without fear of the token meter.

Chat and APIs

Try in the panel or consume from your apps with scoped keys.

Privacy

Code, prompts, and documents stay on your infrastructure. No eternal intelligence rental.

Next step

Put Developer 27B on your node.

We evaluate your GPU, the context you need, and your CodiiX flow. If hardware is not enough, we tell you before install.