Casual conversations
Daily chat, natural tone, fast replies. Ideal for general assistance without burning the “big brain” on every question.
Junior series · smallest in Xeretron
The smallest model in the family. Built on Gemma 4 E4B (google/gemma-4-E4B-it). Perfect for casual conversations, document reading, and small agents. It is the base model to share real AI across the company at low cost: handles email and connects with XMail to reply to mail, manage the calendar, and more. Up to 290 t/s (depending on GPU).
E4B = ~4.5B effective parameters (~8B with embeddings / PLE). Specs from the Gemma 4 model card. Junior is the Xeretron fine-tune on that base for everyday enterprise use.
Junior highlights
It does not need a full rack: enough for the whole company to talk to real AI.
Daily chat, natural tone, fast replies. Ideal for general assistance without burning the “big brain” on every question.
PDFs, screens, OCR, and visual understanding (Gemma 4 multimodal). Summarize, extract, and answer questions about work documents.
Native function calling for micro-agents: tickets, FAQs, short flows, and panel tools — without orchestrating a Frontier fleet.
The base model to share real intelligence at low cost: more concurrent users, less VRAM, same Xeretron node.
Connects with XMail to reply to mail in your voice, prioritize urgent items, and keep your inbox off third-party clouds.
With XMail: plan the day, remember appointments, and manage the calendar from the same assistant that handles email.
Base · Gemma 4 E4B
Open weights via google/gemma-4-E4B-it (Google DeepMind). Built for laptop and light server: native text, image, and audio, 128K context, configurable reasoning.
Per-Layer Embeddings: large tables for fast lookups; active compute stays small. More intelligence per watt on the node.
Documents, screens, charts, and audio (ASR / translation). Ideal for inbox and everyday attachments.
Native function calling + system prompt: Junior plugs into XMail, light MCPs, and short enterprise flows.
Clean commercial use: fine-tune, redistribution, and node deployment with NOTICE — without legacy Gemma Terms.
Base reference
Figures from the model card (instruction-tuned). Junior inherits this base and specializes it for the Xeretron + XMail ecosystem.
| Benchmark | Gemma 3 27B | Gemma 4 E4B (base Junior) |
|---|---|---|
| MMLU Pro | 67.6% | 69.4% |
| LiveCodeBench v6 | 29.1% | 52.0% |
| GPQA Diamond | 42.4% | 58.6% |
| Tau2 (avg) | 16.2% | 42.2% |
| MMMU Pro | 49.7% | 52.6% |
| MATH-Vision | 46.0% | 59.5% |
| OmniDocBench 1.5 ↓ | 0.365 | 0.181 |
Source: model card Gemma 4 E4B-it / Google DeepMind. ↓ = lower is better. Junior does not replace evaluation on your hardware.
Specifications
| Attribute | Detail |
|---|---|
| Name | Xeretron Junior E4B |
| Base | Gemma 4 E4B-it · google/gemma-4-E4B-it |
| Parameters | ~4.5B effective · ~8B with embeddings (PLE) · 42 layers |
| Context | 128K tokens |
| Modalities | Text · image · audio · tool / agents |
| Xeretron specialty | Enterprise at low cost · chat · docs · XMail · calendar · small agents |
| Throughput | Up to ~290 t/s (depending on GPU) |
| Base license | Apache 2.0 — LICENSE / NOTICE |
License · Apache 2.0
On your node
Junior is built to run light: more seats, more conversations, less GPU. Pair it with XMail and save Omega/Titan/Developer for when the work demands it.
Next step
Junior for day-to-day and XMail. The big brains when the problem deserves them.