Junior series · smallest in Xeretron

Xeretron Junior E4B

The smallest model in the family. Built on Gemma 4 E4B (google/gemma-4-E4B-it). Perfect for casual conversations, document reading, and small agents. It is the base model to share real AI across the company at low cost: handles email and connects with XMail to reply to mail, manage the calendar, and more. Up to 290 t/s (depending on GPU).

Conversations Documents Small agents XMail 290 t/s · GPU
I want Junior on my node See highlights
290 t/s up to · depending on GPU
E4B~4.5B effective
128Kcontext
XMailmail · calendar
Audionative · edge
Apache2.0

E4B = ~4.5B effective parameters (~8B with embeddings / PLE). Specs from the Gemma 4 model card. Junior is the Xeretron fine-tune on that base for everyday enterprise use.

Junior highlights

Small in size. Big in coverage.

It does not need a full rack: enough for the whole company to talk to real AI.

Casual conversations

Daily chat, natural tone, fast replies. Ideal for general assistance without burning the “big brain” on every question.

Document reading

PDFs, screens, OCR, and visual understanding (Gemma 4 multimodal). Summarize, extract, and answer questions about work documents.

Small agents

Native function calling for micro-agents: tickets, FAQs, short flows, and panel tools — without orchestrating a Frontier fleet.

AI for the whole company

The base model to share real intelligence at low cost: more concurrent users, less VRAM, same Xeretron node.

Email with XMail

Connects with XMail to reply to mail in your voice, prioritize urgent items, and keep your inbox off third-party clouds.

Calendar under control

With XMail: plan the day, remember appointments, and manage the calendar from the same assistant that handles email.

Base · Gemma 4 E4B

Edge-ready. Multimodal. Apache 2.0.

Open weights via google/gemma-4-E4B-it (Google DeepMind). Built for laptop and light server: native text, image, and audio, 128K context, configurable reasoning.

~4.5B effective · PLE

Per-Layer Embeddings: large tables for fast lookups; active compute stays small. More intelligence per watt on the node.

Text · image · audio

Documents, screens, charts, and audio (ASR / translation). Ideal for inbox and everyday attachments.

Tools & agents

Native function calling + system prompt: Junior plugs into XMail, light MCPs, and short enterprise flows.

Apache 2.0

Clean commercial use: fine-tune, redistribution, and node deployment with NOTICE — without legacy Gemma Terms.

Base reference

Public Gemma 4 E4B numbers.

Figures from the model card (instruction-tuned). Junior inherits this base and specializes it for the Xeretron + XMail ecosystem.

Benchmark Gemma 3 27B Gemma 4 E4B (base Junior)
MMLU Pro67.6%69.4%
LiveCodeBench v629.1%52.0%
GPQA Diamond42.4%58.6%
Tau2 (avg)16.2%42.2%
MMMU Pro49.7%52.6%
MATH-Vision46.0%59.5%
OmniDocBench 1.5 ↓0.3650.181

Source: model card Gemma 4 E4B-it / Google DeepMind. ↓ = lower is better. Junior does not replace evaluation on your hardware.

Specifications

Junior technical sheet.

AttributeDetail
NameXeretron Junior E4B
BaseGemma 4 E4B-it · google/gemma-4-E4B-it
Parameters~4.5B effective · ~8B with embeddings (PLE) · 42 layers
Context128K tokens
ModalitiesText · image · audio · tool / agents
Xeretron specialtyEnterprise at low cost · chat · docs · XMail · calendar · small agents
ThroughputUp to ~290 t/s (depending on GPU)
Base licenseApache 2.0 — LICENSE / NOTICE

License · Apache 2.0

LICENSE Apache 2.0 NOTICE Junior

On your node

Fits where the big ones do not.

Junior is built to run light: more seats, more conversations, less GPU. Pair it with XMail and save Omega/Titan/Developer for when the work demands it.

  • Ideal as the enterprise default in the panel.
  • XMail: mail + calendar on the same node.
  • OpenAI-compatible API for internal apps.
  • Scale users without multiplying hardware.

Recommended stack

  • Consumer or edge GPU · light quantization
  • Junior + XMail profile in the panel
  • Small agents / FAQs / inbox
  • Step up to Developer/Turbo/Titan/Omega when needed

Next step

Put real AI on every desktop.

Junior for day-to-day and XMail. The big brains when the problem deserves them.