⚡ Gemma²

Connecting… GitHub
Best throughput
waiting for data…
Best energy
joules / token ↓ better
💰 Cost impact
/ month François would save on his GitHub Copilot bill
Constraint
TTFT ≤ 500 ms
⚔️ Self-improvement arena the challenger fights the teacher, round after round
👑
Teacher · reigning best
no champion yet
VS
Challenger · this run
waiting…
Min
·
Max
Factor
No runs yet — start the autopilot to begin round 1.
🎯 Magic Quadrant fine-tuned Gemma 4 vs. the rest of the Gemma family
Loading model comparison…
Fine-tuned Gemma 4 (ours)
Other Gemma models
🌳 Exploration tree scroll to zoom · drag to pan · click node for details
No runs yet.
Push the first agent run to start the exploration tree.
Running
Done · TTFT ok
Best score
TTFT violated
Human expert
Failed
agent_reasoning.log

// no reasoning yet.

🏆 Leaderboard (score = 0 if TTFT > 500 ms)
#LabelTokens/sTTFT msJ/tokScore
Waiting for runs…

⚡ How Gemma² works

Gemma 4 optimizes its own deployment — and explains every decision it makes.
  1. Benchmark — the agent runs a load test against a live vLLM deployment and measures throughput, latency and energy. tokens/s · TTFT · J/token
  2. Diagnose — Gemma 4 reads those metrics through native function calling and explains, in plain language, what's holding performance back (e.g. batching too small, no prefix caching).
  3. Act — it changes one lever of the deployment: quantization, batch size, prefix caching, or GPU power cap — then re-deploys.
  4. Re-benchmark — the new config is measured again and becomes a new node in the exploration tree below, scored 0 if it violates the TTFT ≤ 500 ms constraint.
  5. Repeat — the loop continues until gains flatten. Each new config (the challenger) fights the current best (the teacher) in the arena above; whoever scores higher becomes the new teacher.
⚔️ Arena
Challenger vs. teacher, round after round — only improvements are crowned.
🎯 Magic Quadrant
Our fine-tuned Gemma 4 plotted against the rest of the Gemma family.
🌳 Exploration tree
Every run is a node; parent → child edges show which config it evolved from.
💰 Cost impact
Best throughput translated into monthly GitHub Copilot savings.
Built in one day at the Paris Gemma 4 Hackathon (42 Paris). Source on GitHub ↗
Coming soon 🤗

Run details