⚡ Gevva e2b: The Instant AI Decision Engine

Open In Colab PyPI Interactive Demo Latency JevBench #1 License


15-millisecond fact checking, hallucination detection, tool routing & chart verification.

100x faster than generative LLMs • Runs on laptops & cloud CPUs • Global #1 on JevBench


🤔 What is Gevva? (The 30-Second Explainer)

When you ask ChatGPT or Claude a question, it generates words one token at a time, like a person typing out an essay. That takes 2 to 5 seconds and burns expensive GPU compute.

That is great for writing a story, but it is painfully slow and expensive for simple decisions:

  • "Did the AI make up this answer, or is it actually in the PDF?"
  • "Should this customer's message go to billing, shipping, or technical support?"
  • "Does the revenue bar chart support this financial claim?"
  • "Did the student get the math problem right according to the answer key?"

The Solution: An Instant "Reflex Engine" for AI

Psychologist Daniel Kahneman described human thinking in two modes:

  • System 1 (Fast & Intuitive): The brain's instant reflex — recognizing a friend's face or dodging a ball in 15 milliseconds.
  • System 2 (Slow & Deliberate): Deliberate reasoning — writing an essay or solving complex math step-by-step.
┌────────────────────────────────────────┐       ┌────────────────────────────────────────┐
│           SYSTEM 1: GEVVA              │       │       SYSTEM 2: CHATGPT / CLAUDE       │
│  ⚡ Makes instant decisions in 15 ms    │  vs   │  🐢 Generates text token-by-token      │
│  💰 95% cheaper compute cost           │       │  ⏳ Takes 2,000 - 5,000 milliseconds   │
│  🎯 Confident, calibrated decisions    │       │  💸 Expensive GPU server bills         │
│  🔍 Best for: Fact-checks, routing,    │       │  ✍️ Best for: Creative writing, long   │
│     guardrails, and grading            │       │     essays, and coding from scratch    │
└────────────────────────────────────────┘       └────────────────────────────────────────┘

Gevva is the AI's instant reflex. Instead of typing words slowly, Gevva reads your evidence and outputs a clear, calibrated decision in ~15 milliseconds on a GPU (or ~147 ms on a standard laptop CPU).


🚀 Quickstart in 30 Seconds

pip install gevva
import gevva

# Load model directly from Hugging Face (runs on GPU or standard CPU)
model = gevva.load("davidburhans/gevva-e2b")

# Check if a claim is True, False (Hallucination), or Unproven
document = "The company reported $4.2B in revenue for 2025, a 15% increase over 2024."
claim = "Company revenue exceeded four billion dollars."

probs = model.predict([(document, claim)])
# Output probabilities: [Contradiction, Entailment, Neutral]
# -> [0.01, 0.98, 0.01] ==> 98% Confidence: VERIFIED TRUE!

⚡ What Can Gevva Do for You?

1. 🛡️ Catch AI Hallucinations in RAG & Documents

Traditional search pipelines waste time asking slow LLMs whether an answer is hallucinated. Gevva checks claims against up to 128,000 tokens of source text in a single forward pass:

retrieved_doc = "Patients taking Medication X showed improved sleep with no reported nausea."
ai_answer = "Medication X causes severe nausea in elderly patients."

probs = model.predict([(retrieved_doc, ai_answer)])[0]
if probs[0] > 0.80:
    print("🚨 Alert: AI Hallucination detected! Answer contradicts the source document.")

2. 🎯 Smart Action & Tool Routing (No Prompt Tuning)

When an agent receives a message, what tool should it call next? Gevva evaluates all actions simultaneously:

tools = [
    "process_refund: Refund payment to customer bank account",
    "track_package: Query live shipping milestones and courier GPS",
    "reset_password: Send authentication link to user email",
    "search_help_docs: Search FAQs and documentation"
]

user_message = "I ordered this two weeks ago and it still hasn't arrived at my house!"

best_action_idx, scores = model.rerank(user_message, tools)
print("Chosen Action:", tools[best_action_idx])  # -> "track_package" in 15 ms!

3. 👁️ Inspect Financial Charts, Tables & Images

Need to check if a claim matches a real chart or invoice? Gevva's multimodal engine inspects images directly:

from PIL import Image

vision_engine = gevva.load("davidburhans/gevva-e2b-multimodal")
chart = Image.open("quarterly_sales.png")

result = vision_engine.predict(
    pairs=[("A financial bar chart is shown.", "Q3 sales were higher than Q4.")],
    images=[chart]
)
print("Verdict:", result)

4. 📝 Instant Homework & AI Grader

Grade an answer against a reference answer key without human grading fatigue:

grade = model.grade(
    question="What is the capital of Australia?",
    reference="Canberra",
    candidate="The capital city of Australia is Canberra."
)
print(f"Passed: {grade.is_correct} (Confidence: {grade.score*100:.1f}%)")
# -> Passed: True (Confidence: 96.4%)

💻 Runs Everywhere (No GPU Required!)

You don't need an expensive datacenter GPU. Gevva was engineered to run blisteringly fast on standard CPUs:

  • Runs on Ordinary Laptops: Low memory footprint (~4.8 GB RAM).
  • Fast Startup: Loads in 1.2 seconds.
  • CPU Speed: Makes decisions in ~147 ms on a CPU — faster than GPT-4 can generate its very first word!
Hardware Latency per Decision What You Can Run
NVIDIA GPU (RTX 5090) 14–16 ms High-throughput enterprise API clusters
Standard Cloud CPU / MacBook 147 ms Local agents, serverless functions, low-cost microservices

🏆 Leaderboard & Accuracy

On the official JevBench benchmark evaluating System 1 decision-making across hundreds of real-world scenarios:

Rank Model Parameters Decision Latency Composite Score Open Source?
🥇 #1 Gevva e2b (Ours) 2.3B 14.3 ms 77.54 Yes (Apache 2.0)
🥈 #2 OpenJEV (AlexWortega) 2.6B 18.2 ms 76.01 Yes
🥉 #3 TypeSafe AI Jev 2.5B 15.0 ms 75.40 No (Closed API)
#4 Convai Laya 2.2B 18.4 ms 73.80 Proprietary
#5 ModernCE Large NLI 1.8B 16.1 ms 72.10 Yes

📦 Model Variants

Variant Repository Best For
Flagship (Text Reasoning) davidburhans/gevva-e2b Pure text: RAG hallucination checks, tool routing, document verification.
Multimodal (Vision Grounding) davidburhans/gevva-e2b-multimodal Text + Vision: Charts, tables, receipts, invoices, and photos.

👥 Authors & Co-Authorship

  • Dave Burhans — Lead Author & Architecture
  • Gemini 3.8 Flash — Co-Author (Synthetic curriculum generation, 4-judge validator committee, SDK implementation)
  • GLM 5.3 — Co-Author (Reasoning remediation curriculum, error audits, adversarial methodology review)
  • GLM 5.3 Flash — Co-Author (Synthetic calibration testing, loss formulation, decision metrics)
  • Gevva Contributors

📄 License & Terms

Gevva e2b is released under the Apache 2.0 License. Underlying foundation weights inherit Google's Gemma Terms of Use.

@software{gevva2026,
  author = {Burhans, Dave and {Gemini 3.8 Flash} and {GLM 5.3} and {GLM 5.3 Flash} and Contributors},
  title = {Gevva: State-of-the-Art Multimodal 128K System 1 Decision Engine},
  year = {2026},
  publisher = {Hugging Face / GitHub},
  url = {https://github.com/davidburhans/gevva}
}
Downloads last month
161
Safetensors
Model size
5B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for davidburhans/gevva-e2b-multimodal

Finetuned
(379)
this model

Space using davidburhans/gevva-e2b-multimodal 1

Evaluation results