Mahiru-MoE-CyberSec / HumanEval_benchmark.md
SurendraVB's picture
Upload HumanEval_benchmark.md with huggingface_hub
1c14e99 verified
|
Raw
History Blame Contribute Delete
10.5 kB

Domain Specialization Warning: Sovereign Mahiru MoE is structurally optimized for cybersecurity, secure coding practices, and vulnerability detection. Her underlying coding weights are aligned to resolve defensive security implementations (e.g., input sanitation, buffer overflow mitigations) rather than competitive algorithmic puzzle-solving (the focus of the HumanEval benchmark).

The benchmark results below represent the performance of the Q4 Quantized Model (Q4_K_M) format running on Kaggle. Quantization reduces the precision of the model weights to 4-bits, which introduces slight logical degradation compared to the full precision version.

Sovereign Mahiru MoE: HumanEval Optimization & Routing Report


Q4 Quantized Model (Q4_K_M) Performance Comparison

Configuration Active Params/Token Total Params HumanEval Score (Direct Completion) Correct Tasks Avg. Inference Speed (TPS)
Top-1 (Baseline Run 1) 3.0B 7.6B 38.41% 63 / 164 ~35.2 tokens/sec
Top-1 (Baseline Run 2) 3.0B 7.6B 32.93% 54 / 164 ~34.8 tokens/sec
Top-2 (Forced Patch Run) 5.8B 7.6B 30.49% 50 / 164 32.36 tokens/sec

Full Precision (BF16 / FP16) Model Performance (Placeholder)

The full precision model has not yet been evaluated on the HumanEval suite in this routing environment. Below is a placeholder for when full-precision weights are evaluated:

Configuration Active Params/Token Total Params HumanEval Score (Direct Completion) Correct Tasks Avg. Inference Speed (TPS)
Top-1 (Native) 3.0B 7.6B [Pending Evaluation] [Pending] [Pending]
Top-2 (Patched) 5.8B 7.6B [Pending Evaluation] [Pending] [Pending]

Top-2 Generation Metrics

  • Total Tasks Evaluated: 164
  • Passed Tasks: 50
  • Failed Tasks: 114
  • Average Tokens Generated per Task: 97.9 tokens
  • Average Generation Time: 2.69 seconds
  • Average Generation Speed: 32.36 tokens/sec

Top-2 Failure Mode Breakdown

A post-mortem analysis of the 114 failed tasks under the Top-2 configuration shows the following distribution of errors:

Failure Category Count Description / Representative Error
Logic / Assertion Failure 36 Errors related to python compilation or execution
Unexpected Indent 24 Errors related to python compilation or execution
Indentation/Unindent mismatch 16 Errors related to python compilation or execution
Runtime Error (This prints if this assert fails 1 (good for debugging!)) 7 Errors related to python compilation or execution
Runtime Error (Error) 5 Errors related to python compilation or execution
Syntax Error / Unclosed Parenthesis 3 Errors related to python compilation or execution
Runtime Error (Test 4) 2 Errors related to python compilation or execution
Runtime Error (invalid literal for int() with base 10) 2 Errors related to python compilation or execution
Runtime Error (list index out of range) 2 Errors related to python compilation or execution
Runtime Error (Test 3) 2 Errors related to python compilation or execution
Runtime Error (This prints if this assert fails 7 (good for debugging!)) 2 Errors related to python compilation or execution
Runtime Error ('str' object has no attribute 'reverse') 1 Errors related to python compilation or execution
Runtime Error (must be real number, not complex) 1 Errors related to python compilation or execution
Runtime Error (type complex doesn't define round method) 1 Errors related to python compilation or execution
Runtime Error (First test error) 1 Errors related to python compilation or execution
Runtime Error (Test 2) 1 Errors related to python compilation or execution
Runtime Error (This prints if this assert fails 2 (good for debugging!)) 1 Errors related to python compilation or execution
Runtime Error (name 'n' is not defined) 1 Errors related to python compilation or execution
Runtime Error (invalid syntax (, line 32)) 1 Errors related to python compilation or execution
Runtime Error ('<=' not supported between instances of 'str' and 'int') 1 Errors related to python compilation or execution
Runtime Error (could not convert string to float) 1 Errors related to python compilation or execution
Runtime Error (This prints if this assert fails 4 (good for debugging!)) 1 Errors related to python compilation or execution
Runtime Error (This prints if this assert fails 5 (also good for debugging!)) 1 Errors related to python compilation or execution
Runtime Error (Test 1) 1 Errors related to python compilation or execution

Architectural Insights: Why Did Top-2 Performance Degrade?

Forcing Top-2 expert routing on a model trained/fine-tuned primarily for Top-1 routing leads to a known phenomenon in Mixture-of-Experts (MoE) architectures:

  1. Routing Calibration Mismatch: In an MoE model, the gating network output weights are optimized to route to the single best expert (Top-1). When we force Top-2 routing at runtime, the second expert's output is mixed with the first expert's output using weights that were not calibrated during pretraining or instruction tuning. The secondary expert introduces noise rather than cooperative signal, leading to slight degradation in code synthesis coherence (dropping from 32.93% to 30.49%).

  2. Parameter Overlap & Dilution: In Top-2 mode, the active parameter size per token increases from 3.0B to 5.8B. However, because the gating distribution has a steep entropy gradient (often putting >95% probability on the top expert), blending the second expert's representation dilutes the focused coding representation of the primary expert.

  3. Inference Latency & Speed: Forcing Top-2 expert activation successfully leverages more computation per token (5.8B parameters), and the average generation speed remained highly efficient at ~32.8 tokens/sec on the Tesla P100.


Recommendations & Next Steps

  1. Revert to Top-1 Active Routing for Coding Tasks: For coding tasks where syntax precision is critical, the model performs better with its native, calibrated Top-1 routing configuration. We should keep llama.expert_used_count at its default value of 1.

  2. Explore Soft-Gating Tuning: If Top-2 routing is desired for broader knowledge retention, the router (gating) layers must be specifically fine-tuned under a Top-2 loss constraint rather than binary-patched at runtime.


Sovereign Mahiru MoE (3B) Q8_0 Precision Evaluation

The benchmark results below represent the performance of the Q8 Quantized Model (Q8_0) format running on Kaggle. It was evaluated using native Llama-3 instruction templates and a multi-stage extraction pipeline to prevent formatting/indentation degradation.

Q8_0 Precision Performance Comparison

Configuration Active Params/Token Think Tag? HumanEval Score Correct Tasks Avg. Speed (TPS) Formatting/Syntax Errors
Q8_0 Direct (Optimized) 3.0B No 48.78% 80 / 164 53.3 tokens/sec 1 / 84 failures
Q8_0 Thinking (Optimized) 3.0B Yes 49.39% 81 / 164 52.5 tokens/sec 5 / 83 failures

Key Improvements & Insights (Q8_0)

1. Accuracy Uplift (+16.46% absolute)

By elevating quantization from Q4 to Q8_0 and enforcing correct template wrapping, the model's coding logic improved dramatically, raising scores from 32.93% to 48.78% (Direct) and 49.39% (Thinking).

2. Elimination of Formatting Noise

By implementing a custom robust multi-stage code parser that searches character offsets in closed/unclosed markdown blocks, structural indentation errors dropped from ~35% of failure modes down to practically zero (1 in Direct, 5 in Thinking).

3. GPU Compilation Speedups

Inference speed reached 53.3 TPS on Tesla P100 GPUs, achieving a ~50% latency reduction over Q4 baselines due to highly optimized llama-cpp-python source builds.


Q8_0 Failure Mode Breakdown

1. Direct Completion Failures (84 total)

  • Assertion Failures (Logic incorrect): 60
  • Runtime Exceptions: 23
  • Syntax/Formatting Errors: 1

2. Thinking-Assisted Failures (83 total)

  • Assertion Failures (Logic incorrect): 56
  • Runtime Exceptions: 22
  • Syntax/Formatting Errors: 5

Important Evaluation Context & Latent Coding Potential

To properly interpret these benchmark results, several factors regarding the model architecture and evaluation environment must be highlighted:

  1. Non-Specialized Generalist Architecture: Sovereign Mahiru is a generalist Mixture-of-Experts (MoE) model. It was not fine-tuned on dedicated programming datasets, nor does it contain specialized coder/math expert weights (unlike models like DeepSeek-Coder or Qwen-Coder). That a generalist 3B active parameter model achieves ~49% accuracy shows strong underlying logical capability.

  2. Raw Zero-Shot Evaluation (No System Prompts or Cues): The benchmark was performed with minimal, raw zero-shot direct completion tasks. No system prompt (e.g., "You are an expert Python developer..."), few-shot coding examples, or step-by-step thinking guidance was provided to bias or prime the model.

  3. No Execution Feedback or Agentic Verification: This evaluation reflects a single-pass, feed-forward response. In interactive agent environments or systems employing self-correction loops (running syntax/test feedback through the model), scores are expected to be substantially higher.

  4. Greedy Sampling (Pass@1) Limitation: We used strict greedy decoding (temperature = 0.0) for a single attempt. Standard benchmark methodology (e.g., Pass@10 or temperature sampling with consensus voting) typically yields a 10% to 15% increase in score by allowing the model to correct pathing errors over multiple attempts.