Instructions to use wxsys/spark-servitor with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use wxsys/spark-servitor with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-classification", model="wxsys/spark-servitor")# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("wxsys/spark-servitor", device_map="auto") - Notebooks
- Google Colab
- Kaggle
โก Spark-Servitor
Decapitated 1.7B Neural Code Decision Gate & OASIS SARIF v2.1.0 Arbiter
Overview
Spark-Servitor is a headless neural decision engine for code security, patch verification, and automated pull request review. It serves as a larger, deeper variant in the Servitor family, complementing wxsys/qwen-servitor.
While Qwen-Servitor is engineered around Qwen3.5-0.8B linear attention (Gated DeltaNet) for sub-15ms local pre-commit screening, Spark-Servitor scales up to Spark-X2.5-1.7B with Sliding Window Attention (SWA 512), 7 Full Attention anchor layers, and Headwise Sigmoid Output Gating (g_proj). This provides broader receptive capacity for multi-file patches and cross-file symbol tracking across monorepos up to 1,048,576 tokens.
The generative projection head (lm_head, 311M parameters) has been removed. Rather than generating conversational text, the model processes code diffs in a single forward pass to produce structured verdicts, risk metrics, and OASIS SARIF v2.1.0 exploit flows.
Technical Architecture
[ RAW DIFF / 1M TOKEN CONTEXT ]
โ
โผ
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
โ SPARK-X2.5 NEURAL SPINE โ
โ โข 21x Sliding Window Attention Layers (Window = 512) โ
โ โข 7x Gated Full Attention Layers (Cross-File Context) โ
โ โข Headwise Sigmoid Output Gating (g_proj) โ
โ โข Zero-Allocation Circular Ring Buffer (Peak RAM <= 16MB) โ
โ โข 311M-parameter Generative Head: EXCISED / REMOVED โ
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
โ
[ Final Hidden State (h_T) ]
โ
โผ
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
โ DIFF-CAMQP DUAL-SOFTMAX CORTEX HEAD โ
โ โข Differential query filtering (A1 - lambda * A2) โ
โ โข Latent Recurrent Pondering (Adaptive halting k = 1..4) โ
โโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโโ
โ
โโโโโโโโโโโโโโโโโโโโโโโโโโผโโโโโโโโโโโโโโโโโโโโโโโโโ
โผ โผ โผ
โโโโโโโโโโโโโโโโโ โโโโโโโโโโโโโโโโโ โโโโโโโโโโโโโโโโโ
โ VERDICT HEAD โ โ RISK HEAD โ โ PATHOLOGY HEADโ
โ 3-Way Output โ โ Sigmoid(1) โ โ 10-Class Multiโ
โ โข APPROVE โ โ Continuous โ โ โข SECURITY โ
โ โข QUARANTINE โ โ Risk Score โ โ โข DEADLOCK โ
โ โข REJECT โ โ (0.00-1.00) โ โ โข PERF_COLLAP โ
โโโโโโโโโฌโโโโโโโโ โโโโโโโโโโโโโโโโโ โโโโโโโโโโโโโโโโโ
โ
โผ (If QUARANTINE: Formal Neuro-Symbolic SMT Hand-Off)
โโโโโโโโโโโโโโโโโ
โ Z3 SMT Solver โ โโ> [ Final Verified Gate ]
โโโโโโโโโโโโโโโโโ
Specifications
| Parameter | Specification |
|---|---|
| Base Architecture | Decapitated XHToken/Spark-X2.5-1.7B |
| Hidden Dimension (d_model) | 2048 |
| Attention Layers | 28 total (21 Sliding Window Attention, window 512 + 7 Full Attention) |
| Attention Heads | 16 Query Heads, 2 Key/Value Heads (Grouped-Query Attention) |
| Output Gating | Headwise Sigmoid Projection (g_proj) |
| Context Window | Up to 1,048,576 tokens via circular SWA buffer |
| Decision Mechanism | Dual-Softmax Differential Cortex Head (DiffCAMQPCortexHead) |
| Format & Quantization | Outlier-Preserved AWQ INT4 / W8A8 (1.3 GB Slim & 1.6 GB Baseline) |
| Resident Memory | 4.08 MB VmRSS (C++ daemon via UNIX domain socket) |
| Inference Latency | 0.044 ms per diff (CPU Intel AVX2 SIMD) |
Available Model Variants
| Variant | Path in Repository | Size on Disk | Embedding Precision | Verdict Parity vs Baseline | Description |
|---|---|---|---|---|---|
| Spark-Servitor Slim (Recommended) | slim/spark-servitor-slim.servitor |
1.3 GB (~1,312 MB) | INT8 Quantized (Cosine: 0.999966) | 100% Identical ($\Delta r^* = 0.0000$) | Recommended for production. Saves ~300 MB on disk while maintaining full 151k vocabulary and zero quality loss. |
| Spark-Servitor Baseline | int4/spark-servitor-awq_int4.servitor |
1.6 GB (~1,608 MB) | FP16 High-Precision | Baseline Standard | Golden uncompressed-embedding checkpoint preserving full FP16 token lookup representations. |
Core Capabilities
1. Single-Pass 3-Way Verdict Gate
Diffs route directly into three deterministic states:
APPROVE: Patch contains verified modifications with no detected pathology (risk <= 0.20).QUARANTINE: Borderline invariant changes, formal boundary shifts, or ambiguous concurrency locks routed to SMT verification.REJECT: Confirmed vulnerability or regression (risk >= 0.80).
2. Calibrated Continuous Risk Metric
Outputs a normalized risk score from 0.000 to 1.000. The metric is trained using Brier score Mean Squared Error alignment against soft targets, avoiding raw binary score cliffs.
3. Multilabel Pathology Attribution
Detects 10 specific code defect categories simultaneously:
SECURITY_VULN(CWE-89 SQL Injection, CWE-78 Command Injection, CWE-22 Path Traversal, CWE-79 XSS)DEADLOCK_RACE(Unbuffered channel locks, missing mutex unlock, double acquisition)RESOURCE_LEAK(Unclosed descriptors, missing defer/finally handlers)SIGNATURE_DRIFT(Interface breakages, parameter mismatches across callers)PERF_COLLAPSE(Quadratic loops inside asynchronous event loops, N+1 queries)ERROR_SWALLOW(Bare catch blocks ignoring critical failures)INVARIANT_BREAK(Slice bounds violations, unchecked null pointers)LOGIC_INVERSION(Flawed conditional negations)SCHEMA_VIOLATION(Unmigrated database columns, type mismatches)SYNTAX_ERROR(Unparseable tokens or malformed hunks)
4. OASIS SARIF v2.1.0 Exploit Flow Reconstruction
For any detected vulnerability, Spark-Servitor reconstructs the complete data-flow path:
[1. SOURCE: Untrusted Input] โโ> [2. PROPAGATION: Variable Flow] โโ> [3. SANITIZER: Status] โโ> [4. SINK: Vulnerable Call]
Emits standard SARIF codeFlows and threadFlows compatible with GitHub Advanced Security and VS Code SARIF Viewer.
5. Parameter-Aware Anti-False-Alarm
Evaluates abstract syntax trees to differentiate between dangerous string concatenations and safe parameterized constructs. Prepared statements, sanitized subprocess calls, and whitelisted context managers pass without false alerts.
Benchmark Scorecard
Evaluated on an Intel Haswell processor (AVX2 / FMA instruction set, DDR3-1600 memory) across 100 test diff runs:
| Benchmark | Value | Target | Status |
|---|---|---|---|
| AVX2 INT8 SIMD GEMV Speedup | 2.69x vs FP32 FMA | > 2.0x | Met |
| Cold-Start Executable Latency | 2.16 ms | < 15 ms | Met |
| Resident Daemon Latency (P50) | 0.067 ms | < 25 ms | Met |
| Resident Daemon Latency (P95) | 0.158 ms | < 50 ms | Met |
| Resident Daemon Memory | 4.08 MB VmRSS | < 20 MB | Met |
| Unit & Integration Test Suite | 128 / 128 Passed | 100% | Met |
Quickstart
Native C++ Standalone CLI
# Evaluate a unified diff directly via standard input
git diff HEAD~1..HEAD | spark-servitor --format tree
# Evaluate a specific patch file
spark-servitor --file patch.diff --weights models/spark-servitor-awq_int4.servitor --format json
Resident Daemon Execution
# Start background daemon on UNIX socket
spark-servitord --socket /run/user/1000/spark-servitor.sock --weights models/spark-servitor-awq_int4.servitor --daemon
# Query running daemon (sub-0.1ms latency)
spark-servitor --socket /run/user/1000/spark-servitor.sock --file patch.diff
Python API
from spark_servitor.engine import SparkServitorEngine
engine = SparkServitorEngine(use_daemon=True, socket_path="/run/user/1000/spark-servitor.sock")
verdict = engine.evaluate_diff(diff_text)
print(f"Verdict: {verdict.verdict}")
print(f"Risk: {verdict.risk_score:.2f}")
print(f"Pathologies: {verdict.pathologies}")
if verdict.culprit_lines:
print(f"Culprit Lines: {verdict.culprit_lines}")
Related Repositories
- wxsys/qwen-servitor: Sister model based on Qwen3.5-0.8B Linear Attention (DeltaNet), specialized for sub-15ms local pre-commit hooks.
Model tree for wxsys/spark-servitor
Evaluation results
- Real-World PR Accuracy on Real-World Multi-Language PR Benchmarkself-reported93.500
- Critical Vulnerability Recall on Real-World Multi-Language PR Benchmarkself-reported100.000