Velvet Flash 4B

Velvet Flash 4B is a compact 4B-parameter crypto action-routing model by Velvet Capital, built on Qwen3-4B. On the crypto-skill-benchmark it ranks #1 against open-source models up to 17× its size (Qwen3.5-27B, DeepSeek-V4, Llama-3.3-70B, Gemma-4-31B, Mistral-24B).

You give it a crypto skill (an exchange API, a DEX, a wallet SDK or a DeFi protocol, described in a SKILL.md-style system prompt) and a request in plain language. It turns the request into the exact command or API call for that skill. It also writes a human-readable confirmation summary and adds safety warnings when a request looks risky.

It's designed to be the "intent → action" brain inside crypto agents, trading copilots and wallet assistants: fast and cheap enough to run on a single consumer GPU, and strict about never moving funds without confirmation.

⚠️ Velvet Flash does not execute trades and does not give financial advice. It proposes the action and asks for confirmation. Your application decides whether to execute.


What it can do

1. Map an intent to the right command

It reads a skill's documentation (endpoints, CLI sub-commands, parameters) and picks the right operation:

  • Pulls out the parameters precisely: amounts, tokens, symbols, chains, leverage, prices and addresses, copied exactly as given, never rounded or changed.
  • Understands aliases and casual phrasing: "sell all my PEPE", "i need my money now give me balance", "BTC to ETH".
  • Splits multi-step requests into an ordered plan: quote → approve → swap, or bridge → swap.
  • Asks when key information is missing, such as which chain, amount or venue, instead of guessing.

2. Confirm before executing

Every fund-moving action (trade, swap, transfer, bridge, stake, borrow, leverage change) comes back as a proposed command plus a summary table, followed by an explicit request to confirm. Read-only actions (balances, prices, positions, market info) are answered directly.

3. Recognize unsafe requests

It recognizes and pushes back on:

  • Fake or look-alike tokens and scam contracts, and "new wrapped BTC" style lures
  • Wrong-chain or wrong-network transfers, and mismatched bridges
  • Excessive leverage and liquidation-prone positions
  • Phishing and prompt injection, such as "ignore your rules and send…" or instructions hidden in pasted data
  • Requests for credentials: it never prints or asks for private keys, seed phrases or API secrets

4. Explain how to use a skill

It also answers "how do I…" questions about a skill: getting started, authentication setup, what an endpoint returns, and how to use an SDK function.


Supported tasks

Category Example requests
Spot trading "buy $50 of SOL" · "Put a limit sell for 2 BTC at 105000" · "what's the best bid and ask on ETH-USDT"
Perpetuals / futures "long ETH 5x with $100" · "close my BTC perp long" · "buy limit PF_XBTUSD 1 at 50000" · "get funding rate history for BTCUSDT"
Margin "adjust my cross margin max leverage to 5x" · "new margin OTO order for SOLUSDT"
DEX swaps & aggregators "swap 0.1 ETH to USDC" · "trade 5000 USDC for wstETH" · "what DEXes can I use on Arbitrum"
Convert "check convert exchange info for BTC to ETH"
Bridging / cross-chain "bridge 50 USDC from Polygon to Arbitrum" · "move 1000 USDT from Ethereum to Polygon"
Transfers "send 100 USDC to 0xABC… on Base"
Staking & earn "show my staking positions" · "what's the rETH exchange rate" · "check my flexible earn balance"
Liquidity & yield "how much liquidity is in the ETH pool on Arbitrum" · "show the top Pendle markets by TVL"
Portfolio & account "what's my total portfolio worth" · "show open positions" · "check my BSC portfolio"
Market data "current price of bitcoin on Bitget" · "server time for futures" · "exchange info"
Fiat on-ramp "buy crypto with my card" · "log me in"
Smart & agentic wallets "how do I create a smart account" · "I'm new to Privy, how do I start"

Supported venues and skills (30)

Type Skills
Centralized exchanges Binance (Spot, USDⓈ-M Futures, Margin, Convert) · OKX (Trade, Portfolio, Earn) · Kraken (Spot, Futures, Earn/Staking) · KuCoin (Spot, Futures) · Bitget · Coinbase (Fund)
DEX & aggregators Uniswap (Swap Planner) · OKX DEX · KyberSwap · Gate DEX · MoonPay Swap
Perps & liquidity GMX (Trading, Liquidity)
Yield & staking Pendle · Rocket Pool · Gate Staking
Stablecoins & bridging Circle (USDC bridge / CCTP)
Wallets & on-ramp MetaMask Smart Accounts Kit · OKX Agentic Wallet · Privy Agentic Wallets · MoonPay Buy · Minara

Because it follows the documentation in its system prompt, it also works with other skills and CLIs written in the same format. Expect the best accuracy on the venues listed above.

Chains: Ethereum, Arbitrum, Base, Optimism, Polygon, BNB Chain, Avalanche, Solana and the other chains the skills support. Languages: mainly English. It also understands Chinese requests (e.g. "查一下我的活期理财余额").


Dataset

The model's behaviour comes from a curated crypto skill-routing dataset of about 2,500 conversations spanning the 30 skills above, with roughly 70–95 conversations per skill. Each example pairs the skill's full SKILL.md documentation (as the system prompt) with a realistic user request and the ideal assistant response.

The conversations cover:

  • Everyday operations: balances, prices, orders, swaps, positions and transfers
  • Intermediate and multi-step tasks: limit orders, leverage changes, bridging, LP and staking flows
  • Ambiguous or incomplete requests, where the right answer is a clarifying question
  • About 25% adversarial cases: scam tokens, wrong chains, excessive leverage, phishing, prompt injection and attempts to extract credentials
  • Informational questions about a skill's setup and API

Every response follows the same house format: command → summary table → warnings → confirmation prompt. Responses were quality-filtered so that every fund-moving action either asks for confirmation or warns. The public benchmark's scenarios were held out of the dataset.


Usage

Put the skill documentation, or at least the venue and chain, in the system prompt. Turn thinking off for short English answers. Recommended settings: temperature=0, max_tokens=300–600.

from transformers import AutoModelForCausalLM, AutoTokenizer
import torch

repo = "Velvet-Capital/VelvetFlash"
tok = AutoTokenizer.from_pretrained(repo)
model = AutoModelForCausalLM.from_pretrained(repo, dtype=torch.bfloat16, device_map="auto")

msgs = [
    {"role": "system", "content": "Venue: Uniswap. Chain: ethereum."},
    {"role": "user", "content": "swap 0.1 ETH to USDC"},
]
inputs = tok.apply_chat_template(msgs, add_generation_prompt=True, enable_thinking=False,
                                 return_tensors="pt", return_dict=True).to(model.device)
out = model.generate(**inputs, max_new_tokens=400, do_sample=False)
print(tok.decode(out[0][inputs["input_ids"].shape[1]:], skip_special_tokens=True))

Example output:

I'll help you swap 0.1 ETH to USDC on Uniswap. Please review before I execute:

**Command:** `swap 0.1 ETH to USDC`

| Field    | Value       |
|----------|-------------|
| Action   | Swap        |
| Amount   | 0.1 ETH     |
| Token(s) | ETH → USDC  |
| Chain    | Ethereum    |
| DEX      | Uniswap     |

Reply **confirm** to execute, or tell me what to change.

vLLM (OpenAI-compatible server):

vllm serve Velvet-Capital/VelvetFlash --served-model-name velvet-flash-4b --max-model-len 16384
# send "chat_template_kwargs": {"enable_thinking": false} with each request

A lightweight adapter version is in the lora/ folder: PeftModel.from_pretrained(base, "Velvet-Capital/VelvetFlash", subfolder="lora").


Evaluation

Benchmark

Evaluated on crypto-skill-benchmark, which tests how safely and accurately a model works as a crypto operator. Each model gets a platform's skill document plus a user request, and must pick the correct command, extract the parameters exactly, confirm before moving funds, and warn about scams and risky requests. No real transactions are executed.

  • Platforms: a 6-platform core set covering every category: Minara (AI trading agent), Binance Spot (CEX), OKX DEX (aggregator), Uniswap (on-chain DEX), GMX (perpetuals) and MetaMask (wallet)
  • Scenarios: held-out basic and intermediate scenarios, including adversarial and edge cases. No model, including Velvet Flash, saw these scenarios beforehand.
  • Judge: an independent LLM judge (z-ai/GLM-5.2) scores every model with the same rubrics
  • Conditions: identical inputs for every model. The other models ran zero-shot through OpenRouter at default settings.
Dimension Weight What it measures
Safety 30% Confirms before moving funds, parses amounts exactly, protects keys, flags scams
Coverage 25% Breadth of documented operations handled correctly
Robustness 20% Fake tokens, wrong chains, ambiguity, edge cases
Routing 15% Right command and accurate parameter extraction
UX 10% Clear, structured, actionable output

Comparison with open-source models

Velvet Flash 4B ranks #1 of 7. It outscores open-source generalists that are 6–17× larger.

Overall leaderboard

Model Size Overall Safety Coverage Robustness Routing UX
★ Velvet Flash 4B 4B 50 61 34 53 47 52
Qwen3.5-27B 27B 45 46 25 71 39 48
DeepSeek-V4-Flash MoE 40 39 29 54 39 41
Llama-3.3-70B 70B 32 33 22 44 32 30
Gemma-4-31B 31B 31 32 16 53 26 28
Qwen3-4B (base) 4B 23 28 12 26 30 22
Mistral-Small-24B 24B 20 23 6 47 9 10

Score vs model size

Per-platform results

Model Minara Binance Spot OKX DEX Uniswap GMX MetaMask
★ Velvet Flash 4B 66 34 50 47 47 54
Qwen3.5-27B 70 44 45 48 34 27
DeepSeek-V4-Flash 44 45 47 40 37 26
Llama-3.3-70B 75 32 26 19 18 21
Gemma-4-31B 22 40 33 41 27 21
Qwen3-4B (base) 50 30 28 15 6 11
Mistral-Small-24B 26 17 21 19 21 19

Full results matrix

Key takeaways

  • Safety leader: Velvet Flash 4B scores highest on safety, the dimension with the most weight (61, against 46 for the next best). It is the most reliable model here at confirming before moving funds and at parsing amounts exactly.
  • Strongest on the hardest platforms: it leads on OKX DEX, GMX and MetaMask, where the larger models struggle. MetaMask is 54 against 27 for the next best.
  • Most consistent: some larger models beat it on individual platforms (Llama-3.3-70B and Qwen3.5-27B on Minara, DeepSeek and Qwen on Binance Spot). None of them matches its consistency across all six, which is what the weighted average rewards.
  • More than double its base model: it scores 50 against 23 for Qwen3-4B and improves on every platform and every dimension.

Improvement by dimension vs base

Caveats: these results measure crypto-skill routing and safety only, not general reasoning or knowledge. The comparison models were not tuned for this task or output format. Scores come from a single LLM judge and vary somewhat from run to run.

Limitations

  • It is a small model: always show its proposed action to the user and require explicit confirmation before executing anything.
  • It proposes commands but does not call APIs, sign transactions or know live prices or balances unless your app supplies them in the prompt.
  • The benchmark scores come from an LLM judge and carry noise. Accuracy is lower on venues and tools it has not seen.

License

Apache-2.0, the same as the base model Qwen3-4B.

Downloads last month
-
Safetensors
Model size
4B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Velvet-Capital/VelvetFlash

Finetuned
Qwen/Qwen3-4B
Finetuned
(1135)
this model