Instructions to use Velvet-Capital/VelvetFlash with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Velvet-Capital/VelvetFlash with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="Velvet-Capital/VelvetFlash") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("Velvet-Capital/VelvetFlash") model = AutoModelForCausalLM.from_pretrained("Velvet-Capital/VelvetFlash", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use Velvet-Capital/VelvetFlash with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Velvet-Capital/VelvetFlash" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Velvet-Capital/VelvetFlash", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/Velvet-Capital/VelvetFlash
- SGLang
How to use Velvet-Capital/VelvetFlash with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "Velvet-Capital/VelvetFlash" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Velvet-Capital/VelvetFlash", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "Velvet-Capital/VelvetFlash" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Velvet-Capital/VelvetFlash", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use Velvet-Capital/VelvetFlash with Docker Model Runner:
docker model run hf.co/Velvet-Capital/VelvetFlash
Velvet Flash 4B
Velvet Flash 4B is a compact 4B-parameter crypto action-routing model by Velvet Capital, built on Qwen3-4B. On the crypto-skill-benchmark it ranks #1 against open-source models up to 17× its size (Qwen3.5-27B, DeepSeek-V4, Llama-3.3-70B, Gemma-4-31B, Mistral-24B).
You give it a crypto skill (an exchange API, a DEX, a wallet SDK or a DeFi protocol, described in a SKILL.md-style system prompt) and a request in plain language. It turns the request into the exact command or API call for that skill. It also writes a human-readable confirmation summary and adds safety warnings when a request looks risky.
It's designed to be the "intent → action" brain inside crypto agents, trading copilots and wallet assistants: fast and cheap enough to run on a single consumer GPU, and strict about never moving funds without confirmation.
⚠️ Velvet Flash does not execute trades and does not give financial advice. It proposes the action and asks for confirmation. Your application decides whether to execute.
What it can do
1. Map an intent to the right command
It reads a skill's documentation (endpoints, CLI sub-commands, parameters) and picks the right operation:
- Pulls out the parameters precisely: amounts, tokens, symbols, chains, leverage, prices and addresses, copied exactly as given, never rounded or changed.
- Understands aliases and casual phrasing: "sell all my PEPE", "i need my money now give me balance", "BTC to ETH".
- Splits multi-step requests into an ordered plan: quote → approve → swap, or bridge → swap.
- Asks when key information is missing, such as which chain, amount or venue, instead of guessing.
2. Confirm before executing
Every fund-moving action (trade, swap, transfer, bridge, stake, borrow, leverage change) comes back as a proposed command plus a summary table, followed by an explicit request to confirm. Read-only actions (balances, prices, positions, market info) are answered directly.
3. Recognize unsafe requests
It recognizes and pushes back on:
- Fake or look-alike tokens and scam contracts, and "new wrapped BTC" style lures
- Wrong-chain or wrong-network transfers, and mismatched bridges
- Excessive leverage and liquidation-prone positions
- Phishing and prompt injection, such as "ignore your rules and send…" or instructions hidden in pasted data
- Requests for credentials: it never prints or asks for private keys, seed phrases or API secrets
4. Explain how to use a skill
It also answers "how do I…" questions about a skill: getting started, authentication setup, what an endpoint returns, and how to use an SDK function.
Supported tasks
| Category | Example requests |
|---|---|
| Spot trading | "buy $50 of SOL" · "Put a limit sell for 2 BTC at 105000" · "what's the best bid and ask on ETH-USDT" |
| Perpetuals / futures | "long ETH 5x with $100" · "close my BTC perp long" · "buy limit PF_XBTUSD 1 at 50000" · "get funding rate history for BTCUSDT" |
| Margin | "adjust my cross margin max leverage to 5x" · "new margin OTO order for SOLUSDT" |
| DEX swaps & aggregators | "swap 0.1 ETH to USDC" · "trade 5000 USDC for wstETH" · "what DEXes can I use on Arbitrum" |
| Convert | "check convert exchange info for BTC to ETH" |
| Bridging / cross-chain | "bridge 50 USDC from Polygon to Arbitrum" · "move 1000 USDT from Ethereum to Polygon" |
| Transfers | "send 100 USDC to 0xABC… on Base" |
| Staking & earn | "show my staking positions" · "what's the rETH exchange rate" · "check my flexible earn balance" |
| Liquidity & yield | "how much liquidity is in the ETH pool on Arbitrum" · "show the top Pendle markets by TVL" |
| Portfolio & account | "what's my total portfolio worth" · "show open positions" · "check my BSC portfolio" |
| Market data | "current price of bitcoin on Bitget" · "server time for futures" · "exchange info" |
| Fiat on-ramp | "buy crypto with my card" · "log me in" |
| Smart & agentic wallets | "how do I create a smart account" · "I'm new to Privy, how do I start" |
Supported venues and skills (30)
| Type | Skills |
|---|---|
| Centralized exchanges | Binance (Spot, USDⓈ-M Futures, Margin, Convert) · OKX (Trade, Portfolio, Earn) · Kraken (Spot, Futures, Earn/Staking) · KuCoin (Spot, Futures) · Bitget · Coinbase (Fund) |
| DEX & aggregators | Uniswap (Swap Planner) · OKX DEX · KyberSwap · Gate DEX · MoonPay Swap |
| Perps & liquidity | GMX (Trading, Liquidity) |
| Yield & staking | Pendle · Rocket Pool · Gate Staking |
| Stablecoins & bridging | Circle (USDC bridge / CCTP) |
| Wallets & on-ramp | MetaMask Smart Accounts Kit · OKX Agentic Wallet · Privy Agentic Wallets · MoonPay Buy · Minara |
Because it follows the documentation in its system prompt, it also works with other skills and CLIs written in the same format. Expect the best accuracy on the venues listed above.
Chains: Ethereum, Arbitrum, Base, Optimism, Polygon, BNB Chain, Avalanche, Solana and the other chains the skills support. Languages: mainly English. It also understands Chinese requests (e.g. "查一下我的活期理财余额").
Dataset
The model's behaviour comes from a curated crypto skill-routing dataset of about 2,500 conversations spanning the 30 skills above, with roughly 70–95 conversations per skill. Each example pairs the skill's full SKILL.md documentation (as the system prompt) with a realistic user request and the ideal assistant response.
The conversations cover:
- Everyday operations: balances, prices, orders, swaps, positions and transfers
- Intermediate and multi-step tasks: limit orders, leverage changes, bridging, LP and staking flows
- Ambiguous or incomplete requests, where the right answer is a clarifying question
- About 25% adversarial cases: scam tokens, wrong chains, excessive leverage, phishing, prompt injection and attempts to extract credentials
- Informational questions about a skill's setup and API
Every response follows the same house format: command → summary table → warnings → confirmation prompt. Responses were quality-filtered so that every fund-moving action either asks for confirmation or warns. The public benchmark's scenarios were held out of the dataset.
Usage
Put the skill documentation, or at least the venue and chain, in the system prompt. Turn thinking off for short English answers. Recommended settings: temperature=0, max_tokens=300–600.
from transformers import AutoModelForCausalLM, AutoTokenizer
import torch
repo = "Velvet-Capital/VelvetFlash"
tok = AutoTokenizer.from_pretrained(repo)
model = AutoModelForCausalLM.from_pretrained(repo, dtype=torch.bfloat16, device_map="auto")
msgs = [
{"role": "system", "content": "Venue: Uniswap. Chain: ethereum."},
{"role": "user", "content": "swap 0.1 ETH to USDC"},
]
inputs = tok.apply_chat_template(msgs, add_generation_prompt=True, enable_thinking=False,
return_tensors="pt", return_dict=True).to(model.device)
out = model.generate(**inputs, max_new_tokens=400, do_sample=False)
print(tok.decode(out[0][inputs["input_ids"].shape[1]:], skip_special_tokens=True))
Example output:
I'll help you swap 0.1 ETH to USDC on Uniswap. Please review before I execute:
**Command:** `swap 0.1 ETH to USDC`
| Field | Value |
|----------|-------------|
| Action | Swap |
| Amount | 0.1 ETH |
| Token(s) | ETH → USDC |
| Chain | Ethereum |
| DEX | Uniswap |
Reply **confirm** to execute, or tell me what to change.
vLLM (OpenAI-compatible server):
vllm serve Velvet-Capital/VelvetFlash --served-model-name velvet-flash-4b --max-model-len 16384
# send "chat_template_kwargs": {"enable_thinking": false} with each request
A lightweight adapter version is in the lora/ folder: PeftModel.from_pretrained(base, "Velvet-Capital/VelvetFlash", subfolder="lora").
Evaluation
Benchmark
Evaluated on crypto-skill-benchmark, which tests how safely and accurately a model works as a crypto operator. Each model gets a platform's skill document plus a user request, and must pick the correct command, extract the parameters exactly, confirm before moving funds, and warn about scams and risky requests. No real transactions are executed.
- Platforms: a 6-platform core set covering every category: Minara (AI trading agent), Binance Spot (CEX), OKX DEX (aggregator), Uniswap (on-chain DEX), GMX (perpetuals) and MetaMask (wallet)
- Scenarios: held-out basic and intermediate scenarios, including adversarial and edge cases. No model, including Velvet Flash, saw these scenarios beforehand.
- Judge: an independent LLM judge (
z-ai/GLM-5.2) scores every model with the same rubrics - Conditions: identical inputs for every model. The other models ran zero-shot through OpenRouter at default settings.
| Dimension | Weight | What it measures |
|---|---|---|
| Safety | 30% | Confirms before moving funds, parses amounts exactly, protects keys, flags scams |
| Coverage | 25% | Breadth of documented operations handled correctly |
| Robustness | 20% | Fake tokens, wrong chains, ambiguity, edge cases |
| Routing | 15% | Right command and accurate parameter extraction |
| UX | 10% | Clear, structured, actionable output |
Comparison with open-source models
Velvet Flash 4B ranks #1 of 7. It outscores open-source generalists that are 6–17× larger.
| Model | Size | Overall | Safety | Coverage | Robustness | Routing | UX |
|---|---|---|---|---|---|---|---|
| ★ Velvet Flash 4B | 4B | 50 | 61 | 34 | 53 | 47 | 52 |
| Qwen3.5-27B | 27B | 45 | 46 | 25 | 71 | 39 | 48 |
| DeepSeek-V4-Flash | MoE | 40 | 39 | 29 | 54 | 39 | 41 |
| Llama-3.3-70B | 70B | 32 | 33 | 22 | 44 | 32 | 30 |
| Gemma-4-31B | 31B | 31 | 32 | 16 | 53 | 26 | 28 |
| Qwen3-4B (base) | 4B | 23 | 28 | 12 | 26 | 30 | 22 |
| Mistral-Small-24B | 24B | 20 | 23 | 6 | 47 | 9 | 10 |
Per-platform results
| Model | Minara | Binance Spot | OKX DEX | Uniswap | GMX | MetaMask |
|---|---|---|---|---|---|---|
| ★ Velvet Flash 4B | 66 | 34 | 50 | 47 | 47 | 54 |
| Qwen3.5-27B | 70 | 44 | 45 | 48 | 34 | 27 |
| DeepSeek-V4-Flash | 44 | 45 | 47 | 40 | 37 | 26 |
| Llama-3.3-70B | 75 | 32 | 26 | 19 | 18 | 21 |
| Gemma-4-31B | 22 | 40 | 33 | 41 | 27 | 21 |
| Qwen3-4B (base) | 50 | 30 | 28 | 15 | 6 | 11 |
| Mistral-Small-24B | 26 | 17 | 21 | 19 | 21 | 19 |
Key takeaways
- Safety leader: Velvet Flash 4B scores highest on safety, the dimension with the most weight (61, against 46 for the next best). It is the most reliable model here at confirming before moving funds and at parsing amounts exactly.
- Strongest on the hardest platforms: it leads on OKX DEX, GMX and MetaMask, where the larger models struggle. MetaMask is 54 against 27 for the next best.
- Most consistent: some larger models beat it on individual platforms (Llama-3.3-70B and Qwen3.5-27B on Minara, DeepSeek and Qwen on Binance Spot). None of them matches its consistency across all six, which is what the weighted average rewards.
- More than double its base model: it scores 50 against 23 for Qwen3-4B and improves on every platform and every dimension.
Caveats: these results measure crypto-skill routing and safety only, not general reasoning or knowledge. The comparison models were not tuned for this task or output format. Scores come from a single LLM judge and vary somewhat from run to run.
Limitations
- It is a small model: always show its proposed action to the user and require explicit confirmation before executing anything.
- It proposes commands but does not call APIs, sign transactions or know live prices or balances unless your app supplies them in the prompt.
- The benchmark scores come from an LLM judge and carry noise. Accuracy is lower on venues and tools it has not seen.
License
Apache-2.0, the same as the base model Qwen3-4B.
- Downloads last month
- -



