Instructions to use consciousengines/CE-flicker-eco with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use consciousengines/CE-flicker-eco with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="consciousengines/CE-flicker-eco", trust_remote_code=True) messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoModelForCausalLM model = AutoModelForCausalLM.from_pretrained("consciousengines/CE-flicker-eco", trust_remote_code=True, device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use consciousengines/CE-flicker-eco with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "consciousengines/CE-flicker-eco" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "consciousengines/CE-flicker-eco", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/consciousengines/CE-flicker-eco
- SGLang
How to use consciousengines/CE-flicker-eco with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "consciousengines/CE-flicker-eco" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "consciousengines/CE-flicker-eco", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "consciousengines/CE-flicker-eco" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "consciousengines/CE-flicker-eco", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use consciousengines/CE-flicker-eco with Docker Model Runner:
docker model run hf.co/consciousengines/CE-flicker-eco
CE-flicker-eco is an internal, temporary name
Block-diffusion chat with FP8 model weights and learned EccoMoE expert allocation. Completed blocks stream left to right. The current assistant introduces itself as Flicker-Eco; its API model ID is CE-flicker.
Three serving profiles
All use normal BF16 KV, block size 4, confidence threshold 0.95, and up to four denoising steps.
| Profile | Model-weight precision | Expert routing | Tested hardware |
|---|---|---|---|
| Stock BF16 | BF16 | Native top-8 | H200 |
| Stock FP8 | Blocked FP8 | Native top-8 | H200 and L40S |
| EccoMoE FP8 t6 | Blocked FP8 | Learned expert-count allocator, budget 6 | L40S 路 default selected serving profile |
Budget 6 is a compute target; actual expert counts vary by token and layer. The expert allocator is separate from the denoising confidence threshold.
Measured speed, energy and quality
- FP8 stock on H200: 26.5% higher measured peak throughput and 21.6% less GPU energy per output token than BF16 stock.
- EccoMoE t6 on L40S: 9.17% higher throughput and 10.53% less GPU J/token at eight concurrent requests, with an observed 4.67-point MBPP accuracy cost.
H200: stock BF16 versus stock FP8
Same SDAR checkpoint and lmdeploy backend; 1,024 input + forced 1,024 output tokens, no chat/system prompt. Both peaks occur at concurrency 512 and are three-run medians.
| Profile | Peak total output tok/s | Concurrency | GPU J/output token |
|---|---|---|---|
| Stock BF16 | 6,257 | 512 | 0.11027 |
| Stock FP8 | 7,916 | 512 | 0.08640 |
L40S: stock FP8 versus EccoMoE FP8 t6
Same physical GPU and approved prompt. Throughput/energy use forced 256-token outputs; MBPP is a separate paired, natural-stop quality test.
| Profile | 1-user output tok/s | Total output tok/s 路 8 requests | GPU J/token 路 8 requests | Prompted MBPP |
|---|---|---|---|---|
| Stock FP8 | 67.55 | 268.72 | 0.9859 | 74.71% 路 192/257 |
| EccoMoE FP8 t6 | 77.82 | 293.37 | 0.8821 | 70.04% 路 180/257 |
Stock speed/energy pool six repeats; t6 has three. Stock single-user speed drifted 5.1% between phases. MBPP uses one deterministic run per profile: 25 tasks pass only with stock, 13 only with t6; neither run timed out or truncated. The tests preceded the identity-only rename to Flicker-Eco. Prompted quality comparison covers MBPP; current-prompt full MATH/IFEval comparisons remain outside this measured pair.
The default selected serving profile is EccoMoE t6, with this observed tradeoff. These are distinct hardware/workloads; GPU energy excludes CPU and wall-socket power. Full methods, load levels and metric records.
Setup and artifacts
The public checkpoint stores BF16 backbone weights. FP8 in these tables is validated model-weight serving precision via lmdeploy runtime conversion; the KV cache remains BF16.
The package includes the 49 backbone shards, the separate trained allocator, routing source, native-kernel compatibility probe, authenticated OpenAI-compatible streaming gateway and approved prompt. Current served configuration: lmdeploy 0.18.0, FP8 weights / BF16 KV, EccoMoE t6, 8,192 total-context tokens and maximum engine batch 8.
Reproduce the selected profile or either stock profile 路 Allocator integrity check 路 Source provenance and licenses 路 Package checksums.
Provenance
The frozen backbone is SDAR-30B-A3B-Chat, pinned to f5add2a159163a2a8f07e9da7dfcbfaadc73d6d4, under Apache 2.0. Upstream authors retain credit for base-model training and the SDAR algorithm; Conscious Engines supplies the trained allocator and serving/routing adaptation. The original backbone files are preserved.
- Downloads last month
- -
Model tree for consciousengines/CE-flicker-eco
Base model
JetLM/SDAR-30B-A3B-Chat