Models runnable on my workstation (VRAM/RAM friendly). Practical shortlist for local experiments and ICE integration.
🔄 In a Training Loop
Francesco Maiomascio
francescomaiomascio
AI & ML interests
Public reference hub for ML baselines, evaluation resources, and demos.
Publishing is restricted to reproducible artifacts with clear scope, limitations, and documentation.
Recent Activity
liked a model 17 days ago
deepseek-ai/DeepSeek-V4.1-Flash published a model 21 days ago
yailabs/DeepSeek-V4-Flash-GGUF updated a model 21 days ago
yailabs/DeepSeek-V4-Flash-GGUFOrganizations
ICE • Code & Tool-Use
Code-oriented LLMs and tool-use models for agent execution workflows.
Used to evaluate planning, refactoring, and structured action outputs.
-
Qwen/Qwen2.5-Coder-32B-Instruct
Text Generation • 33B • Updated • 1.06M • • 2.16k -
Qwen/Qwen2.5-Coder-7B-Instruct
Text Generation • 8B • Updated • 2.39M • • 809 -
deepseek-ai/deepseek-coder-6.7b-instruct
Text Generation • 7B • Updated • 223k • 511 -
deepseek-ai/deepseek-coder-33b-instruct
Text Generation • 33B • Updated • 4.82k • 584
ICE • Multimodal (Vision + Speech)
Vision-language and speech models for multimodal IO and perception tasks.
Reference set for captioning, OCR-ish flows, ASR, and VLM reasoning.
-
Qwen/Qwen2.5-VL-7B-Instruct
Image-Text-to-Text • 8B • Updated • 6.26M • • 1.72k -
llava-hf/llava-v1.6-mistral-7b-hf
Image-Text-to-Text • 8B • Updated • 498k • 315 -
Salesforce/blip2-opt-2.7b
Image-Text-to-Text • 4B • Updated • 713k • 448 -
google/siglip-so400m-patch14-384
Zero-Shot Image Classification • 0.9B • Updated • 1.42M • 692
Research • Archive
Long-term archive of papers, models, datasets, and tools worth revisiting.
Curated for reference, replication, and future deep dives.
-
The Impact of Hyperparameters on Large Language Model Inference Performance: An Evaluation of vLLM and HuggingFace Pipelines
Paper • 2408.01050 • Published • 8 -
Agent-as-a-Judge: Evaluate Agents with Agents
Paper • 2410.10934 • Published • 22 -
Agent-SafetyBench: Evaluating the Safety of LLM Agents
Paper • 2412.14470 • Published • 12 -
IntellAgent: A Multi-Agent Framework for Evaluating Conversational AI Systems
Paper • 2501.11067 • Published • 12
ICE • Core LLM Baselines
Baseline LLMs for ICE experimentation and regression checks.
Reference set for capability, cost, and stability comparisons.
-
meta-llama/Llama-3.1-8B-Instruct
Text Generation • 8B • Updated • 6.08M • • 8.01k -
meta-llama/Llama-3.1-70B-Instruct
Text Generation • 71B • Updated • 113k • • 995 -
meta-llama/Llama-3.1-8B
Text Generation • 8B • Updated • 566k • • 2.55k -
meta-llama/Llama-3.1-70B
Text Generation • 71B • Updated • 33.6k • • 440
ICE • Retrieval (Embeddings + Rerankers)
Embedding models and rerankers for RAG, search, and semantic indexing.
Reference set for retrieval quality, latency, and multilingual coverage.
ICE • Safety / Guardrails
Safety classifiers and guard models for prompt/content filtering and policy checks.
Reference set for runtime gating, moderation, and risk control lay
Infra • Serving & Optimization
Inference engines, quantization, serving stacks, and perf tooling. Reference list for deployment and latency/cost work.
-
The Impact of Hyperparameters on Large Language Model Inference Performance: An Evaluation of vLLM and HuggingFace Pipelines
Paper • 2408.01050 • Published • 8 -
Seesaw: High-throughput LLM Inference via Model Re-sharding
Paper • 2503.06433 • Published -
PRIMA.CPP: Speeding Up 70B-Scale LLM Inference on Low-Resource Everyday Home Clusters
Paper • 2504.08791 • Published • 141 - Running360
Evaluation Guidebook
📝360Explore LLM benchmark scores over time
Local • Workstation-Ready (≤14B)
Models runnable on my workstation (VRAM/RAM friendly). Practical shortlist for local experiments and ICE integration.
ICE • Core LLM Baselines
Baseline LLMs for ICE experimentation and regression checks.
Reference set for capability, cost, and stability comparisons.
-
meta-llama/Llama-3.1-8B-Instruct
Text Generation • 8B • Updated • 6.08M • • 8.01k -
meta-llama/Llama-3.1-70B-Instruct
Text Generation • 71B • Updated • 113k • • 995 -
meta-llama/Llama-3.1-8B
Text Generation • 8B • Updated • 566k • • 2.55k -
meta-llama/Llama-3.1-70B
Text Generation • 71B • Updated • 33.6k • • 440
ICE • Code & Tool-Use
Code-oriented LLMs and tool-use models for agent execution workflows.
Used to evaluate planning, refactoring, and structured action outputs.
-
Qwen/Qwen2.5-Coder-32B-Instruct
Text Generation • 33B • Updated • 1.06M • • 2.16k -
Qwen/Qwen2.5-Coder-7B-Instruct
Text Generation • 8B • Updated • 2.39M • • 809 -
deepseek-ai/deepseek-coder-6.7b-instruct
Text Generation • 7B • Updated • 223k • 511 -
deepseek-ai/deepseek-coder-33b-instruct
Text Generation • 33B • Updated • 4.82k • 584
ICE • Retrieval (Embeddings + Rerankers)
Embedding models and rerankers for RAG, search, and semantic indexing.
Reference set for retrieval quality, latency, and multilingual coverage.
ICE • Multimodal (Vision + Speech)
Vision-language and speech models for multimodal IO and perception tasks.
Reference set for captioning, OCR-ish flows, ASR, and VLM reasoning.
-
Qwen/Qwen2.5-VL-7B-Instruct
Image-Text-to-Text • 8B • Updated • 6.26M • • 1.72k -
llava-hf/llava-v1.6-mistral-7b-hf
Image-Text-to-Text • 8B • Updated • 498k • 315 -
Salesforce/blip2-opt-2.7b
Image-Text-to-Text • 4B • Updated • 713k • 448 -
google/siglip-so400m-patch14-384
Zero-Shot Image Classification • 0.9B • Updated • 1.42M • 692
ICE • Safety / Guardrails
Safety classifiers and guard models for prompt/content filtering and policy checks.
Reference set for runtime gating, moderation, and risk control lay
Research • Archive
Long-term archive of papers, models, datasets, and tools worth revisiting.
Curated for reference, replication, and future deep dives.
-
The Impact of Hyperparameters on Large Language Model Inference Performance: An Evaluation of vLLM and HuggingFace Pipelines
Paper • 2408.01050 • Published • 8 -
Agent-as-a-Judge: Evaluate Agents with Agents
Paper • 2410.10934 • Published • 22 -
Agent-SafetyBench: Evaluating the Safety of LLM Agents
Paper • 2412.14470 • Published • 12 -
IntellAgent: A Multi-Agent Framework for Evaluating Conversational AI Systems
Paper • 2501.11067 • Published • 12
Infra • Serving & Optimization
Inference engines, quantization, serving stacks, and perf tooling. Reference list for deployment and latency/cost work.
-
The Impact of Hyperparameters on Large Language Model Inference Performance: An Evaluation of vLLM and HuggingFace Pipelines
Paper • 2408.01050 • Published • 8 -
Seesaw: High-throughput LLM Inference via Model Re-sharding
Paper • 2503.06433 • Published -
PRIMA.CPP: Speeding Up 70B-Scale LLM Inference on Low-Resource Everyday Home Clusters
Paper • 2504.08791 • Published • 141 - Running360
Evaluation Guidebook
📝360Explore LLM benchmark scores over time