Models runnable on my workstation (VRAM/RAM friendly). Practical shortlist for local experiments and ICE integration.
🔄 In a Training Loop
Francesco Maiomascio
francescomaiomascio
AI & ML interests
Public reference hub for ML baselines, evaluation resources, and demos.
Publishing is restricted to reproducible artifacts with clear scope, limitations, and documentation.
Recent Activity
liked a model 9 days ago
MiniMaxAI/MiniMax-H3 published a model 10 days ago
yailabs/DeepSeek-V4-Flash-DSpark-YVEX-GGUF liked a model 18 days ago
deepseek-ai/DeepSeek-V4-FlashOrganizations
ICE • Code & Tool-Use
Code-oriented LLMs and tool-use models for agent execution workflows.
Used to evaluate planning, refactoring, and structured action outputs.
-
Qwen/Qwen2.5-Coder-32B-Instruct
Text Generation • 33B • Updated • 1.24M • • 2.1k -
Qwen/Qwen2.5-Coder-7B-Instruct
Text Generation • 8B • Updated • 1.93M • • 771 -
deepseek-ai/deepseek-coder-6.7b-instruct
Text Generation • 7B • Updated • 519k • 508 -
deepseek-ai/deepseek-coder-33b-instruct
Text Generation • 33B • Updated • 3.17k • 581
ICE • Multimodal (Vision + Speech)
Vision-language and speech models for multimodal IO and perception tasks.
Reference set for captioning, OCR-ish flows, ASR, and VLM reasoning.
-
Qwen/Qwen2.5-VL-7B-Instruct
Image-Text-to-Text • 8B • Updated • 8.86M • • 1.68k -
llava-hf/llava-v1.6-mistral-7b-hf
Image-Text-to-Text • 8B • Updated • 696k • 313 -
Salesforce/blip2-opt-2.7b
Image-Text-to-Text • 4B • Updated • 497k • 447 -
google/siglip-so400m-patch14-384
Zero-Shot Image Classification • 0.9B • Updated • 2.19M • 682
Research • Archive
Long-term archive of papers, models, datasets, and tools worth revisiting.
Curated for reference, replication, and future deep dives.
-
The Impact of Hyperparameters on Large Language Model Inference Performance: An Evaluation of vLLM and HuggingFace Pipelines
Paper • 2408.01050 • Published • 9 -
Agent-as-a-Judge: Evaluate Agents with Agents
Paper • 2410.10934 • Published • 23 -
Agent-SafetyBench: Evaluating the Safety of LLM Agents
Paper • 2412.14470 • Published • 12 -
IntellAgent: A Multi-Agent Framework for Evaluating Conversational AI Systems
Paper • 2501.11067 • Published • 13
ICE • Core LLM Baselines
Baseline LLMs for ICE experimentation and regression checks.
Reference set for capability, cost, and stability comparisons.
-
meta-llama/Llama-3.1-8B-Instruct
Text Generation • 8B • Updated • 7.55M • • 6.59k -
meta-llama/Llama-3.1-70B-Instruct
Text Generation • 71B • Updated • 793k • • 942 -
meta-llama/Llama-3.1-8B
Text Generation • 8B • Updated • 1.16M • • 2.36k -
meta-llama/Llama-3.1-70B
Text Generation • 71B • Updated • 41.1k • • 432
ICE • Retrieval (Embeddings + Rerankers)
Embedding models and rerankers for RAG, search, and semantic indexing.
Reference set for retrieval quality, latency, and multilingual coverage.
ICE • Safety / Guardrails
Safety classifiers and guard models for prompt/content filtering and policy checks.
Reference set for runtime gating, moderation, and risk control lay
Infra • Serving & Optimization
Inference engines, quantization, serving stacks, and perf tooling. Reference list for deployment and latency/cost work.
-
The Impact of Hyperparameters on Large Language Model Inference Performance: An Evaluation of vLLM and HuggingFace Pipelines
Paper • 2408.01050 • Published • 9 -
Seesaw: High-throughput LLM Inference via Model Re-sharding
Paper • 2503.06433 • Published -
PRIMA.CPP: Speeding Up 70B-Scale LLM Inference on Low-Resource Everyday Home Clusters
Paper • 2504.08791 • Published • 141 - Running348
Evaluation Guidebook
📝348Explore LLM benchmark scores over time
Local • Workstation-Ready (≤14B)
Models runnable on my workstation (VRAM/RAM friendly). Practical shortlist for local experiments and ICE integration.
ICE • Core LLM Baselines
Baseline LLMs for ICE experimentation and regression checks.
Reference set for capability, cost, and stability comparisons.
-
meta-llama/Llama-3.1-8B-Instruct
Text Generation • 8B • Updated • 7.55M • • 6.59k -
meta-llama/Llama-3.1-70B-Instruct
Text Generation • 71B • Updated • 793k • • 942 -
meta-llama/Llama-3.1-8B
Text Generation • 8B • Updated • 1.16M • • 2.36k -
meta-llama/Llama-3.1-70B
Text Generation • 71B • Updated • 41.1k • • 432
ICE • Code & Tool-Use
Code-oriented LLMs and tool-use models for agent execution workflows.
Used to evaluate planning, refactoring, and structured action outputs.
-
Qwen/Qwen2.5-Coder-32B-Instruct
Text Generation • 33B • Updated • 1.24M • • 2.1k -
Qwen/Qwen2.5-Coder-7B-Instruct
Text Generation • 8B • Updated • 1.93M • • 771 -
deepseek-ai/deepseek-coder-6.7b-instruct
Text Generation • 7B • Updated • 519k • 508 -
deepseek-ai/deepseek-coder-33b-instruct
Text Generation • 33B • Updated • 3.17k • 581
ICE • Retrieval (Embeddings + Rerankers)
Embedding models and rerankers for RAG, search, and semantic indexing.
Reference set for retrieval quality, latency, and multilingual coverage.
ICE • Multimodal (Vision + Speech)
Vision-language and speech models for multimodal IO and perception tasks.
Reference set for captioning, OCR-ish flows, ASR, and VLM reasoning.
-
Qwen/Qwen2.5-VL-7B-Instruct
Image-Text-to-Text • 8B • Updated • 8.86M • • 1.68k -
llava-hf/llava-v1.6-mistral-7b-hf
Image-Text-to-Text • 8B • Updated • 696k • 313 -
Salesforce/blip2-opt-2.7b
Image-Text-to-Text • 4B • Updated • 497k • 447 -
google/siglip-so400m-patch14-384
Zero-Shot Image Classification • 0.9B • Updated • 2.19M • 682
ICE • Safety / Guardrails
Safety classifiers and guard models for prompt/content filtering and policy checks.
Reference set for runtime gating, moderation, and risk control lay
Research • Archive
Long-term archive of papers, models, datasets, and tools worth revisiting.
Curated for reference, replication, and future deep dives.
-
The Impact of Hyperparameters on Large Language Model Inference Performance: An Evaluation of vLLM and HuggingFace Pipelines
Paper • 2408.01050 • Published • 9 -
Agent-as-a-Judge: Evaluate Agents with Agents
Paper • 2410.10934 • Published • 23 -
Agent-SafetyBench: Evaluating the Safety of LLM Agents
Paper • 2412.14470 • Published • 12 -
IntellAgent: A Multi-Agent Framework for Evaluating Conversational AI Systems
Paper • 2501.11067 • Published • 13
Infra • Serving & Optimization
Inference engines, quantization, serving stacks, and perf tooling. Reference list for deployment and latency/cost work.
-
The Impact of Hyperparameters on Large Language Model Inference Performance: An Evaluation of vLLM and HuggingFace Pipelines
Paper • 2408.01050 • Published • 9 -
Seesaw: High-throughput LLM Inference via Model Re-sharding
Paper • 2503.06433 • Published -
PRIMA.CPP: Speeding Up 70B-Scale LLM Inference on Low-Resource Everyday Home Clusters
Paper • 2504.08791 • Published • 141 - Running348
Evaluation Guidebook
📝348Explore LLM benchmark scores over time