Macaron-V1.1
🚀 Hosted API: Mint Recursive (International) · Mint Recursive (Mainland China)
🧩 Artifacts: Macaron Artifacts
🛠️ Serving project: Mixture of LoRA (MoL) serving harness
✉️ Correspondence: contact@mindlab.ltd
Macaron-V1.1 is a 752B-parameter agent model from MindLab Research, post-trained on GLM-5.3. It combines a 744B base model with four 2B LoRA specialists for Chat, Agent, Coding, and Generative UI.
Built with Mint Recursive, Mind Lab's serverless training and inference platform, Macaron-V1.1 follows Macaron-V1, which was based on GLM-5.2. The iteration from V1 to V1.1 was completed in two weeks, using a shared workflow for data preparation, experiments, checkpoint evaluation, and deployment.
Highlights
- 752B release scale: a 744B GLM-5.3 base with four release-labeled 2B LoRA specialists.
- Long-horizon coding: SWE trajectories are organized into reproduction, localization, editing, verification, and recovery to reduce low-relevance exploration. Reported scores are 71.7 on DeepSWE v1.1 and 50.0 on SWE-Marathon.
- Complex office workflows: training emphasizes selecting appropriate tools and avoiding unnecessary actions across documents, messages, and business systems, within the user's intent and authorization. AutomationBench reaches 53.7.
- Chat and Generative UI: expanded real-world scenarios in ChatBench v2 and UI4ABench v2, with reported scores of 66.7 and 83.2 respectively.
- First-attempt UI delivery: 58/60 tasks (96.67%), compared with GLM-5.3's 43/60 (71.67%), a 25-percentage-point increase.
Model Overview
| Field | Value |
|---|---|
| Model name | Macaron-V1.1 |
| Organization | MindLab Research |
| Base model | GLM-5.3 |
| Architecture | GLM-5.3 base + Mixture of LoRA (MoL) specialists |
| Parameter footprint | 752B release label: 744B base + four 2B LoRA specialists |
| Specialists | Chat, Agent, Coding, Generative UI |
| Post-training platform | Mint Recursive |
| Primary domains | Chat, personal-agent and office workflows, coding, Generative UI |
| License | MIT |
Mixture of LoRA (MoL) Architecture
| Specialist | Release-labeled size | Focus |
|---|---|---|
| Chat | 2B | Task progress and interaction quality in open-ended conversations |
| Agent | 2B | Tool selection and execution in complex workflows |
| Coding | 2B | Long-horizon software engineering and terminal tasks |
| Generative UI | 2B | UI delivery, functionality, and visual design |
The specialists share a 744B GLM-5.3 base.
Evaluation
| Category | Benchmark | Macaron V1.1 | Macaron V1 | GLM 5.3 | DeepSeek V4 Pro 0813 | Qwen 3.8 Max | Kimi K3 | Claude Opus 5 |
|---|---|---|---|---|---|---|---|---|
| Chat | ChatBench v2 | 66.7 | 63.6 | 65.8 | 65.7 | 62.2 | 57.7 | 66.2 |
| Agent | AutomationBench | 53.7 | 31.8 | 48.2* | 43.2* | 39.8* | 46.7* | 50.3* |
| Agent | Toolathlon-Verified | 76.0 | 63.0 | 73.0* | 74.1* | 72.5* | 73.2* | 80.6* |
| Coding | DeepSWE v1.1 | 71.7 | 58.4 | 66.9* | 62.7* | 57.0* | 69.0* | 68.8* |
| Coding | SWE-Marathon | 50.0 | 15.0 | 42.5* | 10.6* | — | 48.1* | 50.0* |
| Coding | Terminal-Bench 3.0 | 31.4 | 5.7 | 28.7 | 11.8* | 29.0* | 17.7* | 42.7* |
| GenUI | UI4ABench v2 | 83.2 | 75.8 | 77.6 | 77.7 | 79.3 | 80.6 | 81.4 |
Higher is better. Bold marks the highest displayed score per row; * marks an externally sourced benchmark result; — means unavailable. The table reproduces the announcement chart, including its V1 comparison column. External results retain their original protocols and may differ in harness, task subset, execution budget, and aggregation. The table does not establish a controlled overall model ranking.
Macaron-V1.1 exceeds or matches the displayed Opus 5 scores on ChatBench v2, AutomationBench, DeepSWE v1.1, SWE-Marathon, and UI4ABench v2. It scores below Opus 5 on Toolathlon-Verified and Terminal-Bench 3.0. These comparisons retain the protocol qualifications above.
Evaluation Protocols
| Benchmark | Reported setup and metric |
|---|---|
| ChatBench v2 | Three independent runs per model–case pair using the production system prompt, user persona, and relevant conversation history. A privately deployed GLM-5.2 judge rates responses on a 1–5 scale; ratings are averaged over runs and cases and multiplied by 20. |
| AutomationBench | Public 600-task split of v1.0.6 across six business domains, using the API toolset with at most 50 model-response steps per task. |
| Toolathlon-Verified | Official Toolathlon evaluation service; pass@1. |
| DeepSWE v1.1 | Claude Code agent harness; pass@3, with a task solved if any of up to three attempts succeeds. External official leaderboard results use mini-swe-agent. |
| SWE-Marathon | pass@2. External leaderboard scores retain their respective published settings. |
| Terminal-Bench 3.0 | Claude Code v2.1.207; pass@3. Client-reported output limit of 32,000 tokens per response. Agent timeouts are 10× each task's configured timeout, corresponding to 5–80 hours. |
| UI4ABench v2 | 60 real-world-inspired UI-generation tasks of varying difficulty. Delivery measures first-attempt success; functional completeness is assessed through execution tests and visual quality through rendered screenshots. |
The announcement chart identifies these external sources: DeepSWE, SWE-Marathon, Toolathlon results, the Claude Opus 5 System Card for its Toolathlon score, BenchLM for its Terminal-Bench score, and the DeepSeek blog for Kimi K3 and DeepSeek V4 Pro Terminal-Bench scores. These are source attributions from the supplied chart, not independently revalidated leaderboard snapshots.
The Toolathlon-Verified, DeepSWE v1.1, and Terminal-Bench 3.0 results are also registered in .eval_results/ against Hugging Face benchmark datasets, so Hugging Face can link them to their benchmark leaderboards. These files do not include HF Jobs verifyToken values; they should therefore be displayed as self-reported rather than verified results. ChatBench v2, AutomationBench, SWE-Marathon, and UI4ABench v2 remain model-card self-reported results because their corresponding HF benchmark registrations are not available for this release.
Reproducibility details pending: checkpoint and adapter revisions, task-set revisions where unspecified, evaluation dates, complete generation settings, raw traces, confidence intervals, and train/evaluation overlap analysis. The ChatBench judge belongs to the same GLM family as the base model; cross-family judge calibration is not provided in the announcement.
Training with Mint Recursive
Macaron-V1.1 was trained with Mint Recursive, Mind Lab's serverless model-training platform. The workflow integrates data preparation, experimentation, checkpoint evaluation, and deployment, enabling the iteration from Macaron-V1 to V1.1 to be completed in two weeks.
Training focused on four capability areas:
- Coding: Structuring long-horizon software-engineering trajectories around reproduction, localization, editing, verification, and recovery.
- Agent: Improving tool selection and execution across documents, messaging, and business workflows.
- Chat: Balancing task completion with interaction quality, emotional awareness, and resistance to sycophancy.
- Generative UI: Improving first-attempt delivery, functional completeness, information hierarchy, and visual quality.
Usage
Hosted API
- International users: Mint Recursive.
- Mainland China users: Mint Recursive China.
Consult the platform documentation for available model IDs, authentication, pricing, and rate limits.
Open Weights and Self-Hosted Serving
This repository includes the base checkpoint at the repository root and four LoRA specialists under loras/L0 through loras/L3. Each LoRA directory contains adapter_config.json and adapter_model.safetensors, following the Macaron-V1-Venti release layout.
License
This repository is released under the MIT License. Users should also respect any requirements inherited from the GLM-5.3 base model and from dependencies used by the serving harness.
Contact
- Organization: MindLab Research
- Correspondence: contact@mindlab.ltd
- Downloads last month
- 261
Evaluation results
- datacurve/deep-swe · Deep Swe View evaluation results leaderboard
- harborframework/terminal-bench · Terminalbench 3 View evaluation results leaderboard
- hkust-nlp/Toolathlon · Toolathlon Verified View evaluation results
- Score on ChatBench v2self-reported66.700
- Score on AutomationBenchself-reported53.700
- pass@2 (%) on SWE-Marathonself-reported50.000
- Score on UI4ABench v2self-reported83.200
