--- title: README emoji: ⚡ colorFrom: red colorTo: purple sdk: static pinned: false --- # MetaboLLM [![arXiv](https://img.shields.io/badge/arXiv-2608.06253-b31b1b.svg)](https://arxiv.org/abs/2608.06253) [![GitHub](https://img.shields.io/badge/GitHub-Code-181717.svg)](https://github.com/dohyunku9/MetaboLLM) **MetaboLLM** is a family of metabolomics-specialized large language models designed to integrate biochemical knowledge across heterogeneous resources and support metabolite-, pathway-, reaction-, and enzyme-centered reasoning and description generation. MetaboLLM was developed through continual pretraining, supervised fine-tuning, and structured retrieval using harmonized biochemical knowledge from KEGG, HMDB, PubChem, and SMPDB.

Overview of the MetaboLLM framework

## Highlights - Integrated knowledge covering **237,243 metabolites**, **2,359 pathways**, **12,323 reactions**, and **5,993 enzymes** - Released **four MetaboLLM model variants** across Qwen, Gemma, and Llama backbones - Constructed a metabolomics benchmark containing **17 tasks**, **6,000 training examples**, and **4,200 test examples** - Evaluated factual knowledge, class identification, biochemical relationships, and description generation - Compared MetaboLLM against corresponding base models and five publicly available medical language models - Evaluated transfer on the independently developed MetaBench benchmark ## Resources | Resource | Description | |---|---| | [MetaboLLM-Qwen3-4B](https://huggingface.co/MetaboLLM/MetaboLLM-Qwen3-4B) | Primary MetaboLLM model based on Qwen3-4B | | [MetaboLLM-Qwen3-8B](https://huggingface.co/MetaboLLM/MetaboLLM-Qwen3-8B) | MetaboLLM model based on Qwen3-8B | | [MetaboLLM-Gemma-3-4B](https://huggingface.co/MetaboLLM/MetaboLLM-Gemma-3-4B) | MetaboLLM model based on Gemma-3-4B | | [MetaboLLM-Llama-3.2-3B](https://huggingface.co/MetaboLLM/MetaboLLM-Llama-3.2-3B) | MetaboLLM model based on Llama-3.2-3B | | [MetaboLLM-Benchmark](https://huggingface.co/datasets/MetaboLLM/MetaboLLM-Benchmark) | Training and evaluation benchmark across 17 metabolomics tasks | ## Model Family | Model | Backbone | Release Format | |---|---|---| | MetaboLLM-Qwen3-4B | Qwen3-4B | PEFT LoRA adapter | | MetaboLLM-Qwen3-8B | Qwen3-8B | PEFT LoRA adapter | | MetaboLLM-Gemma-3-4B | Gemma-3-4B | PEFT LoRA adapter | | MetaboLLM-Llama-3.2-3B | Llama-3.2-3B | PEFT LoRA adapter | MetaboLLM-Qwen3-4B served as the primary model in the associated study. All released repositories contain adapter weights and require the corresponding base model. ## Integrated Biochemical Knowledge MetaboLLM was developed from a unified resource integrating complementary information from: - **KEGG** for metabolites, reactions, enzymes, and pathways - **HMDB** for human metabolites, biological roles, chemical taxonomy, and compound descriptions - **PubChem** for chemical structures, identifiers, molecular properties, and compound descriptions - **SMPDB** for curated pathways and physiological descriptions The harmonized resource contains: | Entity type | Count | |---|---:| | Metabolites | 237,243 | | Pathways | 2,359 | | Reactions | 12,323 | | Enzymes | 5,993 | ## MetaboLLM Benchmark The benchmark contains 17 tasks organized into four complementary categories. | Category | Tasks | Train | Test | |---|---:|---:|---:| | Knowledge recall | 5 | 1,000 | 1,000 | | Class identification | 4 | 1,000 | 1,000 | | Relation identification | 3 | 1,000 | 1,000 | | Description generation | 5 | 3,000 | 1,200 | | **Total** | **17** | **6,000** | **4,200** | The benchmark evaluates: - molecular identity and formula knowledge - metabolite and reaction class recognition - metabolite–pathway relationships - metabolite–reaction relationships - reaction–enzyme relationships - metabolite, pathway, and enzyme description generation - structure-rich and structure-poor metabolite description generation ## Benchmark Results ### Biochemical Knowledge and Structured Relationships Mean accuracy across 12 multiple-choice and short-answer tasks. | Category | Model | Mean Accuracy | |---|---|---:| | Medical LLM | MedGemma-1.5-4B | 49.0 | | Medical LLM | FineMedLM-O1 | 57.9 | | Medical LLM | II-Medical-8B-1706 | 63.0 | | Medical LLM | Qwen2.5-Aloe-Beta-7B | 63.9 | | Medical LLM | Meditron3-Qwen2.5-7B | 65.7 | | Base model | Llama-3.2-3B | 51.5 | | Base model | Gemma-3-4B | 52.2 | | Base model | Qwen3-4B | 66.5 | | Base model | Qwen3-8B | 66.8 | | MetaboLLM | MetaboLLM-Llama-3.2-3B | 70.0 | | MetaboLLM | MetaboLLM-Gemma-3-4B | 75.3 | | MetaboLLM | MetaboLLM-Qwen3-8B | 75.9 | | MetaboLLM | **MetaboLLM-Qwen3-4B** | **79.1** | All four MetaboLLM variants outperformed their corresponding unadapted backbones and all evaluated medical language models. MetaboLLM-Qwen3-4B achieved the highest mean accuracy at 79.1%, compared with 65.7% for the strongest evaluated medical model and 66.5% for its corresponding base model. ### Biochemical Description Generation BERTScore-F1 on a 0–100 scale. | Category | Model | Metabolite Description | Pathway Description | Enzyme Description | Structure-Rich Metabolite Description | Structure-Poor Metabolite Description | |---|---|---:|---:|---:|---:|---:| | Medical LLM | MedGemma-1.5-4B | 81.05 | 80.97 | 79.33 | 84.89 | 84.73 | | Medical LLM | FineMedLM-O1 | 83.05 | 82.87 | 81.22 | 84.10 | 83.96 | | Medical LLM | II-Medical-8B-1706 | 82.71 | 82.94 | 81.60 | 86.37 | 85.84 | | Medical LLM | Qwen2.5-Aloe-Beta-7B | 83.34 | 83.39 | 81.68 | 85.47 | 84.52 | | Medical LLM | Meditron3-Qwen2.5-7B | 83.69 | 83.14 | 81.99 | 84.47 | 83.76 | | Base model | Llama-3.2-3B | 82.81 | 82.56 | 81.40 | 84.56 | 84.50 | | Base model | Gemma-3-4B | 81.74 | 81.54 | 80.37 | 85.48 | 85.07 | | Base model | Qwen3-4B | 82.39 | 81.85 | 80.68 | 86.48 | 85.80 | | Base model | Qwen3-8B | 82.30 | 82.28 | 80.79 | 86.20 | 85.61 | | MetaboLLM | MetaboLLM-Llama-3.2-3B | 90.93 | 86.74 | 83.78 | 87.13 | 87.65 | | MetaboLLM | MetaboLLM-Gemma-3-4B | 87.39 | 83.83 | 80.69 | 87.03 | 87.74 | | MetaboLLM | MetaboLLM-Qwen3-8B | 90.94 | 87.26 | 83.70 | 87.66 | 88.02 | | MetaboLLM | **MetaboLLM-Qwen3-4B** | **91.71** | **88.04** | **84.12** | **87.79** | **88.06** | MetaboLLM-Qwen3-4B achieved the highest BERTScore-F1 across all five description-generation tasks. ## External Benchmark Transfer MetaboLLM was also evaluated on MetaBench, an independently developed public metabolomics benchmark. - **MetaboLLM-Qwen3-8B** achieved the highest Knowledge MCQA accuracy at **56.42%** - **MetaboLLM-Qwen3-4B** achieved the highest pathway-description scores, including **85.19 BERTScore-F1**, **25.83 ROUGE-L-F1**, and **17.61 BLEU-2** These results demonstrate transfer beyond the internally constructed MetaboLLM benchmark. ## Quick Start Install the required packages: ```bash pip install -U transformers peft accelerate torch ``` Example using MetaboLLM-Qwen3-4B: ```python import torch from peft import PeftModel from transformers import AutoModelForCausalLM, AutoTokenizer base_model_id = "Qwen/Qwen3-4B-Instruct-2507" adapter_id = "MetaboLLM/MetaboLLM-Qwen3-4B" tokenizer = AutoTokenizer.from_pretrained(adapter_id) base_model = AutoModelForCausalLM.from_pretrained( base_model_id, torch_dtype="auto", device_map="auto", ) model = PeftModel.from_pretrained( base_model, adapter_id, ) messages = [ { "role": "user", "content": "What is the biochemical role of pyruvate?" } ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, return_tensors="pt", ).to(model.device) with torch.no_grad(): outputs = model.generate( inputs, max_new_tokens=256, do_sample=False, ) response = tokenizer.decode( outputs[0][inputs.shape[-1]:], skip_special_tokens=True, ) print(response) ``` Please check the corresponding model card for the exact base-model identifier and model-specific loading instructions. ## Loading the Benchmark ```python from datasets import load_dataset dataset = load_dataset( "MetaboLLM/MetaboLLM-Benchmark", "Knowledge_recall", ) print(dataset["test"][0]) ``` Available configurations: - `Knowledge_recall` - `Class_identification` - `Relation_identification` - `Description_generation` ## Intended Uses MetaboLLM is intended for research involving: - metabolomics-specific question answering - biochemical knowledge recall - metabolite and reaction class identification - metabolite–pathway and reaction–enzyme relationship identification - metabolite, pathway, and enzyme description generation - evaluation of metabolomics-specialized language models ## Limitations - MetaboLLM may generate incorrect or unsupported biochemical statements. - Performance outside metabolomics and biochemical knowledge tasks has not been comprehensively evaluated. - Outputs should be independently verified and should not be used for clinical decision-making. ## Licenses Each MetaboLLM repository is distributed according to its model card and license notice. - Use of each adapter remains subject to the license and terms of its corresponding base model. - Benchmark components created by the MetaboLLM authors are distributed under the terms described in the benchmark repository. - Third-party and source-derived biochemical content remains subject to the licenses and terms of the original providers. Please review the license files in the relevant model and dataset repositories before use. ## Resource Maintainers - Dohyun Ku - Min Gu Kwak