| --- |
| frameworks: |
| - JAX |
| language: |
| - en |
| license: apache-2.0 |
| tags: |
| - OneScience |
| - bioscience |
| - protein-complex-structure-prediction |
| - CombFold |
| - AlphaFold-Multimer |
| - ColabFold |
| tasks: [] |
| datasets: [] |
| --- |
| |
| <p align="center"> |
| <strong> |
| <span style="font-size: 30px;">CombFold</span> |
| </strong> |
| </p> |
| |
| # Model Introduction |
|
|
| CombFold is an open-source pipeline developed by dina-lab3D for predicting the structures of large protein complexes. Starting from the amino acid sequences of the individual chains in a complex, CombFold uses AlphaFold-Multimer to predict multiple candidate subcomplexes and then applies combinatorial assembly algorithms to construct the complete protein complex. |
|
|
| The original publication reports support for complexes containing at least 18,000 amino acids and up to 32 subunits. |
|
|
| Paper: [Assembly of protein complexes by combining AlphaFold and combinatorial optimization](https://www.nature.com/articles/s41592-024-02174-0) |
|
|
| # Model Description |
|
|
| CombFold consists of four main stages: |
|
|
| 1. Define subunits according to protein domains and chain composition and generate `subunits.json`; |
| 2. Generate FASTA files for all subunit pairs and predict pairwise subcomplexes using AlphaFold-Multimer; |
| 3. Optionally predict candidate subcomplexes containing more than two subunits; |
| 4. Extract relative transformations between subunits from the predicted PDB structures and use a C++ combinatorial optimization algorithm to assemble the complete complex. |
|
|
| The main components included in the Hugging Face model package are: |
|
|
| - `model/CombinatorialAssembler/`: C++17 combinatorial assembler and the `AF2trans` structural transformation tool; |
| - `scripts/prepare_fastas.py`: generates FASTA files for pairwise or larger subunit combinations; |
| - `scripts/inference.py`: unified assembly entry point for the Hugging Face model package; |
| - `scripts/run_on_pdbs.py`: original upstream entry point for assembly from predicted PDB structures; |
| - `weight/`: official pretrained AlphaFold-Multimer parameters used for offline ColabFold inference; |
| - `requirements.txt`: additional dependencies required on top of the OneScience base environment. |
|
|
| CombFold itself is not a trainable neural network. Neural-network inference is performed by pretrained AlphaFold-Multimer models, while CombFold is responsible for extracting structural transformations and performing combinatorial assembly. |
|
|
| # Use Cases |
|
|
| | Use Case | Description | |
| | :---: | :--- | |
| | Large protein complex prediction | Combine multiple AlphaFold-Multimer subcomplex predictions into a complete complex structure. | |
| | Homomer prediction | Assemble complexes containing multiple copies of the same unique subunit according to the specified stoichiometry. | |
| | Heteromer prediction | Integrate predictions from different subunit pairs or groups to construct the complete structure. | |
| | Assembly from existing AFM results | Directly use existing AlphaFold-Multimer PDB predictions without rerunning AlphaFold-Multimer. | |
| | Crosslink-guided assembly | Optionally incorporate crosslinking restraints to constrain candidate complex structures. | |
|
|
| # Usage |
|
|
| ## 1. OneCode |
|
|
| You can use the OneCode online environment for an intelligent one-click AI4S programming experience: |
|
|
| [Try OneCode for AI4S Programming](https://web-2069360198568017922-iaaj.ksai.scnet.cn:58043/home) |
|
|
| ## 2. Manual Installation |
|
|
| **Hardware Requirements** |
|
|
| - The CombFold C++ combinatorial assembly stage requires only CPU resources; |
| - Local AlphaFold-Multimer subcomplex prediction typically requires an accelerator; |
| - Memory usage for long-sequence and multi-model inference increases with the total number of residues, MSA depth, number of models, and number of recycles; |
| - PyTorch is not a direct runtime dependency of CombFold or the current ColabFold inference pipeline. |
|
|
| ### Download the Model Package |
|
|
| Install the Hugging Face command-line tool and download the model repository: |
|
|
| ```bash |
| pip install -U huggingface_hub |
| |
| hf download OneScience-Group/CombFold --local-dir ./CombFold |
| cd CombFold |
| ``` |
|
|
| ### Install the Runtime Environment |
|
|
| **OneScience DCU Base Environment** |
|
|
| ```bash |
| conda create -n onescience311 python=3.11 -y |
| conda activate onescience311 |
| |
| python -m pip install onescience[bio-dcu] \ |
| -i http://mirrors.onescience.ai:3141/pypi/simple/ \ |
| --trusted-host mirrors.onescience.ai |
| ``` |
|
|
| Install the additional dependencies: |
|
|
| ```bash |
| python -m pip install --no-deps -r requirements.txt |
| ``` |
|
|
| ### Compile the Combinatorial Assembler |
|
|
| The combinatorial assembly stage of CombFold is implemented in C++17 and requires the following system-level components. These components cannot be installed through `requirements.txt`. |
|
|
| | Component | Purpose | Description | |
| | --- | --- | --- | |
| | C++17 compiler | Compile the C++ source code | `g++` is commonly used on Linux | |
| | GNU Make | Execute the Makefile | GNU Make 4.2.1 has been validated | |
| | Boost headers | Compile-time headers | The directory must contain `boost/algorithm/string.hpp` | |
| | Boost program_options | Link-time runtime library | Usually provided as `libboost_program_options.so` on Linux | |
| |
| Boost headers are generally portable across Linux distributions, but compiled Boost libraries depend on the operating system, CPU architecture, compiler, and `libstdc++` ABI. Therefore, compiled Boost libraries should not be copied directly between Linux, macOS, Windows, or different CPU architectures. |
| |
| It is recommended to use Boost headers and runtime libraries from the same Boost version. |
| |
| The upstream Makefile is configured primarily for macOS Homebrew. On Linux, the following command can be used directly only when Boost is already available in the compiler's default search paths: |
| |
| ```bash |
| cd model/CombinatorialAssembler |
| make |
| cd ../.. |
| ``` |
| |
| #### Specify a Custom Boost Path |
| |
| If Boost is not installed in the default compiler search paths, define the following environment variables: |
| |
| ```bash |
| export COMBFOLD_BOOST_INCLUDE="<Boost source or include directory>" |
| export COMBFOLD_BOOST_LIB="<Boost library directory>" |
| ``` |
| |
| `COMBFOLD_BOOST_INCLUDE` must point to a directory that directly contains the `boost/` subdirectory. |
| |
| `COMBFOLD_BOOST_LIB` must point to a directory that directly contains the `libboost_program_options` library. |
| |
| Before compilation, you can verify the paths using: |
| |
| ```bash |
| test -f "${COMBFOLD_BOOST_INCLUDE}/boost/algorithm/string.hpp" \ |
| && echo "Boost headers OK" |
| |
| find "${COMBFOLD_BOOST_LIB}" -maxdepth 1 \ |
| -name 'libboost_program_options*' -print |
| ``` |
| |
| Compile the assembler: |
| |
| ```bash |
| cd model/CombinatorialAssembler |
| |
| make -j4 \ |
| BOOST_INCLUDE="${COMBFOLD_BOOST_INCLUDE}" \ |
| BOOST_LIB="${COMBFOLD_BOOST_LIB}" |
| |
| cd ../.. |
| ``` |
| |
| ### Weights and Data Preparation |
| |
| The CombFold combinatorial assembler itself does not require model weights. |
| |
| Local generation of AlphaFold-Multimer subcomplex predictions requires the five official AlphaFold-Multimer v3 parameter files: |
| |
| | Asset | Location in the Model Package | Purpose | |
| | --- | --- | --- | |
| | `params_model_1_multimer_v3.npz` | `weight/alphafold/params/` | AFM v3 model 1 | |
| | `params_model_2_multimer_v3.npz` | `weight/alphafold/params/` | AFM v3 model 2 | |
| | `params_model_3_multimer_v3.npz` | `weight/alphafold/params/` | AFM v3 model 3 | |
| | `params_model_4_multimer_v3.npz` | `weight/alphafold/params/` | AFM v3 model 4 | |
| | `params_model_5_multimer_v3.npz` | `weight/alphafold/params/` | AFM v3 model 5 | |
| |
| Use the following ColabFold data path: |
| |
| ```text |
| --data weight/alphafold |
| ``` |
| |
| If you only use existing AlphaFold-Multimer PDB predictions for combinatorial assembly, AlphaFold-Multimer weights and a DCU device are not required. |
| |
| ### Define Subunits |
| |
| The input `subunits.json` is a JSON dictionary keyed by unique subunit names. |
| |
| Each subunit contains the following fields: |
| |
| - `name`: unique subunit name; |
| - `sequence`: amino acid sequence; |
| - `chain_names`: chain names corresponding to this subunit in the complete complex. The number of entries also defines its stoichiometry; |
| - `start_res`: starting residue index of the sequence in the original chain. |
|
|
| Example: |
|
|
| ```json |
| { |
| "A0": { |
| "name": "A0", |
| "chain_names": ["A", "B"], |
| "start_res": 1, |
| "sequence": "MKDILEKLEERRAQARLGGGEKRLEAQHKRGKLTARERIELLLDHGSFEE" |
| } |
| } |
| ``` |
|
|
| The Hugging Face model package provides a complete example: |
|
|
| ```text |
| scripts/example/subunits.json |
| scripts/example/pdbs/ |
| ``` |
|
|
| ### Quick Inference: CPU Assembly from Existing PDB Files |
|
|
| This is the shortest CombFold inference path and does not run AlphaFold-Multimer: |
|
|
| ```bash |
| python scripts/inference.py \ |
| --subunits scripts/example/subunits.json \ |
| --pdbs scripts/example/pdbs \ |
| --output output/example_assembly |
| ``` |
|
|
| The output directory must either not exist or be empty before execution. |
|
|
| After successful execution, the main results are located at: |
|
|
| ```text |
| output/example_assembly/assembled_results/output_clustered_0.pdb |
| output/example_assembly/assembled_results/confidence.txt |
| ``` |
|
|
| You can also use the original upstream positional-argument entry point: |
|
|
| ```bash |
| python scripts/run_on_pdbs.py \ |
| scripts/example/subunits.json \ |
| scripts/example/pdbs \ |
| output/example_assembly |
| ``` |
|
|
| ### Generate Pairwise FASTA Files |
|
|
| Generate FASTA files for every pair of unique subunits defined in `subunits.json`: |
|
|
| ```bash |
| python scripts/prepare_fastas.py \ |
| scripts/example/subunits.json \ |
| --stage pairs \ |
| --output-fasta-folder output/pair_fastas \ |
| --max-af-size 1800 |
| ``` |
|
|
| The output directory must not already exist. |
|
|
| The official example generates files such as: |
|
|
| ```text |
| A0_A0.fasta |
| A0_G0.fasta |
| G0_G0.fasta |
| ``` |
|
|
| ### Minimal DCU Pairwise Inference |
|
|
| For an offline smoke test on a compute node, you can use `single_sequence`, one model, and one recycle: |
|
|
| ```bash |
| colabfold_batch \ |
| output/pair_fastas \ |
| output/colabfold_pairs \ |
| --data weight/alphafold \ |
| --model-type alphafold2_multimer_v3 \ |
| --model-order 1 \ |
| --num-models 1 \ |
| --num-recycle 1 \ |
| --num-relax 0 \ |
| --msa-mode single_sequence \ |
| --disable-unified-memory |
| ``` |
|
|
| The configuration: |
|
|
| ```text |
| single_sequence + 1 model + 1 recycle |
| ``` |
|
|
| is intended only to verify parameter loading, JAX/DCU forward execution, and PDB output generation. |
|
|
| It should not be used to evaluate formal prediction accuracy. |
|
|
| For production-quality predictions, prepare appropriate MSA features and increase the number of models and recycles according to available memory and runtime constraints. |
|
|
| ### End-to-End DCU-to-CPU Inference |
|
|
| At least one predicted PDB structure must be selected for each pair. |
|
|
| To select the top-ranked ColabFold structure: |
|
|
| ```bash |
| mkdir -p output/combfold_pdbs |
| |
| find output/colabfold_pairs -maxdepth 1 \ |
| -type f -name '*rank_001*.pdb' \ |
| -exec cp {} output/combfold_pdbs/ \; |
| ``` |
|
|
| Then run the CombFold combinatorial assembly stage: |
|
|
| ```bash |
| python scripts/inference.py \ |
| --subunits scripts/example/subunits.json \ |
| --pdbs output/combfold_pdbs \ |
| --output output/end2end_assembly |
| ``` |
|
|
| The unified inference entry point also outputs a machine-readable summary, for example: |
|
|
| ```text |
| COMBFOLD_INFERENCE_RESULT={"assembled_structures": 5, "format": "pdb", "status": "PASS", ...} |
| ``` |
|
|
| ### Use Crosslinking Restraints |
|
|
| Use the `--crosslinks` option to provide a crosslink restraint file: |
|
|
| ```bash |
| python scripts/inference.py \ |
| --subunits scripts/example/example_xlinks/subunits.json \ |
| --pdbs scripts/example/example_xlinks/pdbs \ |
| --crosslinks scripts/example/example_xlinks/crosslinks.txt \ |
| --output output/crosslink_assembly |
| ``` |
|
|
| ### Optional Prediction of Larger Subcomplexes |
|
|
| After completing pairwise prediction, larger subcomplex FASTA files can be generated based on the pairwise results: |
|
|
| ```bash |
| python scripts/prepare_fastas.py \ |
| scripts/example/subunits.json \ |
| --stage groups \ |
| --output-fasta-folder output/group_fastas \ |
| --max-af-size 1800 \ |
| --input-pairs-results output/combfold_pdbs |
| ``` |
|
|
| ### Training |
|
|
| CombFold is an inference algorithm that performs combinatorial assembly using predictions generated by pretrained AlphaFold-Multimer models. |
|
|
| CombFold itself does not contain a trainable neural network, training entry point, optimizer, or training-data pipeline. Therefore, this Hugging Face model package does not provide training commands. |
|
|
| Retraining AlphaFold-Multimer is a separate upstream large-scale model-training task and is not part of the CombFold combinatorial assembly workflow. |
|
|
| # OneScience Official Resources |
|
|
| | Platform | OneScience Main Repository | Skills Repository | |
| | --- | --- | --- | |
| | Gitee | https://gitee.com/onescience-ai/onescience | https://gitee.com/onescience-ai/oneskills | |
| | GitHub | https://github.com/onescience-ai/OneScience | https://github.com/onescience-ai/oneskills | |
|
|
| # Citation and License |
|
|
| - Paper: [Assembly of protein complexes by combining AlphaFold and combinatorial optimization](https://www.nature.com/articles/s41592-024-02174-0) |
| - Official implementation: https://github.com/dina-lab3D/CombFold |
| - The CombFold source code is provided under the Apache License 2.0 according to the upstream repository. |
| - AlphaFold, AlphaFold-Multimer, ColabFold, pretrained parameters, and other third-party components are subject to their respective original copyright notices, model terms, and licenses. |