| --- |
| license: other |
| license_name: om-lula-community-license-1.1 |
| license_link: LICENSE |
| pipeline_tag: other |
| tags: |
| - biology |
| - drug-discovery |
| - small-molecule-discovery |
| - protein-ligand |
| - binding-affinity |
| - open-weight |
| - lula-1 |
| - lula-1.1 |
| --- |
| |
| > **License notice.** By downloading, accessing, or using LULA-1.1, you agree to |
| > the [Om LULA Community License 1.1](LICENSE). LULA-1.1 is licensed for |
| > research, evaluation, benchmarking, teaching, and other non-commercial research |
| > uses only. **Any commercial work requires a separate Om commercial license**, |
| > including internal commercial discovery, commercial drug discovery, screening, |
| > hit finding, lead optimization, portfolio decisions, production R&D, product |
| > candidate identification, patent or therapeutic program work, hosted |
| > inference, paid API/SaaS access, resale, support/deployment, product bundling, |
| > and competing model services. For commercial licensing, contact dmc@omtx.ai. |
|
|
| <style> |
| @font-face { |
| font-family: 'OmMontserrat'; |
| font-style: normal; |
| font-weight: 100 900; |
| font-display: swap; |
| src: url('https://huggingface.co/spaces/omtx/README/resolve/main/assets/Montserrat-Latin-Variable.woff2') format('woff2'); |
| } |
| .om-card { background:#0A0A0A; border:1px solid #2A2A2A; border-radius:14px; overflow:hidden; |
| font-family:'OmMontserrat','Montserrat','Helvetica Neue',Arial,sans-serif; color:#fff; |
| margin-bottom:26px; } |
| .om-card p, .om-card h1 { margin-top:0; margin-left:0; margin-right:0; } |
| .om-hero { position:relative; overflow:hidden; background:#0A0A0A; } |
| .om-hero-img { position:absolute; inset:0; width:100%; height:100%; object-fit:cover; |
| object-position:72% 50%; margin:0; } |
| .om-scrim { position:absolute; inset:0; |
| background:linear-gradient(90deg,rgba(10,10,10,0.97) 0%,rgba(10,10,10,0.90) 44%,rgba(10,10,10,0.30) 100%); } |
| .om-cmark { position:absolute; color:#3B3636; font-size:19px; line-height:1; font-weight:300; z-index:5; } |
| .om-copy { position:relative; z-index:2; padding:36px 34px 38px; max-width:660px; } |
| .om-mono { color:#71717A; font-size:11px; font-weight:500; letter-spacing:0.26em; |
| text-transform:uppercase; margin-bottom:14px; } |
| .om-eyebrow { color:#E2E756; font-size:12px; font-weight:600; letter-spacing:0.32em; |
| text-transform:uppercase; margin-bottom:18px; } |
| .om-h1 { font-size:40px; font-weight:700; letter-spacing:0; line-height:1.0; |
| margin-bottom:18px; color:#fff; } |
| .om-accent { color:#E2E756; } |
| .om-lede { color:#D4D4D8; font-size:16px; line-height:1.5; margin-bottom:22px; max-width:54ch; } |
| .om-pill { display:inline-block; border:1px solid #2A2A2A; color:#D4D4D8; font-size:11px; |
| font-weight:700; letter-spacing:0.08em; text-transform:uppercase; padding:8px 15px; |
| border-radius:100px; margin:0 7px 7px 0; } |
| .om-status { border-top:1px solid #2A2A2A; background:#0D0C0C; padding:26px 34px 28px; } |
| .om-chip { display:inline-block; border:1px solid #3B3636; color:#71717A; font-size:11px; |
| font-weight:700; letter-spacing:0.08em; text-transform:uppercase; padding:7px 13px; |
| border-radius:6px; margin:0 7px 7px 0; } |
| .om-chip-live { border-color:#E2E756; color:#E2E756; } |
| </style> |
| |
| <div class="om-card"> |
| <div class="om-hero"> |
| <img class="om-hero-img" src="assets/lula1-hero.jpg" alt="Protein scaffold with a ligand bound in a highlighted pocket" /> |
| <div class="om-scrim"></div> |
| <span class="om-cmark" style="top:15px; left:15px;">+</span> |
| <span class="om-cmark" style="top:15px; right:15px;">+</span> |
| <span class="om-cmark" style="bottom:15px; right:15px;">+</span> |
| <div class="om-copy"> |
| <p class="om-mono">omtx.ai</p> |
| <p class="om-eyebrow">Open-weight release track</p> |
| <h1 class="om-h1">LULA-1.1 <span class="om-accent">sequence-only</span> protein-ligand scoring.</h1> |
| <p class="om-lede"> |
| Protein amino-acid sequence plus ligand SMILES in, binding score out. No structure input, |
| no docking, no folding step. |
| </p> |
| <p> |
| <span class="om-pill">Sequence-only</span> |
| <span class="om-pill">Local inference</span> |
| <span class="om-pill">Open weights</span> |
| </p> |
| </div> |
| </div> |
| <div class="om-status"> |
| <p class="om-eyebrow">Model at a glance</p> |
| <span class="om-chip om-chip-live">1.7M parameters</span> |
| <span class="om-chip om-chip-live">6.8 MB</span> |
| <span class="om-chip om-chip-live">6.14M training pairs</span> |
| <span class="om-chip om-chip-live">13,368 proteins</span> |
| <span class="om-chip om-chip-live">Sequence-only</span> |
| <span class="om-chip om-chip-live">Local inference</span> |
| <span class="om-chip om-chip-live">Open weights</span> |
| </div> |
| </div> |
| |
| # LULA-1.1 |
|
|
| LULA-1.1 is a lightweight, sequence-only protein-ligand binding scorer from Om |
| Therapeutics. It takes a protein amino-acid sequence and ligand SMILES and |
| returns a binding score. There is no structure input, docking, or folding step. |
|
|
| This release uses the same ConPLex-style two-tower scoring architecture as the |
| original LULA-1 open-weight release, with an updated target-balanced training |
| recipe and expanded training coverage. The customer-facing model name is |
| LULA-1.1. |
|
|
| ## What Changed From LULA-1 |
|
|
| LULA-1.1 keeps the LULA-1 two-tower architecture while updating the weights, |
| training coverage, sampling, and protein-context handling. |
|
|
| Compared with LULA-1, LULA-1.1 increases supervised protein-ligand training |
| coverage from 2,763,260 to 6,137,835 pairs, adding 3,374,575 protein-ligand |
| training pairs. |
|
|
| | Coverage | LULA-1 | LULA-1.1 | |
| |---|---:|---:| |
| | Supervised protein-ligand training pairs | 2,763,260 | 6,137,835 | |
| | Binder-labeled training pairs | 2,132,861 | 4,816,392 | |
| | Non-binder-labeled training pairs | 630,399 | 1,321,443 | |
|
|
| The updated sampling recipe is target-balanced to avoid letting high-row-count |
| targets dominate the update stream. LULA-1.1 also uses complete protein-context |
| inference: 1,022-residue ESM windows with 256-residue overlap, C-terminal |
| coverage, overlap-averaged residues, and full-sequence mean pooling excluding |
| BOS/EOS tokens. |
|
|
| The validation evidence for this release is mixed across panels. LULA-1.1 is |
| published as the next open-weight release for research and evaluation; users |
| should benchmark it against their own targets before relying on rank ordering. |
|
|
| ## Protein Coverage |
|
|
| LULA-1.1 represents 13,368 protein source entities across model-ready release |
| inputs. |
|
|
| Example proteins represented include EGFR, JAK2, RET, CDK2, MAPK1, GSK3B, DRD2, |
| OPRM1, CHRM2, HTR2A, ESR1, AR, PPARG, BACE1, and thrombin. |
|
|
| Example protein classes include kinases, GPCRs, nuclear receptors, |
| proteases/peptidases, ion channels and transporters, phosphatases, |
| epigenetic/chromatin regulators, immune/complement/coagulation proteins, and |
| cell-surface receptors. |
|
|
| This release reports aggregate coverage only. It does not include protein-level |
| source manifests, amino-acid sequence tables, ligand rows, per-pair training |
| rows, source object paths, internal private dataset/vintage identifiers, or |
| customer data. |
|
|
| ## What This Release Contains |
|
|
| LULA-1.1 ships as a compact scoring head that runs on top of two public |
| pretrained encoders. Om distributes the LULA-1.1 scoring head in this repository; |
| the third-party encoders remain governed by their own upstream terms. |
|
|
| | Component | Parameters | Source | |
| |---|---:|---| |
| | LULA-1.1 scoring head | 1,705,984 | this repository | |
| | ESM-2 650M protein encoder | 652,358,616 | `facebook/esm2_t33_650M_UR50D` | |
| | ChemBERTa-77M-MTR ligand encoder | about 3,500,000 | `DeepChem/ChemBERTa-77M-MTR` | |
|
|
| ## Usage |
|
|
| ```bash |
| pip install "omtx[lula]>=2.0.14" |
| |
| hf auth login |
| omtx lula download --model lula1.1 |
| omtx lula verify |
| ``` |
|
|
| ### Version Selection |
|
|
| Use the public model selector to choose the release: |
|
|
| ```bash |
| omtx lula download --model lula1 # original LULA-1 open-weight release |
| omtx lula download --model lula1.1 # LULA-1.1 |
| ``` |
|
|
| ```python |
| from omtx.lula import load_model |
| |
| model = load_model("lula1") # original LULA-1 |
| model = load_model("lula1.1") # LULA-1.1 |
| ``` |
|
|
| ```python |
| from omtx.lula import load_model |
| |
| model = load_model("lula1.1") |
| rows = model.score( |
| protein_sequence="MSHHWGYGKHNGPEHWHKDFPIAKGERQSPVDIDTHTAKYDPSLKPLSVSYDQA", |
| smiles=["CCO", "CC(=O)Nc1nnc(s1)S(N)(=O)=O"], |
| ) |
| print(rows) |
| ``` |
|
|
| Batch scoring returns, per molecule: `score`, `rank`, and |
| `top_percentile_in_batch`. Scores are intended for relative prioritization |
| within a candidate set and are not calibrated binding probabilities. |
|
|
| ## Data Locality |
|
|
| Local scoring and fine-tuning run on your machine. Om does not receive your |
| targets, compounds, labels, checkpoints, or scores when you use the local model. |
|
|
| Hosted Om scoring is a separate product surface. |
|
|
| ## Files |
|
|
| | File | Bytes | SHA256 | |
| |---|---:|---| |
| | `model/best.pt` | 6,826,303 | `1933bbdf4aa335d52498c754dc961141d10631e38fb4cc1daa9a908e7e4601ba` | |
| | `model/model_config.json` | 310 | `de22bc8182d7f5f46454455e0e316296b59fb67f3aa420ce78c2728815806ff6` | |
| | `model/inference_config.json` | 436 | `115a343fef54f0d0a108998bb3ce86bd493e29a6518055e5707d896a1ddac36b` | |
|
|
| ## License |
|
|
| LULA-1.1 is distributed under the Om LULA Community License 1.1. It is |
| open-weight, not OSI open source. Commercial use requires a separate written Om |
| commercial license. Contact dmc@omtx.ai for commercial licensing. |
|
|
| ## Attribution |
|
|
| LULA-1.1 uses a ConPLex-style scoring architecture. See the ConPLex reference |
| implementation at https://github.com/samsledje/ConPLex and the publication DOI |
| 10.1073/pnas.2220778120. |
|
|