lula-1.1 / README.md
devcayer's picture
Publish LULA-1.1 open-weight release
34d20a4 verified
|
Raw
History Blame Contribute Delete
9.54 kB
---
license: other
license_name: om-lula-community-license-1.1
license_link: LICENSE
pipeline_tag: other
tags:
- biology
- drug-discovery
- small-molecule-discovery
- protein-ligand
- binding-affinity
- open-weight
- lula-1
- lula-1.1
---
> **License notice.** By downloading, accessing, or using LULA-1.1, you agree to
> the [Om LULA Community License 1.1](LICENSE). LULA-1.1 is licensed for
> research, evaluation, benchmarking, teaching, and other non-commercial research
> uses only. **Any commercial work requires a separate Om commercial license**,
> including internal commercial discovery, commercial drug discovery, screening,
> hit finding, lead optimization, portfolio decisions, production R&D, product
> candidate identification, patent or therapeutic program work, hosted
> inference, paid API/SaaS access, resale, support/deployment, product bundling,
> and competing model services. For commercial licensing, contact dmc@omtx.ai.
<style>
@font-face {
font-family: 'OmMontserrat';
font-style: normal;
font-weight: 100 900;
font-display: swap;
src: url('https://huggingface.co/spaces/omtx/README/resolve/main/assets/Montserrat-Latin-Variable.woff2') format('woff2');
}
.om-card { background:#0A0A0A; border:1px solid #2A2A2A; border-radius:14px; overflow:hidden;
font-family:'OmMontserrat','Montserrat','Helvetica Neue',Arial,sans-serif; color:#fff;
margin-bottom:26px; }
.om-card p, .om-card h1 { margin-top:0; margin-left:0; margin-right:0; }
.om-hero { position:relative; overflow:hidden; background:#0A0A0A; }
.om-hero-img { position:absolute; inset:0; width:100%; height:100%; object-fit:cover;
object-position:72% 50%; margin:0; }
.om-scrim { position:absolute; inset:0;
background:linear-gradient(90deg,rgba(10,10,10,0.97) 0%,rgba(10,10,10,0.90) 44%,rgba(10,10,10,0.30) 100%); }
.om-cmark { position:absolute; color:#3B3636; font-size:19px; line-height:1; font-weight:300; z-index:5; }
.om-copy { position:relative; z-index:2; padding:36px 34px 38px; max-width:660px; }
.om-mono { color:#71717A; font-size:11px; font-weight:500; letter-spacing:0.26em;
text-transform:uppercase; margin-bottom:14px; }
.om-eyebrow { color:#E2E756; font-size:12px; font-weight:600; letter-spacing:0.32em;
text-transform:uppercase; margin-bottom:18px; }
.om-h1 { font-size:40px; font-weight:700; letter-spacing:0; line-height:1.0;
margin-bottom:18px; color:#fff; }
.om-accent { color:#E2E756; }
.om-lede { color:#D4D4D8; font-size:16px; line-height:1.5; margin-bottom:22px; max-width:54ch; }
.om-pill { display:inline-block; border:1px solid #2A2A2A; color:#D4D4D8; font-size:11px;
font-weight:700; letter-spacing:0.08em; text-transform:uppercase; padding:8px 15px;
border-radius:100px; margin:0 7px 7px 0; }
.om-status { border-top:1px solid #2A2A2A; background:#0D0C0C; padding:26px 34px 28px; }
.om-chip { display:inline-block; border:1px solid #3B3636; color:#71717A; font-size:11px;
font-weight:700; letter-spacing:0.08em; text-transform:uppercase; padding:7px 13px;
border-radius:6px; margin:0 7px 7px 0; }
.om-chip-live { border-color:#E2E756; color:#E2E756; }
</style>
<div class="om-card">
<div class="om-hero">
<img class="om-hero-img" src="assets/lula1-hero.jpg" alt="Protein scaffold with a ligand bound in a highlighted pocket" />
<div class="om-scrim"></div>
<span class="om-cmark" style="top:15px; left:15px;">+</span>
<span class="om-cmark" style="top:15px; right:15px;">+</span>
<span class="om-cmark" style="bottom:15px; right:15px;">+</span>
<div class="om-copy">
<p class="om-mono">omtx.ai</p>
<p class="om-eyebrow">Open-weight release track</p>
<h1 class="om-h1">LULA-1.1 <span class="om-accent">sequence-only</span> protein-ligand scoring.</h1>
<p class="om-lede">
Protein amino-acid sequence plus ligand SMILES in, binding score out. No structure input,
no docking, no folding step.
</p>
<p>
<span class="om-pill">Sequence-only</span>
<span class="om-pill">Local inference</span>
<span class="om-pill">Open weights</span>
</p>
</div>
</div>
<div class="om-status">
<p class="om-eyebrow">Model at a glance</p>
<span class="om-chip om-chip-live">1.7M parameters</span>
<span class="om-chip om-chip-live">6.8 MB</span>
<span class="om-chip om-chip-live">6.14M training pairs</span>
<span class="om-chip om-chip-live">13,368 proteins</span>
<span class="om-chip om-chip-live">Sequence-only</span>
<span class="om-chip om-chip-live">Local inference</span>
<span class="om-chip om-chip-live">Open weights</span>
</div>
</div>
# LULA-1.1
LULA-1.1 is a lightweight, sequence-only protein-ligand binding scorer from Om
Therapeutics. It takes a protein amino-acid sequence and ligand SMILES and
returns a binding score. There is no structure input, docking, or folding step.
This release uses the same ConPLex-style two-tower scoring architecture as the
original LULA-1 open-weight release, with an updated target-balanced training
recipe and expanded training coverage. The customer-facing model name is
LULA-1.1.
## What Changed From LULA-1
LULA-1.1 keeps the LULA-1 two-tower architecture while updating the weights,
training coverage, sampling, and protein-context handling.
Compared with LULA-1, LULA-1.1 increases supervised protein-ligand training
coverage from 2,763,260 to 6,137,835 pairs, adding 3,374,575 protein-ligand
training pairs.
| Coverage | LULA-1 | LULA-1.1 |
|---|---:|---:|
| Supervised protein-ligand training pairs | 2,763,260 | 6,137,835 |
| Binder-labeled training pairs | 2,132,861 | 4,816,392 |
| Non-binder-labeled training pairs | 630,399 | 1,321,443 |
The updated sampling recipe is target-balanced to avoid letting high-row-count
targets dominate the update stream. LULA-1.1 also uses complete protein-context
inference: 1,022-residue ESM windows with 256-residue overlap, C-terminal
coverage, overlap-averaged residues, and full-sequence mean pooling excluding
BOS/EOS tokens.
The validation evidence for this release is mixed across panels. LULA-1.1 is
published as the next open-weight release for research and evaluation; users
should benchmark it against their own targets before relying on rank ordering.
## Protein Coverage
LULA-1.1 represents 13,368 protein source entities across model-ready release
inputs.
Example proteins represented include EGFR, JAK2, RET, CDK2, MAPK1, GSK3B, DRD2,
OPRM1, CHRM2, HTR2A, ESR1, AR, PPARG, BACE1, and thrombin.
Example protein classes include kinases, GPCRs, nuclear receptors,
proteases/peptidases, ion channels and transporters, phosphatases,
epigenetic/chromatin regulators, immune/complement/coagulation proteins, and
cell-surface receptors.
This release reports aggregate coverage only. It does not include protein-level
source manifests, amino-acid sequence tables, ligand rows, per-pair training
rows, source object paths, internal private dataset/vintage identifiers, or
customer data.
## What This Release Contains
LULA-1.1 ships as a compact scoring head that runs on top of two public
pretrained encoders. Om distributes the LULA-1.1 scoring head in this repository;
the third-party encoders remain governed by their own upstream terms.
| Component | Parameters | Source |
|---|---:|---|
| LULA-1.1 scoring head | 1,705,984 | this repository |
| ESM-2 650M protein encoder | 652,358,616 | `facebook/esm2_t33_650M_UR50D` |
| ChemBERTa-77M-MTR ligand encoder | about 3,500,000 | `DeepChem/ChemBERTa-77M-MTR` |
## Usage
```bash
pip install "omtx[lula]>=2.0.14"
hf auth login
omtx lula download --model lula1.1
omtx lula verify
```
### Version Selection
Use the public model selector to choose the release:
```bash
omtx lula download --model lula1 # original LULA-1 open-weight release
omtx lula download --model lula1.1 # LULA-1.1
```
```python
from omtx.lula import load_model
model = load_model("lula1") # original LULA-1
model = load_model("lula1.1") # LULA-1.1
```
```python
from omtx.lula import load_model
model = load_model("lula1.1")
rows = model.score(
protein_sequence="MSHHWGYGKHNGPEHWHKDFPIAKGERQSPVDIDTHTAKYDPSLKPLSVSYDQA",
smiles=["CCO", "CC(=O)Nc1nnc(s1)S(N)(=O)=O"],
)
print(rows)
```
Batch scoring returns, per molecule: `score`, `rank`, and
`top_percentile_in_batch`. Scores are intended for relative prioritization
within a candidate set and are not calibrated binding probabilities.
## Data Locality
Local scoring and fine-tuning run on your machine. Om does not receive your
targets, compounds, labels, checkpoints, or scores when you use the local model.
Hosted Om scoring is a separate product surface.
## Files
| File | Bytes | SHA256 |
|---|---:|---|
| `model/best.pt` | 6,826,303 | `1933bbdf4aa335d52498c754dc961141d10631e38fb4cc1daa9a908e7e4601ba` |
| `model/model_config.json` | 310 | `de22bc8182d7f5f46454455e0e316296b59fb67f3aa420ce78c2728815806ff6` |
| `model/inference_config.json` | 436 | `115a343fef54f0d0a108998bb3ce86bd493e29a6518055e5707d896a1ddac36b` |
## License
LULA-1.1 is distributed under the Om LULA Community License 1.1. It is
open-weight, not OSI open source. Commercial use requires a separate written Om
commercial license. Contact dmc@omtx.ai for commercial licensing.
## Attribution
LULA-1.1 uses a ConPLex-style scoring architecture. See the ConPLex reference
implementation at https://github.com/samsledje/ConPLex and the publication DOI
10.1073/pnas.2220778120.