adminspec's picture
Explain Laxmi research ambition and staged goals
80758d5 verified
|
Raw History Blame Contribute Delete
10.4 kB
---
license: apache-2.0
library_name: pytorch
tags:
- software-state-prediction
- experimental
- negative-results
- custom-model
- real-patch-evaluation
---
# Laxmi — experimental software-state prediction
**Predict the consequences of a software change before executing it.**
Created and maintained by **Navneet Prabhakar**, creator and sole contributor to Laxmi. Third-party libraries retain their authorship and licenses.
Laxmi investigates whether observed software state, history, an action and a proposed code change can predict what happens next. This release keeps the bounded controller experiment inspectable and adds local real-patch evaluation tooling to the **same research project**. It is not a chat model, a general code-review assistant or an autonomous software engineer.
> **Version 0.1.1 · Experimental research release**
>
> Controller weights and runtime are unchanged from v0.1.0. The latest model repair failed its improvement test. The new real-patch exact-execution tool is a baseline, **not a trained model**. None of the four controller candidates is promoted as an accepted model.
## The goal
Laxmi's long-term ambition is a **testable software world model**: a system that maintains a predictive view of software state as evidence arrives. Given the current state, execution history, a proposed action or code change, and declared operating conditions, it should forecast possible consequences **before execution**, express uncertainty, and revise its view when new observations arrive. A useful forecast would help an engineer decide what to inspect, test, change or leave alone.
The research path has three stages:
1. **Predictive state:** learn action- and source-conditioned outcomes, including bounded multi-step effects and differences between proposed changes. Measure forecasts against actual outcomes and strong methods given the same information.
2. **Prospective advice:** save forecasts before real controller actions, then test whether the advice improves a human review decision at acceptable error, coverage, latency and cost.
3. **Broader engineering assessment:** use source and pull-request context, observed history and explicit constraints to assess likely consequences, compare a finite set of alternatives, and explain findings with traceable evidence.
The ultimate aspiration is useful engineer-like work within a **defined role and measurable constraints**. Progress is measured by prospective outcomes and decisions, including honest abstention when evidence is insufficient. This v0.1.1 release is an early, bounded controller experiment toward the first stage. It has not established a general software world model, a qualified advisory, or autonomous engineering capability.
## The experiment
A controller records progress and retries failed operations. An edit to its recovery logic can delay completion or repeat work. The research question is whether pre-action evidence can predict that consequence.
```text
State + history + source/edit + declared conditions
↓
Save a prediction
↓
Execute a controlled experiment
↓
Compare with observed outcomes
```
The released consumer forecasts admitted synthetic controller conditions. It can observe, predict, compare proposals, save/restore state and score separately recorded outcomes. It never dispatches native controller actions. Recording a forecast with this consumer alone does not authenticate prospective ordering.
## Results, including the failures
| Seed | Original H1 exact outcomes | Normalized H1 exact outcomes | Changed pairs both correct, either readout |
|---|---:|---:|---:|
| 17 | 8 / 24 | 8 / 24 | 0 / 6 |
| 43 | 4 / 24 | 4 / 24 | 0 / 6 |
Readout normalization did not improve the task. These are exposed development results from September 15–16, 2026, not independent acceptance scores. Longer-horizon likelihood comparisons remain unresolved. Probabilities are not established as calibrated, and independent implementation transfer remains open.
The packaged example checks installation and replay. It must not be presented as a fresh benchmark. See [verification.json](verification.json) for the exact smoke scope and [BENCHMARK.md](BENCHMARK.md) for the next comparison design.
### Real-patch development screen
This release also contains [local real-patch evaluation tooling](evaluation/real-patch/README.md). Six exposed historical changes were inventoried; only one currently has a valid forecast packet, and five have oracle-only scenarios excluded from forecast accuracy. On the one scored development case, the **same-information exact-execution baseline** matched both branch observations and the patch effect. The scenario was selected after its outcome was known. There are no blind/acceptance-eligible cases, no learned real-patch model scored, and no released case data. The result shows that this fully runnable case offers no measured headroom over exact execution; it does not validate generalization.
## Download and run
**Controller runtime verified platform: Apple Silicon macOS, Python 3.12.11.** That wheel is macOS ARM64 tagged. Other Python 3.12 patch builds and platforms have not been verified for the controller runtime. The source archive is provided for inspection and experimentation; it does not establish Linux controller support. No GPU is required for the example.
With the Hugging Face `hf` CLI installed:
```bash
hf download Executespec/laxmi-controller-experimental --revision v0.1.1 --local-dir laxmi-release
cd laxmi-release
shasum -a 256 -c SHA256SUMS
python3.12 -m venv .venv
.venv/bin/python -m pip install runtime/laxmi_readout_runtime_capsule-0.1.0-py3-none-macosx_11_0_arm64.whl
.venv/bin/python -I -B verify_release.py --release . --python "$PWD/.venv/bin/python" --output smoke-17-original --candidate 17-original
```
Expected completion: `17-original: all six operations passed`. The output directory must be new. The check observes the example, predicts, compares, restores, saves and scores an outcome. Omit `--candidate` to check all four candidates. The verifier makes candidate/example files private (`0600`) as required by the runtime. It writes local snapshots and logs; it does not run the controlled native software or train a model.
The wheel declares exact dependency versions, also recorded in `requirements-macos-arm64-py312.txt`. Dependencies are installed separately; they are not bundled. Version pins are not a dependency artifact hash lock. The source package includes the runtime source and license files. No optimizer checkpoint is shipped.
The separate real-patch package has its own [local installation instructions](evaluation/real-patch/README.md). It installed and **validated** its input in an offline Linux ARM64 container, but Linux exact execution and any hosted API remain unqualified. Its OpenAPI file is a design draft only.
## Package layout
| Path | Contents |
|---|---|
| `candidates/17-original/`, `17-normalized/`, `43-original/`, `43-normalized/` | Original retained safetensors weights, configuration, manifests and matching consumer metadata |
| `runtime/` | Reviewed runtime wheel and source ZIP |
| `evaluation/real-patch/` | Local exact-baseline evaluator wheel, source, contract and design-only API draft; no model weights or case data |
| `examples/` | One retained synthetic observation, proposal comparison and outcome template |
| `verify_release.py` | Installation/replay check using only this package |
| `release.json`, `SHA256SUMS` | Candidate identity, runtime identity and file integrity |
| `verification.json` | Sanitized installation/replay evidence |
| `LICENSE`, `NOTICE`, `LICENSE_SCOPE.md`, `RELEASE_SCOPE.md` | Apache 2.0, attribution and exact distribution scope |
| `CITATION.cff` | Suggested research citation |
Weights, configuration, readout metadata and runtime must match. Original and normalized readouts share tensor layouts but differ in computation. Do not mix them. Existing candidate metadata intentionally retains `acceptance=false` and `release_qualified=false`; those fields record scientific qualification, not whether the experimental package has been uploaded.
## Known boundaries
- The input contract is bounded synthetic controller evidence. Do not relabel real executions to pass admission.
- No arbitrary-PR understanding, patch generation, autonomous engineering or production suitability is established.
- Missing observations remain unknown rather than fabricated zero counts.
- Internal hashes verify consistency, not source ownership, event absence or wall-clock ordering.
- Longer rollouts may contain unresolved/pruned probability mass; it is not silently renormalized away.
- Supported consumer horizons are 1, 2, 4 and 8; this release's installation smoke covers H1 only.
- The learned probabilities are not proven calibrated and do not authorize external actions.
- The real-patch exact comparator is a baseline, not learned prediction. Only one exposed development forecast was scored; no public API or Linux exact-execution qualification exists.
## Benchmark and iteration
A future benchmark will compare identical pre-action evidence against native outcomes, with persistence/prior/source-rule baselines and Jev/Laya on shared questions. **No Jev/Laya comparison has been run.** Base versus fitted checkpoints will be separate arms. Public example replay is distinct from fresh prospective accuracy or independent transfer.
Later versions should pin source, schema, runtime, weights, data exposure and results together. This version remains available as a baseline, including its negative result. See [CHANGELOG.md](CHANGELOG.md).
## Author, attribution and license
**Navneet Prabhakar** — creator, maintainer and sole contributor to Laxmi.
[GitHub profile](https://github.com/navneetprabhakar)
Original components explicitly listed in [RELEASE_SCOPE.md](RELEASE_SCOPE.md) are released under [Apache License 2.0](LICENSE). Preserve applicable attribution and license notices when redistributing; see [NOTICE](NOTICE). Third-party dependencies retain their own terms. This release does not distribute the full corpus or private operational archives.
If you use Laxmi in research, please cite [CITATION.cff](CITATION.cff). Citation is requested, not an additional license condition.