- AgentFEM DENIM
- Which artifact should I use?
- Prospective sealed evaluation (v2.2.0, no retraining)
- Capability v4 balanced model (v2.2.0)
- Material-conditioned DENIM v2
- AgentFEM integration
- Historical boundary-extension checkpoint
- Why this checkpoint is an incomplete-physics test
- Research progression
- Use the current v2.2 checkpoint
- Scope and limitations
- Article evidence v3
- Which artifact should I use?
AgentFEM DENIM
Which artifact should I use?
| Purpose | Artifact | Status |
|---|---|---|
| Current research checkpoint and new inference | capability_v4/ |
Recommended; conditional v2.2, 8,477 parameters |
| AgentFEM execution of the current checkpoint | capability_v4/model.json, weights.safetensors, and the denim.conditional.v2 adapter |
Three 3D implicit global cases verified |
| Reproduce the earlier conditional study | conditional_v2/ |
Historical v2.0 checkpoint |
| Reproduce the original fixed-material study | denim.pt, denim-expanded.pt, agentfem_bundle/ |
Legacy 918-parameter line |
| Inspect ablation and few-shot experiments | article_evidence_v3/ |
Research evidence, not the recommended production checkpoint |
The current v2.2 model was trained on the pre-sealed 3,660-trajectory protocol. The additional 128 trajectories are evaluation-only and do not change the weights.
Prospective sealed evaluation (v2.2.0, no retraining)
The unchanged capability-v4 weights were evaluated once on 128 trajectories
published first at dataset revision
712756b9fb98b16376d8fd413709f09299a9fd1a.
The eight sealed path families have no family-name overlap with training, and
the repository audit found zero exact cross-role collisions in full model
inputs or input/target pairs.
| Sealed cohort | Trajectories | Macro RMSE | Relative L2 | Micro RMSE / yield |
|---|---|---|---|---|
| Globally unseen paths | 64 | 0.333 MPa | 0.099% | 0.127% |
| 150/200-cycle horizons | 32 | 0.365 MPa | 0.478% | 0.156% |
| Amplitude + mean shift | 32 | 0.912 MPa | 0.255% | 0.339% |
| All sealed tests | 128 | 0.486 MPa | 0.246% | 0.182% |
Across 92,928 states, micro RMSE is 0.509 MPa and R2 is 0.999994. The
trajectory-macro bootstrap 95% interval is 0.432--0.543 MPa; median, p95 and
maximum per-trajectory RMSE are 0.324, 1.184 and 1.359 MPa. The maximum return-
mapping yield residual is 96 Pa. Full metrics, data binding, evaluation code
and checksums are in sealed_test_v1/.
How to read these errors
The metrics answer different questions and should be read together:
- RMSE in MPa gives the absolute engineering size of the stress error.
- Relative L2 divides the total prediction error by the total reference- stress magnitude. It remains well behaved when cyclic stresses cross zero.
- RMSE / yield compares state-point micro RMSE with the 280 MPa reference yield stress. The overall value of 0.182% means an RMS error of roughly two thousandths of the yield stress; it is not a pointwise percentage guarantee.
Over all sealed trajectories, the reference stress RMS is 206.814 MPa. The 0.509 MPa micro RMSE therefore gives 0.246% relative L2. The worst individual trajectory RMSE is 1.359 MPa, or 0.485% of yield stress. The maximum peak-stress error is 3.196 MPa, or 1.142% of yield stress. These normalized values clarify the error scale without replacing the original MPa values.
MAPE is intentionally omitted: stress reversals pass through zero, so dividing by instantaneous reference stress would create unstable and misleading percentages. Exact definitions and denominators are included in the machine- readable sealed evaluation JSON.
This is a prospective synthetic constitutive benchmark, not experimental
validation. sealed_test_v1 is permanently excluded from training and model
selection; using it to develop a future checkpoint turns it into a development
benchmark and requires a new sealed test version.
Capability v4 balanced model (v2.2.0)
capability_v4/ is the recommended research checkpoint for long-cycle and
full-3D path studies. It retains the 8,477-parameter conditional architecture
and adds balanced replay over the pre-sealed 3,660-trajectory protocol. The 100-cycle
OOD RMSE improves from 57.634 to 5.599 MPa; known-family strict-test RMSE is
3.450 MPa versus 3.308 MPa for v2.1. Full metrics, limitations, safe weights,
training code and AgentFEM global validation evidence are included in that folder.
The two long-cycle results represent different regimes and must not be ranked
as one ladder. The historical 100-cycle stress test reaches PEEQ 0.330--5.582,
outside the declared PEEQ <= 0.10 application domain. The sealed 150/200-
cycle test deliberately stays inside that domain (maximum PEEQ 0.0933) and
obtains 0.365 MPa trajectory-macro RMSE. Together they show strong in-domain
state propagation and a separately reported extreme-accumulation boundary.
Material-conditioned DENIM v2
The conditional_v2/ directory adds an 8,477-parameter material-conditioned
DENIM trained with strict roles for all 2,660 trajectories: 1,484 train, 407
validation and 769 test. It reaches 3.308 MPa RMSE on 537 strict held-out J2/
Chaboche trajectories and reduces the incomplete-material long-history RMSE
from 10.906 to 4.263 MPa, while retaining 0.798 MPa on the original frozen test.
This checkpoint is a material-point model. The fixed-material
agentfem_bundle/ remains the globally verified AgentFEM runtime artifact.
AgentFEM integration
For the current conditional model, use capability_v4/ as the bundle directory.
Its manifest selects architecture adapter denim.conditional.v2; the safe
weights are loaded locally from weights.safetensors. The top-level
agentfem_bundle/ is retained only for reproducing the earlier fixed-material,
918-parameter runtime study.
The legacy bundle has been checked with AgentFEM commit
40c965225fc3b8d12a93a68e0d6b01dd35d47d82 and AgentFEM-learning commit
06d454e4ee66a7ee4f4c8a705d5958ff9577d18c:
| Runtime gate | Result |
|---|---|
| Legacy-to-safe material-point maximum stress difference | 2.68e-7 Pa |
| 121-step maximum absolute stress | 292.547138 MPa |
| Final PEEQ | 0.0095270883 |
| Plastic bar, serial | 4/4 increments, converged |
| Plastic bar, two MPI ranks | 4/4 increments, converged |
| Serial/two-rank maximum-stress difference | 5.96e-8 Pa |
The elastic tangent passes a strict finite-difference check. Across the
audited plastic states, the current automatic-differentiation tangent differs
from fixed-old-state central differences by approximately 0.05–0.80%. The
reported global tests converge without cutback, while an exact consistent
plastic tangent remains an open numerical improvement. See
artifacts/agentfem/RUNTIME_VALIDATION.md and
artifacts/agentfem/runtime_validation.json.
Historical boundary-extension checkpoint
This package retains the originally published denim.pt checkpoint and adds
denim-expanded.pt, trained after a 500-trajectory capability-boundary
extension. The original 32-trajectory test split remains frozen.
The expanded checkpoint and all reported boundary metrics are tied to the
immutable dataset revision
c84f416e5a71daa157e406c50afc3fc73509b9ca.
| Test | Original checkpoint | Expanded checkpoint |
|---|---|---|
| Published held-out paths | 1.136 MPa | 0.714 MPa |
| New path OOD | 1.163 MPa | 0.743 MPa |
| Amplitude OOD | 3.091 MPa | 2.294 MPa |
| Long-history stress test | 11.261 MPa | 10.906 MPa |
The long-history result is the present capability boundary: additional
ordinary trajectories substantially improve path and amplitude tests but do
not remove long-horizon drift. Full metrics and data-quality evidence are in
artifacts/model_metrics.json and artifacts/QUALITY_REPORT.md.
DENIM means Discrete-Energy Neural Internal-variable Model. It is a compact gray-box constitutive model for path-dependent small-strain J2 plasticity. Elasticity, yield geometry, associative flow, plastic incompressibility, non-negative plastic increments and the implicit return map remain explicit. Neural components represent only the unknown isotropic and kinematic hardening closure.
The name does not currently imply a proved discrete-energy theorem. The
published bundle declares energy: false: return mapping and the listed
kinematic/plasticity constraints are implemented, while a formal global
energy/dissipation guarantee remains outside the present claim.
Why this checkpoint is an incomplete-physics test
The AgentFEM reference material has three kinematic-memory channels and a piecewise-linear isotropic-hardening table. This checkpoint has only two memory channels and receives neither the reference equations nor parameters. The second learned state must close two unresolved reference time scales.
The model has 918 trainable parameters. On two completely held-out non-proportional path families:
| Model | RMSE | R2 |
|---|---|---|
| Incomplete J2 | 59.541 MPa | 0.730208 |
| GRU (49,254 parameters) | 76.988 MPa | 0.548933 |
| DENIM | 1.136 MPa | 0.999902 |
Research progression
DENIM follows an earlier multiaxial benchmark in the same dataset repository. The results form one research progression, but not one raw-score leaderboard:
| Stage | Physics available to the model | Path-OOD stress RMSE | Meaning |
|---|---|---|---|
| MLP / GRU / LSTM / TCN / physics-state GRU | weak or soft physics | 60.44–105.81 MPa | sequence fitting generalizes poorly to unseen paths |
| Physics-integrator NN | correct J2/Chaboche equations and state structure | 0.0237 MPa | white-box fusion ceiling when the governing form is known |
| DENIM | J2 skeleton, but hardening law and one reference memory scale are hidden | 1.136 MPa | closes genuinely missing evolution terms and remains FE-deployable |
The Physics-integrator and DENIM rows use different frozen protocols. Their absolute RMSE values must not be ranked directly. The progression asks a more useful question: how far can architecture-level physics be retained when the true evolution equations or internal-variable structure are not known?
Coarse/fine path RMSE was 0.944/0.958 MPa. In 12-element notched-bar tests, cyclic and monotonic reaction relative-L2 errors were 0.636% and 0.500%. A severe cyclic case passed with one global trust-region fallback.
Use the current v2.2 checkpoint
import torch
from safetensors.torch import load_file
from src.t2_denim_conditional import ConditionalDENIM, rollout
model = ConditionalDENIM(channels=2, embedding=32, hidden=48)
model.load_state_dict(
load_file("capability_v4/weights.safetensors", device="cpu")
)
model.eval()
# strain: [batch, steps, 6]; descriptor: [batch, 13]
# tensor-shear Voigt order: xx, yy, zz, xy, yz, xz
result = rollout(strain, young, poisson, yield_stress, descriptor, model)
stress = result["stress"]
Use material_descriptor(...) from src/t2_denim_conditional.py to construct
the descriptor. The capability_v4/model.json manifest records the exact
parameter/state contract. The top-level inference.py, config.json and
denim.pt belong to the legacy fixed-material model.
For current finite-element use, download capability_v4/ and select the
denim.conditional.v2 AgentFEM-learning adapter. The following commands are
legacy fixed-material examples retained for reproduction:
python examples/learned_constitutive_denim/case.py \
--bundle /path/to/agentfem_bundle
python examples/learned_constitutive_denim/global_bar.py \
--bundle /path/to/agentfem_bundle --displacement 0.008
The fixed source implementation is available at AgentFEM-learning commit 06d454e.
Scope and limitations
- controlled synthetic J2/Chaboche families; not a universal metals model or experimental calibration;
- small-strain, rate-independent, isotropic J2 plasticity;
- trained with privileged high-fidelity internal-state supervision;
- material-family indicators and parameter descriptors are required; arbitrary unseen constitutive families are not automatically supported;
- AgentFEM/PyTorch implicit execution is verified, but this is not a compiled or production-certified UMAT;
- implicit return mapping is slower than direct GRU inference;
- the severe cyclic structure test still needed one global fallback;
- the current plastic autodiff tangent is numerically useful but does not meet the strict material-point consistency tolerance used in the runtime audit.
The associated data, source trajectories and evidence are available at AgentFEM-Material-Loading-Memory.
Article evidence v3
See article_evidence_v3/ for ablation, material-family transfer, and AgentFEM global-solve evidence.
- Downloads last month
- 97

