NASA IMS Test Set 2 six-gate benchmark
A frozen chronological evaluation of the MCIFT six-gate condition-monitoring policy on 984 NASA IMS bearing recordings.
- Evidence status
- Public run report — benchmark PR under review
- Protocol
- mcift.gates.ims-six-gate.v1
- Last verified
- 2026-07-21
- Benchmark commit
7ae4383fd642
Executive result
The six-gate run produced broad screening activity but only five persistent and three high-confidence warnings. The first persistent and high-confidence warning appeared 49 recordings, or 8 hours 10 minutes, before the terminal recording. This timing is not necessarily lead time before physical fault onset.
The complete six-gate system produced no positives on the 96-recording healthy holdout. G6 alone activated on 8 recordings, but the remaining MCIFT gates prevented those isolated activations from becoming screening, persistent, or high-confidence decisions.
Why six gates
The gates separate broad change screening from temporal persistence, progression, directional consistency and conventional vibration corroboration.
- G1 — Global deformation
- G2 — Local relationship damage
- G3 — Persistence
- G4 — Directional consistency
- G5 — Robust progression
- G6 — Conventional vibration agreement
G6 uses centred RMS, excess kurtosis and crest factor as conventional vibration corroboration. The basic any-feature/any-channel G6 rule was sensitive on the healthy holdout, producing 8 positives in 96 recordings.
Decision policy
G1 OR G2G3 AND (G1 OR G2) AND (G4 OR G5)G3 AND G6 AND (G1 OR G2) AND (G4 OR G5) AND at least 5 of 6 available gatesUnavailable gates are reported separately from failed gates.
Frozen protocol
- Dataset
- NASA IMS Bearings · Test Set 2
- Recordings
- 984
- Sampling rate
- 20,000 Hz
- Samples per recording
- 20,480
- Channels
- 4 bearing channels
- Terminal failure
- Bearing 1 outer race
- Seed · history
- 20260101 · 32
- Processing · schema
- mcift.exchange.vibration.v1 · 4
| Partition | Interval | Count | Role |
|---|---|---|---|
| Reference | [0,96) | 96 | Fit channel scales and relationships |
| Calibration | [96,352) | 256 | Fit thresholds |
| Healthy holdout | [352,448) | 96 | Fresh history; monitor not updated |
| Evaluation | [448,984) | 536 | Chronological order; monitor not updated |
Labels did not influence preprocessing or threshold selection. Evaluation data did not influence fitting.
Main outcomes
The terminal recording is used as the timing endpoint because independently documented physical fault-onset time is not available.
Healthy holdout and false-positive context
Healthy-holdout activity
No full-MCIFT positives were observed in this frozen 96-recording holdout. This is not proof of a zero false-positive rate in the wider population.
Negative controls
The tested disrupted controls produced no persistent or high-confidence warnings under this frozen protocol. Negative controls reduce some simple alternative explanations, but they do not prove that MCIFT identified a causal physical mechanism.
Localization did not succeed
The documented terminal failure concerned bearing 1. The run did not correctly localize that bearing.
Candidate nodes and edges are associations, not causal localization.
Runtime and reproducibility
- Complete run ID
- 20260721T092816Z-9f97efa9
- Full wall-clock runtime
- 1,904.3 s · 31 min 44 s
- Primary evaluation
- 147.7 s · 3.63 recordings/s
- Controls
- 997.7 s
- Peak resident memory
- 468,983,808 B
- Dataset parsing
- 984 / 984
- Result rows
- 536 · 96 · 536/control
- Deterministic figures
- 13
The public repository contains the runner, frozen configuration, schemas and methodology. Raw NASA recordings and local run directories are not committed.
The benchmark was executed against the recorded MCIFT development source. The public 0.1.0a1 package contains the reviewed public implementation line, but the exact benchmark provenance remains tied to the recorded source and benchmark commits.
MCIFT 27663fab8a9b164fbe8143537af24b988c5e109c · benchmark 7ae4383fd642f1ab99e64bebdb1e7047b39496dcWhat this result supports
- The six-gate software can be executed reproducibly on the frozen IMS Set 2 protocol.
- The implementation produced separate screening, persistent and high-confidence evidence.
- The run produced strong warnings before the terminal recording.
- The frozen healthy holdout had no full-MCIFT positives.
- The tested disrupted controls had no persistent or high-confidence warnings.
- The run is computationally practical on ordinary CPU hardware.
What this result does not support
- Proof of generalization, production reliability or safety certification.
- Proof of predictive superiority or independently established physical-fault-onset lead time.
- Reliable causal localization.
- Validation on other IMS sets; Exathlon evaluation has not yet been completed.