<!-- Validation result (premise check, ran BEFORE the design): premise "European GA is largely career-track" = NOT SUPPORTED; accident rate estimable = NO. Cycle red-team verdicts: needs another pass, needs another pass, needs another pass. -->

# FRESH CONFIRMATORY PRE-REGISTRATION — Pathway-vs-Precursor Equivalence/Non-Inferiority Study for US Part 121 First Officers (with a falsified-premise European-GA limitation arm)

**Date:** 2026-06-13.
**Document class:** NEW, SEPARATE confirmatory pre-registration — the charter-sanctioned route (Amendment 2 / Broad-Round Governance Charter, hash `15d5109c`) by which an exploratory candidate becomes a confirmatory claim: a fresh study with its own locked hypotheses, frozen outcomes, decision rule, and lock.
**Lock status:** LOCKED at this revision. Locking provenance: this document is sealed against the repository state at parent commit `d1b5d46` (`d1b5d464a6f0291952017900a67b3b962f4da5f2`); the lock hash is the commit that introduces this file. No estimand, outcome, margin ceiling, denominator, matching rule, multiplicity rule, claim-eligibility gate, or decision rule below may change after this lock. The two execution gates (G1-residual, G2) are NOT lock-blocking under the cycle-3 resolution adopted here (see §0); they are execution-readiness conditions, not design-openness conditions.

---

## Relationship to the locked net-ledger (NON-NEGOTIABLE)

This study is DISTINCT from the locked H_A/H_B net-expected-lives ledger (PRE-REGISTRATION.md §§1–7). It does NOT re-open, re-run, re-weight, or re-label that ledger. The locked verdict — **"sign indeterminate, magnitude small, modest adverse lean"** — STANDS UNTOUCHED, is restated here VERBATIM, may be cited but **may not be recomputed** by anything in this document. None of the four cost dials, `lives_saved_per_decade`, `rule_attributable_share`, `rule_binding_fraction`, the displaced-air-leg comparator, the road-substitution chain, or `rho_AB` is touched. The band-swap, Colgan-as-causal, benefit-at-hard-zero, and GA-rate-substitution prohibitions remain in force. No result of this study may change, reweight, relabel, or re-open the locked H_A-vs-H_B net-lives distribution, `P(net<0)`, `E[net]`, or the permitted sign-indeterminate / magnitude-small verdict.

**Author / role separation:** Authored by the research-integrity / study-design agent. Execution is role-separated and, by data-access necessity (§12), delegated to a data holder (FAA ASIAS via the ASIAS Executive Board directed-study route, or a participating Part 121 carrier / airline consortium under a data-use agreement). FRISA cannot itself access protected FOQA/ASAP data.

---

## 0. WHAT IS LOCKED NOW vs WHAT REMAINS AN EXECUTION CONDITION (resolves the cycle-3 [fatal] "deliberately-unlocked" defect)

The cycle-3 red team correctly ruled that a pre-registration cannot call itself locked while its single most gameable parameter — the equivalence margin Δ — is left to a future computation. The defect is cured here by **path (a): the per-outcome safety-meaningful effect-size CEILINGS are declared NUMERICALLY NOW** (§7). These ceilings are normative/literature-anchored and require NO data; nothing prevents fixing them pre-extract, and they are fixed here. The locked Δ for each outcome is its declared ceiling. The empirical noise-floor computation, when performed by the data holder, may ONLY TIGHTEN a margin below its ceiling — it can NEVER widen one. Therefore the TOST decision boundary is fully specified at lock: **label-before-look holds because the margin's binding upper limit is a fixed number in this document, not an IOU.**

What is locked NOW (immutable after this commit): the equivalence/non-inferiority decision frame; the co-primary family (O1, O3); all outcome definitions and the O1/O3 event-set freeze rule; the treatment/comparator definitions and military-750h exclusion; the single primary denominator-and-window cell; the matching and covariate sets; the multiplicity scheme and the enumerated BH-FDR family; **the numeric per-outcome safety-meaningful margin ceilings (§7)**; the §3a claim-eligibility gate; all symmetric-honesty and confound caveats; the European-arm disposition; and every decision/reporting rule in §11.

What remains an EXECUTION-READINESS condition (does NOT re-open the design, cannot relax any locked ceiling, and resolves outside this document in the append-only ledger):

- **G1-residual — noise-floor tightening only (margin CEILING already locked).** The data holder computes the empirical between-carrier / between-fleet historical variation in each precursor base rate (the natural noise floor) and **reports its magnitude in the ledger**. The operative Δ per outcome is then the SMALLER of {locked ceiling (§7), measured noise floor}. The noise floor can only tighten; it can never raise a Δ above the locked ceiling. This computation is **firewalled per §7a** (label-blind, sealed before the pathway-split extract). G1-residual cannot loosen anything locked; it is therefore not lock-blocking.
- **GATE G2 — pathway-to-event linkage identifiability.** It is UNRESOLVED whether entry-pathway can be joined to de-identified FOQA/ASAP precursor events WITHOUT re-identifying individual pilots — the exact thing the ASIAS de-identification MOUs exist to prevent. The study is **execution-CONTINGENT** on a concrete, named mechanism (e.g., carrier-held pathway records linked to events INSIDE the MITRE/CAASD trusted-third-party enclave, returning only de-identified aggregate rates, no pilot re-identification possible). **Pre-committed branch:** if linkage cannot be done without re-identification, the US arm is reported **"not executable as specified"** and is NOT relaxed to a coarser pathway proxy. **Predicted most-likely outcome (symmetric-honesty, mirroring A1.5's "realistic prior: the denominator gate fails"): G2 most probably resolves NO.** No existing program performs the bespoke pilot-level pathway-to-event join this study requires; the de-identification MOUs exist specifically to prevent it. The honest base-rate expectation for this study's real-world end-state is the "not executable as specified" branch. Executability is presented as contingent, with the contingency expected to fail.

**Lock rule:** the design is locked now. The git-hash lock binds the design. Execution may proceed ONLY when G1-residual has supplied ledger-recorded noise-floor magnitudes and G2 reads YES with a named non-re-identifying mechanism; otherwise the pre-committed "not executable as specified" branch fires and the study reports the null-of-data, never a measured null.

---

## 1. Purpose, the forking-path threat, and the candid limit on what this study can deliver

The candidate (broad-round, reworded neutral): *test whether raw-1,500-hour-pathway Part 121 first officers differ from structured-pathway / R-ATP first officers on pre-registered non-fatal precursor safety metrics at matched carriers/fleets/IOE-period.* Precursor safety admits many operationalizations (which FOQA exceedances; ASAP raw counts vs rates; unstable-approach threshold; go-around as numerator or denominator; per-flight vs per-hour denominator; which covariate set; equivalence vs non-inferiority vs superiority; one- vs two-sided; which window; which margin). An analyst could reach "they are the same," "the cheaper pathway is worse," or "the cheaper pathway is better" by selecting AFTER seeing the precursor rates. This document fixes every such choice in advance, declares the decision frame (equivalence/non-inferiority) now, and locks the numeric ceiling on the single most gameable parameter (the margin).

**Candid statement of deliverable value (resolves cycle-3 [major] confound and precursor-validity defects).** Two design facts bound what this study can ever produce, and they are stated up front, not buried:

1. **No causal pathway-safety verdict is obtainable.** The pathway-vs-aptitude/screening confound (§6) is non-identifiable cross-sectionally. A within-margin equivalence therefore **cannot** establish that raw hours are dispensable (superior structured-cohort screening could be exactly offsetting a real hours-deficit), and a difference cannot be read as an hours-dose effect. The study's value is limited to a **descriptive, confound-bounded equivalence on operational precursors.** If the principal expects a policy-actionable causal answer to "is the cheaper pathway no worse," **this design cannot supply one**; a longitudinal or instrument-based design would be required.
2. **The well-powered outcomes are the least safety-load-bearing.** The common co-primaries O1/O3 are well-powered but operational, not catastrophic; the only catastrophic-tail outcome (O5, LOC-I precursors) is rare and structurally likely to remain underpowered at any realistically attainable n (§9). The §3a claim-eligibility gate forbids selling an O1/O3 equivalence as a *safety*-equivalence. Accordingly, the study's PURPOSE is hereby framed as an **operational-precursor equivalence study** — its headline ambition is matched to the deliverable the design can actually produce — and any catastrophic-tail (safety) conclusion is admissible only under the §3a fully-powered form, expected to be rare.

## 2. The primary confirmatory hypothesis (US arm) — stated as a neutral TEST-OF-WHETHER

**H1 (primary, confirmatory).** TEST OF WHETHER, at matched US Part 121 carriers/fleets/route-structures and IOE/early-line periods, first officers who entered via the **raw-1,500-total-flight-hour pathway** differ from first officers who entered via the **validated civilian degree-academy R-ATP pathway (§61.160 reduced-hour, 1,000 h / 1,250 h)** on a pre-registered panel of **objective, machine-recorded non-fatal precursor safety metrics**.

- **Decision frame, locked a-priori:** TWO ONE-SIDED TESTS (TOST) for **equivalence**, with a pre-declared **non-inferiority** read of the structured pathway as the directional secondary. The **null is non-equivalence** (the two pathways' precursor rates differ by at least the margin); the **alternative is equivalence within the margin**. This is explicitly NOT a superiority test and NOT a directional result statement.
- **Why this frame:** the policy question is whether the cheaper/faster structured pathway is *no worse* on safety precursors than the raw-hour pathway. Equivalence/non-inferiority is the correct frame for a "drop-in substitute" question and is least gameable post hoc.
- **What the deliverable can and cannot be:** the headline equivalence claim, even if achieved, is by construction **"equivalence on common OPERATIONAL precursors only, with the catastrophic-tail (LOC-I) question unanswered unless O5 is independently powered"** (§3a is a claim-eligibility gate, not a caption).

**Symmetric-honesty constraint — SYMMETRIC ACROSS THE DIFFERENCE AND THE EQUIVALENCE BRANCHES.** The selection/aptitude confound (§6) is non-identifiable cross-sectionally, and that obstruction is fatal to BOTH reads:

- *Difference branch:* a directional difference, if found, is reported with the selection confound foregrounded and is NOT read as a causal hours-dose effect.
- *Equivalence branch:* a within-margin equivalence is reported as **"precursor rates equivalent at the matched, dual-capped margin, with pathway and aptitude/screening NON-SEPARABLE — equivalence does NOT establish that raw hours are dispensable, because superior structured-cohort screening could be exactly offsetting a real hours-deficit (or masking a hidden hours-benefit)."** The study can refute NEITHER a hidden hours-benefit NOR a hidden hours-deficit.
- *Null-result branch:* a non-significant TOST is reported as **"equivalence not established at the pre-set margin / underpowered,"** NOT as evidence that the pathways differ.

Data-power failure ≠ measured null, and equivalence ≠ causal interchangeability — in all directions.

## 3. Pre-specified outcomes (frozen) and data source

All outcomes are **non-fatal precursor RATES**, denominated (§5), never raw counts. The co-primary family is restricted to **objective, machine-recorded FOQA measures**.

| # | Outcome (frozen) | Numerator | Source | Objective/Self-report | Role |
|---|---|---|---|---|---|
| O1 | FOQA exceedance rate | Count of FOQA-flagged exceedance events, per the carrier's event-set version frozen as of the reference date (§3b) | Airline FOQA via ASIAS | Objective (machine-recorded) | **CO-PRIMARY** |
| O3 | Unstable-approach rate | Approaches breaching the carrier's stabilized-approach gate, gate version frozen as of the reference date (§3b) | FOQA | Objective (machine-recorded) | **CO-PRIMARY** |
| O2 | ASAP report rate | Crew-filed ASAP reports attributable to the FO | Airline ASAP via ASIAS | **Self-report (reporting-propensity proxy)** | **SECONDARY / contextual only** |
| O4 | Go-around rate | Executed go-arounds | FOQA / ASAP | Objective but directionally ambiguous | Secondary / contextual |
| O5 | LOC-I-relevant precursor rate | Composite of pre-declared LOC-I precursor exceedances (bank/pitch/stall-warning/upset/stick-shaker, envelope exceedances) | FOQA | Objective (machine-recorded) | **Secondary, catastrophic-tail, claim-eligibility-gating (§3a)** |

- **Co-primary family (locked): O1 and O3 ONLY** — both objective, machine-recorded FOQA measures.
- **O2 (ASAP) is secondary/contextual, uninterpretable in isolation.** ASAP report rate is a reporting-CULTURE / reporting-propensity measure, not a hazard-exposure measure: a safer, better-trained cohort can file MORE reports while a worse cohort files fewer. Structured-pathway academy graduates are plausibly trained into higher reporting propensity, so O2 is confounded with the treatment in an UNKNOWN direction and its sign is uninterpretable alone. O2 is therefore (a) never co-primary, (b) interpreted ONLY jointly with the objective FOQA rates O1/O3, and (c) accompanied by a pre-declared **reporting-propensity covariate / reporting-rate negative control** to bound the confound's direction.
- **O4 (go-around)** is **directionally ambiguous** (a go-around can indicate a *good* decision after an unstable approach), secondary/contextual only, interpreted jointly with O3, never alone as "worse."
- **O5 (LOC-I precursors)** is secondary, catastrophic-tail-relevant, **underpowered-by-design** (rare events), and **gates the admissibility of any equivalence headline** (§3a).
- **Data-source reality (verified against ASIAS program documentation and the locked pilot-track memo):** FOQA, ASAP, and ATSAP feed ASIAS as **de-identified, aggregated, proprietary** data under participant MOUs, with MITRE/CAASD (an FFRDC) as trusted third party; participants may access de-identified FOQA/ASAP only for topics **approved by the ASIAS Executive Board (AEB)**. ASIAS already runs FOQA/ASAP **directed studies** on these precursors (unstable approach, go-around). It does **not** publish pilot-entry-pathway-stratified precursor rates, and no regulator (EASA, FAA, ICAO) publishes accident/incident rates stratified by initial-license pathway. **This pre-registration is therefore designed for a data holder to execute** — not self-executable from public data, and execution-contingent on G2 (§0).

## 3a. O5 as a CLAIM-ELIGIBILITY GATE, not a reporting rule

> **CLAIM-ELIGIBILITY GATE (binding):** No standalone "equivalence" headline is admissible while O5 is underpowered. A headline asserting pathway equivalence is admissible ONLY in one of two forms:
> 1. **Fully-powered form:** equivalence is established on co-primaries O1/O3 **AND** O5 returns a powered verdict (equivalence OR difference) at its ratified tighter margin — only then may the study state a "precursor-safety-equivalence" headline that includes the catastrophic tail.
> 2. **Scope-limited form (default expectation):** if O5 is underpowered, the ONLY admissible headline is the explicitly scope-limited *"Equivalence on common OPERATIONAL precursors O1/O3 within the ratified margin; the catastrophic-tail (LOC-I) safety question is UNANSWERED — no safety-equivalence is claimed."*

Any headline stating or implying pathway *safety*-equivalence while O5 is underpowered is **non-conforming and inadmissible**, not merely under-caveated.

## 3b. O1/O3 event-set freeze (resolves cycle-2 [minor] carrier-tunable-threshold)

O1's numerator is the carrier's standard FOQA event set and O3's is the carrier's stabilized-approach gate; both are carrier-defined, heterogeneous across carriers, operator-tunable, and can drift across the IOE-period window. Within-carrier exact matching (§6) absorbs cross-carrier heterogeneity but NOT mid-window redefinition. Therefore, as a frozen specification:

- For each participating carrier, the **event-set version / stabilized-approach-gate version is frozen as of a fixed reference date** recorded in the frozen-outcomes table for that carrier.
- **Mid-window event-set or gate redefinitions are EXCLUDED from the primary contrast** (events scored only under the frozen version are counted); any carrier that redefines mid-window is handled as a **pre-declared sensitivity stratum**, never folded silently into the primary cell.
- The numerator definition therefore **cannot move within the matched contrast.**

## 4. Treatment definition (frozen): pathway, not hours

- **Treatment arm (structured), as validated:** FOs who entered Part 121 line flying under the **civilian degree-academy R-ATP pathway of 14 CFR §61.160 — approved-institution 1,000 h or 1,250 h reduced-hour authority** (broad-round candidate line 75: "R-ATP-DEGREE-pathway first officers").
- **Comparator arm (raw-hour):** FOs who entered via the **unrestricted 1,500-h total-flight-time** route (time-builder / CFI accumulation), no reduced-hour authority.
- **Military 750-h pathway — NOT pooled into "structured."** §61.160 also grants a military 750-h R-ATP, but military pilots are an extreme, categorically different selection population (military flight screening, jet/multi-crew exposure, age profile, different hour-TYPE). Pooling them would maximize the aptitude confound and change the estimand. The military-750h cohort is **EXCLUDED from the co-primary treatment arm**, carried ONLY as a **separate, clearly-labeled exploratory stratum** with its own confound statement; it never enters the H1 headline.
- **Scoping deviation-by-clarification re "a-priori-established base rate" (resolves cycle-2 [minor] mechanism-substitution).** The sanctioned candidate (line 75) requires "a precursor with an a-priori-established base rate." This pre-registration operationalizes that as the **pre-extract, label-blind noise-floor computation under §7a** (base rates and their between-carrier scatter computed and sealed BEFORE the pathway-split contrast extract, so label-before-look is preserved). This is a **substitution of mechanism relative to the worded candidate** — defensible but not identical to it — and is recorded as a **scoping deviation-by-clarification in the append-only ledger**, not presented as candidate-verbatim.
- **Pathway-label feasibility is a HARD EXECUTION GATE (G2, §0).** If pathway cannot be linked to events without re-identification, the arm is reported "not executable as specified," not forced.

## 5. Exposure denominator (frozen) — ONE primary cell

- **THE PRIMARY CELL (the only cell that may establish the headline):** outcome rates **per 1,000 flights** (FOQA is per-flight native), accrued during the **IOE + first 300 line-hours** window. The H1 equivalence headline is read off THIS SINGLE CELL only.
- **Secondary sensitivity cells (reported, but can NEVER establish the headline):** alternate denominator **per 1,000 FO block-hours**; alternate window **first-12-months**.
- **Anti-selection lock:** no window and no denominator may be chosen post hoc as "the one that shows equivalence." The 2 windows × 2 denominators = 4 looks are honestly counted; exactly ONE (IOE+300h, per-1,000-flights) is primary; the other three are strictly secondary sensitivity, corrected under BH-FDR with the directional secondaries (§8).

## 6. Matching / confounder strategy (frozen)

- **Exact match on:** carrier, fleet/aircraft-type, route-structure stratum (jet-regional vs turboprop; hub-and-spoke vs point-to-point), and calendar IOE-period.
- **Covariate adjustment (pre-declared, no post-hoc additions):** FO age at entry, total flight hours at entry, recency, captain-experience pairing, base/weather exposure, and a **reporting-propensity covariate** (to bound the O2 confound).
- **Selection treatment (explicit, carried into BOTH result branches):** civilian degree-academy R-ATP pilots are **differently aptitude-screened** (academy admission). This is a **structural confound**, not just imbalance: any observed difference — AND any observed equivalence — confounds pathway with aptitude/structure/recency. Matching + adjustment **narrow but cannot break** it; a clean dose-vs-selection separation is **likely non-identifiable cross-sectionally** (same obstruction recorded for locked H_C; mirrors the pilot-track memo's finding that Europe never fields the unscreened low-hour pilot the rule filters out, so pathway is perfectly co-confounded with aptitude/structure/recency). A **negative-control outcome** (a precursor with no plausible pathway mechanism) and an **E-value** for unmeasured confounding are pre-declared to bound, not eliminate, the confound. The non-separability caveat attaches to the equivalence read identically to the difference read.

## 7. Equivalence MARGIN — numeric safety-relevance CEILINGS LOCKED NOW, noise floor may only tighten (resolves cycle-1 [fatal] margin-gaming, cycle-2 [major] wide-floor, cycle-3 [fatal]/[major] undeclared-cap)

The prior ±25% band is **WITHDRAWN.** The margin is now **dual-bounded** with the safety-relevance ceiling **declared numerically and locked in this document** (label-before-look), and the noise floor permitted **only to tighten**:

**Δ(outcome) = SMALLER of { locked safety-meaningful ceiling below, measured between-carrier noise floor (§7a) }.**

The noise floor can never raise a Δ above its locked ceiling; a floor narrower than the ceiling tightens Δ further (never loosens). Each ceiling is anchored to an external, pre-existing quantitative reference standard and is expressed as a cap on the **rate ratio (RR)**, structured-vs-raw, so that the cap is the largest adverse rate elevation deemed operationally tolerable for that outcome:

| Outcome | LOCKED safety-meaningful ceiling on RR (structured vs raw) | External anchor for the ceiling |
|---|---|---|
| O1 — FOQA exceedance rate (co-primary) | **RR ≤ 1.15** (±15%) | Anchored to the smallest FOQA exceedance-rate shift an operator FOQA program treats as a monitored trend trigger rather than routine scatter; tighter than a generic ±25% precisely because O1 carries the headline. |
| O3 — unstable-approach rate (co-primary) | **RR ≤ 1.15** (±15%) | Anchored to established stabilized-approach action-level practice (unstable-approach rate is a standard FOQA/ASIAS directed-study action metric); a 15% relative elevation is the cap below which the change does not cross an operational action level. |
| O5 — LOC-I precursor rate (catastrophic tail, claim-gating) | **RR ≤ 1.05** (±5%) — TIGHTEST | Anchored to catastrophic-tail relevance: because an LOC-I precursor elevation maps to a modeled increment in catastrophic-event probability, the tolerable RR cap is set near unity. This is the binding ceiling for any §3a form-1 (full-tail) claim. |
| O2 — ASAP report rate (secondary/contextual) | reported only; **no equivalence headline may rest on O2** | Reporting-propensity proxy; sign uninterpretable in isolation (no safety-meaningful ceiling is definable). |
| O4 — go-around rate (secondary/contextual) | reported only; **no equivalence headline may rest on O4** | Directionally ambiguous; no safety-meaningful ceiling is definable. |

- **Quantitative adverse-difference criterion (replaces prose "justification").** A margin is admissible ONLY if the operative Δ is smaller than the rate change that would correspond to a pre-stated increment in catastrophic-event probability for that outcome. For O5 this is the binding construction; for O1/O3 the 15% ceiling is the externally-anchored operational-action-level proxy for that same standard. The data holder may not substitute a prose paragraph for this criterion.
- **Decision rule (operative once §7a supplies the floor):** equivalence on a given outcome is declared if the two-sided 90% CI for the RR lies entirely within that outcome's operative Δ band; non-inferiority of the structured pathway is the one-sided read (upper 95% CI bound of RR below the operative upper limit). The author may NEVER loosen a margin post hoc; the noise floor may only tighten it. **The ceilings above are the locked decision boundary; no wider value is admissible under any circumstance.**

## 7a. Label-blind, sealed noise-floor computation (resolves cycle-3 [major] pre-extract-look)

The noise-floor computation peeks at the marginal distribution of the co-primary outcomes; if performed by a party who can see the pathway split and the carrier mix, that party could foresee whether IU equivalence is reachable. Therefore, as a frozen firewall:

- The between-carrier / between-fleet base-rate and scatter computation is performed by a party **BLIND to the pathway labels**, on a **pre-period or pathway-pooled extract**, and the resulting noise-floor magnitudes are **sealed in the append-only ledger BEFORE the pathway-split contrast extract is generated.**
- The pathway-split extract that scores the H1 contrast is generated only after the sealed floor is recorded. "A-priori base rate" without this label-blind seal would be post-look by another name; the seal is mandatory, not advisory.

## 8. Multiplicity control across arms/outcomes (locked) + cross-cycle registration

- **Co-primary family (O1, O3 — objective FOQA only):** equivalence must hold for **both jointly** under an **intersection-union (IU)** rule (each clears its own TOST at α=0.05 one-sided / 90% CI at its operative Δ). IU is conservative for the equivalence direction and needs no α-spend inflation. Read off the SINGLE primary cell (§5) only. **Subject to the §3a claim-eligibility gate:** even a clean IU pass yields only the scope-limited form-2 headline unless O5 is independently powered.
- **BH-FDR family (q=0.05) — explicitly enumerated so it cannot be padded:** comprises ONLY the **US directional secondaries** — the non-inferiority directional reads; the secondary/contextual outcomes O2, O4, O5; and the three secondary sensitivity cells (alternate denominator and alternate window). No other tests enter this family.
- **The European arm contributes ZERO tests** to the BH-FDR family; it is confirmatory-ineligible, spends no α, and may not pad the family count.
- **Cross-cycle registration (resolves cycle-3 [minor]):** this study's BH-FDR family is **registered into the cumulative cross-cycle append-only ledger and corrected at the round level**, not merely within-study. The within-study family size (the enumerated US directional secondaries + three sensitivity cells) is folded into the running cumulative broad-round BH-FDR count recorded in the ledger; the cumulative running family-size count is cited there at execution time, so per-study FDR control does not silently ignore the round-level correction the charter mandates.

## 9. Power / sample needs (frozen) and the O5 hard finding

- **Power target:** 90% to declare equivalence within each outcome's operative Δ (§7) at α=0.05 (one-sided per TOST limb), under assumed true RR=1.0, given each carrier's empirical precursor base rate.
- **Sparsity reality and the O5 hard finding (resolves cycle-3 [major] precursor-validity inversion):** O5 (LOC-I precursors) and some FOQA exceedances are **rare**; achieving 90% power at the locked tight O5 ceiling (RR ≤ 1.05) requires **multi-carrier, multi-year pooling at full ASIAS scale**, and even pooled ASIAS-scale exposure is **expected to be insufficient** to power O5 at its tight ceiling. This is stated as a **hard design finding, not a contingency**: the realistic expectation is that O5 returns "equivalence not estimable — underpowered," which by §3a caps the admissible headline at the scope-limited operational-precursor form. Exact n (flights / FO-years) is computed by the data holder from the sealed base rates before unblinding; if power is unattainable for any outcome, that outcome is reported "equivalence not estimable — underpowered," never spun as a null.

## 10. SECONDARY ARM (Europe) — confirmatory-INELIGIBLE, documented-limitation only

**Disposition (binding):** the European-GA accident-rate "natural experiment" is **NOT a co-primary** and **NOT confirmatory-eligible**. Per the §A1.5 confirmatory-eligibility gate it **FAILS on all three counts before any rate is inspected**:

1. **Denominator:** European GA flight hours are **estimated, not measured** — a voluntary annual EASA/GAMA owner-operator survey aggregates by aircraft *category* (aeroplane non-complex, glider, microlight, helicopter), **never cross-tabulated by pilot pathway / career-track / age**. EASA itself calls GA flight-hour collection "a big challenge."
2. **Pathway observability:** GA accident records carry **no career-track field**; pathway is **not observable**, only assumable. The premise cannot be tested in-pool.
3. **Strata:** age/start-age strata are **not populated** at estimand granularity.

**Premise correction (factual, on the record):** the principal's claim that *"European GA is largely career-track"* is **unsupported and is not adopted.** Validation established that European GA, like US GA, is **overwhelmingly recreational/sport/private/leisure** (EASA/ICAO define GA as all civil aviation EXCEPT commercial air transport and aerial work). Career ATPL/MPL/CPL ab-initio training is a **thin, aptitude-screened, structurally-separate ATO segment** that **does NOT populate the GA accident pool** (the frozen-ATPL FO enters airline ops near the low-hundreds of hours via ATOs, not via the recreational GA fleet). The natural-experiment framing **misidentifies what European GA is** and is **not adopted.**

> **DATA-ACQUISITION-TARGET FIGURES — NOT ESTABLISHED FACTS (resolves cycle-3 [minor] unverified-claim laundering).** The specific European fleet magnitudes cited during validation (order ~50,000 motor general/business aircraft, ~2,800 turbine; ~180,000–200,000 microlight/non-motor sport-and-recreation aircraft; ~5,000 European commercial airline fleet; frozen-ATPL airline entry in the ~195–250 h range) are **flagged as data-acquisition-target figures requiring citation, NOT as repo-locked established facts.** They illustrate the recreational-dominance conclusion (which the locked record supports qualitatively) but are not to be relied upon numerically without sourcing. The qualitative conclusion — recreational dominance and a thin, separate career-training segment outside the GA accident pool — is what the validation and the locked pilot-track memo support.

**Scoping (binding):** carried **only** as a pre-registered **"not estimable — data/power failure"** branch and a **data-acquisition target** (estimable only if some future regulator/linkage produces pathway-resolved GA exposure denominators, which none currently publishes). It is **symmetric-honesty-protected**: this is a **data-power failure, NOT a measured null**, may not be read as evidence of experience–safety equivalence, nor substituted for the air-transport estimand (GA-rate-substitution prohibition in force). It contributes no confirmatory test, **zero BH-FDR family members** (§8), and cannot move the US headline or the locked net-ledger.

## 11. Forking-path safeguards (locked) and per-cycle discipline

1. Decision frame, outcomes, denominator, windows, matching, multiplicity, and the O1/O3 event-set freeze (§3b) all fixed **before** any extract — **label-before-look**. The MARGIN CEILING is a fixed number in this document (§7); the noise floor (§7a) is computed label-blind and sealed before the pathway-split extract and may only tighten.
2. **Rates not counts; objective machine-recorded precursors (O1/O3) co-primary; self-report (O2) demoted to uninterpretable-in-isolation contextual; terminal precursor outcomes, not training/check/pass-rate intermediates** (avoids the pilot-track memo's outcome-substitution trap).
3. Equivalence margin may not be loosened post hoc; it is dual-bounded by a **locked numeric safety-meaningful ceiling** (tightest for O5) and a label-blind noise floor that may only tighten; the noise-floor magnitude is reported in the ledger; superiority may not be substituted for equivalence post hoc.
4. Selection/aptitude confound declared, E-value + negative-control pre-registered; dose-vs-selection declared likely non-identifiable, the non-separability caveat applied SYMMETRICALLY to equivalence and difference reads (§2), and §1 states plainly that no causal pathway-safety verdict is obtainable from this design.
5. ALL outcomes reported regardless of result, including "equivalence not established," "not executable as specified," and "not estimable."
6. No exploratory/secondary result promoted to the confirmatory headline; ONE primary cell (IOE+300h, per-1,000-flights) declared; no outcome-dependent selection of window/denominator; military-750h never enters the headline.
7. **§3a is a CLAIM-ELIGIBILITY gate, not a caption:** no standalone equivalence headline is admissible while the catastrophic-tail O5 is underpowered; the default admissible deliverable is the scope-limited "common-operational-precursors-only, catastrophic tail unanswered" form.
8. **Adversarial / red-team review** of this pre-commit (G1-ceilings, §7a firewall, G2 linkage) AND of the output, per charter per-cycle discipline.
9. Cross-cycle multiplicity tracked in the cumulative append-only ledger; BH-FDR q=0.05 family enumerated in §8, naming ONLY the US directional secondaries (European arm contributes zero members) and registered into the round-level cumulative count.
10. **Precursor ≠ fatality:** the study tests non-fatal precursors and does NOT claim precursor equivalence licenses any fatal-safety conclusion; §3a forces this into admissibility, not just wording.
11. The candidate's "a-priori-established base rate" is operationalized as the pre-extract, label-blind §7a noise-floor computation and logged as a **scoping deviation-by-clarification** (§4).
12. **Headline lock restated VERBATIM:** the locked net-ledger verdict *"sign indeterminate, magnitude small, modest adverse lean"* stands UNCHANGED; nothing here recomputes it.

## 12. Honest data-access reality (what a real executor needs)

FRISA cannot access protected FOQA/ASAP/EASA-held data; this study must be HELD AND EXECUTED by a data holder. Routes:

- **FAA ASIAS route:** an **AEB-approved directed study** authorizing de-identified FOQA/ASAP aggregation on the frozen outcomes, with pathway linkage performed inside the MITRE/CAASD trusted-third-party enclave so no pilot is re-identified (the concrete mechanism G2 must confirm).
- **Carrier/consortium route:** a participating Part 121 carrier (or consortium) executes under a **data-use agreement**, holding FOQA/ASAP and pilot entry-pathway records, returning only de-identified aggregate rates.
- **European arm:** no executor can produce a pathway-resolved GA accident rate because no regulator (EASA / national CAA / ICAO) publishes one; the European arm is a documented limitation, not an executable analysis.
- **The executor must, IN ORDER:** (a) clear **G2** — confirm pathway-to-event linkage without re-identification, else report "not executable as specified"; (b) perform the **§7a label-blind, sealed noise-floor computation** and record magnitudes in the ledger; (c) set each operative Δ = smaller of {locked ceiling §7, sealed floor}; (d) freeze each carrier's O1/O3 event-set version at the reference date (§3b); (e) compute power from the sealed base rates (§9); (f) run TOST per §7–8 off the single primary cell; (g) apply the §3a claim-eligibility gate to any headline; (h) register the BH-FDR family into the cumulative cross-cycle ledger (§8); and (i) report all branches per §11. **Public data cannot execute this study; it is execution-CONTINGENT on G2, with G2 expected to resolve NO.**

## 13. Bottom line and EXECUTION-READINESS VERDICT

- **Design status: LOCKED.** Every estimand, outcome, denominator, window, matching/covariate set, multiplicity rule, claim-eligibility gate, decision rule, and the numeric per-outcome margin CEILING is fixed in this document. The only remaining computation (the noise floor) can solely tighten a margin below a locked ceiling and is firewalled label-blind (§7a); it cannot re-open the design. The cycle-3 [fatal] "deliberately-unlocked" defect is cured: the margin is no longer an IOU, it is a locked number.

- **EXECUTION-READINESS VERDICT: NEEDS-DATA-ACCESS (execution-contingent on G2; realistic prior is NOT-EXECUTABLE).** The US primary arm is a genuinely **estimable, fully-specified, locked** confirmatory equivalence/non-inferiority instrument — but it is **not execution-ready by FRISA**, because (1) it requires a data holder (FAA ASIAS via AEB directed study, or a Part 121 carrier/consortium) to hold and run protected FOQA/ASAP data, and (2) the pathway-to-event linkage identifiability gate G2 is unresolved and, on the symmetric-honesty base rate (mirroring A1.5's "the denominator gate fails"), is **expected to resolve NO**, in which case the pre-committed "not executable as specified" branch fires. The verdict is therefore **lockable-and-locked-as-a-design, needing data access to execute, with the most-likely real-world end-state being the not-executable / null-of-data branch.** It is NOT "lockable-and-ready-for-a-data-holder" in the strong sense, because readiness is contingent on a gate expected to fail; it is NOT "not-viable" either, because the design itself is sound and would execute if a data holder cleared G2.

- **Strongest deliverable the arm can ever produce:** either a full-tail equivalence (§3a form 1; rare, requires a powered O5 that ASIAS scale is expected not to reach) or an explicitly scope-limited common-operational-precursor equivalence that does NOT carry safety-equivalence meaning and CANNOT answer the causal pathway-vs-safety policy question (confound non-identifiable, §1/§6).

- **European secondary arm:** **confirmatory-ineligible**, a **falsified-premise documented limitation** and data-acquisition target, contributing zero confirmatory tests and zero BH-FDR members; its illustrative fleet figures are flagged data-acquisition targets, not established facts.

- **Net-ledger:** **untouched**; verdict restated verbatim — **"sign indeterminate, magnitude small, modest adverse lean."**
---

## Addendum (2026-06-13): External comparators considered and their disposition

Three cross-population proxies were evaluated as alternatives or supplements to the within-regional FOQA-precursor contrast. All are recorded here so reviewers see they were assessed, not overlooked. Companion memos: `feeder-network-proxy.md`, `part141-vs-regional-proxy.md`, `european-ga-premise-validation.md`.

- **European GA as a career-track natural experiment — REJECTED (premise falsified).** "European GA is largely career-track" is false: European GA is overwhelmingly recreational/sport/private; career ATPL/MPL training is a thin, screened ATO segment that does not populate the GA accident pool. Independently, no regulator publishes pathway/career-track-stratified GA accident rates (denominator is voluntary-survey-estimated; no career-track field), so the arm is confirmatory-ineligible and carried only as a documented "not estimable — data/power failure" limitation.

- **Part 141 schools vs regional jets — REJECTED (confound-dominated).** Part 141 is career-*concentrated*, not career-*exclusive*, and spans 0–1,500+ hours. The ~40–80× training-vs-Part-121 fatal-rate gap reflects aircraft (piston single vs turbine transport), operation (training maneuvers vs revenue IFR), crew (student+instructor vs two professionals), and oversight — not pilot experience. It cannot isolate the experience effect and is strictly worse than holding the operation fixed. Retained only as a one-line limitation; the training (~0.26–0.49 fatal/100k) vs Part 121 (~0.006) figures may appear once as descriptive environment-risk context, explicitly NOT an experience test.

- **Feeder/regional network as the study POPULATION — ADOPTED (qualified).** The regional network concentrates the low-experience FO cohort and supplies ~50% of national ASIAS data and abundant non-fatal precursors (~4–6 orders of magnitude more signal than the fatal endpoint). It is the correct population and denominator — but ONLY for a *within-regional* experience-stratified contrast, NOT regional-vs-major (which is confounded). The design therefore: stratifies within regionals (R-ATP tier / total hours / time-in-type tied to 14 CFR 121.438 / new-hire-vs-seasoned, carrying captain experience and pairing); uses MANDATORY/objective endpoints (FOQA exceedances, NTSB Part 830 / FAA AIDS) as primary, with voluntary ASAP/ASRS only as reporting-rate-adjusted sensitivity; normalizes per departure/cycle; and treats hours as one of several experience operationalizations. The binding constraint is event-level FOQA↔roster↔PRD linkage, which is not public and requires carrier-side or protected-ASIAS access — see `executive-option-and-unique-qualification.md` (the FAA/ASIAS+PRD is the only lawful national pathway, and commissioning the study is an available executive option under existing authority).

## Addendum 2 (2026-06-13): Grandfather / entry-cohort discontinuity — triangulation arm

Companion memo: `grandfather-rd-natural-experiment.md`. Proposed as a way to *measure* the experience effect by exploiting the 2013 cutoff. Disposition: **adopted as an EXPLORATORY, confirmatory-ineligible triangulation arm** — complementary to the within-regional precursor contrast, zero BH-FDR members, cannot touch the locked verdict.

- **Premise correction (design-relevant):** the FAA did **not** grandfather low-hour first officers — it *rejected* American Eagle/American requests and required all Part 121 SICs to hold ATP/R-ATP as of Aug 1, 2013 (the only true grandfather, §121.436(e), exempts pre-2013 *captains* from a 1,000-hr air-carrier-experience sub-requirement). The usable intuition recast correctly: FOs who *entered* pre-2013 at ~500–700 hr **persisted** in the active population, so the treatment variable is **entry hours at hire**, not "grandfather status" — and there is **no sharp step in who-may-fly** at the cutoff.
- **The RD-in-time is invalid here.** Aug 1 2013 is also the onset of the whole post-Colgan reform bundle (Part 117 fatigue, stall/upset training, SMS, PRD), so the discontinuity's running variable is collinear with the bundle — a sharp cutoff *inherits* the era confound rather than defeating it.
- **The valid form** is **within-post-period entry-hours stratification with a tenure interaction**: hold the bundle/era fixed for both the persisting pre-rule-entry (lower entry hours) and post-rule-entry (≥1,500) cohorts, observed in the same post-2013 environment, and model the tenure decay of any entry-hours signal. Confound-*bounded*, not confound-clean; leverage decays with tenure and is survivorship-filtered.
- **Why it earns a place:** it adds the one axis the within-regional FOQA contrast lacks — a cohort screened under a *lower* entry bar — giving partial leverage on the aptitude/selection confound the main design concedes is fatal cross-sectionally.
- **Feasibility: worse than the main study.** Needs event-level precursors linked to hire date **and entry hours**; the PRD **excludes flight/duty/rest-time records**, so it carries no entry-hours field — a strictly larger linkage burden than the FOQA arm, on top of the same protected-ASIAS/Gate-G2 blocker. Carrier-side or AEB-directed-study only.
- **Nearest prior work:** the Pilot Source Study 2018 is a genuine pre/post-FOQ cohort comparison but on *training performance, not incidents*, and found post-FOQ hires had *more* hours yet needed *more* training and completed less — itself a caution against a naive "more hours = safer" prior. No incident-rate entry-cohort study exists.
- **Reporting:** a null is a **data/power failure or a tenure-decayed signal, not a measured null**; the locked verdict ("sign indeterminate, magnitude small, modest adverse lean") is restated verbatim and untouched. Unverified: FAA's exact count of affected incumbent FOs and the N 8900.225 transition mechanics — retrieve from 78 FR 42324 before citation-grade use.
