# Final Audit Report — "Host type governs influenza evolutionary strategy across reservoir and spillover hosts"

- **Authors:** Maltepes, Markin, **Shank**, Ort, Sabre, Damodaran, Park, Kistler, Anderson, **Moncla**
- **Venue:** bioRxiv preprint, posted 2026-09-17 · **DOI:** 10.64898/2026.09.15.751820
- **Local copy:** `preprint_moncla_h3nx.pdf` / `preprint_text.txt` · **Existing review:** `PEER_REVIEW.md` · **Reproduction:** `REPRODUCTION/RESULTS.md`
- **Code audited (authors' own, cloned):** `REPRODUCTION/_repos/h3nx-paper` (`github.com/moncla-lab/h3nx-paper`)
- **COI (declared, and applied AGAINST the paper):** Shank, Markin & Anderson are our own group. Every finding below — including the favorable ones — was held to an exact-file / exact-number bar, and every reviewer suspicion that fell in the paper's favor was re-attacked before being retired. "Be more skeptical, not less" was the operating instruction throughout.

---

## 1. Executive overview

**What the paper claims.** From 6,104 subsampled H3Nx genomes, the paper argues that *host type — not the virus — sets the evolutionary strategy*. Birds reassort heavily and show little directional (sweep-like) selection; mammals reassort less but carry detectable HA/NA adaptation; swine uniquely do both. Reassortment is neutral-to-deleterious in birds (47.4% NAm-avian / 29.8% Eurasian reassortant lineages purged <1 yr), marginally beneficial in swine, and enriched on mammal→mammal (chiefly human→swine) host switches (OR≈5.49) but not on avian switches. Two methodological engines carry the paper: a modified McDonald–Kreitman (Bhatt/Kistler) adaptive-rate estimator, and TreeSort run through a novel 1000-replicate uncertainty pipeline that emits ≥95%-support summary trees.

**Does it hold?** **Yes.** This is a strong, unusually self-critical paper whose central thesis survived a hostile reproduction against the authors' own code and outputs. We reproduced its load-bearing numbers directly — the host-switch enrichment (OR=5.4946, p=9.34×10⁻⁸), the purge fractions (47.4 / 29.8 / 34.7%), the swine persistence-FET (6mo OR=2.44, p=1.6×10⁻⁶), the between-subtype NA binomial (p=3.07×10⁻⁷), the clock-rate table, and the reassortment-rate ranking — each to the stated precision from the authors' shipped artifacts. Every reviewer hypothesis that would have undercut a conclusion (reassortant exclusion manufacturing the avian contrast; persistence–rate circularity; clock-scaling flipping the avian>swine ordering; coin-flip branch calls destroying the OR; an estimator artifact masquerading as "no avian selection") was tested and **fell in the paper's favor**. The paper looks *better* after reproduction, not worse.

**Strongest remaining concern (first).** The one substantive gap is **reproducibility, not correctness**: three *load-bearing* per-segment reassortment-enrichment claims (C24–C26: NA/PA over- and NS/MP under-representation in avian; NA/PB1 under-representation in swine, Bonferroni p = 0.007 / 0.02 / 0.03) rest on a Monte-Carlo null whose **generating code is absent from the released repository**, and whose direction is *not* obvious from the raw shipped per-segment rates (in `na_avian`, NS is the highest raw-rate segment yet is claimed *under*-represented — enrichment is relative to an unshipped null expectation). Adjacent to this, two load-bearing per-host reassortment rates (**Eurasian-avian C14 = 0.1248 ± 0.0018; canine C17**) have **no shipped per-replicate log**, so they cannot be verified the way the swine/human/equine/NAm-avian rates were. None of these gaps produces a wrong conclusion in what we could check — but they mean a reader cannot re-derive several headline numbers end-to-end from the released materials.

**Recommendation: Accept with MINOR revision.** No conclusion requires new experiments. The revisions are (a) archive the segment-enrichment null generator and the missing per-replicate logs / null trees; (b) fix one abstract-tier p-value transcription (9×10⁻⁷ → 9×10⁻⁸); (c) reconcile the avian Fig-4 FET numbers with the figure's own source JSON; (d) tighten a handful of Methods sentences (segment-specific reassortant exclusion; the ≥95%-support filter's load-bearing role; the flyway p-value formula); (e) soften "no detectable adaptive signal" for *Eurasian* avian HA, whose own bootstrap CI excludes zero.

---

## 2. Findings by severity

No **CRITICAL** or **MAJOR** findings. All upheld findings are MINOR, NIT, or SOUND (verified-correct). Reproducibility/release gaps that block end-to-end re-derivation of *load-bearing* claims are collected first because they are the most actionable.

### MINOR

**F1 · Segment-enrichment Monte-Carlo null (C24–C26) is untested and its code is absent from the repo.** *(dim: completeness; load-bearing)*
Three load-bearing claims — NAm-avian NA/PA over- & NS/MP under-representation, Eurasian-avian, and swine NA/PB1 under-representation, with Bonferroni p = 0.007 / 0.02 / 0.03 — rest on a per-segment over/under-representation permutation null. A repo-wide search of every `.ipynb`/`.py` for `overrepresent|underrepresent|per-segment|monte-carlo|shuffle|permut|(r+1)/(n+1)` returns **no notebook that implements this test**; the only Bonferroni hits are the *geographic* `flyway_binomial.ipynb` (66 "flyway" references, zero "segment") and the FET/clock notebooks that cover different analyses. Tension worth flagging: the shipped `pipeline_runs/na_avian/log.csv` per-segment mean rates rank **NS highest (0.156)**, NA second, MP lowest — yet C24 calls NS *under*-represented and NA *over*-represented, so the entire direction of C24 is a property of the missing null (not a contradiction — enrichment is relative to expectation — but it is unverifiable from shipped numbers).
**Fix:** release the segment-enrichment null script (or a seed) so C24–C26 are reproducible end-to-end; failing that, report the observed and expected per-segment counts alongside the p-values so the direction can be checked.

**F2 · Two load-bearing per-host reassortment rates have no shipped per-replicate log (Eurasian-avian C14, canine C17).** *(dim: stats/completeness; load-bearing)*
Per-replicate `log.csv` exists only for `na_avian, swine, human, equine`; `pipeline_runs/eurasian_avian/` and `canine/` ship only `summary.json`/`summary.nwk` (confirmed by `find … -name log.csv`). The swine/human/equine rates reproduce to 4 decimals from their logs (F13); the NAm-avian mean is 1.7% off (F7); but **Eurasian-avian 0.1248 ± 0.0018 — which underpins the "415 events > swine 307 yet lower *rate*" clock-scaling argument (C20) and the whole avian-vs-swine ordering M4 defends — cannot be checked against any shipped replicate log.** A single-summary-tree recompute (415 is_reassorted nodes / 6.41 subs tree length ÷ clock 0.001002) gives the *clock-free density* 0.0649, not the 1000-replicate per-year mean 0.1248 — consistent with the rate being a replicate-mean statistic whose log was not shipped, not with an error.
**Fix:** ship `eurasian_avian/log.csv` and `canine/log.csv` (or the aggregation script) so C14/C17 are verifiable like the other four hosts.

**F3 · Headline host-switch p-value mis-stated by one order of magnitude (9×10⁻⁷ vs actual 9×10⁻⁸).** *(dim: stats)*
The load-bearing global host-switch enrichment (C42) is reported at line 438 and in the Fig 5A discussion as "OR=5.49, p-value = 9×10⁻⁷, 95% CI = 3.08, 9.81." The OR and CI are correct, but the authors' own notebook (`within-between_FET.ipynb` cell 9) stores `p = 9.34×10⁻⁸`, and an independent scipy recompute from their own 2×2 (21/776/26/5279) gives p = 9.341×10⁻⁸. This is a single Fisher test — no multiple-testing correction converts 9e-8 to 9e-7 — so it is a pure transcription error in the exponent. Conclusion (highly significant enrichment) is unaffected.
**Fix:** change 9×10⁻⁷ → 9.3×10⁻⁸ in the Results text.

**F4 · Avian Fig-4 FET OR/p values in the text disagree with the figure's own source JSON.** *(dim: stats)*
For the avian persistence-fitness FET (C39/C40), the text (lines ~414–418) gives NAm 10yr OR=0.63/p=0.04, Eurasian 6yr OR=0.52/p=0.008, Eurasian 10yr OR=0.56/p=0.046. The object Fig 4 is plotted from, `FET_per_clade.json` — reproduced exactly by independent scipy recompute from `counts_per_clade.json` — gives OR=0.608/p=0.026, 0.544/**0.017**, 0.547/0.035. These are not rounding artifacts (0.608→0.61 not 0.63; p=0.017 cannot round to 0.008 — a >2× understatement). All qualitative conclusions survive (every cell p<0.05, OR<1). *Note:* the finding's original "swine values match exactly" framing is wrong — the text quotes no swine OR/p numbers at all — so drop that selectivity argument; the numeric text-vs-figure discrepancy itself is real and confirmed.
**Fix:** reconcile the Results text with `FET_per_clade.json`, or state which data version the text numbers come from.

**F5 · "No detectable adaptive signal in avian HA / equivalent to PB1" is contradicted by the authors' own bootstrap for *Eurasian* avian.** *(dim: stats; load-bearing wording)*
Claim C9 holds for NAm avian but not Eurasian. In the authors' own `eurasian_avian/ha_3_3_adaptation_bootstrapped.json`, Eurasian-avian HA rate = 2.26×10⁻⁴ with a percentile 95% CI **[3.78×10⁻⁵, 4.33×10⁻⁴] that excludes zero** (only 1% of bootstraps ≤0), ~9× the Eurasian PB1 rate. Their own plotting notebook (`plot_MK.ipynb`) computes CIs by the identical percentile method and renders them as the Fig-1B error bars — so the CI that excludes zero is the authors' own, plotted in their own figure, while the prose says "no detectable signal." The "distinguishable from zero" leg is firm; the "distinguishable from PB1" leg is only marginal (CIs overlap; paired diff frac≤0 = 0.03). NAm-avian HA (frac≤0 = 0.04) is itself borderline, so the true picture is a gradient, not a clean detect/no-detect split. An independent site-level test (our MEME reproduction, R1c) finds ~0 avian-HA episodic sites, so the *biology* stands; only the one-lineage verbal characterization is imprecise.
**Fix:** soften C9 for Eurasian avian — "a low but nonzero adaptive rate (~9× PB1, bootstrap CI excludes zero)," distinct from the NAm case.

**F6 · Methods overstate the reassortant exclusion as blanket; it was correctly segment-specific and was NOT applied to HA.** *(dim: M1)*
Methods (lines 726–728) read as a blanket per-strain exclusion ("Reassortant strains identified through TreeSort were excluded…"). In fact the HA MK inputs equal the *full* TreeSort tree tip counts exactly (na_avian 1194=1194, eurasian 887=887, human 1078=1078 from `input_data_ha_all_3.json` vs a dendropy leaf count of the reference trees), because HA is the TreeSort reference segment — a strain flagged `is_reassorted=1` has some *other* segment discordant with HA, so its HA is correctly retained. PB1 instead tracks the non-PB1-reassortant set (human PB1 = 1072 exactly). The scheme actually run (per-segment exclusion) is the *methodologically correct* one, and it means the low-avian-HA-adaptation result uses *all* tips, not a reduced non-reassortant sample — strengthening the result. The wording could mislead a reader into thinking otherwise.
**Fix:** state that the exclusion is segment-specific and that HA (the reference segment) retains all strains.

**F7 · NAm-avian reassortment rate (0.3472) does not reproduce from the shipped per-replicate log (0.3413).** *(dim: stats)*
The `simple_rate` column is provably authoritative — swine (0.1303), human (0.0952), equine (0.0096) reproduce their mean *and* SD to 4 decimals — yet `pipeline_runs/na_avian/log.csv` (n=1000) gives mean 0.3413 (SD 0.0129), 1.7% below the reported 0.3472 ± 0.0131, and 0.3472 is not recoverable by any trimming (max replicate 0.4116). Both moments differ slightly, consistent with the reported value coming from a re-run/re-subsampled log not committed. Harmless — NAm avian remains by far the highest host — but the shipped log does not match the paper.
**Fix:** confirm 0.3472 against the exact figure log, or ship the matching log.

**F8 · Rate-induced censoring of the persistence metric is real and measurable, but does not touch a load-bearing claim.** *(dim: M3)*
The persistence metric (Methods L924–926; notebook CELL 8) truncates a reassortant lineage at the *next* downstream reassortment, so a higher-rate host self-censors sooner. Confirmed empirically: 26.1% of NAm-avian reassorted lineages are terminated by a downstream reassortment vs 14.7% in swine (1.8×). Re-doing the metric as a Kaplan–Meier curve with reassortment-terminations *censored* shrinks the naive cross-host "% purged <1yr" gap from 12.7 pts to 6.5 pts — ~half the descriptive turnover gap is censoring-driven. **But** the paper explicitly declines a raw cross-host persistence claim (L367–368, "did not differ significantly between avian and swine, Supp Fig 15") and anchors the fitness inference on within-host observed-vs-null and reassorted-vs-nonreassorted contrasts, which are censoring-matched or censoring-immune (see F9, and checked-and-sound §3). The confound exists but bites only the descriptive turnover numbers.
**Fix:** optional caveat in the persistence paragraph / Supp Fig 15–16 legend that the descriptive cross-host turnover fractions are partly a function of reassortment rate itself.

**F9 (release) · The null-tree generator and the 1000 `randomized_tree_{i}.nwk` files are absent from the repo.** *(dim: M3; reproducibility)*
The fitness null consumes 1000 pre-generated shuffled trees per clade (`fet_fitness.ipynb` CELL 4), but neither the trees nor the branch-reassignment/shuffle generator are present — only the consuming notebooks and the resulting JSONs. `.gitattributes` declares a `randomized_tree/hosts.zip` LFS object that was never committed (dangling entry; not on disk, not in `git ls-files`). Everything verifiable (methods text, output consumption, conserved event counts — null means na_avian 312.7, eurasian 207.3, swine 153.5 all within 1 SD of observed) is internally consistent; this is a completeness gap, not evidence of error. *Housekeeping:* `RESULTS.md` L96 ("R3/M3 skipped — no summary trees") is stale — the trees and null JSONs are present and M3 was fully testable.
**Fix:** ship the null-tree generation script (and/or seed).

### NIT

**F10 · The ≥95%-support threshold is genuinely load-bearing for the OR=5.49 host-switch result — the paper should say so.** *(dim: m1)*
All 21 reassorted host-switch internal nodes have their sibling on the parent's side of the host boundary, so a literal "50/50 coin-flip" resolution of uncertain events would collapse OR from 5.49 to a median 1.83 (only ~49% of reps significant) and mammal-mammal from 18 to a median 9. The result is safe *only because* the ≥95%-support filter removes near-coin-flip events (min surviving support 0.95; zero events in the [0.45,0.55] band; empirical-support perturbation keeps OR=5.49 at 100% significance — see F16). The Methods describe the random per-replicate resolution but never connect it to why the threshold makes host-switch branch assignment safe.
**Fix:** one Methods sentence noting that because only ≥95%-support events are retained, the random resolution of uncertain events does not materially affect host-switch branch assignments; optionally report the 21 events' support values.

**F11 · Per-event reassortment support for the *global* h3nx tree is not archived.** *(dim: m1; reproducibility)*
The host-switch analysis consumes only `host-switching/trees/h3nx/{summary.nwk, traits.json}`. `summary.nwk` carries only binary `[&is_reassorted, rea=…]`; `traits.json` carries host/country/region/subtype/order (+ their confidences) but **no `reassorted_confidence`**. Per-node reassortment support exists only for the per-host trees (`pipeline_runs/*/summary.json`), which cannot be node-mapped to the global tree. So a reader cannot verify how many of the specific 21 host-switch events were near the 95% threshold vs support=1.0 — the F16 robustness argument relies on transferring the per-host support distribution to the global tree.
**Fix:** archive the global per-node reassortment support, or report the support value for each of the 21 (35 incl. leaves) host-switch events in Supp Table 1.

**F12 · Config/format defects in the released MK pipeline (metafile_sep and missing per-host FASTAs).** *(dim: M1; reproducibility)*
`adaptive_evo_config_h3nx.json` declares `metafile_sep="\t"` and a `{virus}_{gene}.tsv` metadata path, but the shipped metadata are comma-separated `.csv`; `rate_of_adaptation.py` line 153 passes the config separator straight to `pd.read_csv`, so as-shipped it hits `FileNotFoundError` on the nonexistent `.tsv` (and would misparse on a `.csv` with a tab sep). The per-host alignment FASTAs the config references (`./mk-test-files/alignments/host/*.fasta`) are absent entirely — only derived intermediate JSONs ship — so the MK pipeline is not runnable end-to-end from the shipped config. No published result is affected (the authors evidently ran an interactive-notebook-derived variant; M1 used independent per-host alignments).
**Fix:** set `metafile_sep=","` (or rename to true `.tsv`) and include the per-host alignment FASTAs or a script to regenerate them.

**F13 · Clock-notebook event counts (647/442) disagree with the Fig-2 legend counts (415/307).** *(dim: stats)*
`clock_rate.ipynb` cell 7's illustrative "tree-length check" hardcodes `rea_events={eurasian_avian:647, swine:442}` (executed output 100.96 vs 37.50), while the Fig-2 legend states 415 vs 307 — which match `is_reassorted=1` counts in `summary_baltic.nwk` exactly and reproduce at the 95% threshold. The 647/442 figures reproduce nowhere in the repo (conf>0 gives 749/539, ≥0.5 gives 558/363, ≥0.95 gives 415/307). The qualitative direction (Eurasian events-per-tree-length > swine) holds under both, so no headline number is wrong; two shipped artifacts simply disagree on the underlying counts.
**Fix:** use the 95%-support 415/307 in the diagnostic cell, or annotate what 647/442 represent.

**F14 · Flyway Monte-Carlo p-value uses r/n, not the Methods-stated (r+1)/(n+1).** *(dim: stats)*
`randomized_tree/flyway_binomial.ipynb` computes `p = sum(d>=obs)/n` with no pseudocount, contradicting the Methods empirical-p formula `p=(r+1)/(n+1)`. Immaterial — every flyway p came out 1.0 after Bonferroni. (Minor: the segment-nonrandomness p-values 0.007/0.020/0.030 are *consistent* with (r+1)/1001 for r=6/19/30, but that code is not in the clone — ties to F1.)
**Fix:** align the flyway notebook with the stated formula, or note in Methods that the flyway test used the uncorrected empirical p.

**F15 (housekeeping, our repo) · `REPRODUCTION/RESULTS.md` R2b table is stale vs the committed autoclock JSONs.** *(dim: M4)*
Several per-community rates in the R2b table were transcribed from an earlier run (e.g. na_avian "1.30e-3, 1.57e-3" vs JSON 1.48e-3 / 7.62e-4; swine 2nd community 5.45e-3 vs 3.61e-3; canine 1.84e-3 vs 2.31e-3; human comm0 also off by ~0.5e-3). All committed values remain in the plausible HA range and none approach the flip threshold, so the M4 conclusion is unaffected — but our own notes should be regenerated from the JSONs before any of these numbers are quoted.
**Fix:** regenerate the R2b table directly from `results/autoclock_ha_*.json`.

*(A non-load-bearing text-vs-artifact mismatch surfaced during coverage: C6 states HA TMRCA 95% CI ≈ 1910-05 → 1917-09, but the committed `auspice/h3nx_ha.json` root gives num_date CI [1908.853, 1916.195] ≈ 1908-11 → 1916-03 — shifted ~1.5 yr later and narrower. Likely the paper's CI comes from a refined/human-inclusive build; worth reconciling in the reproducibility ledger, not a conclusion-level issue.)*

---

## 3. Checked and sound / refuted suspicions (balance)

The COI instruction was to attack the paper harder, so the balance below matters: the majority of load-bearing checks came back **sound**, and every major reviewer suspicion was refuted against the authors' own artifacts.

**Reviewer majors that were tested and fell in the paper's favor:**

- **M1 (reassortant exclusion manufactures the avian contrast) — REFUTED.** Re-running the authors' own Bhatt/MK estimator on the same HA alignments *with* vs *without* reassortants shows the effect on the rate is negligible and runs *opposite* to the confound hypothesis: excluding reassortants slightly *lowers* the avian rate (na_avian 0.052→0.017; eurasian 0.404→0.385, ×10⁻³/codon/yr). Both with-reassortant avian rates stay far below every mammal (human 3.09, swine 1.50, canine 1.11, equine 0.64). The mammal–avian gap is 1–3 units; the exclusion moves avian by ~0.02–0.04. **M1 is not a real confound.** (Optional one-line sensitivity note would preempt the concern.)

- **M2 ("no avian adaptation" is an estimator artifact) — WITHDRAWN.** Our independent site-level `hyphaeon meme` (frame-corrected; an initial contrary result was a reading-frame bug on *our* side, caught by the verification workflow) finds **0 FDR episodic sites in avian HA**, robust to tree-vs-TN93 (R1c, ≤1 site in any host). Agrees with the paper's "little directional selection in birds." (See F5 for the one-lineage bootstrap nuance, which is a wording issue, not a reversal.)

- **M3 (persistence–rate circularity) — DOES NOT BITE THE CONCLUSIONS.** Beyond F8: three rate-aware recomputes all uphold the paper. (1) Observed-vs-null purged-1yr (censoring matched both sides): swine purges far *less* than its null (34.7% vs 46.2%, p=0.003, beneficial); NAm avian ≈ null (47.4% vs 48.0%, p=0.60, neutral). (2) Kaplan–Meier log-rank with reassortment-terminations censored (the exact M3 fix): swine reassortants persist significantly *longer* (χ²=10.67, p=0.0011, 4.22y vs 3.04y); NAm avian no difference (p=0.41). (3) The long-term fit/unfit FET metric is **structurally immune** to M3 — it measures distance to *any* descendant leaf, not truncated at the next reassortment (confirmed in `fet_fitness.ipynb` CELL 5, `max over k.leaves`; 28–31% of reassorted nodes have an intervening reassortment the metric spans). Swine OR≈2.0 p<0.001, avian ns, reproduced from the authors' JSON. *Refinement:* "avian neutral" should read "NAm avian neutral, Eurasian avian deleterious" (Eurasian FET OR<1 significant, KM p=0.0008 rea *shorter*) — which sharpens the swine-vs-avian contrast rather than weakening it.

- **M4 (clock-scaling flips the avian>swine ordering) — REFUTED.** The rate is linear in the reference-HA clock (Methods L775–780; `clock_rate.ipynb` cell 7). Clock-independent incongruence density (events/substitution) already puts avian on top *before any clock*: na_avian 245.3 > eurasian 124.5 > swine 33.2 > human 22.8 > canine 20.5 > equine 5.2 (independently corroborated by a from-scratch tree parse: 161 > 65 > 26 > 25 > 19 > 5). A flip would require the avian HA clock < 5.3×10⁻⁴; every defensible clock set (authors' 2.665; our autoclock dominant-community 4.75; adversarial worst case 1.09) keeps the ordering. The *only* value that flips it is our single-clock pooled estimate 1.4×10⁻⁴ — a documented pathology (tMRCA 1332 CE). Notably the clock works *against* the ordering (avian clock 2.77× *lower* than swine), which survives anyway. Our independent avian HA clock (0.00148, covariation-aware autoclock) agrees with the authors' (0.001415, ~4%); the authors use covariation-aware TreeTime, not the naive RTT that broke on our side; our lower mammalian clocks would *raise* mammalian rates, cutting against the headline, and the ordering still holds.

**Reviewer minor m1 (coin-flip branch calls) — DOES NOT BITE.** The nodes-only Fisher test reproduces exactly (a=21, b=776, c=26, d=5279, OR=5.4946, p=9.34×10⁻⁸; switches {avian→mammal:2, mammal→mammal:18, mammal→avian:1}). Across all per-host surviving events (n=2463) min support = 0.95 and *zero* fall in the coin-flip band; calibrated Monte-Carlo perturbation drawing each event's flip probability from the empirical support keeps OR at 5.49 with 100% of reps significant. The ≥95%-support pipeline is exactly the filter that neutralizes the coin-flip concern (F10 asks only that the paper *say* so).

**Load-bearing numbers reproduced exactly from the authors' own artifacts (SOUND):**

- Purge fractions 47.4 / 29.8 / 34.7% — reproduced from `summary_baltic.nwk` with a from-scratch parser matching baltic's `is_reassorted` handling (node counts 306/198/150).
- Swine/human/equine reassortment rates (0.1303 ± 0.0052, 0.0952 ± 0.0063, 0.0096 ± 0.0001) and **all** swine persistence-FET values (6mo OR=2.438 p=1.6×10⁻⁶; 2yr OR=2.478 p=1.1×10⁻⁷; 6yr OR=1.881 p=0.0024; 10yr OR=1.343 p=0.26) — reproduced to full precision from `counts_per_clade.json`.
- Between-subtype NA binomial (C28–C33): within-probabilities 0.382 / 0.412, counts 35/138 and 39/109, headline p = 3.07×10⁻⁷ (Eurasian 1.19×10⁻⁴), and the underdetection robustness bounds (11.6% / 7.4%) — all match the authors' notebook and independent recompute. (One cosmetic mislabeled code comment; computation correct.)
- Avian-order host-switch FET (OR=0.49/p=0.69; 1.08/0.84), host-switch counts (35/99 = 35%; 21 internal = 18 mammal-mammal + 2 + 1), all six clock rates, and the dataset arithmetic (5023 + 1078 + 3 = 6104) — all verified.
- Dataset-curation audit (our reproduction): 0 wrong-class hosts, 0 duplicate headers, 100% dated 1963–2024, 0 sequences >10% N/gap. Curation is sound.

**Refuted findings NOT carried forward** (checked, and rejected as problems): a claim that the M3(a) critique should be *withdrawn* (over-reached — it conflated the within-host null, which is appropriately censoring-matched, with the cross-host claim the review actually targets; the within-host null is a feature, not the flaw); and one placeholder/stub finding with no substantive content.

---

## 4. Coverage statement

**Reproduced (recomputed or exactly cross-read against the authors' own outputs):** MK adaptive rates with/without reassortants (M1); persistence + purge fractions and the shuffled-tree-null / within-host / long-term-FET fitness analyses (M3); clock rates and clock-independent reassortment-rate ordering (M4); host-switch enrichment OR + mammal-mammal specificity + coin-flip robustness (m1); between-subtype NA binomial; swine/human/equine/NAm-avian per-replicate rates; the global host-switch and avian-order FET tables; dataset arithmetic. Independent tooling: site-level MEME episodic selection (R1/R1c) and autoclock HA clocks (R2b).

**Read but not independently recomputed:** the paper text and PDF figures/tables in full; the authors' notebooks (`within-between_FET*`, `reassortment_persistence-cumulative`, `fet_fitness`, `clock_rate`, `subtype_switching_analyses`, `flyway_binomial`); the config and MK pipeline source; the committed auspice builds (root tMRCA CI checked — F15 note).

**Still open (the actionable revision list):**
1. **Segment-enrichment null (C24–C26)** — untested *and* unreproducible: the Monte-Carlo generator is not in the repo (F1). Highest priority given three load-bearing claims whose direction is not obvious from raw shipped rates. An independent recompute is possible in principle (observed per-segment counts from the summary trees + the stated (r+1)/(n+1) null) if the authors supply the null construction.
2. **Eurasian-avian (C14) and canine (C17) rates** — no shipped per-replicate log (F2); recompute from the aggregation if released.
3. **Null-tree generator + 1000 randomized trees** (F9) and **global per-node reassortment support** (F11) — release for end-to-end reproducibility.
4. Textual corrections: F3 (p-value exponent), F4 (avian Fig-4 FET vs JSON), F5 (Eurasian avian wording), F6/F10/F14 (Methods sentences), F13 (diagnostic-cell counts).

**Bottom line.** The reassortment-rate / adaptation / persistence / host-switch spine of the paper is thoroughly covered and holds up; every load-bearing check we could run reproduced, and every reviewer suspicion that would have overturned a conclusion was refuted against the authors' own code. The residual issues are release/reproducibility gaps and text-vs-artifact corrections. **Accept with minor revision.**
