Evidence-first notes on bioscience and deep tech, at the edge of the lab and the market. Information only — not investment advice, not medical advice.
The 30-second version
- What. In a prospective, multicenter cohort with centrally read liver biopsy as the reference standard (LITMUS Imaging Study, N = 357, Nature Medicine, published online July 24, 2026), imaging-based fibrosis markers (MRE, VCTE-LSM, Agile 3+, Agile 4) statistically met pre-specified minimum acceptable performance criteria for advanced fibrosis and cirrhosis (AUC 0.84–0.91, P < 0.01 to P = 0.03). The best-performing serum marker, NIS2+, “did not significantly exceed the pre-specified minimum acceptable criteria” for MASH (AUC 0.83, P = 0.15) or at-risk MASH (AUC 0.82, P = 0.19) — the study’s own language.
- So what. At-risk MASH is precisely the phenotype (MASH plus stage F2–F3 fibrosis) that current labeling requires for resmetirom and semaglutide. This inverts the usual prediction-to-action gap seen elsewhere in this field, where accuracy is mature but no treatment exists to act on it: here the treatments already exist and are guideline-endorsed, but the accuracy shortfall sits exactly at the gate that defines who qualifies for them.
- Now what. One candidate explanation is that the ceiling on MASH-detection accuracy reflects noise in the biopsy-based reference standard itself (inter-reader disagreement on inflammation and ballooning), not a limit of the biomarkers — but this is a hypothesis, not a finding, because the study does not report inter-reader agreement (kappa). No 95% confidence intervals are reported for any AUC in the abstract, and the cohort is a tertiary hepatology population with 47% advanced fibrosis (F3+F4), so positive predictive value in lower-prevalence primary-care screening populations would be expected to differ substantially even at an identical AUC.
The five-minute read
Two axes, one biopsy reference
All markers in this study were validated against the same reference standard: centrally read liver biopsy. But biopsy grades two different things — fibrosis stage (a structural, scarring axis) and steatohepatitis activity, i.e., ballooning degeneration and lobular inflammation (a cellular, inflammatory axis). The LITMUS Imaging Study is the first prospective, multicenter, pre-registered test of both axes side by side in the same 357 participants (fibrosis stages F0–F4 at 12%, 16%, 25%, 32%, and 15%). The result is a clean dissociation: elastography and composite scores (MRE, VCTE-LSM, Agile 3+, Agile 4) statistically cleared the bar for advanced fibrosis (AUC up to 0.91, P < 0.01) and cirrhosis (AUC 0.87–0.91, all P < 0.01), while the top serum marker, NIS2+, did not statistically clear the bar for MASH or at-risk MASH. A 2023 retrospective LITMUS meta-cohort (Vali et al., Lancet Gastroenterology & Hepatology) had already suggested this pattern; this study confirms it prospectively, which is itself a meaningful evidentiary upgrade even though the underlying scientific conclusion is not new.
The gate that does not clear
Why does the miss matter more than a single AUC comparison suggests? Because “at-risk MASH” is not just a research label — it is the phenotype that current U.S. labeling for resmetirom (accelerated approval, March 2024) and semaglutide 2.4 mg (MASH indication, August 2025) requires: MASH confirmed plus stage F2–F3 fibrosis (secondary-source attribution, not independently cross-checked against the primary approval documents in this analysis). Fibrosis staging is already well served by elastography in clinical guidelines. The unmet target is the one that gates access to the two newest approved drugs for this disease. In effect, biopsy or an unvalidated surrogate remains the practical path to determining treatment eligibility for the inflammatory component of MASLD.

What this study does not settle
Three limits should travel with every number above. First, the abstract reports no 95% confidence intervals for any AUC — in a design where AUC 0.83 (P=0.15) is read as a miss and AUC 0.84 (P=0.03) is read as a pass, interval estimates matter and are absent. Second, whether the lower MASH-detection accuracy reflects a limit of the biomarkers or noise in the biopsy-based reference standard (inter-reader disagreement on ballooning and inflammation) is not determined by this paper, because inter-reader agreement (kappa) is not reported; the “reference-standard noise” explanation below is presented as a hypothesis, not a conclusion. Third, this cohort is a tertiary-referral biopsy population with 47% advanced fibrosis, not a primary-care screening population where advanced fibrosis prevalence is typically in the low single digits — positive predictive value would differ substantially even at an identical AUC.
Deep dive
1. Background
Non-invasive tests (NITs) for MASLD (metabolic dysfunction-associated steatotic liver disease) have been converging on a two-step algorithm in guidelines such as EASL-EASD-EASO (2024, secondary summary): a low-cost risk score (FIB-4) to triage, followed by elastography to stage fibrosis. The unresolved question has been whether NITs can also identify steatohepatitis (MASH) and, more specifically, “at-risk MASH” — MASH plus clinically significant fibrosis — without biopsy. A 2023 retrospective LITMUS meta-cohort study (Vali et al., Lancet Gastroenterology & Hepatology) found that no single marker or composite score significantly exceeded a pre-specified AUROC of 0.80 for at-risk MASH. The LITMUS Imaging Study (ClinicalTrials.gov NCT05479721, sponsor University of Oxford, prospective observational cohort, enrollment starting 2019-09-04) was designed to test the same question prospectively, in a multicenter cohort with centrally read biopsy, against pre-specified minimum acceptable performance criteria (MAC) set before results were seen.
2. What this study newly shows
Among 357 participants (fibrosis stages F0–F4 at 12%, 16%, 25%, 32%, 15%), the following AUCs and significance tests against the pre-specified MAC are reported in the abstract:
| Target | Best marker | AUC | P vs. MAC | Verdict (abstract’s own language) |
|---|---|---|---|---|
| MASH | NIS2+ (serum) | 0.83 | 0.15 | Did not significantly exceed MAC |
| At-risk MASH (MASH + ≥F2) | NIS2+ (serum) | 0.82 | 0.19 | Did not significantly exceed MAC |
| MASH | MRE (imaging, best) | 0.71 | not stated | Lower than the serum marker; not separately tested against MAC in the abstract |
| At-risk MASH | MRE (imaging, best) | 0.75 | not stated | Lower than the serum marker; not separately tested against MAC in the abstract |
| Advanced fibrosis | MRE | 0.91 | <0.01 | Met MAC |
| Advanced fibrosis | Agile 3+ | 0.84 | 0.03 | Met MAC |
| Cirrhosis | MRE | 0.91 | <0.01 | Met MAC |
| Cirrhosis | VCTE-LSM | 0.87 | <0.01 | Met MAC |
| Cirrhosis | Agile 3+ | 0.89 | <0.01 | Met MAC |
| Cirrhosis | Agile 4 | 0.88 | <0.01 | Met MAC |
The pattern the paper itself draws out (paraphrased, not a direct quote): serum biomarkers tend to do relatively better at identifying at-risk MASH than imaging does, while elastography and composite scores perform well for staging advanced fibrosis and cirrhosis — but “relatively better” for serum markers still fell short of the pre-specified statistical bar. This is a regulatory-grade confirmation of a pattern the 2023 retrospective meta-cohort had already flagged; the novelty is in the strength of the design (prospective, multicenter, pre-registered, pre-specified MAC), not in a new biological finding.
3. Strengths and limits of the methodology
The design strengths are real: prospective enrollment, multiple centers, centrally read biopsy, and — critically — a MAC that was pre-specified before results were seen, which removes the researcher-degrees-of-freedom problem common to retrospective diagnostic-accuracy studies. Three limits, required reading alongside every number above:
- No 95% confidence intervals are reported for any AUC in the abstract. In a structure where AUC 0.83 (P=0.15) is read as a miss and AUC 0.84 (P=0.03) is read as a pass, the absence of interval estimates is a material gap, not a footnote.
- Whether the lower MASH-detection accuracy (0.71–0.83) relative to fibrosis-staging accuracy (0.84–0.91) reflects a true limit of the biomarkers, or instead reflects noise in the biopsy-based reference standard itself — ballooning degeneration and lobular inflammation are known to have lower inter-reader agreement than fibrosis staging — cannot be determined from this paper, because inter-reader agreement (kappa) for this cohort is not reported. This reattribution (marker vs. label noise) is a hypothesis this analysis finds worth tracking, not a conclusion this study established.
- The cohort is a tertiary hepatology referral population with 47% advanced fibrosis (F3+F4) and only 12% F0. AUC is prevalence-invariant, but positive and negative predictive value are not: applying these AUCs to a primary-care or diabetes/obesity screening population, where advanced fibrosis prevalence is typically single digits, would be expected to yield materially different predictive values even though the discrimination (AUC) is unchanged. Multiple-comparison correction across the markers and four targets tested is not confirmed in the abstract, which should temper confidence in the single P=0.03 result (Agile 3+, advanced fibrosis) in particular.
4. Connections to neighboring domains
AI and digital pathology. If the ceiling on MASH-detection AUC is set by reference-standard noise rather than by the biomarkers, then the more promising lever may not be a better serum or imaging marker but a more reproducible label — for example, machine-learning-based digital pathology scoring that turns ballooning/inflammation grading into a continuous, reproducible score. That would be a case of AI functioning as a label generator rather than a predictor, a role distinct from most AI-in-diagnostics examples this outlet has covered to date. This connection is plausible and structurally motivated but has zero direct support within this paper (kappa is not reported); it should be treated as a hypothesis to watch, not a result.
Regulatory and reimbursement pathways. In the United States, the biomarker regulatory-qualification pathway (e.g., FDA surrogate-endpoint qualification) and the laboratory-developed test (LDT) reimbursement pathway are, from the outset, separate gates; clearing one does not imply clearing the other. This case is one instance where that separation is visible for the same underlying test category — details in section 5 below.
Imaging throughput. MRE is the single best-performing marker across every fibrosis-related target in this study (AUC 0.91), yet it is not the first-line tool in most clinical algorithms. The constraint is not diagnostic accuracy but MRI scanner slot availability and cost — a systems-engineering and throughput problem, not a biology problem.
5. Commercialization and investment angle
TRL by test (two distinct gates — regulatory qualification and LDT/reimbursement — are listed as parallel facts, not ranked against each other):
| Marker | TRL | Basis |
|---|---|---|
| VCTE-LSM (FibroScan) | 9 | FDA clearance, multi-guideline inclusion, reimbursement, routine clinical use; this study adds prospective supporting evidence for cirrhosis. |
| Agile 3+ / Agile 4 | 8 | Software add-on to commercial FibroScan hardware, guideline-referenced, met MAC for cirrhosis (both) and advanced fibrosis (Agile 3+) in this study. |
| MRE | 8 | Integrated into commercial MRI platforms, clinical use; best AUC in this study for both fibrosis targets; throughput/cost, not accuracy, is the adoption constraint. |
| NIS2+ / NASHnext | 7–8 | Commercial laboratory-developed test (LDT) available, with a reimbursement pathway reported separately (see below); in this study, the pre-specified statistical criterion for MASH and at-risk MASH was not met. Regulatory qualification and LDT/reimbursement are two different gates. |
| cT1 (Perspectum) | not assessable here | Result not reported in the abstract; unverified in this analysis. |
| Core thesis: “NITs can substitute for biopsy to select MASH treatment candidates” | 7 | Reached prospective, multicenter, centrally-read system-level validation (meets the definition of TRL 7) but did not cross into TRL 8 for this specific indication, because the pre-specified performance criterion was not met for the target that defines treatment eligibility. |
The serum-marker and reimbursement case. NIS2+ (developer: Genfit) showed the highest accuracy among serum biomarkers evaluated in this study for both MASH (AUC 0.83) and at-risk MASH (AUC 0.82), but — in the study’s own language — “did not significantly exceed the pre-specified minimum acceptable criteria” (P = 0.15 and P = 0.19, respectively). This is the abstract’s own statistical verdict, not a product evaluation by this outlet.
In the United States, the biomarker regulatory-qualification pathway and the laboratory-developed test (LDT) reimbursement pathway are, from the outset, separate gates that do not imply one another. This case makes that separation concretely visible: the consortium’s own pre-specified performance criterion was not met for this marker in this paper, while — per the test developer’s July 2, 2026 press release (secondary source, not independently verified in this analysis) — Medicare coverage for NASHnext (the commercial laboratory-developed test built on this marker class) was reported to take effect around August 10, 2026. Separately, the FDA reportedly accepted a Letter of Intent in September 2025 to consider VCTE-LSM as a reasonably likely surrogate endpoint in MASH drug development (secondary trade-press report, not independently verified in this analysis) — a qualification-track development that, notably, concerns a different marker (elastography, not the serum marker discussed above).
More generally: when a particular measurement platform becomes qualified as a regulatory surrogate endpoint for a disease area, clinical trial design in that area can converge on a single measurement approach. This is a structural feature of surrogate-endpoint regulatory pathways in general, not a characteristic unique to any single company or device.
Fact list (relationships stated as facts, not evaluations):
- NASHnext is a laboratory-developed test offered by Labcorp under license of Genfit’s NIS4 technology.
- VCTE-LSM, Agile 3+, and Agile 4 are products of Echosens (privately held).
- The MRE hardware driver is a product of Resoundant Inc. (privately held, a Mayo Clinic spinoff); MRE is integrated as an option on commercial MRI platforms from multiple manufacturers.
- cT1/LiverMultiScan is a product of Perspectum Ltd. (privately held); this study’s abstract does not report a cT1 result.
- On the treatment-demand side: resmetirom (Rezdiffra) received U.S. accelerated approval in March 2024 (secondary source); semaglutide 2.4 mg (Wegovy) received a MASH indication in August 2025 (secondary source), with efficacy figures from the ESSENCE trial attributed to secondary reporting and not cross-checked against the primary trial publication in this analysis.
- Several industry partners in the LITMUS consortium hold MASH drug pipelines, per the ClinicalTrials.gov collaborator list for NCT05479721: Pfizer, Novartis, Takeda, Intercept Pharmaceuticals, and Boehringer Ingelheim. For these companies, the practical value of this study is less about diagnostics per se and more about patient enrichment and surrogate endpoints for their own trials.
Investment relevance here is indirect. This single study does not resolve which measurement approach ultimately anchors MASH treatment-eligibility determination in practice.
6. The counterargument (skeptic block)
The firm’s skeptic gate returned proceed-with-caveats (conditional): the numeric layer is verified-clean (all ten AUC/P-value pairs cross-checked against the abstract, zero discrepancies), but the interpretive layer is conditional. The following caveats are carried into this article verbatim in substance, as required by that gate:
- No 95% confidence intervals are reported anywhere in the abstract. In a structure where AUC 0.83 is read as a miss and AUC 0.84 is read as a pass, reading “met/did not meet” as a clean binary without interval estimates is a real risk of over-reading the data.
- Whether the lower MASH-detection performance relative to fibrosis-staging performance reflects a limit of the markers or noise in the biopsy-based reference standard cannot be determined by this paper — inter-reader agreement (kappa) for this cohort is not reported. The “reference-standard noise” explanation in this article is a hypothesis, not an established conclusion.
- These figures come from a tertiary hepatology referral cohort in which advanced fibrosis accounts for 47% of participants; in a primary-care population with single-digit advanced-fibrosis prevalence, positive predictive value would be expected to be substantially lower even at an identical AUC.
One additional note from the same review: multiple-comparison correction across markers and targets is not confirmed in the abstract, which should temper confidence in isolated borderline results (e.g., Agile 3+ for advanced fibrosis, P = 0.03).
7. Metrics to watch going forward
- Inter-reader agreement (kappa) for MASH/ballooning/inflammation grading in this cohort, if the full text becomes accessible — this single value would resolve whether the “reference-standard noise” hypothesis in sections 4 and 6 is supported.
- 95% confidence intervals for the reported AUCs, and the paper’s operational definition of the pre-specified MAC, from the full text.
- Whether the reported NASHnext Medicare coverage carries conditions such as coverage-with-evidence-development, once independently confirmed beyond the developer’s press release.
- Progress of the FDA’s reported surrogate-endpoint qualification review for VCTE-LSM, if independently confirmed.
- Whether any prospective trial registers patients using NIT-defined at-risk MASH status versus standard-of-care biopsy-defined status, which would be the first direct test of whether the accuracy gap identified here changes clinical outcomes.
- cT1 (Perspectum) performance, not reported in the abstract, if disclosed in the full text or a follow-up publication.
References
- LITMUS Imaging Study Consortium. 2026. “Prospective validation of imaging and serum diagnostic biomarkers of steatohepatitis and fibrosis in MASLD: the LITMUS Imaging Study.” Nature Medicine, published online July 24, 2026. https://doi.org/10.1038/s41591-026-04496-2 (Full text paywalled; all figures used here are from the abstract.)
- Vali, et al. 2023. “Biomarkers for staging fibrosis and non-alcoholic steatohepatitis in non-alcoholic fatty liver disease (LITMUS) individual participant data meta-analysis.” The Lancet Gastroenterology & Hepatology. (DOI/URL not confirmed in the source analysis; cited for its finding that no marker significantly exceeded a pre-specified AUROC of 0.80 for at-risk MASH in a retrospective meta-cohort.)
- LITMUS Imaging Study protocol. Study protocol paper, PubMed ID 37802221. https://pubmed.ncbi.nlm.nih.gov/37802221/
- ClinicalTrials.gov. “NCT05479721.” U.S. National Library of Medicine, National Institutes of Health. https://clinicaltrials.gov/study/NCT05479721 (Direct registry query; sponsor University of Oxford, prospective observational cohort, collaborator list includes Perspectum, Antaros Medical, Resoundant, Pfizer, Novartis, Takeda, Intercept Pharmaceuticals, and Boehringer Ingelheim.)
- Innovative Health Initiative (IHI, formerly IMI2). “LITMUS Project Factsheet.” https://www.ihi.europa.eu/projects-results/project-factsheets/litmus (Direct query; EU public funding approximately €15,797,881, EFPIA industry contribution approximately €25,427,538, coordinator Newcastle University, lead EFPIA partner Pfizer Ltd.)
Secondary-source claims cited above in-text (the developer’s July 2, 2026 press release on Medicare coverage; the reported September 2025 FDA Letter of Intent acceptance for VCTE-LSM; EASL-EASD-EASO 2024 algorithm summary; resmetirom and semaglutide approval dates and ESSENCE trial figures) are attributed as secondary and not independently verified in this analysis. Specific outlet names and URLs for these were not confirmed in the underlying analysis file and are therefore omitted here rather than approximated, to avoid fabricating citations. One additional industry-trade figure identified during source review was found inconsistent with the abstract’s own MRE figures and is deliberately excluded from this article.
Disclosure
This post is for information only and is not investment advice, and not medical advice. Treatment decisions should always be made with your own clinician.
COI note: The LITMUS Imaging Study was funded through the Innovative Health Initiative (IHI, formerly IMI2), a European public-private partnership. Public (EU) funding was approximately €15.8 million; industry (EFPIA) contribution was approximately €25.4 million — industry contribution was roughly 1.6 times the public contribution (source: IHI project factsheet). Manufacturers of several of the evaluated tests — Echosens (VCTE-LSM, Agile 3+/4), Genfit (NIS2+), Resoundant (MRE hardware), and Perspectum (cT1) — are consortium partners, as are several pharmaceutical companies with MASH drug programs (per the NCT05479721 collaborator list). Despite this structure, the result that the consortium’s own commercial serum test did not significantly exceed the pre-specified minimum acceptable criteria was reported as-is: pre-registration (NCT05479721) and pre-specified performance criteria set before results were seen removed the degrees of freedom that would otherwise allow selective reporting after the fact. Individual author-level conflict-of-interest disclosures were not accessible (the full text is paywalled) and are noted as unverified. Companies and tests named above are described factually and are not a solicitation to buy or sell any security. The author holds no position in, and has no financial interest in, the companies named.
Leave a comment