HealthTechCrunch

Continuous Glucose Monitor Accuracy Compared Across Devices

Staff Writer · · 11 min read
Cover illustration for “Continuous Glucose Monitor Accuracy Compared Across Devices”
Therapeutics & Wearables · August 9, 2026 · 11 min read · 2,426 words

Here is something every CGM on the market shares, and it matters before any accuracy number makes sense: none of them actually measure blood glucose. Every current CGM measures glucose in interstitial fluid, the fluid bathing the cells between your capillaries and your tissues. That compartment is never in real-time equilibrium with your blood. Think of it like a river and a lake connected by a narrow channel — the lake's water level reflects the river's, but always a few minutes behind, and more so when the river is flooding.

The lag is clinically consequential. CGM readings trail capillary blood glucose by five to fifteen minutes, and that gap widens during rapid shifts: a post-meal spike, an intense workout, the rebound from a hypoglycemia treatment. A 2013 study in Diabetes isolated the underlying transport delay in fasted adults and found a mean lag of 5.3 to 6.2 minutes just for glucose to cross from the vascular space into the interstitial compartment, before any sensor noise or signal processing enters the picture. That range is a hard biological floor. No engineering improvement eliminates it.

This creates a regulatory awkwardness worth holding onto. Accuracy evaluations compare CGM interstitial readings against venous blood samples drawn at the same moment. The lag makes any CGM look less accurate than it actually is relative to true interstitial glucose, because the two compartments are sampling different things at any given instant. The device is being graded against a reference it is physiologically incapable of perfectly matching. Every MARD figure you encounter downstream carries that embedded distortion.

In your day-to-day use, the directional arrow deserves as much attention as the number itself. A reading of 85 mg/dL with a steep downward arrow is a fundamentally different clinical situation than 85 mg/dL with a flat line. During rapid changes, the number misleads; the trend is what you act on.

Study design compounds this further. The timing of calibration and comparator measurements in clinical trials can mask or amplify the lag effect. A trial that staggers reference blood draws slightly relative to CGM readings, intentionally or not, produces a different MARD than one with tightly synchronized measurements. Manufacturer-reported MARDs and independently generated MARDs are not always capturing the same phenomenon, even when the methodology looks identical on paper.

How the major prescription CGMs compare on MARD

The figures here come from manufacturer pivotal trials and FDA clearance data. Independent head-to-head results arrive in the next section, and the divergence is not trivial.

Abbott FreeStyle Libre 3 reported a MARD of 7.9% in its clearance studies, the lowest of any currently available prescription sensor in that context. It delivers readings every minute rather than the five-minute intervals most competing devices use. Abbott has been transitioning users to the Libre 3 Plus through 2025, extending wear to 15 days. One real-world caveat worth knowing: Abbott issued a Medical Device Correction for select Libre 3 and Libre 3 Plus batches in late 2025 due to a manufacturing issue. No headline MARD captures batch-level quality variability, and that number shouldn't make you forget it.

Dexcom G7 (10-day) reported a MARD of 8.2% on the upper arm and 9.1% on the abdomen in its clearance trial. Placement matters measurably, and that spread is worth factoring into practical decisions. It is FDA-cleared for children as young as age two.

Dexcom G7 15-Day, cleared by the FDA in April 2025 and launched in the U.S. in December 2025, reported a MARD of 8.0% in its pivotal study. That trial enrolled 130 adults with diabetes across six U.S. sites; more than 80% of individual sensors achieved a MARD below 10%. Accuracy held consistent regardless of age, sex, diabetes type, HbA1c, or BMI within that cohort. The 15-day version carries an adults-only indication.

Medtronic Guardian 4 reported a MARD of 10.78% on the abdomen and 10.64% on the arm. Those figures sit at or just above the accepted clinical threshold for non-adjunctive use, depending on insertion site. The Guardian 4 functions within Medtronic's closed-loop system ecosystem, a context that meaningfully changes how its accuracy profile should be interpreted. A number that looks weak in isolation can look quite different when the system around the sensor is designed to compensate.

Medtronic Simplera showed a reported MARD range of 8.2% to above 10% across studies. A range that wide signals meaningful variability in how consistently the device performs across conditions. It pairs with Medtronic's Smart MDI system and delivers readings every five minutes.

Senseonics Eversense 365 occupies an entirely different category. It is subcutaneously implanted, FDA-cleared in September 2024, and designed to last up to one year. The ENHANCE study, published in Diabetes Technology and Therapeutics in May 2025, reported an overall MARD of 8.8% through 365 days across 110 participants using primarily one calibration per week. Confirmed alert detection was 96.6% at the 70 mg/dL hypoglycemia threshold and 97.9% at 180 mg/dL. Ninety percent of sensors survived the full 365-day wear period. Insertion requires a minor outpatient procedure, so it is not interchangeable with a peel-and-stick patch. If you want to eliminate the cognitive load of frequent sensor changes, an 8.8% MARD sustained across a full year is a serious accuracy profile, not a consolation prize for the implant inconvenience.

What independent head-to-head studies reveal that manufacturer trials don't

Diagram: Manufacturer vs. Independent MARD: The Gap That Changes Everything. Visualizes: Show the contrast between manufacturer-reported MARD figures and independently measured MARD figures for three CGM devices side by side.

Every device's MARD rises when tested outside a manufacturer-sponsored protocol. The relevant question is how much it rises and what the clinical implications actually are.

The Hanson et al. 2024 study compared the Dexcom G7 against the FreeStyle Libre 3 in a single-arm, prospective, multicenter trial involving 55 adult participants with type 1 or type 2 diabetes. The publication itself acknowledges that study design, data analysis, and editorial support involved a manufacturer of one of the products studied, which limits how far it can function as an independent reference.

The Freckmann et al. 2025 study is the most important independent three-way comparison currently available, and the results are worth sitting with. Twenty-four adults with type 1 diabetes wore the FreeStyle Libre 3, Dexcom G7, and Medtronic Simplera simultaneously for up to 15 days. Three seven-hour in-clinic sessions used YSI reference measurements, with deliberate glucose excursions into hyperglycemia and hypoglycemia induced. The independent MARDs: FreeStyle Libre 3 at 11.6%, Dexcom G7 at 12.0%, Medtronic Simplera at 11.6%. All three sit substantially higher than the manufacturer-reported figures for the same devices. The authors explicitly flag that all three devices underperformed previous independent studies, a finding the manufacturers have not adequately addressed publicly.

The gap between what manufacturers report and what independent researchers find is not a fluke. It is a consistent pattern, and if you are using a CGM to make treatment decisions without knowing this, you are working with an incomplete picture. You could say the manufacturer number is the highlight reel, and the independent study is the game tape.

The FreeStyle Libre 3 and Dexcom G7 outperformed Simplera across most comparators, though the margin was not dramatic.

A companion paper in Diabetes Care (Freckmann et al., July 2025), involving 23 participants, examined CGM-derived glycemic metrics specifically: time in range and time in tight range. The G7 and Libre 3 produced essentially identical results for both metrics. Simplera diverged meaningfully. Commentary by Roy Beck in Diabetes Care confirmed the equivalence between G7 and Libre 3 on those outcomes and corroborated the Simplera divergence.

For clinicians and users who rely on TIR and TITR to guide therapy adjustments, the choice between G7 and Libre 3 is largely a matter of preference and ecosystem fit. The choice to use or avoid Simplera, at least in the mid-to-high glucose range, is a clinical decision with documented consequences.

Where each device performs better or worse within the glucose range

Overall MARD averages performance across the entire glucose spectrum. That is useful as far as it goes, but it can mislead you if you stop there. The clinical situations where accuracy matters most, hypoglycemia and post-meal peaks, are exactly where average figures obscure the most important differences.

Normal to high glucose ranges

The Freckmann 2025 data show that G7 and Libre 3 track post-meal hyperglycemia more closely than Simplera across mid-to-high glucose zones. If you are managing postprandial excursions or titrating basal-bolus insulin, both generally outperform Simplera in the range where most of those decisions get made.

Hypoglycemia detection

This is where Simplera changes the conversation. In the Freckmann 2025 head-to-head, Simplera detected 93% of low glucose events, outperforming both G7 and Libre 3 in that specific zone. That is a meaningful finding for any user with hypoglycemia unawareness. The tradeoff is that Simplera reads consistently lower across the mid-to-high range, driving a higher rate of unnecessary low alerts. Alert fatigue erodes the urgency that makes alarms useful, and that erosion compounds over time. If your primary clinical risk is severe hypoglycemia, Simplera's sensitivity can be a legitimate counterweight to its weaker overall MARD. The hierarchy here is not obvious; it depends entirely on the patient in front of you.

Warm-up and the first 12 hours of wear

Simplera's MARD in the first 12 hours of wear was approximately 20.0% in the Freckmann study. That is a qualitatively different accuracy category, not a rounding error. G7's first-12-hour MARD was around 12.8%. Libre 3 showed particularly strong performance after that initial period. The practical consequence is that sensor change timing matters far more for Simplera users than for users of the other two. If you change a Simplera before a night planned for dose-adjustment decisions, or before a meal you are tracking postprandially, you are introducing a window of substantially degraded reliability that you may not have been warned about at the point of prescription. That is a gap in the informed consent conversation, and it is worth naming.

During exercise

Exercise degrades CGM accuracy across all devices, and the drop is large enough to change how readings should be used during physical activity. A 2026 prospective study covering both laboratory and real-world conditions found an overall mean MARD of 13.63% during exercise across devices, with a 95% confidence interval of 11.41% to 15.84%. That accuracy drop is universal, not brand-specific. The interstitial lag widens during the rapid glucose shifts exercise induces, and no sensor architecture currently available closes that gap. If you are managing athletic performance or exercise-induced hypoglycemia, you should treat CGM readings during activity as directional guidance, not precise values. The arrow and the trend matter more during a hard interval session than the number on the screen.

How CGM accuracy changes in hospital and critically ill settings

Diagram: Accuracy by Condition: How MARD Shifts Across Contexts. Visualizes: Show how a CGM's effective MARD escalates as clinical conditions move further from the controlled outpatient ideal.

CGM use in hospitals is expanding rapidly. The accuracy degradation in critically ill patients is severe enough to change how the technology should be interpreted in that environment entirely, and the clinical teams deploying it should know the numbers before relying on them.

A nonrandomized prospective trial studying the Dexcom G6 Pro and FreeStyle Libre Pro in critically ill patients requiring continuous intravenous insulin infusion found G6 Pro MARD of 22.7% and FreeStyle Libre Pro MARD of 25.2%. Both are far above the sub-10% clinical threshold that defines outpatient non-adjunctive use. That is not a marginal overshoot. That is a fundamental change in what the device can reliably communicate.

A separate observational study from St. Vincent's Hospital Melbourne, involving 103 patients and published in October 2025, reported a MARD of 11.23%, considerably better than the IV insulin trial. The contrast is instructive: ICU context is not monolithic. Severity of illness, vasopressor use, and the presence of edema likely explain much of the variation between studies, and the MARD you get in a general ward patient is not the MARD you get in someone on norepinephrine.

The mechanisms are well understood even if they remain unsolvable at the sensor level. Three mechanisms are at work: (i) poor peripheral perfusion slows interstitial glucose equilibration, (ii) vasopressors and corticosteroids alter glucose dynamics in ways that can outpace the sensor's calibration assumptions, and (iii) edema disrupts the physical contact between the sensor filament and the tissue it is sampling. None of these are engineering problems with near-term solutions.

CGM can still serve a meaningful surveillance role in hospital settings, particularly for detecting trends and reducing fingerstick burden in non-critical patients. But the MARD values that define outpatient accuracy do not transfer to the ICU. Confirmatory blood glucose checks remain essential, and if you are using CGM readings for insulin dosing in critically ill patients, you should treat the interstitial number as one input among several, not as the ground truth.

Reading CGM accuracy claims without being misled by them

Before any MARD figure earns your trust, three questions need answers, and skipping any of them leaves a meaningful blind spot.

First: was the study manufacturer-sponsored or independent? The Freckmann 2025 data make the gap explicit. Independent MARDs for all three devices studied were substantially higher than manufacturer-reported figures for the same products. That gap is not fraud; it reflects the difference between controlled pivotal trial conditions and the messier reality of head-to-head testing with deliberate glucose excursions. But you should assume it exists for every device on the market until independent data say otherwise.

Second: what glucose range and clinical context were studied? A strong overall MARD can coexist with weak hypoglycemia detection. Strong hypoglycemia sensitivity can coexist with mid-range divergence that skews every TIR calculation. Accuracy is not a single number; it is a profile, and that profile needs to match the actual use case.

Third: what was the comparator? YSI reference blood glucose, venous laboratory draws, and capillary fingerstick testing each produce different baselines. Because of the interstitial lag, all CGMs will appear less accurate against venous blood than they are relative to true interstitial glucose. Two studies of the same device using different comparators can produce legitimately different MARDs without either being wrong.

Here is where the evidence lands. For the lowest overall MARD in clearance studies, Libre 3 at 7.9% and Dexcom G7 15-Day at 8.0% are near-equivalent, and independent trials confirm they produce essentially identical TIR and TITR outcomes. If your primary concern is hypoglycemia unawareness, Simplera's superior low-glucose detection can be a real advantage, contingent on your tolerance for mid-range divergence and the first-12-hour reliability gap. If you want to eliminate the recurring burden of sensor changes, the Eversense 365's year-long implant with an 8.8% MARD through its full lifespan is a genuinely different value proposition. None of these is the universally correct answer. The right choice for you is the one whose accuracy profile maps onto the glucose range and clinical stakes that actually matter in your specific situation.

Sources

  1. hcplive.com

More in Therapeutics & Wearables