Research

AI-generated text

GlucoFM: a dual‑stream foundation model that separates slow and fast CGM dynamics

GlucoFM is a self‑supervised foundation model for continuous glucose monitor (CGM) data that explicitly separates slower baseline glycemic trends from short‑term deviations.

GlucoFM: a dual‑stream foundation model that separates slow and fast CGM dynamics

Consumer wearables offer motion and physiological signals useful for estimating activity and sleep, but they provide only an indirect view of glucose regulation. Continuous glucose monitors (CGMs) complement these by sampling interstitial glucose every few minutes via a small subcutaneous sensor, capturing fasting, overnight and post‑meal patterns. Interpreting these traces remains difficult, especially because high‑quality clinical labels are sparse and expensive to collect.

Why GlucoFM?

Many existing CGM foundation models — including CGMformer, GluFormer, and CGM‑JEPA — process glucose through a single representation stream instead of explicitly separating slow baseline trends from transient events. CGM data are not homogeneous: they show relatively slow baseline patterns punctuated by short deviations that may reflect meals, activity, or sensor artifacts. This motivated the design of GlucoFM, a self‑supervised foundation model with a dual‑stream architecture that separates lower‑frequency glycemic state from short‑term residual events while preserving time‑of‑day information and a mask for missing observations. Latent‑prediction objectives teach the model daily context and temporal evolution.

Pre‑training: scale and handling real CGM issues

GlucoFM was pre‑trained on 109,066 hours of unlabeled CGM data drawn from Wear‑CGM and four published datasets, totaling 477 participant/session records. CGM recordings commonly contain gaps, varying sampling intervals, and sensor artifacts. The model aligns each recording to a 24‑hour, five‑minute grid and keeps an observation mask to distinguish measured from unobserved positions.

The dual‑stream encoder learns a lower‑frequency state component (slower glycemic trends) and a residual event component (short‑term deviations from physiology, behavior, or sensing). Rather than reconstructing noisy raw glucose values, GlucoFM uses two complementary latent predictive pre‑training tasks:

  • Contextual prediction: mask parts of a daily glucose sequence and predict their latent representations from surrounding context. Predicting in latent space captures broader daily patterns without forcing reconstruction of every sensor reading.
  • Temporal dynamics: predict how a person’s steady baseline and short‑term deviations shift from hour to hour, encouraging capture of continuous glucose dynamics instead of isolated snapshots.

CGM‑aware augmentations (baseline drift, compression‑like drops, sparser sampling, short disconnections) expose the model to the kinds of variation and missingness encountered in real CGM recordings.

Practical capabilities and evaluation

GlucoFM was evaluated across four cohorts (CGMacros, Stanford, Hall and ShanghaiT2DM) and seven clinical prediction tasks — diabetes risk, insulin resistance, beta‑cell dysfunction, hyperlipidemia, hypoglycemia, obesity, and glucotype — for a total of 14 cohort–task evaluations. Separately, the model was assessed on two‑hour postprandial glycemic response (PPGR) forecasting.

The experiments tested whether frozen day‑level representations are informative for unseen participants, whether they add historical context for postprandial trajectory prediction, whether combining multiple days improves subject‑level prediction, how well representations transfer across cohorts, and how effectively they adapt when labeled data are scarce.

Accuracy across metabolic tasks

Using subject‑disjoint window‑level linear probing (freeze encoder, train a linear classifier on 24‑hour representations with no participant overlap between train and test), GlucoFM achieved the highest task‑averaged PR‑AUC among evaluated methods. Across the 14 cohort–task evaluations, average PR‑AUC rose from 54.7 for the strongest CGM‑specific baseline retrained on the same data to 58.8 with GlucoFM — an absolute gain of 4.1 percentage points (~7.5% relative). When compared to the best‑performing GluFormer variant pre‑trained on the same corpus, GlucoFM’s PR‑AUC was on average 5.8 percentage points higher.

GlucoFM attained the top PR‑AUC in all diabetes‑risk and beta‑cell‑dysfunction evaluations and in three of four insulin‑resistance evaluations.

Predicting postprandial glycemic responses

For PPGR forecasting, 874 paired meal events from 34 participants were evaluated using subject‑disjoint cross‑validation. Dexcom and Libre devices were modeled separately under identical splits. Models received progressively richer inputs: a frozen model representation combined with one hour of pre‑meal CGM, meal nutrition (energy, carbohydrate, fat, protein, fiber), and participant‑level data (fasting glucose, BMI, diabetes status).

With full context, GlucoFM achieved the lowest mean absolute error (MAE) of 21.88 mg/dL, compared with 22.90 mg/dL for the best baseline and 27.69 mg/dL for the train‑fold mean baseline. This indicates GlucoFM provides complementary historical context useful for predicting postprandial glucose changes.

Using multiple days to improve subject prediction

Single 24‑hour traces may not capture a person’s full glucose behavior. The team encoded each day separately and averaged representations across up to seven days, weighting each participant equally. Additional days improved PR‑AUC in most settings and datasets: for example, Stanford beta‑cell dysfunction improved by 9.6 points and Hall diabetes prediction by 14.0 points. CGMacros showed mostly positive gains across Dexcom, Libre and fused sensor data. ShanghaiT2DM insulin‑resistance was an exception under simple averaging, suggesting optimal aggregation strategies may vary by task. Overall, frozen daily representations can be combined to strengthen subject‑level prediction without retraining the encoder.

Cross‑cohort transfer

The study tested cross‑dataset transfer: training a downstream classifier on one cohort and evaluating it on another. GlucoFM outperformed the second‑best method in 11 of 12 cross‑cohort evaluations by margins of 0.5–8.6 PR‑AUC points and trailed once by 0.6 points. Absolute PR‑AUCs ranged from 61.6% (Stanford→Hall tasks) to 90.0% (Hall→CGMacros insulin resistance), indicating that emphasizing physiological structure helps frozen representations generalize beyond cohort‑specific noise.

Few‑shot learning with limited labels

Two few‑shot scenarios were tested: varying the number of labeled participants per class, and varying the fraction of observations available per participant. GlucoFM achieved the highest task‑averaged PR‑AUC at every evaluated data budget, including the most limited settings (one labeled subject per class and 1% of observations). The advantage was most pronounced when labeled subjects were scarce, demonstrating efficient learning from very few examples.

Does the dual‑stream design matter?

The team compared the full dual‑stream encoder to alternatives emphasizing raw input, state‑only (slow trends) or event‑only (fast deviations). The event‑only model performed worst, showing transient fluctuations alone are insufficient. Raw‑input and state‑only variants were competitive, but the full dual‑stream model consistently performed best, supporting the idea that separating slower and faster glucose dynamics into complementary streams improves learned representations.

Conclusions and limitations

These results suggest CGM models benefit from explicitly modeling multiscale glucose dynamics — slower trends, short‑term deviations, daily timing, and sensor missingness. By learning reusable patterns from unlabeled CGM, GlucoFM produced representations that performed strongly on prediction, cross‑cohort transfer and few‑shot tasks, offering a way to leverage limited labeled clinical data more effectively.

Limitations include variation in metabolic responses across individuals, cohorts and devices, and a pre‑training population of modest size. Next steps include pre‑training on larger, more diverse populations, extending GlucoFM beyond independently processed 24‑hour windows toward native multi‑day modeling to capture trends over weeks or months, and exploring real‑time adaptation.

Acknowledgements

Contributors: Zechen Li, Keerthana Natarajan, Weizhi Zhang, Simon A. Lee, Yuwei Zhang, Maxwell A Xu, Menglian Zhou, Zeinab Esmaeilpour, Flora D. Salim (University of New South Wales), Mark Malhotra, Lindsey Sunden, Shwetak Patel, Yuzhe Yang, and Ahmed A. Metwally. The authors also acknowledge Bobak J. Mortazavi and Ricardo Gutierrez‑Osuna (Texas A&M University) for providing the CGMacros dataset used in the study.