TL;DR for operators

A forest-monitoring team is not choosing between sensors on accuracy alone. It is choosing between a structurally rich measurement source that is expensive and intermittent and a cheaper annual representation that can keep far more field observations usable across time and geography.

On the same restricted set of 589 plots, tuned LiDAR-only and combined models reach $R^2$ up to 0.79, while the annual learned representation reaches 0.67. But when annual availability expands the usable dataset to 6,801 observations, the best model using those annual representations reaches $R^2 = 0.82$.

That reversal is the point. The highest headline $R^2$ does not show that the annual representation is a better structural sensor; it shows that coverage and temporal availability can change the effective training set enough to change model performance. These annual learned Earth-observation representations—Google Satellite Embeddings, or GSE—therefore compete with LiDAR partly through information supply, not just information quality.

For operators, the decision is whether repeated LiDAR acquisition is worth its structural detail, whether scalable annual embeddings provide enough coverage, or whether the two should be combined according to where structural accuracy versus geographic and temporal continuity matters most.

The best sensor can lose when it leaves too much data unusable

A forest-monitoring program rarely chooses between two measurement systems under laboratory conditions. One source may describe canopy structure well but arrive only when an aircraft flies. Another may contain less directly interpretable information but update every year across the region.

That difference changes more than collection logistics. It determines which field measurements can actually enter the model.

Shashika Lamahewage and Chandi Witharana test this problem across seven northeastern U.S. states using continuous forest inventories, airborne LiDAR, and annual AlphaEarth Google Satellite Embeddings, or GSE.1 Their target is aboveground biomass: the living plant mass above the ground used here as a forest-carbon monitoring quantity.

The key comparison has two stages.

Data condition Best reported result What it establishes
589 plots with comparable LiDAR and GSE availability LiDAR-only: $R^2 = 0.79$ Direct structural sensing remains strong when observations are held to the same restricted sample
Same 589-plot setting GSE-only: $R^2 = 0.67$ GSE is not intrinsically more predictive than LiDAR in the common-sample comparison
Expanded annual GSE setting, 6,801 observations GSE-only: $R^2 = 0.82$ Annual availability can compensate through much greater usable training coverage

The highest $R^2$ therefore comes from the representation that performs worse when the data comparison is constrained to the same 589 plots. The change is not evidence that satellite embeddings suddenly contain better canopy measurements. The training environment changed.

Annual coverage turns temporal alignment into a modeling advantage

The study starts with heterogeneous forest inventories collected at different times. Remote-sensing observations also arrive on their own schedules. Matching the two strictly can discard large amounts of otherwise informative field data.

LiDAR is particularly exposed to this constraint because the paper integrates multiple airborne missions acquired between roughly 2011 and 2025. Requiring inventory plots to overlap the satellite period and fall within three years of a LiDAR acquisition leaves only 589 observations for the direct comparison.

GSE changes that constraint because the AlphaEarth representation is available annually from 2017 through 2025 at 10-meter resolution. The researchers estimate mean annual biomass change between consecutive plot inventories,

$$ \Delta \mathrm{AGB} = \frac{\mathrm{AGB}\ast{t_2}-\mathrm{AGB}\ast{t_1}}{t_2-t_1}, $$

and use that rate to align eligible field measurements to remote-sensing years within the study’s stated three-year window.

With annual GSE observations, that procedure expands the reported usable sample from 589 to 6,801 observations. The best expanded GSE-only configuration—using dimensionality reduction, spatial predictors, tuning, and XGBoost—reaches $R^2 = 0.82$ and a reported MAE of 245.23 Mg ha$^{-1}$.

For an operational monitoring program, availability is therefore part of measurement quality. A source that updates consistently can preserve more supervision, cover more geographic variation, and reduce losses caused by acquisition timing. Those gains can outweigh weaker per-observation structural information.

LiDAR still provides information the embeddings do not clearly replace

The restricted comparison prevents a replacement interpretation.

On 589 plots, tuned GSE-only models top out at $R^2 = 0.67$. Tuned LiDAR-only Random Forest reaches $0.79$, while the best tuned combined LiDAR-GSE configuration also reaches $0.79$. The combined models additionally show low residual geographic dependence after spatial correction; for combined XGBoost, Moran’s $I$ falls to $-0.0095$ with $p = 0.482$.

That spatial test has a specific purpose. Nearby prediction errors can resemble one another, allowing a geographic model to exploit location structure rather than generalize from the intended predictors. Adding coordinate-derived features and checking residual Moran’s $I$ is therefore a diagnostic and corrective step. It is evidence that residual spatial structure was reduced; it is not equivalent to testing the model on an entirely unseen geographic region.

The study’s permutation sensitivity analysis points in the same direction. LiDAR plus spatial information remains a substantial contributor to performance. GSE principal components on their own show relatively small sensitivity, although their contribution improves when combined with spatial information.

The embeddings appear to encode ecologically relevant structure: some features have moderate relationships with upper-canopy LiDAR metrics. But an individual GSE dimension cannot be read as a named ecological measurement. The representation remains largely a black box.

A practical monitoring stack uses the two sources differently

For a forest-carbon MRV team, regional agency, or land manager, the evidence supports a tiered acquisition strategy rather than a single winner.

Use annual GSE as the coverage layer. Where repeated airborne acquisition is impractical, annual embeddings can keep more inventory observations temporally relevant and support frequently refreshed regional biomass estimates.

Use LiDAR where structural detail has higher marginal value. High-value project areas, calibration zones, uncertain forests, or locations where biomass estimates affect consequential accounting decisions may justify the additional acquisition cost.

Keep field inventories as the measurement anchor. Neither remote-sensing source removes dependence on representative plot data. The study’s gain from annual embeddings works precisely because more inventory observations can be aligned to those annual representations.

A further Cognaptus inference follows: organizations could use an annual GSE-based model to identify regions where uncertainty, disturbance, or economic exposure warrants subsequent LiDAR collection. The paper does not test such an acquisition policy directly, so its return on investment remains an operational hypothesis rather than a reported result.

Deployment still needs a harder geographic test

Several boundaries materially affect how the reported accuracy should be used.

The inventory network combines programs with different plot shapes, sizes, objectives, and geolocation quality. Biomass labels also inherit uncertainty from allometric equations, while very high observations above roughly 600 Mg ha$^{-1}$ contribute disproportionately to prediction error. The annual growth adjustment assumes that mean change between inventories is a reasonable approximation over the alignment interval, an assumption that disturbance or silvicultural intervention can violate.

More importantly, the study uses random 80/20 train-test splitting with spatial covariates rather than explicitly holding out geographic regions. Before deploying a model into forests materially different from those represented in training, an operator would want spatially blocked or out-of-region validation.

The manuscript also contains minor reporting inconsistencies that make directional conclusions safer than some secondary point estimates. Scenario II is reported as 6,801 observations in Table 2 and the results but as 6,802 in one methods sentence; the retained GSE principal-component count is reported as both 53 and 54; and some MAE and relative-bias values differ between Table 5 and later prose. These conflicts do not reverse the central sample-size result, but they argue against treating every reported secondary metric as exact.

The procurement decision is about information supply, not a leaderboard

This paper changes the forest-monitoring decision in a useful way. LiDAR remains the stronger structural measurement under the common 589-plot comparison. Annual GSE becomes compelling because it changes how much training evidence survives the alignment process.

For organizations designing regional biomass systems, that separates two forms of value: structural information per observation and continuity of observations across space and time.

A defensible operating model is consequently a portfolio: annual embeddings for broad, repeatable monitoring; LiDAR for targeted structural measurement; and field inventories for calibration and validation. The study provides predictive evidence that this architecture can work across a large northeastern U.S. region. Whether it transfers reliably to unseen geographies, different inventory regimes, and operational carbon-accounting thresholds still requires validation designed around those deployment conditions.

Cognaptus: Automate the Present, Incubate the Future.


  1. Shashika Lamahewage and Chandi Witharana (2026). Foundation-Model Earth Representations Enable Regional-Scale Forest Aboveground Biomass Monitoring Across the Northeastern United States. arXiv:2607.27217. https://arxiv.org/abs/2607.27217 ↩︎