What GEODE is trying to fix

Most “scientific foundation models” assume all problems can be pushed through a common format: grids, tokens, or symbolic encodings. In that setup, the messy parts—boundary conditions, monitor histories, specific geometries—get translated by an external “scientific interface” layer before the model ever sees them. That interface is both:

  • Outside the pretrained model, and
  • Outside the audit trail of what is actually reused across tasks.

The GEODE work takes the complementary approach: keep the real scientific interfaces and physical domains in place, and let the foundation model adapt around them instead of the other way around [\[1\]](http://arxiv.org/abs/2609.38067v1).

Concretely, the paper:

  • Keeps boundary histories, sparse monitor records, and loading histories in their native inference classes.
  • Keeps outputs on Cartesian, latitude–longitude, and unstructured domains instead of forcing everything onto a common grid [\[1\]](http://arxiv.org/abs/2609.38067v1).

The core object is GEODE: a foundation model built as a shared routed library of wavelet operators, with task-specific scientific interfaces plugged into that library [\[1\]](http://arxiv.org/abs/2609.38067v1).

That library is jointly pretrained to represent:

  • Cavity flow
  • Radiation dose
  • Elastoplastic stress

After that pretraining, they add a heat exchanger and a reactor subchannel by training only a private interface containing 2.1% of the model parameters [\[1\]](http://arxiv.org/abs/2609.38067v1).

For anyone building agents around existing simulation stacks, that’s a very different foundation-model story: the model is explicitly not the “single interface to everything.” Instead, it is a shared operator core that plugs into many heterogeneous interfaces.

How GEODE is structured

From the abstract, you can think of GEODE in three layers [\[1\]](http://arxiv.org/abs/2609.38067v1):

  1. Task-specific scientific interfaces
  • These handle the native inputs: boundary condition histories, monitor readings, loads.
  • They also emit outputs on the native mesh/coordinate types: Cartesian grids, latitude–longitude grids, or unstructured meshes.
  • Every new physical system (e.g., a reactor subchannel) gets its own private interface module.
  1. Shared routed library of wavelet operators
  • This is the foundation part: a common bank of wavelet operators that can be reused across tasks.
  • “Routed” implies there’s some learned mechanism that decides which operators to apply where for each task and input; the exact routing mechanism is not detailed in the abstract, but the point is selective reuse, not a single monolithic block.
  1. Routing and parameter isolation
  • When adding new tasks, only the task interface is trained—2.1% of total parameters—while the shared operator library stays frozen [\[1\]](http://arxiv.org/abs/2609.38067v1).
  • This parameter isolation is contrasted with “unrestricted fine-tuning,” where the full network is allowed to update.

The training and evaluation story has three main phases:

  1. Joint pretraining on multi-physics
  • Cavity flow, radiation dose, elastoplastic stress are represented within a single jointly pretrained GEODE model [\[1\]](http://arxiv.org/abs/2609.38067v1).
  • Here, both the shared wavelet library and the task interfaces for those three problems are trained.
  1. Incremental acquisition of new systems
  • A heat exchanger and a reactor subchannel are added by training only the new private interfaces, which collectively represent 2.1% of model parameters [\[1\]](http://arxiv.org/abs/2609.38067v1).
  • The shared wavelet library is reused but not updated.
  1. Comparative fine-tuning setups
  • Parameter-isolated setup: only the new interface is trained; shared library is frozen.
  • Unrestricted fine-tuning: the whole model, including shared operators, is updated on the new task.

The key difference between these setups is what happens to the original tasks (cavity flow, dose, stress) after you train on the new ones.

What happens when you add new tasks

When GEODE learns the heat exchanger and reactor subchannel tasks via parameter isolation:

  • Earlier predictions remain unchanged [\[1\]](http://arxiv.org/abs/2609.38067v1).

Under unrestricted fine-tuning, where you allow the whole model to adapt to the new tasks:

  • Predictions on the earlier tasks degrade by factors of 14–29 [\[1\]](http://arxiv.org/abs/2609.38067v1).

This is a critical point for anyone building long-lived systems:

  • If you want a single foundation model to power multiple engineering workflows, you cannot assume that standard fine-tuning will respect earlier calibrations.
  • Isolating most parameters and only training small per-interface heads (2.1% here) can dramatically cut catastrophic forgetting relative to letting the whole model move.

But the paper also emphasizes that:

“Preservation alone does not establish reuse” [\[1\]](http://arxiv.org/abs/2609.38067v1).

Even if freezing the shared library successfully prevents regressions on old tasks, that doesn’t prove those tasks are benefiting from the pretrained computation when you add new ones—or that the new tasks are actually using the old knowledge.

Measuring actual “reuse” vs just not breaking things

To test whether the shared wavelet library is actually doing useful work for a given task and data regime, the authors introduce a norm-matched randomized-library control [\[1\]](http://arxiv.org/abs/2609.38067v1).

The idea at a high level:

  • Build a version of the model where the shared wavelet library is replaced by randomized operators, but norm-matched so that the raw magnitudes/scale match the real library.
  • Compare performance between:
  • The model with the pretrained library, and
  • The model with the randomized, norm-matched library.

What they find is:

  • The contribution of pretrained computation (the real, learned wavelet library) is conditional on the task and data regime [\[1\]](http://arxiv.org/abs/2609.38067v1).

In other words:

  • There are regimes where the pretrained library clearly adds value over randomized operators.
  • There are also regimes where you cannot assume that having the shared library improves performance simply because it’s there and frozen.

For practitioners, this is a caution about overclaiming transfer:

  • A model that doesn’t regress on old tasks when you add new ones is not automatically a good “foundation” for those new tasks.
  • You need explicit controls—here, randomized norm-matched operators—to quantify what is actually being reused rather than just co-located.

Why full-field error can lie to you

The paper also looks at how you evaluate errors on scientific fields and finds a subtle but important issue:

  • A full-field relative L2 error can substantially understate error relative to spatial variation when the field level dominates the norm [\[1\]](http://arxiv.org/abs/2609.38067v1).

Intuitively:

  • If a field is dominated by a large mean level and the model gets that level roughly right, the overall L2 norm will look small even if the spatial structure (gradients, local variations) is badly wrong.

So they introduce a separate decomposition of error that can expose this effect [\[1\]](http://arxiv.org/abs/2609.38067v1). While the abstract does not spell out the exact decomposition, the upshot is:

  • You should not rely on a single full-field norm when assessing scientific models; you need metrics that distinguish level from variation.

They also find that:

  • Task-specific operators remain more accurate on three of the five problems studied [\[1\]](http://arxiv.org/abs/2609.38067v1).

So even with a shared wavelet library and multi-task pretraining, bespoke operators tuned for a given task can still outperform the foundation model on a majority of problems in this benchmark.

This cuts against a naive “one model to rule them all” narrative:

  • Multi-task coverage and shared operators don’t automatically beat carefully crafted task-specific systems across the board.

How this fits in an AI+simulation stack

Although the paper doesn’t describe a production stack, the pieces described in the abstract map naturally onto an agentic science/engineering setup:

  • Existing simulation codes and data systems stay wrapped in their native interfaces:
  • They produce boundary and loading histories, sparse monitor records, and fields on mixed domains (Cartesian, lat–lon, unstructured) [\[1\]](http://arxiv.org/abs/2609.38067v1).
  • GEODE’s task-specific interfaces are the bridge:
  • They ingest those native representations.
  • They project queries into the shared wavelet operator library and decode outputs back into the domain-native form.
  • The shared wavelet library becomes a reusable “physics kernel”:
  • Jointly pretrained across cavity flow, radiation dose, elastoplastic stress.
  • Reused when acquiring heat exchanger and reactor subchannel tasks with only 2.1% new parameters trained [\[1\]](http://arxiv.org/abs/2609.38067v1).

For an AI engineer building agents over heterogeneous physical systems, this suggests a pattern:

  • Do not force a single grid or tokenization layer across physics codes if that means throwing away domain-specific interfaces.
  • Instead, maintain rich, per-system adapters and share only a relatively small, highly structured operator library.
  • Parameter isolate that shared core when adding new systems so you can:
  • Preserve earlier calibrations (avoiding 14–29x degradation from unrestricted fine-tuning), and
  • Run explicit controls to see whether the shared computation is helping [\[1\]](http://arxiv.org/abs/2609.38067v1).

Why this matters now

The GEODE paper makes three conceptual distinctions that are easy to blur:

  1. Multi-task coverage
  • Can a single pretrained model represent many different problems at once (here: cavity flow, dose, stress, then heat exchanger and reactor subchannel) [\[1\]](http://arxiv.org/abs/2609.38067v1)?
  1. Preservation
  • When you add or adapt to new tasks, do you preserve performance on the old ones?
  • Parameter isolation here avoids degradations of 14–29x that appear under unrestricted fine-tuning [\[1\]](http://arxiv.org/abs/2609.38067v1).
  1. Pretrained reuse
  • Do new tasks actually benefit from previously learned computation, beyond what a random-but-norm-matched operator bank would give you?
  • The randomized-library controls show this is conditional on task and data regime, not automatic [\[1\]](http://arxiv.org/abs/2609.38067v1).

For production builders, these distinctions map to different tests you should plan for:

  • Coverage tests: Can the architecture even express the families of problems you care about?
  • Preservation tests: Before and after each incremental adaptation, measure old-task performance; avoid unrestricted fine-tuning when regressions appear.
  • Reuse tests: Use targeted ablations—like norm-matched random operators—to check that your “shared core” is actually contributing, not just sitting inert.

GEODE also pushes against the idea that a foundation model must dictate the singular interface for everything. Instead, it treats the scientific interfaces themselves as first-class, and makes the foundation piece a structured library that can serve them.

What to watch next

From the abstract alone, several questions for future work stand out:

  • Scaling behavior across more tasks
  • The paper reports five problems total. It will matter how the routed wavelet library scales in effectiveness—both in coverage and in true reuse—when the number of physical systems grows significantly [\[1\]](http://arxiv.org/abs/2609.38067v1).
  • Interface design principles
  • We know task-specific scientific interfaces exist and can be small (2.1% of parameters when adding two new tasks) [\[1\]](http://arxiv.org/abs/2609.38067v1).
  • It remains to be seen what general recipes emerge for designing such interfaces for entirely new domains.
  • Better error decompositions
  • The paper shows that standard full-field relative L2 error can mislead when field level dominates the norm, and provides a decomposition that clarifies this [\[1\]](http://arxiv.org/abs/2609.38067v1).
  • A broader toolkit of decomposed metrics will likely be needed as these models enter safety-critical settings.
  • Operator sharing vs. task-specificity trade-offs
  • Task-specific operators beat the shared approach on three out of five problems [\[1\]](http://arxiv.org/abs/2609.38067v1).
  • Understanding when to invest in shared libraries vs. keeping things bespoke will be an ongoing design choice for engineering organizations.

For now, GEODE’s main contribution is a mechanism and an evaluation mindset: couple heterogeneous interfaces through a routed operator library, isolate most parameters when extending, and rigorously separate “we didn’t break previous tasks” from “we are truly reusing what we pretrained.”

What is not documented

The abstract and metadata do not document:

  • The exact architecture of the routed wavelet operator library (e.g., number of operators, routing mechanism, network depths).
  • Training details: datasets, loss functions, optimization procedures, compute budget, or training schedule.
  • Quantitative error numbers beyond relative factors (e.g., the precise metrics before and after fine-tuning, or per-task performance tables).
  • Implementation details of the task-specific interfaces (input/output tensor shapes, how native domains are encoded, etc.).
  • The precise mathematical form of the error decomposition used to separate field level from spatial variation.
  • Any runtime, memory, or deployment characteristics of GEODE in real-world pipelines.