Why composing LoRAs is hard in practice
LoRA is attractive because it lets you fine-tune a large language model cheaply, once per task: you freeze the base and learn a low‑rank “adapter” delta instead of touching the full weights paper.
The trouble starts when you have many such task‑specific adapters and you want:
- a *single* combined model,
- without retraining on all task data,
- and without per‑task routing logic at inference.
The paper “New LoRA Skills Should Read but Never Write” identifies three standard strategies and their problems paper:
- Merging in weight space (e.g., summing deltas into the base): causes *interference* between skills.
- Retraining on all task data: expensive; essentially re‑doing a multi‑task fine‑tune.
- Routing between separate adapters: you abandon the “single model” goal and add task‑specific runtime logic.
They trace these difficulties to two usually‑implicit choices that every composition method makes paper:
- Factorization choice inside each LoRA update.
A LoRA update has *infinitely many equivalent factorizations*; each factorization implements the same effective weight update, but it determines *what a learned interaction between adapters can see* paper. While an adapter is used alone, the choice is invisible; when you try to compose, it suddenly matters.
- Direction of coupling between skills.
When you couple an “old” skill with a new one, the coupling can “point” in either direction. The direction decides *whether the old skills keep computing what they computed before* [paper](http://arxiv.org/abs/2609.31600v1]. A bad direction lets new skills overwrite old ones.
For builders, that second point is the real failure mode you’ve probably seen: you add a new LoRA and some older behavior silently regresses, even when their domains seem unrelated. READ’s core idea is to fix both the factorization and the coupling direction.
READ: Read-only Expansion of Adapter Deltas
The method introduced in the paper is called READ: Read-only Expansion of Adapter Deltas paper.
High‑level design goals:
- Preserve each skill’s original behavior *exactly* when you add it.
- Allow new skills to leverage the representational subspaces of existing skills.
- Prevent new skills from *modifying* what old skills compute.
- Keep inference cost identical to a single fine‑tuned model: no routing, no extra passes.
READ does this via two key mechanisms:
- Canonicalizing each adapter.
Every LoRA adapter is rewritten into a balanced canonical form that preserves its update exactly [paper](http://arxiv.org/abs/2609.31600v1]. This fixes the factorization ambiguity: instead of arbitrary low‑rank factors, you enforce a canonical coordinate system for each adapter’s “input” and “output” subspaces.
- Since the update is preserved exactly, plugging the canonicalized adapter back into the base produces the *same* behavior as the original adapter when used alone [paper](http://arxiv.org/abs/2609.31600v1].
- The benefit is that all further operations—especially cross‑skill coupling—operate in a consistent basis, so the semantics of “reading from” or “writing to” a subspace are well‑defined.
- One‑directional coupling: new skills read, never write.
READ then introduces a coupling matrix between skills, but constrains it so that the coupling grows in one direction only [paper](http://arxiv.org/abs/2609.31600v1].
Specifically:
- A new skill can read the input subspaces of old skills.
- A new skill cannot write into their output subspaces [paper](http://arxiv.org/abs/2609.31600v1].
This “read‑only” design makes the coupling asymmetric:
- Old skills keep computing *exactly* what they computed before; their outputs are not perturbed by the presence of the new skill.
- New skills can condition on richer features (the input subspaces of prior adapters) to implement more complex or composite behaviors.
The only trainable object at each append is the new skill's row of the coupling matrix [paper](http://arxiv.org/abs/2609.31600v1]. That is structurally important:
- You never revisit or adjust the couplings of older skills.
- Append operations are localized: add one adapter, train one row.
Once you have a stack of skills and couplings, the composed update folds into the base weights [paper](http://arxiv.org/abs/2609.31600v1]. The result is:
- No inference‑time routing.
- No task‑specific rules.
- Same forward pass structure and cost as a “single” fine‑tuned model.
For agent engineers, that means you get multi‑skill behavior without changing your serving stack.
Step-by-step: appending a new LoRA skill with READ
Based on the description in the paper, an append looks conceptually like this paper:
- Train a new LoRA adapter for a new task.
- Use standard LoRA fine‑tuning on your frozen base model for the new task.
- The result is a low‑rank update specific to that task.
- Canonicalize the new adapter.
- Rewrite the adapter into the balanced canonical form used by READ.
- This transformation keeps the effective update *exactly the same* as the original adapter [paper](http://arxiv.org/abs/2609.31600v1].
- Integrate into the coupling scheme.
- Introduce a new row in the coupling matrix corresponding to this skill.
- Set up the coupling so that this new skill can read input subspaces of existing skills.
- Enforce the one‑directional constraint so it cannot write into their output subspaces [paper](http://arxiv.org/abs/2609.31600v1].
- Train only the new coupling row.
- Freeze:
- The base model,
- All existing adapters,
- All existing rows of the coupling matrix.
- Train *only* the new skill’s row of the coupling matrix [paper](http://arxiv.org/abs/2609.31600v1].
- Fold the composed update into the base.
- Combine the base weights, all canonicalized adapters, and their couplings into a single effective weight update.
- Deploy a model that carries all the skills with no additional inference cost and no routing [paper](http://arxiv.org/abs/2609.31600v1].
For a production pipeline, you can imagine this as a skill “registry” of LoRA deltas, each stored in canonical form, along with a persistent coupling matrix that grows by one row per skill.
Where READ fits in an agentic stack
From the paper’s setup, READ is primarily a model‑side solution to multi‑task capability, not an orchestration trick [paper](http://arxiv.org/abs/2609.31600v1]. In a typical agentic stack:
- Today’s pattern
- Use one base model,
- Load different LoRAs per tool or workflow,
- Implement routing via:
- prompt instructions,
- tool selection policies,
- or explicit adapter switches.
- With READ
- All those LoRAs are *composed once* into a single set of weights.
- Orchestration simplifies:
- No need to choose an adapter per call.
- No “leaking” of task boundaries into prompts.
- You still decide which tools or external APIs to call, but the language model behaves as if it had been multi‑task trained on your skill set.
The claim that the composed update folds into the base weights with no inference cost, routing, or task-specific rules matters here [paper](http://arxiv.org/abs/2609.31600v1]:
- You don’t pay latency for extra adapter passes.
- You don’t maintain runtime logic to choose “which LoRA” for each request.
- You reduce failure modes where the router picks the wrong adapter.
From an operations standpoint, READ shifts complexity from runtime to an offline append process that is strictly local to each new skill.
Why “read-only” matters for skill survival
The authors emphasize that factor coordinates and coupling direction, which are never exposed by a lone adapter, decide whether composed skills survive [paper](http://arxiv.org/abs/2609.31600v1].
Two insights for practitioners:
- Factorization is a hidden interface.
When you train a LoRA in isolation, you only care that its overall delta works. But under composition:
- Different factorizations that are equivalent in isolation become *inequivalent* because they present different bases to any cross‑skill coupler.
- Arbitrary factorization choices can starve the coupler of semantically meaningful directions to connect skills.
READ’s balanced canonical form is essentially a standardization of this hidden interface so that couplings can be learned reliably [paper](http://arxiv.org/abs/2609.31600v1].
- Write access is what kills old skills.
If you let a new adapter *write* into the output subspaces of older ones, then:
- Small updates to the new skill’s parameters can propagate backwards and alter established behaviors.
- You lose the guarantee that “adding a new skill will not break old ones.”
By forcing couplings to grow only in a read-only direction—new skill reads old input subspaces but cannot write into their outputs—READ ensures old skills keep computing exactly what they computed before [paper](http://arxiv.org/abs/2609.31600v1].
For developers shipping systems with long‑lived skills (e.g., a set of regulatory or safety behaviors), this is particularly important: you can append new capabilities while preserving core, previously‑validated behaviors.
Empirical picture: how much does READ buy you?
The paper evaluates READ across four benchmark suites and two model families, adding skills one at a time [paper](http://arxiv.org/abs/2609.31600v1].
Key reported results:
- Across several families, READ improves every suite average over the strongest published baselines built from the same adapters [paper](http://arxiv.org/abs/2609.31600v1].
- On SuperGLUE, READ improves the suite average by more than twenty points over these baselines [paper](http://arxiv.org/abs/2609.31600v1].
- On a domain suite, READ improves the suite average by more than seven points [paper](http://arxiv.org/abs/2609.31600v1].
- Nearly all complete addition sequences end above every direct baseline [paper](http://arxiv.org/abs/2609.31600v1].
The important details for deployment:
- These baselines are built from the same adapters. So gains are coming from the composition method, not better single‑task fine‑tuning.
- Skills are added one at a time, which aligns with how teams typically grow capability over months.
This suggests that if you already have a library of per‑task LoRAs, READ can potentially turn that pile of one‑off adapters into a more coherent multi‑skill model without redoing task training.
Practical implications for builders
Given only the information in the paper, here are some grounded implications for engineering multi‑skill systems:
- Prefer composition over routing when latency and simplicity matter.
READ gives you a way to retain the cheapness of LoRA fine‑tuning while moving back to a unified model that doesn’t need adapter routing at inference [paper](http://arxiv.org/abs/2609.31600v1].
- Treat LoRA factorization as part of the API.
The balanced canonical form is not just a math trick; it makes the interaction between skills well‑posed by fixing the coordinate system the coupler sees [paper](http://arxiv.org/abs/2609.31600v1]. If you roll your own composition scheme, you’ll want similar thinking about factorization.
- Make append operations local and monotone.
READ’s design—only the new skill’s row of the coupling matrix is trained—gives a template for safe continual learning where old skills can be assumed stable [paper](http://arxiv.org/abs/2609.31600v1]. Architectures that keep revisiting older parameters will have a harder time preserving behavior.
- Plan for evaluation sequences, not just endpoints.
The paper reports not only final performance but how “complete addition sequences” behave, noting that nearly all end above every baseline [paper](http://arxiv.org/abs/2609.31600v1]. In production, you care about the model after every incremental append; READ is explicitly designed for that use case.
In terms of integration into your stack, READ is a training‑time and offline‑composition technique; there is no new runtime primitive you must support beyond serving a standard model with modified weights [paper](http://arxiv.org/abs/2609.31600v1].
What is not documented
The paper’s abstract states the high‑level method and results, but several important details are *not* documented in the provided text:
- The exact mathematical form and algorithm to construct the balanced canonical form of a LoRA adapter.
- How the coupling matrix is parameterized layer‑wise or module‑wise, and how it interacts with multiple adapters across a deep network.
- Training hyperparameters or data used when fitting the new skill’s coupling row (optimizer, learning rate, data mixture).
- The specific model families, architectures, and parameter scales used in the evaluations.
- The identities and compositions of the four benchmark suites beyond the mention of SuperGLUE and a “domain suite”.
- Any implementation details or reference code for integrating READ into existing frameworks.