From “hit the docs again” to source-specific competence
When you wire an LLM agent to an external source—API docs, policy manuals, a codebase, a wiki—the default pattern is:
- index it once,
- run retrieval on every query,
- maybe keep some interaction-level memory.
All repeated use of that source is still just repeated access: every task treats the source as if it’s new.
The paper “From Knowledge Access to Source Learning: Developing Source-Specific Competence” proposes a different object in your stack: a persistent source model for each authoritative source your agent relies on, and a process called source learning to build it over time 1.
Instead of just getting better at generic retrieval, the agent is meant to develop reusable source-specific competence:
- how this particular source is structured,
- how its content is interpreted,
- how its knowledge is applied in tasks.
This is orthogonal to user-level memory. A tool like Breadcrumb records what *you* do on your Mac—screens, meetings, AI chats—and exposes that history to your AI as a searchable memory layer 2. Source learning is about what the *agent* learns about an external source it keeps calling into, not about the agent’s egocentric experience.
What “source learning” actually is
The core definition from the paper:
Source learning: developing reusable source-specific competence over a persistent authoritative source 1.
The representation of this competence is a persistent source model:
- It is attached to a specific source (e.g., a particular API’s docs).
- It captures reusable understanding of:
- how the source’s knowledge is structured,
- how it should be interpreted,
- and how it is applied in downstream tasks 1.
Critically:
- This model is persistent: it survives across tasks and sessions.
- It is refined progressively as the agent works with that source.
In other words, instead of each task building ad-hoc, per-query views of the same docs, the system maintains a stable, improving representation of the source itself.
SourceLearn: two learning loops over the same source
To implement source learning, the authors introduce SourceLearn, which combines two complementary mechanisms 1:
- Self-Directed Source Learning (SDSL)
- The agent identifies what remains incompletely understood in the source.
- It adaptively revisits the source to close those gaps 1.
- Task-Guided Source Learning (TGSL)
- The agent uses downstream experience—its performance on actual tasks—to reveal:
- local representational gaps in the source model,
- and recurring needs in how knowledge should be organized 1.
In both cases, there are two key constraints:
- Where the learning signal comes from
“Learning signals determine what should be reconsidered” 1.
- In SDSL, the signal is: “these parts of the source are still unclear.”
- In TGSL, the signal is: “these failures or friction points in tasks point to missing or misorganized source understanding.”
- Where the updated content comes from
“Persistent updates are reconstructed from the authoritative source” 1.
This is important: they do *not* simply store hallucinated interpretations; any persistent change is re-derived from the original source, which is treated as ground truth.
This is a different stance from memory systems that preserve whatever the model said earlier. For instance, experience-based memory systems (the baselines they compare against) preserve knowledge from prior interactions, but still “largely treat repeated use of the same source as repeated access” instead of as an opportunity for deeper understanding 1.
Step-by-step, conceptually
The paper doesn’t spell out the full algorithmic sequence, but the documented structure implies this kind of cycle:
- Initial source model construction
- Start from the authoritative source.
- Build an initial persistent source model that organizes the content somehow (details are not given in the text we have).
- Run tasks against the source
- Use this persistent source model (plus direct access to the source) to solve knowledge-intensive tasks involving that source 1.
- Generate learning signals
- Self-directed: identify which parts of the source remain poorly understood.
- Task-guided: detect representational gaps and recurring organization needs based on how tasks succeeded or failed 1.
- Revisit the authoritative source
- For parts flagged by the learning signals, revisit the source itself (docs, manual, etc.) to reconstruct better representations.
- Persistently update the source model
- Apply updates to the persistent source model, creating new reusable competence specific to that source 1.
- Repeat across many tasks
- Over time, as tasks accumulate, the persistent source model should increasingly encode the “right way” to read and use that source.
The key control-flow idea for builders: separate “where should I look again?” from “how do I re-encode it?”. Learning signals come from performance; updates are re-derived from the authoritative text.
How this fits into an agentic stack
The paper situates SourceLearn relative to three families of systems:
- Improved source access and organization
- Retrieval and indexing methods that make it easier to get the right chunk from a large source 1.
- Hybrid RAG is explicitly named as a baseline; SourceLearn’s gains are measured relative to it, but the paper text here does not describe Hybrid RAG’s design 1.
- Agent memory systems
- Systems that preserve reusable knowledge from prior interactions, often user- or episode-centric 1.
- Breadcrumb is an example of this kind of memory for developer workflows: it records everything you do on a Mac (screen, meetings, AI transcripts, decisions) and turns that into local, encrypted memory that your AI can search via >30 MCP tools 2.
- These systems help the agent remember *episodes*; SourceLearn is specifically about understanding *an external source*.
- Source-specific learning
- SourceLearn is explicitly “over a persistent authoritative source” 1.
- The learner’s object isn’t the user, the environment, or generic skills; it’s “this API’s docs”, “this legal codebase”, etc.
In a production agent stack, SourceLearn would logically sit:
- Downstream of raw storage/indexing: it still needs the authoritative source to reconstruct updates.
- Alongside or under your vector store / RAG layer: it refines how that source is structured and used.
- Parallel to user- or task-centric memories: Breadcrumb-like systems keep track of “what you did”; SourceLearn keeps track of “what this source really means.”
You can think of the persistent source model as a long-lived, specialized adapter around one corpus, not a generic long-term memory.
Why it matters: the documented empirical gains
The authors evaluate SourceLearn across five benchmarks and three LLM backends 1. Within the text we have, the important documented facts are:
- SourceLearn “achieves the best performance in 13 of 15 settings” 1.
- It obtains “gains of up to 22.6 points over Hybrid RAG” 1.
- It also provides “substantial overall improvements over static source representations and experience-based memory baselines” 1.
So empirically, at least for the tasks they measure:
- Treating repeated source use as an opportunity to *learn the source* yields measurable gains over:
- better hybrid retrieval alone,
- static (non-learning) source representations,
- and generic “experience-based” memory that doesn’t explicitly model the source.
The exact tasks, metrics, and backends are not described in the provided text, but the pattern is consistent: learning a persistent, source-specific model pays off.
Design implications and trade-offs for builders
Even without the internal details, the paper’s framing suggests several concrete design implications.
1. Distinct memory objects: source vs. experience
Most builders today conflate at least two things:
- Experience memory: what the agent or user did (sessions, decisions, workflows).
- Breadcrumb makes this first-class: egocentric logs, meeting notes, AI sessions, and rules that apply automatically by folder scope 2.
- Source understanding: how to interpret and apply a persistent external source (docs, manuals, standards) 1.
SourceLearn’s framing is explicit that source-specific competence should be its own target. Architecturally, that means:
- one layer that organizes user/agent history (Breadcrumb-like),
- another layer that organizes each persistent external source, with its own life cycle.
2. Learning signals vs. update content
The paper draws a hard line:
- Learning signals (from self-directed and task-guided loops) tell you what to reconsider 1.
- But all persistent updates are reconstructed from the authoritative source 1.
For actual systems, that translates to:
- Avoid storing “I think section 4 means X” as ground truth. Instead, store “section 4 is important and needs a better representation” and always pull from section 4 itself when constructing that representation.
- This guards against model drift/hallucination in long-lived source models by anchoring every update in the original source.
3. Task-guided reorganizations of a source
Task-Guided Source Learning is explicitly about:
- “local representational gaps” and
- “recurring needs in how source knowledge should be organized” 1.
For builders, this suggests a pattern:
- Use failure and friction in real tasks to drive *schema evolution* of how you represent a source.
- If the same awkward cross-reference keeps tripping the agent, that’s a signal to reorganize that part of the persistent source model, not just tweak retrieval scores.
The paper doesn’t document what concrete reorganizations SourceLearn performs, so any particular schema design would be your own.
4. Self-directed revisiting as a scheduled background job
Self-Directed Source Learning “identifies what remains incompletely understood and adaptively revisits the source” 1.
For a deployed agent, that implies a concrete scheduling decision:
- Don’t only touch the source when there is a user query.
- Run background processes that try to close perceived gaps in understanding of that source, using idle time or explicit maintenance windows.
The paper doesn’t provide resource costs or policies, so trade-offs around compute vs. benefit would be up to you.
Relationship to other long-horizon work
The sources include several other long-horizon and memory-focused systems; they help contextualize SourceLearn, even if they tackle different problems:
- AutoCompact trains a coding agent to decide *when* to compact context, *what* state to preserve, and *how* to continue, improving pass rates on SWE-bench Verified and SWE-PolyBench Verified by 9.2% and 5.0% over a base model 3. This is context management over trajectories, not source-specific learning.
- Persistent Context Graphs / ReCAP compact long interaction histories by storing attention-derived importance and dependencies, then reassessing relevance using new user requests 6. Again, history-centric rather than source-centric.
- MemLife compacts long-term egocentric video into text episodes and uses a time-indexed reader to retrieve them; this improves over baselines without reprocessing raw video at query time 5. That’s multimedia personal memory, distinct from text source learning.
Taken together:
- AutoCompact/ReCAP/MemLife/Breadcrumb show how to curate and compact experience (what happened in a repo, a conversation history, your life, your desktop).
- SourceLearn focuses on curating and refining understanding of a persistent external source (the authoritative text your agent depends on).
These are complementary axes. A robust agent stack will likely need both.
What to watch next
From the limited text, a few open directions follow naturally, though the paper doesn’t document solutions:
- Operationalizing self-directed “incomplete understanding”
The paper states SDSL “identifies what remains incompletely understood” 1 but doesn’t document the detection mechanism. Designing good proxies for “confusion” over a source is an open engineering question.
- Benchmarks and source types
We know there are five benchmarks and three LLM backends 1, but not what kinds of sources they represent (APIs, legal texts, textbooks, etc.). Seeing where SourceLearn works best would inform whether to adopt similar patterns for your own sources.
- Interfaces between source models and other memories
The paper conceptually distinguishes source representations, static representations, and experience-based memories 1, but doesn’t specify APIs between them. In your stack, you’ll need concrete contracts between:
- persistent source models,
- interaction histories,
- and environment logs like Breadcrumb’s.
- Safety and auditing of source models
Since updates are reconstructed from authoritative sources 1, there’s a natural hook for audit tools, but the paper doesn’t describe any.
For now, the main actionable idea is conceptual: treat each critical external source as an object your agent is allowed to *learn*, not just *query*, and enforce that every “understanding” you persist traces back to the authoritative text.
What is not documented
Based on the provided sources, the following are *not* established:
- The internal architecture, data structures, or APIs of the persistent source model used by SourceLearn.
- The exact algorithms or prompts used for Self-Directed Source Learning and Task-Guided Source Learning, including how “incomplete understanding” and “representational gaps” are detected.
- The specific benchmarks, datasets, or domains used in the five-benchmark evaluation, and the identities or sizes of the three LLM backends.
- Details of Hybrid RAG’s implementation, beyond being a baseline that SourceLearn outperforms by up to 22.6 points in some settings.
- Any compute costs, latency impacts, or resource trade-offs of running SourceLearn in practice.
- Concrete integration patterns of SourceLearn with tools like Breadcrumb, MemLife, or ReCAP; these systems are documented separately and no integration is described.