Mehdi Akiki
Published on

How rust-analyzer Finds a Definition Through Macros and Modules

Authors
  • Mehdi Akiki avatar
    Name
    Mehdi Akiki
    Twitter

Article · Through the layers

Pressing go to definition looks like a text-search feature. The cursor is on Client, another file contains struct Client, so perhaps the editor searches for the same word.

That model fails quickly in Rust. A name can be renamed by use, re-exported through several modules, created inside a macro expansion, disabled by cfg, or present in the same source file under more than one crate configuration.

While contributing around compiler tooling, I found a more useful model:

Go to definition is a round trip from source syntax to semantic identity and back to source.

The return trip matters. A compiler can stop after it knows what a name means. An IDE must also find the file and range a person can open.

One small example already needs semantics

Consider this workspace:

// src/model.rs
#[derive(Default)]
pub struct Client;

// src/api.rs
pub use crate::model::Client as ApiClient;

// src/main.rs
mod api;
mod model;

macro_rules! construct {
    ($ty:path) => {
        $ty::default()
    };
}

fn main() {
    let _client = construct!(crate::api::ApiClient);
}

Ask for the definition of ApiClient at the call site. A useful answer may be the re-export in api.rs, or the underlying Client definition in model.rs, depending on the navigation command. The token also sits inside a macro invocation whose expansion contains another path.

A string search cannot decide any of this. rust-analyzer needs the crate graph, module tree, imports, macro expansion, and the meaning of the path at this exact location.

The five-stage trace I use

I reduce the operation to five stages:

cursor position
    ↓
syntax token and enclosing construct
    ↓
source or expanded syntax context
    ↓
resolved semantic definition
    ↓
original file and navigation range

This is not a literal list of five function calls. Internal APIs evolve. It is a stable way to reason about the work.

1. Turn an offset into syntax

The editor sends a document and byte or character position. rust-analyzer parses each file into a lossless syntax tree. Whitespace, comments, malformed code, and incomplete constructs remain representable because an IDE must work while I am typing.

The token at the cursor is only the beginning. Client may be:

  • a name being declared;
  • a path segment being referenced;
  • a field or method;
  • part of a use tree;
  • an associated type;
  • a token passed to a macro.

The enclosing syntax tells the IDE which semantic question to ask. This separation is important: the syntax layer knows the shape of one file, but it intentionally does not store name resolution or type information.

2. Keep the file context with the node

A syntax node is not globally unique enough for semantic work. The same physical file can participate in more than one crate or module context. Different cfg settings can also change which items exist.

Conceptually, semantic lookup needs both pieces:

(file or expansion identity, syntax node)

rust-analyzer uses file-aware wrappers and semantic APIs for this job. This prevents a common mental mistake: assuming that one source file always maps to one module definition.

For example, a shared source file can be included from two crate roots. The text is identical, but a path beginning with crate:: has different meaning in each crate. Source-to-definition mapping can therefore be one-to-many at the file boundary.

3. Cross the syntax-to-definition boundary

For a declaration node, rust-analyzer needs to discover its semantic owner. Its architecture guide describes the core source_to_def idea as a recursive lookup:

  1. find a containing syntax item whose semantic definition is known;
  2. resolve that parent definition;
  3. ask the parent for its semantic children;
  4. map those children back to source;
  5. select the child whose source matches the node.

The direction feels inverted at first. Why not store a direct pointer from every syntax node to a definition?

Because syntax is cheap, local, and replaceable after an edit. Semantic definitions are derived from crate-wide facts. Keeping them separate lets the syntax tree remain a value and lets incremental queries recompute only the affected meaning.

The parent-first search also gives the child a context. A function named parse means little by itself. A parse item inside a resolved module has a stable semantic owner.

4. Resolve references, not only declarations

Most go-to-definition requests start on a reference, not on the declaration itself. The IDE classifies the syntax and asks the semantic model what it resolves to.

For crate::api::ApiClient, resolution uses facts such as:

  • which crate owns this file context;
  • which module contains the expression;
  • what crate means there;
  • which names the api module exports;
  • whether ApiClient is a definition, import, or re-export;
  • which conditional items are active.

Methods add type inference and trait lookup. An operator such as + can navigate to an implementation of Add. The ? operator can involve From conversion. This is why rust-analyzer's go-to-definition feature is better understood as semantic navigation than identifier lookup.

Macros create virtual code with real meaning

The macro in the example expands roughly into:

crate::api::ApiClient::default()

The expanded expression must participate in parsing, name resolution, and type inference. But there is no editable .rs file containing that exact generated expression.

I think about expanded code as a virtual file with a source map:

real call-site token ──expand──> virtual token
real macro definition ──expand──> virtual generated syntax
virtual result         ──map────> original source range

Tokens captured from $ty have an origin at the call site. Tokens written by the macro have an origin around the macro definition. Procedural macros make this mapping harder because generated tokens carry spans, and those spans determine how diagnostics and navigation relate back to source.

Navigation through a macro therefore needs two answers:

  1. What does the token mean in the expanded program?
  2. Which original source range is the best place to show?

If a token has no precise editable origin, the IDE has to choose a useful fallback. This is also why preserving meaningful spans in procedural macros improves much more than error underlines.

Modules are not folders with nicer syntax

The file tree helps rust-analyzer build the module tree, but the two are not identical.

filesystem path  ≠  module identity  ≠  crate identity

mod payments; can resolve to one of Rust's supported module-file layouts. A file can be reached from a crate root only when module declarations connect it. A dependency source file belongs to another crate even if it is visible in the editor. Generated code and macro expansions have file identities without ordinary filesystem paths.

For go to definition, the important structure is the semantic module tree derived from the crate graph and active configuration. Treating directories as namespaces gives plausible but wrong answers around multiple targets, build scripts, integration tests, and conditional modules.

Re-exports make “the definition” a product choice

Suppose a library intentionally exposes:

pub use internal::v2::Client;

There are now at least two useful destinations:

  • the public re-export that explains the supported API;
  • the original struct that explains the implementation.

Go to declaration and go to definition can make different choices here. Navigation may also skip a blanket standard-library implementation when a more specific destination is useful. The semantic model supplies possible identities; the IDE feature still applies a product policy to choose the navigation target.

This distinction matters when debugging. A correct resolver can still produce an unhelpful jump if the final target-selection policy is wrong.

Why edits do not require full recomputation

rust-analyzer holds source and project structure as input and derives semantic facts on demand. An edit replaces syntax for the changed file and invalidates dependent queries. Unrelated crate facts can remain reusable.

The syntax/semantics split is what makes this practical:

  • syntax answers what text structure exists here;
  • the crate and module model answers where the node lives;
  • name resolution answers which definition a name denotes;
  • type inference answers method and associated-item questions;
  • source maps answer where the result came from.

I avoid saying that go to definition “walks an AST until it finds the name.” The expensive and interesting work is not tree walking. It is moving between these identities without losing context.

A debugging checklist for wrong navigation

When navigation fails, I test the layers in order:

  1. Project loading: Is the correct Cargo target and crate graph loaded?
  2. Configuration: Are the expected features, target, and cfg flags active?
  3. Syntax: Is the code parsed as the construct I think it is?
  4. Macro expansion: Does the macro expand, and do captured tokens keep useful origins?
  5. Name resolution: Does the path resolve in this module and hygiene context?
  6. Type inference: Is enough type information available for a method or associated item?
  7. Source mapping: Does the semantic result map to an original editable range?
  8. Target policy: Is the IDE choosing the declaration, re-export, impl, or underlying definition I expect?

This checklist is more effective than beginning inside the LSP handler. The protocol usually transports a location request and returns locations. The hard problem lives in the semantic layers beneath it.

What I learned from compiler-tooling work

The useful lesson is larger than one IDE feature. Syntax and meaning should not be fused too early, but every derived fact needs a path back to user-visible evidence.

That pattern appears in compilers, data systems, and developer tools: preserve the origin, compute a stable internal identity, and make the reverse mapping explicit. When the reverse mapping is missing, the system may know the answer and still be unable to explain it.

The examples here are deliberately small and do not describe private project code. They reflect the same kind of source-to-semantics boundary I have had to understand when working in Rust tooling.

Further reading

Go to definition is fast because the architecture keeps layers separate. It is accurate because the request carries context through them. And it is useful because the final semantic identity can still be mapped back to the source a developer can inspect.