RFA-186 · Case file with fixtures · Case 158 of 694 · Runtime evidence
Why BufRead::fill_buf Repeats Bytes Until consume
fill_buf exposes currently buffered bytes without marking them read. Processing and advancing are separate operations: inspect the slice, remember the accepted length, release its borrow, and call consume exactly once.
- Reviewed
- Rust
- Rust 1.98.1, edition 2024
- Targets
- all targets
- Profiles
- dev, release, test
Direct answer
What this Rust failure means
- Why it happens
- fill_buf only exposes currently available buffered data; the reader advances after the caller explicitly marks accepted bytes with consume.
- First discriminating check
- Call fill_buf twice without consume, record both slices, then repeat while consuming the exact first-slice length between calls.
fill_buf sounds like an instruction to fetch the next block. It is better understood as a request to view the unread block currently available. Viewing it does not advance the reader.
The failing program wraps abcdef in a three-byte BufReader. The first fill_buf returns abc. The second call also returns abc, not def.
Observation and commitment are separate
BufRead::fill_buf returns a borrowed slice of the internal buffer, filling it from the underlying reader only when necessary. BufRead::consume marks a number of those bytes as read.
The low-level protocol is:
fill_buf observe unread bytes
process decide how many bytes were accepted
consume(n) commit that progress
fill_buf observe what remains
The repaired program copies the first slice for its assertion, consumes its length, and then receives the next three bytes.
The split supports parsers that need to look ahead
A parser may need to inspect bytes before deciding whether it has a complete token. If fill_buf advanced automatically, merely checking for a delimiter could lose input.
With the split API, I can search the available slice, consume only a complete prefix, and leave an incomplete suffix for the next decision. This is useful for zero-copy scanning because accepted bytes do not need to be copied into an intermediate buffer.
“Zero-copy” here still includes memory loads. It means the parser can borrow existing storage instead of writing the bytes into another allocation.
The borrow determines the code shape
The returned slice borrows the reader mutably. I cannot call consume while that slice remains in use. A common pattern puts inspection in a smaller scope and keeps only the accepted length:
let accepted = {
let available = reader.fill_buf()?;
inspect(available)
};
reader.consume(accepted);
This is not ceremony added by the borrow checker. It prevents the internal buffer from being advanced or refilled while a reference into it remains active.
If data must survive consumption, I transform or copy only what the next stage owns.
Partial consumption changes the next view
I do not have to consume the whole slice. If fill_buf returns abcdef and the parser accepts two bytes, consume(2) makes the next view begin at c. An implementation may return cdef from existing storage without reading again.
The exact buffer size and refill boundary are not protocol boundaries. BufReader::with_capacity makes a deterministic fixture, but production parsing must work for every possible split.
This is the same rule used for streaming pattern matching: chunk boundaries must not change meaning.
Over-consuming is a logic error
The amount passed to consume must not exceed the unread bytes supplied by fill_buf. The documentation calls exceeding it a logic error. I calculate progress from the exact slice just inspected rather than from a stale counter or expected record size.
I also call consume exactly once. Consuming the same accepted length again skips unprocessed data. Retrying after a downstream error therefore needs a defined commit point.
An empty slice means EOF for that reader
fill_buf returns an empty slice at EOF. A short non-empty slice does not mean EOF; it may reflect buffer capacity, underlying read behaviour, or currently available data.
I process every non-empty slice and stop only when the returned slice is empty. For nonblocking or async designs, readiness and EOF may have different representations, but the same observation-versus-consumption principle remains useful.
Higher-level methods may be a better choice
For ordinary lines, read_line or lines is clearer. For fixed-size records, read_exact may express the success contract better. I use fill_buf when borrowing, delimiter search, or partial consumption produces a real benefit.
Low-level control also means owning limits, error recovery, and progress invariants. A loop that repeatedly calls fill_buf without consuming anything can spin forever on the same bytes.
My buffered-parser checklist
When a BufRead loop repeats or skips data, I check:
- Which call observes bytes and which call commits progress?
- Is the consumed length derived from the current borrowed slice?
- Can the parser accept only a prefix?
- Is
consumeaccidentally missing or called twice? - Does the algorithm work at every buffer split?
- Is an empty slice the only EOF signal being used?
- Does retained data own its bytes before the reader advances?
The core principle reaches beyond Rust: look-ahead and acknowledgment are different state transitions. fill_buf lets me inspect without losing data. Progress happens only when I explicitly call consume.