Files
erofs-freebsd-out-tree/docs/pre9-lzma-cache-implementation.md
T
2026-08-13 10:44:59 +02:00

143 lines
6.5 KiB
Markdown

# Pre9 LZMA decoded extent cache
## Decision
Pre7, Pre8, and Pre9 all reproduced the same LZMA read-liveness symptom. The
manual evidence shows that a single large read completes quickly while a full
file read made up of successive small reads remains in the LZMA decode path.
The source path creates and destroys a decoded extent for every mapped read;
the existing FreeBSD buffer cache only retains compressed block buffers.
A cache local to `z_erofs_read_uio()` would not cross the VOP/read boundary
between successive small user reads. Commit `61f4709` therefore kept one
decoded extent on each compressed-file vnode. Review found that retained
memory could then grow with the number of open vnodes. This follow-up moves
the cache to `struct erofs_mount`: each mount retains at most one decoded
extent until unmount, so ordinary open-file count cannot expand the cache.
A single entry can thrash between files and concurrent readers; that is an
accepted first-stage tradeoff.
The cache is deliberately limited to complete, non-partial MicroLZMA extents.
Partial references keep the existing path because their decoded length and
reference semantics are different. Other codecs are also unchanged in this
phase: the measured failure is LZMA-specific, and broadening the change would
make the validation and regression attribution less precise.
## Data and lifecycle
`struct erofs_zextent_cache` stores the decoded allocation and the mapping
identity needed to prove that it can be reused:
```c
struct erofs_zextent_cache {
void *data;
erofs_nid_t m_nid;
erofs_off_t m_pa;
erofs_off_t m_la;
uint64_t m_plen;
uint64_t m_llen;
unsigned int m_deviceid;
unsigned int m_flags;
unsigned char m_algorithmformat;
};
```
The cache mutex is initialized immediately after allocating
`struct erofs_mount`. Every mount failure path reaches `erofs_sb_free()`, and
normal unmount calls the same helper after `vflush()`. There is no cache field
or cache lifecycle dependency in `struct erofs_node`. The cache key includes
the inode NID because one mount entry is shared by all regular file vnodes.
NID zero remains a valid key; `data != NULL` is the validity bit.
EROFS is read-only, so a decoded extent does not need write invalidation.
`vflush()` completes before unmount frees the mount object, and
`z_erofs_extent_cache_fini()` releases the single mount-owned allocation.
`EROFS_MAP_META` is explicitly excluded, so metadata-backed tailpacking keeps
the previous path. `packed_inode` and `metabox_en` are also excluded because
they are mount-private backing objects, not user file vnodes.
## Concurrency
Decompression and block reads occur without holding `z_extent_cache_lock`. A reader
first takes the lock only to compare the complete key and copy a cache hit.
On a miss, it builds a private decoded extent. It then takes the lock again:
1. If another reader published the same key, copy that published extent and
discard the duplicate private allocation.
2. Otherwise replace the one cached extent, copy the requested range, unlock,
and free the old allocation.
This permits duplicate construction under concurrent misses but keeps all
shared pointer access under the mutex. The lock is never held across
`bread()`, allocation, or decompression, all of which may sleep. The cache
allocation is never freed while a reader is copying it because replacement,
hit copying, and pointer clearing are serialized by the same mutex.
The helpers are reached only when `want <= MAXPHYS` and contain `KASSERT`
checks for that contract. Thus the largest lock-held copy is the FreeBSD
`MAXPHYS` request size, not the 12 MiB on-disk extent cap. Calls that can pass
more than `MAXPHYS` use the existing uncached copy path. This bounded mutex
copy is accepted for the first stage; a refcounted immutable entry is deferred
until runtime evidence shows that this copy is materially contended.
## Bounds and unchanged paths
The existing mapping sanity checks cap a mapped compressed extent at
`Z_EROFS_PCLUSTER_MAX_DSIZE` before the read path. `z_erofs_read_extent()` also
checks the compressed and decoded lengths before allocation. The read path now
explicitly checks `mapoff` and the extent length before converting them to
`size_t`; this protects the cache offset arithmetic on platforms where
`size_t` is narrower than the on-disk fields.
Cache use requires all of the following:
```text
compressed mapped extent
not EROFS_MAP_PARTIAL_REF
Z_EROFS_COMPRESSION_LZMA
initialized mount cache
not `EROFS_MAP_META`
not a mount-private backing inode
request length no greater than `MAXPHYS`
```
Fragments, holes, partial references, non-LZMA codecs, uncompressed files,
metadata reads, and mount-private backing inodes retain their previous code
paths and error handling.
## Rejected alternatives for Pre9
- A function-local cache was rejected because it cannot span successive VOP
reads that caused the observed amplification.
- A per-vnode cache was rejected after the `61f4709` review: open file count
could retain one decoded extent per vnode without a system-wide bound.
- A larger cross-vnode cache was rejected because it would require an eviction
policy and larger memory accounting. The one-entry per-mount cache is the
controlled compromise: it has a fixed mount-scoped bound and simple teardown.
- A Linux page/folio/XArray/workqueue port was rejected because those are not
FreeBSD vnode/buf primitives and would create an unnecessary compatibility
layer.
- A decoder stream pool was deferred: it may reduce allocator overhead but
does not remove repeated full extent decompression.
- Changing `MAXPHYS`, changing the disk format, bypassing the buffer cache, or
adding decoder retries was rejected because none addresses the demonstrated
decoded-result reuse and each changes unrelated behavior or resource bounds.
## Static validation and required runtime matrix
This change is static-only in the source phase. The follow-up test agent must
run the unchanged LZMA fixture against Pre9 and require:
1. bounded single reads at 4 KiB, 16 KiB, 64 KiB, and 1 MiB with source-range
hash equality;
2. repeated small reads spanning the same extent, with completion and hash
equality;
3. complete sequential SHA-256 with recorded exit status and elapsed time;
4. concurrent reads of the same file and close/reopen reads;
5. partial-reference, plain, LZ4, DEFLATE, and ZSTD regression coverage;
6. clean unmount, md detach, and module unload after every case.
The automation verdict and DUT verdict must remain separate. A timeout or
missing final hash remains a failure or blocked result; it must not be promoted
to PASS because bounded reads succeed.