143 lines
6.5 KiB
Markdown
143 lines
6.5 KiB
Markdown
# Pre9 LZMA decoded extent cache
|
|
|
|
## Decision
|
|
|
|
Pre7, Pre8, and Pre9 all reproduced the same LZMA read-liveness symptom. The
|
|
manual evidence shows that a single large read completes quickly while a full
|
|
file read made up of successive small reads remains in the LZMA decode path.
|
|
The source path creates and destroys a decoded extent for every mapped read;
|
|
the existing FreeBSD buffer cache only retains compressed block buffers.
|
|
|
|
A cache local to `z_erofs_read_uio()` would not cross the VOP/read boundary
|
|
between successive small user reads. Commit `61f4709` therefore kept one
|
|
decoded extent on each compressed-file vnode. Review found that retained
|
|
memory could then grow with the number of open vnodes. This follow-up moves
|
|
the cache to `struct erofs_mount`: each mount retains at most one decoded
|
|
extent until unmount, so ordinary open-file count cannot expand the cache.
|
|
A single entry can thrash between files and concurrent readers; that is an
|
|
accepted first-stage tradeoff.
|
|
|
|
The cache is deliberately limited to complete, non-partial MicroLZMA extents.
|
|
Partial references keep the existing path because their decoded length and
|
|
reference semantics are different. Other codecs are also unchanged in this
|
|
phase: the measured failure is LZMA-specific, and broadening the change would
|
|
make the validation and regression attribution less precise.
|
|
|
|
## Data and lifecycle
|
|
|
|
`struct erofs_zextent_cache` stores the decoded allocation and the mapping
|
|
identity needed to prove that it can be reused:
|
|
|
|
```c
|
|
struct erofs_zextent_cache {
|
|
void *data;
|
|
erofs_nid_t m_nid;
|
|
erofs_off_t m_pa;
|
|
erofs_off_t m_la;
|
|
uint64_t m_plen;
|
|
uint64_t m_llen;
|
|
unsigned int m_deviceid;
|
|
unsigned int m_flags;
|
|
unsigned char m_algorithmformat;
|
|
};
|
|
```
|
|
|
|
The cache mutex is initialized immediately after allocating
|
|
`struct erofs_mount`. Every mount failure path reaches `erofs_sb_free()`, and
|
|
normal unmount calls the same helper after `vflush()`. There is no cache field
|
|
or cache lifecycle dependency in `struct erofs_node`. The cache key includes
|
|
the inode NID because one mount entry is shared by all regular file vnodes.
|
|
NID zero remains a valid key; `data != NULL` is the validity bit.
|
|
|
|
EROFS is read-only, so a decoded extent does not need write invalidation.
|
|
`vflush()` completes before unmount frees the mount object, and
|
|
`z_erofs_extent_cache_fini()` releases the single mount-owned allocation.
|
|
`EROFS_MAP_META` is explicitly excluded, so metadata-backed tailpacking keeps
|
|
the previous path. `packed_inode` and `metabox_en` are also excluded because
|
|
they are mount-private backing objects, not user file vnodes.
|
|
|
|
## Concurrency
|
|
|
|
Decompression and block reads occur without holding `z_extent_cache_lock`. A reader
|
|
first takes the lock only to compare the complete key and copy a cache hit.
|
|
On a miss, it builds a private decoded extent. It then takes the lock again:
|
|
|
|
1. If another reader published the same key, copy that published extent and
|
|
discard the duplicate private allocation.
|
|
2. Otherwise replace the one cached extent, copy the requested range, unlock,
|
|
and free the old allocation.
|
|
|
|
This permits duplicate construction under concurrent misses but keeps all
|
|
shared pointer access under the mutex. The lock is never held across
|
|
`bread()`, allocation, or decompression, all of which may sleep. The cache
|
|
allocation is never freed while a reader is copying it because replacement,
|
|
hit copying, and pointer clearing are serialized by the same mutex.
|
|
|
|
The helpers are reached only when `want <= MAXPHYS` and contain `KASSERT`
|
|
checks for that contract. Thus the largest lock-held copy is the FreeBSD
|
|
`MAXPHYS` request size, not the 12 MiB on-disk extent cap. Calls that can pass
|
|
more than `MAXPHYS` use the existing uncached copy path. This bounded mutex
|
|
copy is accepted for the first stage; a refcounted immutable entry is deferred
|
|
until runtime evidence shows that this copy is materially contended.
|
|
|
|
## Bounds and unchanged paths
|
|
|
|
The existing mapping sanity checks cap a mapped compressed extent at
|
|
`Z_EROFS_PCLUSTER_MAX_DSIZE` before the read path. `z_erofs_read_extent()` also
|
|
checks the compressed and decoded lengths before allocation. The read path now
|
|
explicitly checks `mapoff` and the extent length before converting them to
|
|
`size_t`; this protects the cache offset arithmetic on platforms where
|
|
`size_t` is narrower than the on-disk fields.
|
|
|
|
Cache use requires all of the following:
|
|
|
|
```text
|
|
compressed mapped extent
|
|
not EROFS_MAP_PARTIAL_REF
|
|
Z_EROFS_COMPRESSION_LZMA
|
|
initialized mount cache
|
|
not `EROFS_MAP_META`
|
|
not a mount-private backing inode
|
|
request length no greater than `MAXPHYS`
|
|
```
|
|
|
|
Fragments, holes, partial references, non-LZMA codecs, uncompressed files,
|
|
metadata reads, and mount-private backing inodes retain their previous code
|
|
paths and error handling.
|
|
|
|
## Rejected alternatives for Pre9
|
|
|
|
- A function-local cache was rejected because it cannot span successive VOP
|
|
reads that caused the observed amplification.
|
|
- A per-vnode cache was rejected after the `61f4709` review: open file count
|
|
could retain one decoded extent per vnode without a system-wide bound.
|
|
- A larger cross-vnode cache was rejected because it would require an eviction
|
|
policy and larger memory accounting. The one-entry per-mount cache is the
|
|
controlled compromise: it has a fixed mount-scoped bound and simple teardown.
|
|
- A Linux page/folio/XArray/workqueue port was rejected because those are not
|
|
FreeBSD vnode/buf primitives and would create an unnecessary compatibility
|
|
layer.
|
|
- A decoder stream pool was deferred: it may reduce allocator overhead but
|
|
does not remove repeated full extent decompression.
|
|
- Changing `MAXPHYS`, changing the disk format, bypassing the buffer cache, or
|
|
adding decoder retries was rejected because none addresses the demonstrated
|
|
decoded-result reuse and each changes unrelated behavior or resource bounds.
|
|
|
|
## Static validation and required runtime matrix
|
|
|
|
This change is static-only in the source phase. The follow-up test agent must
|
|
run the unchanged LZMA fixture against Pre9 and require:
|
|
|
|
1. bounded single reads at 4 KiB, 16 KiB, 64 KiB, and 1 MiB with source-range
|
|
hash equality;
|
|
2. repeated small reads spanning the same extent, with completion and hash
|
|
equality;
|
|
3. complete sequential SHA-256 with recorded exit status and elapsed time;
|
|
4. concurrent reads of the same file and close/reopen reads;
|
|
5. partial-reference, plain, LZ4, DEFLATE, and ZSTD regression coverage;
|
|
6. clean unmount, md detach, and module unload after every case.
|
|
|
|
The automation verdict and DUT verdict must remain separate. A timeout or
|
|
missing final hash remains a failure or blocked result; it must not be promoted
|
|
to PASS because bounded reads succeed.
|