| Age | Commit message (Collapse) | Author |
|
# Conflicts:
# fs/coredump.c
# fs/f2fs/f2fs.h
# fs/fuse/dax.c
# fs/xfs/libxfs/xfs_btree.c
|
|
https://git.kernel.org/pub/scm/linux/kernel/git/vfs/vfs.git
# Conflicts:
# fs/smb/server/smb2pdu.c
# fs/smb/server/vfs.c
# fs/smb/server/vfs.h
|
|
page-based ->private flag helpers are used in the compression path, where
large folios are not enabled. They can use folio versions with
page_folio(). The two remaining users in data.c and segment.c can use
fio->folio instead of fio->page (two are in a union).
Drop page-based helpers after the conversion and rename
PAGE_PRIVATE_{GET,SET,CLEAR}_FUNC() and the PAGE_PRIVATE_* flags to
F2FS_FOLIO_PRIVATE_* to match. Convert the folio/page union from
f2fs_io_info union to folio only, since no page user is left.
The folio helpers do a plain read-modify-write where the page ones used
set_bit()/clear_bit(). It is fine because the converted code either holds
folio lock or, in f2fs_compress_write_end_io(), matches what the
non-compressed code does in f2fs_write_end_bio().
Link: https://lore.kernel.org/20260920-remove-pg_private-v5-7-bb68b6a21869@nvidia.com
Signed-off-by: Zi Yan <ziy@nvidia.com>
Co-developed-by: David Hildenbrand (Arm) <david@kernel.org>
Signed-off-by: David Hildenbrand (Arm) <david@kernel.org>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Suggested-by: Tal Zussman <tz2294@columbia.edu>
Reviewed-by: Tal Zussman <tz2294@columbia.edu>
Reviewed-by: Lance Yang <lance.yang@linux.dev>
Reviewed-by: Chao Yu <chao@kernel.org>
Assisted-by: LLM
Cc: Jaegeuk Kim <jaegeuk@kernel.org>
|
|
f2fs sets its PAGE_PRIVATE_* flags in page->private and checking
page->private != NULL is equivalent to checking PG_private. Change
PagePrivate() to page_private(). Meanwhile, in set_page_private_##name(),
page->private is first set to 0/NULL before an PAGE_PRIVATE_* flag is set,
but it can cause confusion when PG_private is removed and
page->private != NULL is used instead. Change it to initialize
page->private to PAGE_PRIVATE_NOT_POINTER instead and retain the original
semantics.
It prepares for a future commit that removes PG_private.
No functional change intended.
Link: https://lore.kernel.org/20260920-remove-pg_private-v5-6-bb68b6a21869@nvidia.com
Signed-off-by: Zi Yan <ziy@nvidia.com>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Acked-by: Usama Arif <usama.arif@linux.dev>
Acked-by: Chao Yu <chao@kernel.org>
Reviewed-by: David Hildenbrand (Arm) <david@kernel.org>
Assisted-by: LLM
Cc: Jaegeuk Kim <jaegeuk@kernel.org>
|
|
A malformed F2FS image can enable F2FS_FEATURE_DEVICE_ALIAS without
providing a multi-device configuration. For a regular single-device
filesystem, f2fs_scan_devices() returns without allocating sbi->devs.
In this state, f2fs_dev_is_alloc_blocked() passes the device alias
feature check and dereferences FDEV(0), resulting in a NULL pointer
dereference during segment allocation.
Device aliasing requires at least one secondary device. Reject
superblocks that enable device aliasing without entries for both the
main and secondary devices in sanity_check_raw_super().
Fixes: eae3faf210bd ("f2fs: support dynamic reserve/release for device aliasing")
Reported-by: syzbot+ae5b8eb92ed40411ce16@syzkaller.appspotmail.com
Closes: https://syzkaller.appspot.com/bug?extid=ae5b8eb92ed40411ce16
Signed-off-by: Seongjae Jeong <jsjlee1020@gmail.com>
Reviewed-by: Chao Yu <chao@kernel.org>
Signed-off-by: Jaegeuk Kim <jaegeuk@kernel.org>
|
|
The newly added label causes a warning when QUOTA is turned off:
fs/f2fs/super.c: In function '__f2fs_remount':
fs/f2fs/super.c:3142:1: error: label 'restore_holder' defined but not used [-Werror=unused-label]
3142 | restore_holder:
| ^~~~~~~~~~~~~~
Fixes: 3f140a60c92a ("f2fs: quota: fix stale lock holder on remount failure")
Signed-off-by: Arnd Bergmann <arnd@arndb.de>
Acked-by: Randy Dunlap <rdunlap@infradead.org>
Tested-by: Randy Dunlap <rdunlap@infradead.org>
Reviewed-by: Chao Yu <chao@kernel.org>
Signed-off-by: Jaegeuk Kim <jaegeuk@kernel.org>
|
|
Now that everything takes a mnt_idmap as const store a const pointer in
struct vfsmount, struct mount_kattr and struct kstatmount and return one
from mnt_idmap(), file_mnt_idmap() and ovl_upper_mnt_idmap(). Finally,
also make mnt_idmap_get() return a const pointer. Also convert the
remaining local variables that are initialized from the accessors.
alloc_mnt_idmap() keeps returning a non-const pointer. It is the only
place where an idmapping is actually written to.
Link: https://patch.msgid.link/20260901-work-idmap-const-v1-26-54ccd48e100b@kernel.org
Reviewed-by: Jan Kara <jack@suse.cz>
Reviewed-by: Seth Forshee <sforshee@kernel.org>
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
|
|
Convert to const struct mnt_idmap.
A mount's idmapping is immutable. The only thing that is allowed to be
modified afterwards is the reference count and that is hidden behind
mnt_idmap_get() and mnt_idmap_put(). Everything else only ever reads
from the idmapping. This is the same model that struct cred uses and the
idmapping is also rather sensitive.
So make the idmap argument const wherever we can. The conversion is done
from the bottom up so callers can continue to pass a non-const pointer
to a const parameter until the conversion is finished.
No functional changes.
Link: https://patch.msgid.link/20260901-work-idmap-const-v1-24-54ccd48e100b@kernel.org
Acked-by: Paul Moore <paul@paul-moore.com>
Reviewed-by: Jan Kara <jack@suse.cz>
Reviewed-by: Seth Forshee <sforshee@kernel.org>
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
|
|
Convert to const struct mnt_idmap.
A mount's idmapping is immutable. The only thing that is allowed to be
modified afterwards is the reference count and that is hidden behind
mnt_idmap_get() and mnt_idmap_put(). Everything else only ever reads
from the idmapping. This is the same model that struct cred uses and the
idmapping is also rather sensitive.
So make the idmap argument const wherever we can. The conversion is done
from the bottom up so callers can continue to pass a non-const pointer
to a const parameter until the conversion is finished.
No functional changes.
Link: https://patch.msgid.link/20260901-work-idmap-const-v1-23-54ccd48e100b@kernel.org
Reviewed-by: Jan Kara <jack@suse.cz>
Reviewed-by: Seth Forshee <sforshee@kernel.org>
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
|
|
Convert to const struct mnt_idmap.
A mount's idmapping is immutable. The only thing that is allowed to be
modified afterwards is the reference count and that is hidden behind
mnt_idmap_get() and mnt_idmap_put(). Everything else only ever reads
from the idmapping. This is the same model that struct cred uses and the
idmapping is also rather sensitive.
So make the idmap argument const wherever we can. The conversion is done
from the bottom up so callers can continue to pass a non-const pointer
to a const parameter until the conversion is finished.
No functional changes.
Link: https://patch.msgid.link/20260901-work-idmap-const-v1-22-54ccd48e100b@kernel.org
Reviewed-by: Jan Kara <jack@suse.cz>
Reviewed-by: Seth Forshee <sforshee@kernel.org>
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
|
|
Convert to const struct mnt_idmap.
A mount's idmapping is immutable. The only thing that is allowed to be
modified afterwards is the reference count and that is hidden behind
mnt_idmap_get() and mnt_idmap_put(). Everything else only ever reads
from the idmapping. This is the same model that struct cred uses and the
idmapping is also rather sensitive.
So make the idmap argument const wherever we can. The conversion is done
from the bottom up so callers can continue to pass a non-const pointer
to a const parameter until the conversion is finished.
No functional changes.
Link: https://patch.msgid.link/20260901-work-idmap-const-v1-21-54ccd48e100b@kernel.org
Reviewed-by: Jan Kara <jack@suse.cz>
Reviewed-by: Seth Forshee <sforshee@kernel.org>
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
|
|
Convert to const struct mnt_idmap.
A mount's idmapping is immutable. The only thing that is allowed to be
modified afterwards is the reference count and that is hidden behind
mnt_idmap_get() and mnt_idmap_put(). Everything else only ever reads
from the idmapping. This is the same model that struct cred uses and the
idmapping is also rather sensitive.
So make the idmap argument const wherever we can. The conversion is done
from the bottom up so callers can continue to pass a non-const pointer
to a const parameter until the conversion is finished.
No functional changes.
Link: https://patch.msgid.link/20260901-work-idmap-const-v1-20-54ccd48e100b@kernel.org
Reviewed-by: Jan Kara <jack@suse.cz>
Reviewed-by: Seth Forshee <sforshee@kernel.org>
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
|
|
Convert to const struct mnt_idmap.
A mount's idmapping is immutable. The only thing that is allowed to be
modified afterwards is the reference count and that is hidden behind
mnt_idmap_get() and mnt_idmap_put(). Everything else only ever reads
from the idmapping. This is the same model that struct cred uses and the
idmapping is also rather sensitive.
So make the idmap argument const wherever we can. The conversion is done
from the bottom up so callers can continue to pass a non-const pointer
to a const parameter until the conversion is finished.
No functional changes.
Link: https://patch.msgid.link/20260901-work-idmap-const-v1-19-54ccd48e100b@kernel.org
Reviewed-by: Jan Kara <jack@suse.cz>
Reviewed-by: Seth Forshee <sforshee@kernel.org>
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
|
|
Convert to const struct mnt_idmap.
A mount's idmapping is immutable. The only thing that is allowed to be
modified afterwards is the reference count and that is hidden behind
mnt_idmap_get() and mnt_idmap_put(). Everything else only ever reads
from the idmapping. This is the same model that struct cred uses and the
idmapping is also rather sensitive.
So make the idmap argument const wherever we can. The conversion is done
from the bottom up so callers can continue to pass a non-const pointer
to a const parameter until the conversion is finished.
No functional changes.
Link: https://patch.msgid.link/20260901-work-idmap-const-v1-18-54ccd48e100b@kernel.org
Acked-by: Paul Moore <paul@paul-moore.com>
Reviewed-by: Jan Kara <jack@suse.cz>
Reviewed-by: Seth Forshee <sforshee@kernel.org>
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
|
|
Convert to const struct mnt_idmap.
A mount's idmapping is immutable. The only thing that is allowed to be
modified afterwards is the reference count and that is hidden behind
mnt_idmap_get() and mnt_idmap_put(). Everything else only ever reads
from the idmapping. This is the same model that struct cred uses and the
idmapping is also rather sensitive.
So make the idmap argument const wherever we can. The conversion is done
from the bottom up so callers can continue to pass a non-const pointer
to a const parameter until the conversion is finished.
No functional changes.
Link: https://patch.msgid.link/20260901-work-idmap-const-v1-17-54ccd48e100b@kernel.org
Acked-by: Paul Moore <paul@paul-moore.com>
Reviewed-by: Jan Kara <jack@suse.cz>
Reviewed-by: Seth Forshee <sforshee@kernel.org>
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
|
|
Convert to const struct mnt_idmap.
A mount's idmapping is immutable. The only thing that is allowed to be
modified afterwards is the reference count and that is hidden behind
mnt_idmap_get() and mnt_idmap_put(). Everything else only ever reads
from the idmapping. This is the same model that struct cred uses and the
idmapping is also rather sensitive.
So make the idmap argument const wherever we can. The conversion is done
from the bottom up so callers can continue to pass a non-const pointer
to a const parameter until the conversion is finished.
No functional changes.
Link: https://patch.msgid.link/20260901-work-idmap-const-v1-15-54ccd48e100b@kernel.org
Reviewed-by: Jan Kara <jack@suse.cz>
Reviewed-by: Seth Forshee <sforshee@kernel.org>
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
|
|
Convert to const struct mnt_idmap.
A mount's idmapping is immutable. The only thing that is allowed to be
modified afterwards is the reference count and that is hidden behind
mnt_idmap_get() and mnt_idmap_put(). Everything else only ever reads
from the idmapping. This is the same model that struct cred uses and the
idmapping is also rather sensitive.
So make the idmap argument const wherever we can. The conversion is done
from the bottom up so callers can continue to pass a non-const pointer
to a const parameter until the conversion is finished.
No functional changes.
Link: https://patch.msgid.link/20260901-work-idmap-const-v1-14-54ccd48e100b@kernel.org
Reviewed-by: Jan Kara <jack@suse.cz>
Reviewed-by: Seth Forshee <sforshee@kernel.org>
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
|
|
Convert to const struct mnt_idmap.
A mount's idmapping is immutable. The only thing that is allowed to be
modified afterwards is the reference count and that is hidden behind
mnt_idmap_get() and mnt_idmap_put(). Everything else only ever reads
from the idmapping. This is the same model that struct cred uses and the
idmapping is also rather sensitive.
So make the idmap argument const wherever we can. The conversion is done
from the bottom up so callers can continue to pass a non-const pointer
to a const parameter until the conversion is finished.
No functional changes.
Link: https://patch.msgid.link/20260901-work-idmap-const-v1-13-54ccd48e100b@kernel.org
Acked-by: Paul Moore <paul@paul-moore.com>
Reviewed-by: Jan Kara <jack@suse.cz>
Reviewed-by: Seth Forshee <sforshee@kernel.org>
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
|
|
Convert to const struct mnt_idmap.
A mount's idmapping is immutable. The only thing that is allowed to be
modified afterwards is the reference count and that is hidden behind
mnt_idmap_get() and mnt_idmap_put(). Everything else only ever reads
from the idmapping. This is the same model that struct cred uses and the
idmapping is also rather sensitive.
So make the idmap argument const wherever we can. The conversion is done
from the bottom up so callers can continue to pass a non-const pointer
to a const parameter until the conversion is finished.
No functional changes.
Link: https://patch.msgid.link/20260901-work-idmap-const-v1-10-54ccd48e100b@kernel.org
Reviewed-by: Jan Kara <jack@suse.cz>
Reviewed-by: Seth Forshee <sforshee@kernel.org>
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
|
|
Convert to const struct mnt_idmap.
A mount's idmapping is immutable. The only thing that is allowed to be
modified afterwards is the reference count and that is hidden behind
mnt_idmap_get() and mnt_idmap_put(). Everything else only ever reads
from the idmapping. This is the same model that struct cred uses and the
idmapping is also rather sensitive.
So make the idmap argument const wherever we can. The conversion is done
from the bottom up so callers can continue to pass a non-const pointer
to a const parameter until the conversion is finished.
No functional changes.
Link: https://patch.msgid.link/20260901-work-idmap-const-v1-8-54ccd48e100b@kernel.org
Reviewed-by: Jan Kara <jack@suse.cz>
Reviewed-by: Seth Forshee <sforshee@kernel.org>
Signed-off-by: Christian Brauner (Amutable) <brauner@kernel.org>
|
|
The skipped count does not indicate the number of pages anymore.
Signed-off-by: Chao Yu <chao@kernel.org>
Signed-off-by: Jaegeuk Kim <jaegeuk@kernel.org>
|
|
The count does not indicate the number of pages anymore.
Signed-off-by: Chao Yu <chao@kernel.org>
Signed-off-by: Jaegeuk Kim <jaegeuk@kernel.org>
|
|
This patch exposes per-cache memory usage (entry struct memory vs.
cached block buffer memory) for meta, node, and compress caches under
the memory section in debugfs:
- meta entry: xxx KB, meta cache: xxx KB
- node entry: xxx KB, node cache: xxx KB
- compress entry: xxx KB, compress cache: xxx KB
Reviewed-by: Zhiguo Niu <zhiguo.niu@unisoc.com>
Signed-off-by: Chao Yu <chao@kernel.org>
Signed-off-by: Jaegeuk Kim <jaegeuk@kernel.org>
|
|
This patch introduces ftrace tracepoints to observe and profile metadata
cache operations:
- trace_f2fs_cache_set_dirty to trace marking a cached block dirty
- trace_f2fs_write_cache to trace single block writeback submission
- trace_f2fs_write_caches to tracesbatch writeback and sync sessions
Reviewed-by: Zhiguo Niu <zhiguo.niu@unisoc.com>
Signed-off-by: Chao Yu <chao@kernel.org>
Signed-off-by: Jaegeuk Kim <jaegeuk@kernel.org>
|
|
This patch adds fault injection support for metadata cache allocations
to improve error-path test coverage.
- it integrates entry allocations with FAULT_KMALLOC
Reviewed-by: Zhiguo Niu <zhiguo.niu@unisoc.com>
Signed-off-by: Chao Yu <chao@kernel.org>
Signed-off-by: Jaegeuk Kim <jaegeuk@kernel.org>
|
|
This patch migrates compressed cluster caching from the fake VFS inode
page cache (sbi->compress_inode) to compress cache (sbi->compress_blocks).
It converts compression caching and decompression paths to use
compress cache APIs, uses entry->ino for per-inode invalidation, and
removes sbi->compress_inode.
Reviewed-by: Zhiguo Niu <zhiguo.niu@unisoc.com>
Signed-off-by: Chao Yu <chao@kernel.org>
Signed-off-by: Jaegeuk Kim <jaegeuk@kernel.org>
|
|
This patch introduces compress_blocks in f2fs_sb_info structure, initializes
and destroys the compress cache during filesystem mount and unmount.
It adds an ino union field in struct f2fs_cached_block (sharing space
with writeback linkage for zero memory overhead) to track per-inode
cached blocks, and registers compress cache into the memory shrinker.
Reviewed-by: Zhiguo Niu <zhiguo.niu@unisoc.com>
Signed-off-by: Chao Yu <chao@kernel.org>
Signed-off-by: Jaegeuk Kim <jaegeuk@kernel.org>
|
|
This patch migrates F2FS node block caching from the fake VFS inode
page cache (sbi->node_inode) to node cache (sbi->node_blocks).
It updates node related helpers to use node cache APIs and reference to
struct f2fs_cached_block, and removes sbi->node_inode, especially, unifies
inline and regular dentry block handling via struct f2fs_dentry_block_ref.
Reviewed-by: Zhiguo Niu <zhiguo.niu@unisoc.com>
Signed-off-by: Chao Yu <chao@kernel.org>
Signed-off-by: Jaegeuk Kim <jaegeuk@kernel.org>
|
|
This patch introduces node_blocks in f2fs_sb_info structure, initializes
and destroys the node cache during filesystem mount and unmount.
It also introduces helper wrappers for node cache operation, and registers
node cache into the memory shrinker and writeback kthread.
Reviewed-by: Zhiguo Niu <zhiguo.niu@unisoc.com>
Signed-off-by: Chao Yu <chao@kernel.org>
Signed-off-by: Jaegeuk Kim <jaegeuk@kernel.org>
|
|
This patch migrates F2FS meta block caching from the fake VFS inode
page cache (sbi->meta_inode) to meta cache (sbi->meta_blocks).
It converts CP, SIT, NAT, SSA, recovery, and GC metadata I/O paths to
operate on struct f2fs_cached_block instead of folio, and removes
sbi->meta_inode.
Reviewed-by: Zhiguo Niu <zhiguo.niu@unisoc.com>
Signed-off-by: Chao Yu <chao@kernel.org>
Signed-off-by: Jaegeuk Kim <jaegeuk@kernel.org>
|
|
This patch introduces a background writeback kthread (f2fs_writeback-x:y)
to periodically flush dirty metadata cache entries with a default
interval of 5 seconds.
It manages thread lifecycle across mount, unmount, and remount (rw/ro)
transitions.
It introduces a sysfs entry /sys/fs/f2fs/<disk>/cache_wb_interval to control
writeback interval.
Reviewed-by: Zhiguo Niu <zhiguo.niu@unisoc.com>
Signed-off-by: Chao Yu <chao@kernel.org>
Signed-off-by: Jaegeuk Kim <jaegeuk@kernel.org>
|
|
This patch integrates the metadata cache into the F2FS memory shrinker
subsystem to reclaim clean, unreferenced cached blocks under memory
pressure.
It implements f2fs_shrink_cache() using a 3-phase cache reclamin method:
1. isolate clean entries from lru list
2. truncate from radix tree under lock
3. splice un-reclaimed entries back
And hooks the new interface into f2fs_shrink_count() and f2fs_shrink_scan().
Reviewed-by: Zhiguo Niu <zhiguo.niu@unisoc.com>
Signed-off-by: Chao Yu <chao@kernel.org>
Signed-off-by: Jaegeuk Kim <jaegeuk@kernel.org>
|
|
This patch introduces meta_blocks in f2fs_sb_info structure, initializes
and destroys the meta cache during filesystem mount and unmount.
It also introduces helper wrappers for meta cache operation.
Reviewed-by: Zhiguo Niu <zhiguo.niu@unisoc.com>
Signed-off-by: Chao Yu <chao@kernel.org>
Signed-off-by: Jaegeuk Kim <jaegeuk@kernel.org>
|
|
This patch introduces the core metadata block caching infrastructure to
manage f2fs metadata independently of the page cache.
It implements:
- core cache APIs: get, create, put, drop, backed by a radix tree and
a single global LRU list.
- support multiple status of cached block: LOCKED, UPTODATE, DIRTY,
WRITEBACK, INLINE.
- internal bio based read/write helpers with adjacent block vector merging.
Reviewed-by: Zhiguo Niu <zhiguo.niu@unisoc.com>
Signed-off-by: Chao Yu <chao@kernel.org>
Signed-off-by: Jaegeuk Kim <jaegeuk@kernel.org>
|
|
fsstress reports a kernel BUG in f2fs_truncate_partial_cluster():
kernel BUG at fs/f2fs/compress.c:1238!
RIP: 0010:f2fs_truncate_partial_cluster+0x292/0x2a0
Call Trace:
<TASK>
f2fs_truncate+0xf6/0x210
f2fs_setattr+0x6b7/0x770
notify_change+0x33b/0x520
do_truncate+0xc2/0xf0
vfs_truncate+0x153/0x1d0
ksys_truncate+0x78/0xd0
__x64_sys_truncate+0x16/0x20
do_syscall_64+0xbe/0x540
The root cause is that Thread A (truncate) and Thread B (background
writeback or fsync) can race as follows:
Thread A Thread B
- f2fs_setattr
- f2fs_truncate
- f2fs_truncate_blocks
- f2fs_truncate_partial_cluster
- f2fs_is_compressed_cluster
return 1
- f2fs_write_cache_pages
- f2fs_write_multi_pages
- f2fs_write_raw_pages
- f2fs_write_single_data_page
dn.data_blkaddr != COMPRESS_ADDR
(cluster converted to normal)
- f2fs_prepare_compress_overwrite
- f2fs_is_compressed_cluster
return 0
- return 0
- f2fs_bug_on(sbi, err == 0): BUG!
Writeback path does not acquire i_gc_rwsem or filemap_invalidate_lock.
When a compressed cluster fails compression during writeback, it is
overwritten with raw data blocks. If Thread A checked
f2fs_is_compressed_cluster() before the conversion, but calls
f2fs_prepare_compress_overwrite() after the conversion,
f2fs_prepare_compress_overwrite() returns 0 because the cluster is no
longer a compressed cluster.
To fix this, remove the f2fs_bug_on() and retry checking the cluster status
when f2fs_prepare_compress_overwrite() returns 0, so that it can fall back
to f2fs_do_truncate_blocks() to handle it as a normal cluster.
Cc: stable@kernel.org
Fixes: 3265d3db1f16 ("f2fs: support partial truncation on compressed inode")
Reviewed-by: Zhiguo Niu <zhiguo.niu@unisoc.com>
Signed-off-by: Chao Yu <chao@kernel.org>
Signed-off-by: Jaegeuk Kim <jaegeuk@kernel.org>
|
|
Otherwise, it may cause NULL pointer dereference in race condition:
Thread A Thread B
- f2fs_setattr
- f2fs_truncate
- f2fs_truncate_blocks
- f2fs_do_truncate_blocks
- f2fs_truncate_inode_blocks
- truncate_dnode
- truncate_node
- invalidate_mapping_pages
- folio->mapping = NULL
- folio_lock
- is_node_folio
- F2FS_F_SB(folio): dereference on folio->mapping
Cc: stable@kernel.org
Fixes: 019a8912425e ("f2fs: introduce is_{meta,node}_folio")
Reported-by: Zhiguo Niu <zhiguo.niu@unisoc.com>
Closes: https://lore.kernel.org/linux-f2fs-devel/c5d29b31-e764-4782-9639-be84b1b2bfc2@kernel.org
Signed-off-by: Chao Yu <chao@kernel.org>
Reviewed-by: Zhiguo Niu <zhiguo.niu@unisoc.com>
Signed-off-by: Jaegeuk Kim <jaegeuk@kernel.org>
|
|
In gang-lookup loops such as last_fsync_dnode(), f2fs_sync_node_pages(),
and f2fs_fsync_node_pages(), candidate dirty node folios returned by
filemap_get_folios_tag() are inspected and filtered before acquiring
their folio locks:
- IS_DNODE
- is_cold_node
- ino_of_node
- ofs_of_node (called by IS_DNODE)
These helpers previously relied on F2FS_F_SB(folio) which dereferences
folio->mapping to obtain sbi. This can suffer a race condition with
concurrent node truncation, causing a NULL pointer dereference panic:
Thread A Thread B
- f2fs_sync_node_pages
- filemap_get_folios_tag
- truncate_node
- invalidate_mapping_pages
- filemap_remove_folio
- folio->mapping = NULL
- IS_DNODE
- F2FS_F_SB
- folio->mapping->host (panic)
To resolve this race condition, pass struct f2fs_sb_info *sbi explicitly
to these helpers.
All these helpers fundamentally rely on F2FS_NODE_FOOTER(), which
dynamically locates the node footer at:
folio_address(folio) + F2FS_BLKSIZE(sbi) - sizeof(struct node_footer).
Consequently, F2FS_NODE_FOOTER() and its direct sibling footer accessors:
- IS_INODE
- nid_of_node
- cpver_of_node
- next_blkaddr_of_node
are also parameterized with struct f2fs_sb_info *sbi.
In get_dnode_base() and get_dnode_addr(), derive sbi safely via
`inode ? F2FS_I_SB(inode) : F2FS_F_SB(node_folio)` to support callers
passing NULL inode (e.g. is_alive()).
Signed-off-by: Chao Yu <chao@kernel.org>
Signed-off-by: Jaegeuk Kim <jaegeuk@kernel.org>
|
|
For compressed file, compressed write may fail and fall back to raw
write.
-Thread A - Thread B
- f2fs_write_multi_pages - f2fs_down_write(&sbi->cp_rwsem);
- f2fs_write_compressed_pages - ...
- f2fs_trylock_op - ...
- f2fs_down_read_trylock(&sbi->cp_rwsem); - ...
- f2fs_write_raw_pages - ...
Thread B acquires the lock first, which causes Thread A to fail lock
acquisition and fall back to raw-data write.
IPU_FORCE is enabled on small-capacity storage devices (< 16GB),
so f2fs_write_raw_pages() overwrites the original compressed data
in-place.
The in-memory node has been updated with raw-data addresses, while the
node metadata stored on eMMC still remains in compressed state.
If a power-cut occurs before the node is flushed to disk, on-disk
inconsistency arises: disk data is raw, yet metadata treats it as a
compressed cluster, leading to decompression failure.
To eliminate this risk completely, force out-place update for all
write operations on compressed file.
Cc: stable@kernel.org
Fixes: 4c8ff7095bef ("f2fs: support data compression")
Signed-off-by: Jiucheng Xu <jiucheng.xu@amlogic.com>
Reviewed-by: Chao Yu <chao@kernel.org>
Signed-off-by: Jaegeuk Kim <jaegeuk@kernel.org>
|
|
In f2fs_post_evict_inode(), record_bits is assigned with BIT(APPEND_INO)
and BIT(UPDATE_INO) respectively, so the second assignment overwrites
the first one: when both FI_APPEND_WRITE and FI_UPDATE_WRITE are set,
only the UPDATE_INO entry is recorded.
The two flags are not mutually exclusive.Once the APPEND_INO entry is
lost after the inode eviction, the shortcut in f2fs_do_sync_file() can
take the flush_out path in the next fsync(), skipping f2fs_fsync_node_pages().
As a result, the dnodes holding the newly allocated blocks are not
fsync-marked, and after a sudden power loss, the data appended before the
eviction is silently lost.
This restores the previous behavior where the two ino entries were
added independently.
Fixes: 314c9e476ffc ("f2fs: call __add_ino_entry out of the eviction path")
Signed-off-by: Zhiguo Niu <zhiguo.niu@unisoc.com>
Reviewed-by: Chao Yu <chao@kernel.org>
Signed-off-by: Jaegeuk Kim <jaegeuk@kernel.org>
|
|
During checkpoint, f2fs_sync_dirty_inodes() and f2fs_sync_inode_meta()
iterate over dirty inodes in their respective lists. If igrab() fails on
an inode (e.g. because it is in the freeing state), the loop continues
without moving the current inode to the tail of the list. As a result,
subsequent iterations pick the same inode repeatedly, preventing other
ready dirty inodes in the list from making forward progress and leading
to a livelock.
Fix this by moving the current inode to the tail of the list
(list_move_tail(&fi->{dirty_list,gdirty_list}, head)) before attempting
igrab() in both f2fs_sync_dirty_inodes() and f2fs_sync_inode_meta().
Additionally, if igrab() fails, yield the CPU with cond_resched() to
allow the evicting thread to finish eviction. Remove the redundant
f2fs_submit_merged_write() call, since .writepages already submits cached
bios via f2fs_submit_merged_write_cond().
v2:
- Also apply list_move_tail() to f2fs_sync_dirty_inodes().
- Remove redundant f2fs_submit_merged_write() calls from both functions,
keeping only cond_resched().
- Update commit title and description.
Signed-off-by: Daeho Jeong <daehojeong@google.com>
Reviewed-by: Chao Yu <chao@kernel.org>
Signed-off-by: Jaegeuk Kim <jaegeuk@kernel.org>
|
|
The node_change lock serializes block reservation in the PRE_AIO path
against checkpoint preparation, since block reservation can create
dirty node pages and update checkpoint accounting.
However, writes that remain within the inline data area return before
the block reservation path. Thus, they do not call
inc_valid_block_count(), change a node mapping from NULL_ADDR to
NEW_ADDR, or create a dirty node page as a result of block reservation.
They also do not update total_valid_block_count or
alloc_valid_block_count.
The inline path only copies the existing inline data to the data folio,
sets FI_DATA_EXIST, and marks the inode folio for deferred inline data
flushing. FI_DATA_EXIST can dirty inode metadata, but it does not
reserve a block or update the node mapping and checkpoint accounting
that node_change is intended to serialize.
Skip f2fs_map_lock() for writes that fit within MAX_INLINE_DATA. Keep
the existing locking for inline conversion, which can update filesystem
metadata and requires checkpoint serialization.
Signed-off-by: Seongjae Jeong <jsjlee1020@gmail.com>
Reviewed-by: Chao Yu <chao@kernel.org>
Signed-off-by: Jaegeuk Kim <jaegeuk@kernel.org>
|
|
F2FS remount can fail while duplicating a journalled quota file name.
f2fs_remount()
- sbi->umount_lock_holder = current
- kstrdup(F2FS_OPTION(sbi).s_qf_names[i])
: fails
- free previously duplicated names
- return -ENOMEM
The direct return leaves umount_lock_holder pointing at the failed
remount context. A later checkpoint may then wrongly treat another
context as the s_umount lock holder.
Add a restore_holder cleanup path and use it before returning -ENOMEM.
Fixes: eb85c2410d6f ("f2fs: quota: fix to avoid warning in dquot_writeback_dquots()")
Signed-off-by: Jianan Huang <jnhuang95@gmail.com>
Reviewed-by: Chao Yu <chao@kernel.org>
Signed-off-by: Jaegeuk Kim <jaegeuk@kernel.org>
|
|
Block size constants, shift counts, masks, and page-to-block ratios were
globally defined assuming a fixed 4KB block size.
Parameterize F2FS_BLKSIZE, F2FS_BLKSIZE_BITS, F2FS_BLKSIZE_MASK,
F2FS_BLKS_PER_PAGE, and CP_CHKSUM_OFFSET with struct f2fs_sb_info *sbi
to reference sbi->blocksize, sbi->log_blocksize, and runtime masks.
Update call sites across metadata operations, data I/O, file operations,
garbage collection, inline data, and sysfs information.
Signed-off-by: Kelvin Zhang <zhangxp1998@gmail.com>
Reviewed-by: Chao Yu <chao@kernel.org>
Signed-off-by: Jaegeuk Kim <jaegeuk@kernel.org>
|
|
The file indexing tree boundaries separating direct, indirect, and
double-indirect node blocks depend on the number of data block addresses
and node IDs contained in each node block.
Parameterize index boundary macros (NODE_DIR1_BLOCK,
NODE_DIR2_BLOCK, NODE_IND1_BLOCK, NODE_IND2_BLOCK, and NODE_DIND_BLOCK) and
address capacity helpers (DEF_ADDRS_PER_BLOCK, NIDS_PER_BLOCK,
cur_addrs_per_inode, addrs_per_page, and addrs_per_folio) with
struct f2fs_sb_info *sbi.
Update file mapping, block allocation, truncate, garbage collection,
and recovery paths to compute indexing tree offsets from the runtime
geometry.
Signed-off-by: Kelvin Zhang <zhangxp1998@gmail.com>
Reviewed-by: Chao Yu <chao@kernel.org>
Signed-off-by: Jaegeuk Kim <jaegeuk@kernel.org>
|
|
Byte-to-block and block-to-byte conversion helpers previously relied on
global PAGE_SIZE, PAGE_SHIFT, or F2FS_BLKSIZE_BITS.
Parameterize F2FS_BYTES_TO_BLK, F2FS_BLK_TO_BYTES, F2FS_BLK_END_BYTES,
F2FS_BLK_ALIGN, and max_file_blocks with struct f2fs_sb_info *sbi using
sbi->log_blocksize.
Update all call sites across data mapping, file operations, fiemap queries,
fsverity checks, NAT bitmap allocations, and truncate paths.
Signed-off-by: Kelvin Zhang <zhangxp1998@gmail.com>
Reviewed-by: Chao Yu <chao@kernel.org>
Signed-off-by: Jaegeuk Kim <jaegeuk@kernel.org>
|
|
Sector-to-block and block-to-sector conversion helpers previously
assumed a fixed sector count per block derived from 4KB pages.
Parameterize F2FS_LOG_SECTORS_PER_BLOCK, SECTOR_FROM_BLOCK, and
SECTOR_TO_BLOCK with struct f2fs_sb_info *sbi, using
sbi->log_sectors_per_block.
Update all call sites across metadata I/O, data mapping, discard
operations, and zoned block device reporting.
Signed-off-by: Kelvin Zhang <zhangxp1998@gmail.com>
Reviewed-by: Chao Yu <chao@kernel.org>
Signed-off-by: Jaegeuk Kim <jaegeuk@kernel.org>
|
|
The usable capacity of dedicated on-disk extended attribute blocks and
inline xattr regions scales with the filesystem block size.
Parameterize VALID_XATTR_BLOCK_SIZE and MAX_INLINE_XATTR_SIZE to
calculate usable xattr limits dynamically from sbi->blocksize rather than
hardcoding PAGE_SIZE or DEF_ADDRS_PER_INODE.
Update mount option consistency validation for inline_xattr_size to
evaluate allowed boundaries dynamically against the runtime block size.
Signed-off-by: Kelvin Zhang <zhangxp1998@gmail.com>
Reviewed-by: Chao Yu <chao@kernel.org>
Signed-off-by: Jaegeuk Kim <jaegeuk@kernel.org>
|
|
An inode block ends with five i_nid entries followed by a node footer.
Data block address pointers (i_addr[]) precede them. Similarly, direct and
indirect node blocks contain data addresses or node IDs followed by a node
footer at the end of the block.
Describe struct f2fs_inode, struct direct_node, and
struct indirect_node using flexible array members. Compute the locations
of i_nid and node footers dynamically from the filesystem block size.
Introduce F2FS_INODE_NIDS() and F2FS_NODE_FOOTER() helpers to access these
tail fields.
Compute sbi->addrs_per_inode, sbi->addrs_per_block, and
sbi->nids_per_block in init_sb_info(), and update node management, file
mapping, and recovery paths accordingly.
Signed-off-by: Kelvin Zhang <zhangxp1998@gmail.com>
Reviewed-by: Chao Yu <chao@kernel.org>
Signed-off-by: Jaegeuk Kim <jaegeuk@kernel.org>
|
|
An on-disk directory block contains a bitmap, reserved padding, an
array of directory entries (struct f2fs_dir_entry), and matching filename
slots. A fixed compile-time structure couples their offsets to 4KB blocks
and relies on static reserved-space definitions.
Remove struct f2fs_dentry_block and compute region offsets (bitmap
bytes, directory-entry count, and filename slots) dynamically from the
filesystem block size. Access directory blocks through
struct f2fs_dentry_ptr views initialized with the runtime geometry.
Update directory operations, inline dentry handling, and recovery paths
to use the dynamic block layout. No functional change is introduced for
4KB blocks.
Signed-off-by: Kelvin Zhang <zhangxp1998@gmail.com>
Reviewed-by: Chao Yu <chao@kernel.org>
Signed-off-by: Jaegeuk Kim <jaegeuk@kernel.org>
|
|
An on-disk orphan block contains a variable-length array of 32-bit
inode numbers followed by a fixed footer at the end of the block. A
compile-time whole-block structure cannot represent the footer position
when the block size varies at runtime.
Remove struct f2fs_orphan_block, introduce
struct f2fs_orphan_block_footer, and compute sbi->orphans_per_block
dynamically in init_sb_info(). Add helpers to access the inode entry array
and footer from a block buffer.
Update orphan inode recovery, checkpointing, and mount paths to use the
parameterized helpers. This preserves the on-disk format while decoupling
orphan handling from compile-time constants.
Signed-off-by: Kelvin Zhang <zhangxp1998@gmail.com>
Reviewed-by: Chao Yu <chao@kernel.org>
Signed-off-by: Jaegeuk Kim <jaegeuk@kernel.org>
|