diff options
| author | Mark Brown <broonie@kernel.org> | 2026-09-30 12:48:37 +0100 |
|---|---|---|
| committer | Mark Brown <broonie@kernel.org> | 2026-09-30 12:48:37 +0100 |
| commit | c0c20ac78811bbb321b9f1b1b7ee3c00c847e95d (patch) | |
| tree | 66376e23a2e88beaa9ffcfb9420f6ed8885578a4 | |
| parent | 19cb490f7b0c1e5808a81ac385dff32045aa98aa (diff) | |
| parent | bf234c28d9e24e3d6c42a6202e3f92fba514c2ae (diff) | |
| download | linux-next-c0c20ac78811bbb321b9f1b1b7ee3c00c847e95d.tar.gz linux-next-c0c20ac78811bbb321b9f1b1b7ee3c00c847e95d.zip | |
Merge branch 'fs-next' of linux-next
# Conflicts:
# fs/coredump.c
# fs/f2fs/f2fs.h
# fs/fuse/dax.c
# fs/xfs/libxfs/xfs_btree.c
861 files changed, 27227 insertions, 15840 deletions
@@ -58,7 +58,7 @@ S: Longford, Ireland S: Sydney, Australia N: Tigran A. Aivazian -E: tigran@aivazian.fsnet.co.uk +E: aivazian.tigran@gmail.com W: http://www.moses.uklinux.net/patches D: BFS filesystem D: Intel IA32 CPU microcode update support diff --git a/Documentation/ABI/testing/sysfs-fs-f2fs b/Documentation/ABI/testing/sysfs-fs-f2fs index 0cebc89799dd..c4746c416ac2 100644 --- a/Documentation/ABI/testing/sysfs-fs-f2fs +++ b/Documentation/ABI/testing/sysfs-fs-f2fs @@ -1020,3 +1020,10 @@ Contact: "Daeho Jeong" <daehojeong@google.com> Description: This is a read-only entry to show the upper bound section number for pinned files. Pinned files will only be allocated within sections 0 to pinned_area_max_secno - 1. + +What: /sys/fs/f2fs/<disk>/cache_wb_interval +Date: August 2026 +Contact: "Chao Yu" <chao@kernel.org> +Description: This is a writable entry to control writeback interval of + f2fs_writeback-x:y, the range is [100, 30000], by default the value + is 5000, unit is ms. diff --git a/Documentation/admin-guide/binfmt-misc.rst b/Documentation/admin-guide/binfmt-misc.rst index d26b63a27c25..9e84b877d06d 100644 --- a/Documentation/admin-guide/binfmt-misc.rst +++ b/Documentation/admin-guide/binfmt-misc.rst @@ -19,6 +19,9 @@ To actually register a new binary type, you have to set up a string looking like ``:name:type:offset:magic:mask:interpreter:flags`` (where you can choose the ``:`` upon your needs) and echo it to ``/proc/sys/fs/binfmt_misc/register``. +The first character of the string is its field delimiter and can be any +ASCII punctuation character other than the backslash ``\``. + Here is what the fields mean: - ``name`` diff --git a/Documentation/admin-guide/nfs/pnfs-scsi-server.rst b/Documentation/admin-guide/nfs/pnfs-scsi-server.rst index b202508d281d..a3074450ac15 100644 --- a/Documentation/admin-guide/nfs/pnfs-scsi-server.rst +++ b/Documentation/admin-guide/nfs/pnfs-scsi-server.rst @@ -16,7 +16,7 @@ addition to the MDS. As of now the file system needs to sit directly on the exported LUN, striping or concatenation of LUNs on the MDS and clients is not supported yet. -On a server built with CONFIG_NFSD_SCSI, the pNFS SCSI volume support is +On a server built with CONFIG_NFSD_SCSILAYOUT, the pNFS SCSI volume support is automatically enabled if the file system is exported using the "pnfs" option and the underlying SCSI device support persistent reservations. On the client make sure the kernel has the CONFIG_PNFS_BLOCK option diff --git a/Documentation/filesystems/befs.rst b/Documentation/filesystems/befs.rst index a22f603b2938..c1dbfe9c95d2 100644 --- a/Documentation/filesystems/befs.rst +++ b/Documentation/filesystems/befs.rst @@ -44,8 +44,8 @@ implementation. Which is it, BFS or BEFS? ========================= Be, Inc said, "BeOS Filesystem is officially called BFS, not BeFS". -But Unixware Boot Filesystem is called bfs, too. And they are already in -the kernel. Because of this naming conflict, on Linux the BeOS +But the UnixWare Boot Filesystem is called bfs, too, and it was already +in the kernel. Because of this naming conflict, on Linux the BeOS filesystem is called befs. How to Install diff --git a/Documentation/filesystems/bfs.rst b/Documentation/filesystems/bfs.rst deleted file mode 100644 index ce14b9018807..000000000000 --- a/Documentation/filesystems/bfs.rst +++ /dev/null @@ -1,60 +0,0 @@ -.. SPDX-License-Identifier: GPL-2.0 - -======================== -BFS Filesystem for Linux -======================== - -The BFS filesystem is used by SCO UnixWare OS for the /stand slice, which -usually contains the kernel image and a few other files required for the -boot process. - -In order to access /stand partition under Linux you obviously need to -know the partition number and the kernel must support UnixWare disk slices -(CONFIG_UNIXWARE_DISKLABEL config option). However BFS support does not -depend on having UnixWare disklabel support because one can also mount -BFS filesystem via loopback:: - - # losetup /dev/loop0 stand.img - # mount -t bfs /dev/loop0 /mnt/stand - -where stand.img is a file containing the image of BFS filesystem. -When you have finished using it and umounted you need to also deallocate -/dev/loop0 device by:: - - # losetup -d /dev/loop0 - -You can simplify mounting by just typing:: - - # mount -t bfs -o loop stand.img /mnt/stand - -this will allocate the first available loopback device (and load loop.o -kernel module if necessary) automatically. If the loopback driver is not -loaded automatically, make sure that you have compiled the module and -that modprobe is functioning. Beware that umount will not deallocate -/dev/loopN device if /etc/mtab file on your system is a symbolic link to -/proc/mounts. You will need to do it manually using "-d" switch of -losetup(8). Read losetup(8) manpage for more info. - -To create the BFS image under UnixWare you need to find out first which -slice contains it. The command prtvtoc(1M) is your friend:: - - # prtvtoc /dev/rdsk/c0b0t0d0s0 - -(assuming your root disk is on target=0, lun=0, bus=0, controller=0). Then you -look for the slice with tag "STAND", which is usually slice 10. With this -information you can use dd(1) to create the BFS image:: - - # umount /stand - # dd if=/dev/rdsk/c0b0t0d0sa of=stand.img bs=512 - -Just in case, you can verify that you have done the right thing by checking -the magic number:: - - # od -Ad -tx4 stand.img | more - -The first 4 bytes should be 0x1badface. - -If you have any patches, questions or suggestions regarding this BFS -implementation please contact the author: - -Tigran Aivazian <aivazian.tigran@gmail.com> diff --git a/Documentation/filesystems/exfat.rst b/Documentation/filesystems/exfat.rst new file mode 100644 index 000000000000..ce5c9344a7b2 --- /dev/null +++ b/Documentation/filesystems/exfat.rst @@ -0,0 +1,117 @@ +.. SPDX-License-Identifier: GPL-2.0 + +================================== +The Linux exFAT filesystem driver +================================== + + +.. Table of contents + + - Overview + - Utilities support + - Supported mount options + + +Overview +======== + +exFAT is a filesystem designed for removable storage and other devices that +need to store large files. The Linux exFAT filesystem driver provides read +and write support for exFAT volumes. + +To mount an exFAT volume, use the ``exfat`` filesystem type:: + + mount -t exfat /dev/sdX1 /mnt + + +Utilities support +================= + +The exfatprogs project provides userspace utilities for creating, checking, +repairing, inspecting, and tuning exFAT filesystems. Use exfatprogs when +creating or checking an exFAT filesystem. For example, use ``mkfs.exfat`` +to create a filesystem and ``fsck.exfat`` to check or repair one. + +The project is available at: + + https://github.com/exfatprogs/exfatprogs + + +Supported mount options +======================= + +The exFAT driver supports the following mount options: + +======================= ==================================================== +uid= +gid= Set the owner and group of all files and + directories. The default is the uid and gid of + the process mounting the filesystem. + +umask= Set the permission mask for files and directories. + The default is the umask of the process mounting + the filesystem. + +dmask= Set the permission mask for directories. + +fmask= Set the permission mask for files. + +allow_utime= Control the permission check for changing file + timestamps. Only permission bits 0022 are used. + Permission bit 0020 allows members of the file's + group to change timestamps, and permission bit 0002 + allows other users to change timestamps. The + default is derived from dmask (``~dmask & 0022``). + +iocharset=name Character set used to convert between user-visible + filenames and the UTF-16 character encoding used by + exFAT. The default is + CONFIG_EXFAT_DEFAULT_IOCHARSET, which is ``utf8`` + unless changed at kernel configuration time. Use + ``iocharset=utf8`` for UTF-8 filename handling. + +errors= Specify exFAT behavior on filesystem errors. The + value must be ``panic``, ``continue``, or + ``remount-ro``. These respectively panic, continue + without changing the filesystem, or remount the + filesystem read-only. The default is + ``remount-ro``. + +discard Issue discard/TRIM requests to the block device + when clusters are freed. This is disabled by + default. ``nodiscard`` disables it explicitly. + +keep_last_dots Keep trailing periods in path components during + lookup. Without this option, trailing periods are + stripped. Existing entries with trailing periods + can be accessed when this option is enabled, but + creating new entries with trailing periods is + rejected. + +sys_tz Use the system timezone as the UTC offset when an + exFAT timestamp does not contain a valid timezone + offset. This takes precedence over time_offset. + +time_offset=minutes Set the UTC offset, in minutes, used when an exFAT + timestamp does not contain a valid timezone offset. + Values from -1440 to 1440 are accepted. The default + is 0. This option is ignored when sys_tz is set. + +zero_size_dir Create directories with zero size and without + allocating a cluster. This is disabled by default; + the default behavior allocates a cluster for a new + directory. ``nozero_size_dir`` disables it + explicitly. +======================= ==================================================== + + +Deprecated mount options +------------------------ + +The following options are accepted for compatibility but should not be used: + +``utf8`` + Deprecated. Use ``iocharset=utf8`` instead. + +``debug``, ``namecase=``, ``codepage=`` + Deprecated and ignored by the exFAT driver. diff --git a/Documentation/filesystems/index.rst b/Documentation/filesystems/index.rst index 734a45e51667..1100130ccf0a 100644 --- a/Documentation/filesystems/index.rst +++ b/Documentation/filesystems/index.rst @@ -75,7 +75,6 @@ Documentation for filesystem implementations. autofs autofs-mount-control befs - bfs btrfs ceph coda @@ -87,6 +86,7 @@ Documentation for filesystem implementations. ecryptfs efivarfs erofs + exfat ext2 ext3 ext4/index diff --git a/Documentation/filesystems/locking.rst b/Documentation/filesystems/locking.rst index 844d65eb47a5..6330653287d5 100644 --- a/Documentation/filesystems/locking.rst +++ b/Documentation/filesystems/locking.rst @@ -61,23 +61,23 @@ inode_operations prototypes:: - int (*create) (struct mnt_idmap *, struct inode *,struct dentry *,umode_t); + int (*create) (const struct mnt_idmap *, struct inode *,struct dentry *,umode_t); struct dentry * (*lookup) (struct inode *,struct dentry *, unsigned int); int (*link) (struct dentry *,struct inode *,struct dentry *); int (*unlink) (struct inode *,struct dentry *); - int (*symlink) (struct mnt_idmap *, struct inode *,struct dentry *,const char *); - struct dentry *(*mkdir) (struct mnt_idmap *, struct inode *,struct dentry *,umode_t); + int (*symlink) (const struct mnt_idmap *, struct inode *,struct dentry *,const char *); + struct dentry *(*mkdir) (const struct mnt_idmap *, struct inode *,struct dentry *,umode_t); int (*rmdir) (struct inode *,struct dentry *); - int (*mknod) (struct mnt_idmap *, struct inode *,struct dentry *,umode_t,dev_t); - int (*rename) (struct mnt_idmap *, struct inode *, struct dentry *, + int (*mknod) (const struct mnt_idmap *, struct inode *,struct dentry *,umode_t,dev_t); + int (*rename) (const struct mnt_idmap *, struct inode *, struct dentry *, struct inode *, struct dentry *, unsigned int); int (*readlink) (struct dentry *, char __user *,int); const char *(*get_link) (struct dentry *, struct inode *, struct delayed_call *); void (*truncate) (struct inode *); - int (*permission) (struct mnt_idmap *, struct inode *, int, unsigned int); + int (*permission) (const struct mnt_idmap *, struct inode *, int, unsigned int); struct posix_acl * (*get_inode_acl)(struct inode *, int, bool); - int (*setattr) (struct mnt_idmap *, struct dentry *, struct iattr *); - int (*getattr) (struct mnt_idmap *, const struct path *, struct kstat *, u32, unsigned int); + int (*setattr) (const struct mnt_idmap *, struct dentry *, struct iattr *); + int (*getattr) (const struct mnt_idmap *, const struct path *, struct kstat *, u32, unsigned int); ssize_t (*listxattr) (struct dentry *, char *, size_t); int (*fiemap)(struct inode *, struct fiemap_extent_info *, u64 start, u64 len); void (*update_time)(struct inode *inode, enum fs_update_time type, @@ -86,12 +86,12 @@ prototypes:: int (*atomic_open)(struct inode *, struct dentry *, struct file *, unsigned open_flag, umode_t create_mode); - int (*tmpfile) (struct mnt_idmap *, struct inode *, + int (*tmpfile) (const struct mnt_idmap *, struct inode *, struct file *, umode_t); - int (*fileattr_set)(struct mnt_idmap *idmap, + int (*fileattr_set)(const struct mnt_idmap *idmap, struct dentry *dentry, struct file_kattr *fa); int (*fileattr_get)(struct dentry *dentry, struct file_kattr *fa); - struct posix_acl * (*get_acl)(struct mnt_idmap *, struct dentry *, int); + struct posix_acl * (*get_acl)(const struct mnt_idmap *, struct dentry *, int); struct offset_ctx *(*get_offset_ctx)(struct inode *inode); locking rules: @@ -148,7 +148,7 @@ prototypes:: struct inode *inode, const char *name, void *buffer, size_t size); int (*set)(const struct xattr_handler *handler, - struct mnt_idmap *idmap, + const struct mnt_idmap *idmap, struct dentry *dentry, struct inode *inode, const char *name, const void *buffer, size_t size, int flags); diff --git a/Documentation/filesystems/netfs_library.rst b/Documentation/filesystems/netfs_library.rst index ddd799df6ce3..0c9786ffe192 100644 --- a/Documentation/filesystems/netfs_library.rst +++ b/Documentation/filesystems/netfs_library.rst @@ -195,7 +195,7 @@ structure is defined:: struct inode inode; const struct netfs_request_ops *ops; struct fscache_cookie * cache; - loff_t remote_i_size; + loff_t _remote_i_size; unsigned long flags; ... }; @@ -229,11 +229,14 @@ filesystem: Local caching cookie, or NULL if no caching is enabled. This field does not exist if fscache is disabled. - * ``remote_i_size`` + * ``_remote_i_size`` The size of the file on the server. This differs from inode->i_size if local modifications have been made but not yet written back. + Use netfs_read_remote_i_size() and netfs_write_remote_i_size() to access + this field. Hold inode->i_lock when writing it. + * ``flags`` A set of flags, some of which the filesystem might be interested in: diff --git a/Documentation/filesystems/ntfs3.rst b/Documentation/filesystems/ntfs3.rst index 2b86a9b3a6de..e1d86dcc3316 100644 --- a/Documentation/filesystems/ntfs3.rst +++ b/Documentation/filesystems/ntfs3.rst @@ -109,6 +109,29 @@ this table marked with no it means default is without **no**. Kernel. Not to be confused with NTFS ACLs. The option specified as acl enables support for POSIX ACLs. + * - ads + - Enable access to alternate data streams, i.e. named $DATA + attributes. A stream is addressed by appending a colon and the + stream name to the file name. Reading the pseudo-stream + ``query_streams`` returns the names of the streams attached to a + file, one per line:: + + cat file:query_streams # list stream names + cat file:ads1 # read a stream + touch file:ads2 # create a stream on an existing file + rm file:ads1 # remove a stream + + Creating a file together with a stream in a single call, and + renaming a stream, are not supported. Enabled by default. + + * - nocase + - Perform file name lookups case-insensitively. + + * - delalloc + - Delay block allocation for buffered writes until writeback, + rather than allocating as each write is issued. This lets the + driver allocate larger contiguous runs and reduces fragmentation. + Todo list ========= - Full journaling support over JBD. Currently journal replaying is supported diff --git a/Documentation/filesystems/porting.rst b/Documentation/filesystems/porting.rst index 60880eb0c49d..e666edab789f 100644 --- a/Documentation/filesystems/porting.rst +++ b/Documentation/filesystems/porting.rst @@ -348,7 +348,7 @@ simply of return 1. Note that all actual eviction work is done by caller after As before, clear_inode() must be called exactly once on each call of ->evict_inode() (as it used to be for each call of ->delete_inode()). Unlike before, if you are using inode-associated metadata buffers (i.e. -mark_buffer_dirty_inode()), it's your responsibility to call +mmb_mark_buffer_dirty()), it's your responsibility to call invalidate_inode_buffers() before clear_inode(). NOTE: checking i_nlink in the beginning of ->write_inode() and bailing out @@ -1203,16 +1203,16 @@ will fail-safe. --- -** mandatory** +**mandatory** lookup_one(), lookup_one_unlocked(), lookup_one_positive_unlocked() now take a qstr instead of a name and len. These, not the "one_len" versions, should be used whenever accessing a filesystem from outside -that filesysmtem, through a mount point - which will have a mnt_idmap. +that filesystem, through a mount point - which will have a mnt_idmap. --- -** mandatory** +**mandatory** Functions try_lookup_one_len(), lookup_one_len(), lookup_one_len_unlocked() and lookup_positive_unlocked() have been @@ -1229,7 +1229,7 @@ already been performed such as after vfs_path_parent_lookup() --- -** mandatory** +**mandatory** d_hash_and_lookup() is no longer exported or available outside the VFS. Use try_lookup_noperm() instead. This adds name validation and takes @@ -1370,7 +1370,7 @@ similar. --- -** mandatory** +**mandatory** lock_rename(), lock_rename_child(), unlock_rename() are no longer available. Use start_renaming() or similar. @@ -1409,3 +1409,16 @@ use only if you have no alternative. The .create inode_operation no longer receives the 'excl' arg. It must always assume the file does not already exist. If the filesystem needs to be involved in non-exclusive create, it should provide atomic_open. + +--- + +**mandatory** + +All struct mnt_idmap pointers handed to filesystems are const now. +->create(), ->mkdir(), ->mknod(), ->symlink(), ->rename(), ->setattr(), +->getattr(), ->permission(), ->tmpfile(), ->get_acl(), ->set_acl() and +->fileattr_set() as well as the xattr ->set() handler and the vfs_*() +helpers take a const struct mnt_idmap *. mnt_idmap() and file_mnt_idmap() +return one. The idmapping is immutable so nothing should have modified it +anyway. References are taken and dropped via mnt_idmap_get() and +mnt_idmap_put() as before, both accept a const pointer. diff --git a/Documentation/filesystems/proc.rst b/Documentation/filesystems/proc.rst index c102b62023cd..fc59c98acca1 100644 --- a/Documentation/filesystems/proc.rst +++ b/Documentation/filesystems/proc.rst @@ -1963,6 +1963,10 @@ For example:: $ echo 0x7 > /proc/self/coredump_filter $ ./some_program +If the coredump socket protocol is used a coredump server can select memory +types to include dynamically. See COREDUMP_MEMORY_TYPES in +include/uapi/linux/coredump.h. + 3.5 /proc/<pid>/mountinfo - Information about mounts -------------------------------------------------------- diff --git a/Documentation/filesystems/sharedsubtree.rst b/Documentation/filesystems/sharedsubtree.rst index 8b7dc9159083..8bae6d9a7e04 100644 --- a/Documentation/filesystems/sharedsubtree.rst +++ b/Documentation/filesystems/sharedsubtree.rst @@ -564,8 +564,8 @@ f) Unmount semantics where 'A' is a mount mounted on mount 'B' at dentry 'b'. If mount 'B' is shared, then all most-recently-mounted mounts at dentry - 'b' on mounts that receive propagation from mount 'B' and does not have - sub-mounts within them are unmounted. + 'b' on mounts that receive propagation from mount 'B' are unmounted as + well, if every mount below them is also unmounted. Example: Let's say 'B1', 'B2', 'B3' are shared mounts that propagate to each other. @@ -584,10 +584,18 @@ f) Unmount semantics So all 'C1', 'C2' and 'C3' should be unmounted. - If any of 'C2' or 'C3' has some child mounts, then that mount is not - unmounted, but all other mounts are unmounted. However if 'C1' is told - to be unmounted and 'C1' has some sub-mounts, the umount operation is - failed entirely. + If any of 'C2' or 'C3' has a child mount that cannot be unmounted + then that mount is not unmounted. But all other mounts are unmounted. + A child mount that is itself unmounted by the same unmount + propagation does not keep its parent mounted. However if 'C1' is + supposed to be unmounted and 'C1' has some sub-mounts, the unmount + fails. + + A lazy umount (MNT_DETACH) takes a whole tree. Every mount of the + tree then propagates its unmount from its own parent as described + above, so the mounts that receive propagation lose the corresponding + trees as well. Documentation/filesystems/propagate_umount.txt has the + precise rules, including the ones for locked mounts. g) Clone Namespace diff --git a/Documentation/filesystems/smb/ksmbd.rst b/Documentation/filesystems/smb/ksmbd.rst index 672c5d3892ff..728cfca13205 100644 --- a/Documentation/filesystems/smb/ksmbd.rst +++ b/Documentation/filesystems/smb/ksmbd.rst @@ -82,10 +82,12 @@ Signing Update Supported. Pre-authentication integrity Supported. SMB3 encryption(CCM, GCM) Supported. (CCM/GCM128 and CCM/GCM256 supported) SMB direct(RDMA) Supported. -SMB3 Multi-channel Partially Supported. Planned to implement - replay/retry mechanisms for future. +SMB3 Multi-channel Supported. Receive Side Scaling mode Supported. SMB3.1.1 POSIX extension Supported. +MSDFS Planned for future. +Continuous Availability (CA) Under development. +AD/DC Under development. ACLs Partially Supported. only DACLs available, SACLs (auditing) is planned for the future. For ownership (SIDs) ksmbd generates random subauth @@ -98,8 +100,9 @@ ACLs Partially Supported. only DACLs available, SACLs member. Kerberos Supported. Durable handle v1,v2 Supported. -Persistent handle Planned for future. -SMB2 notify Planned for future. +Persistent handle Under development. +Resilient handle Planned for future. +SMB2 notify Under development. Sparse file support Supported. DCE/RPC support Partially Supported. a few calls(NetShareEnumAll, NetServerGetInfo, SAMR, LSARPC) that are needed @@ -113,7 +116,8 @@ ksmbd/nfsd interoperability Planned for future. The features that ksmbd support are Leases, Notify, ACLs and Share modes. SMB3.1.1 Compression Supported. SMB3.1.1 over QUIC Planned for future. -Signing/Encryption over RDMA Planned for future. +Signing over RDMA Under development. +Encryption over RDMA Supported. SMB3.1.1 GMAC signing support Planned for future. ============================== ================================================= diff --git a/Documentation/filesystems/squashfs.rst b/Documentation/filesystems/squashfs.rst index 45653b3228f9..d6397961d16b 100644 --- a/Documentation/filesystems/squashfs.rst +++ b/Documentation/filesystems/squashfs.rst @@ -177,9 +177,9 @@ or if the compressed block was larger than the uncompressed block. Inodes are packed into the metadata blocks, and are not aligned to block boundaries, therefore inodes overlap compressed blocks. Inodes are identified -by a 48-bit number which encodes the location of the compressed metadata block -containing the inode, and the byte offset into that block where the inode is -placed (<block, offset>). +by a 64-bit number: the upper 48 bits encode the location of the compressed +metadata block containing the inode, and the lower 16 bits give the byte offset +into that block where the inode is placed (<block, offset>). To maximise compression there are different inodes for each file type (regular file, directory, device, etc.), the inode contents and length diff --git a/Documentation/filesystems/vfs.rst b/Documentation/filesystems/vfs.rst index dec7816303c6..1824462a4371 100644 --- a/Documentation/filesystems/vfs.rst +++ b/Documentation/filesystems/vfs.rst @@ -117,7 +117,7 @@ members are defined: const struct fs_parameter_spec *parameters; void (*kill_sb) (struct super_block *); struct module *owner; - struct file_system_type * next; + struct hlist_node list; struct hlist_head fs_supers; struct lock_class_key s_lock_key; @@ -415,33 +415,33 @@ As of kernel 2.6.22, the following members are defined: .. code-block:: c struct inode_operations { - int (*create) (struct mnt_idmap *, struct inode *,struct dentry *, umode_t); + int (*create) (const struct mnt_idmap *, struct inode *,struct dentry *, umode_t); struct dentry * (*lookup) (struct inode *,struct dentry *, unsigned int); int (*link) (struct dentry *,struct inode *,struct dentry *); int (*unlink) (struct inode *,struct dentry *); - int (*symlink) (struct mnt_idmap *, struct inode *,struct dentry *,const char *); - struct dentry *(*mkdir) (struct mnt_idmap *, struct inode *,struct dentry *,umode_t); + int (*symlink) (const struct mnt_idmap *, struct inode *,struct dentry *,const char *); + struct dentry *(*mkdir) (const struct mnt_idmap *, struct inode *,struct dentry *,umode_t); int (*rmdir) (struct inode *,struct dentry *); - int (*mknod) (struct mnt_idmap *, struct inode *,struct dentry *,umode_t,dev_t); - int (*rename) (struct mnt_idmap *, struct inode *, struct dentry *, + int (*mknod) (const struct mnt_idmap *, struct inode *,struct dentry *,umode_t,dev_t); + int (*rename) (const struct mnt_idmap *, struct inode *, struct dentry *, struct inode *, struct dentry *, unsigned int); int (*readlink) (struct dentry *, char __user *,int); const char *(*get_link) (struct dentry *, struct inode *, struct delayed_call *); - int (*permission) (struct mnt_idmap *, struct inode *, int); + int (*permission) (const struct mnt_idmap *, struct inode *, int); struct posix_acl * (*get_inode_acl)(struct inode *, int, bool); - int (*setattr) (struct mnt_idmap *, struct dentry *, struct iattr *); - int (*getattr) (struct mnt_idmap *, const struct path *, struct kstat *, u32, unsigned int); + int (*setattr) (const struct mnt_idmap *, struct dentry *, struct iattr *); + int (*getattr) (const struct mnt_idmap *, const struct path *, struct kstat *, u32, unsigned int); ssize_t (*listxattr) (struct dentry *, char *, size_t); void (*update_time)(struct inode *inode, enum fs_update_time type, int flags); void (*sync_lazytime)(struct inode *inode); int (*atomic_open)(struct inode *, struct dentry *, struct file *, unsigned open_flag, umode_t create_mode); - int (*tmpfile) (struct mnt_idmap *, struct inode *, struct file *, umode_t); - struct posix_acl * (*get_acl)(struct mnt_idmap *, struct dentry *, int); - int (*set_acl)(struct mnt_idmap *, struct dentry *, struct posix_acl *, int); - int (*fileattr_set)(struct mnt_idmap *idmap, + int (*tmpfile) (const struct mnt_idmap *, struct inode *, struct file *, umode_t); + struct posix_acl * (*get_acl)(const struct mnt_idmap *, struct dentry *, int); + int (*set_acl)(const struct mnt_idmap *, struct dentry *, struct posix_acl *, int); + int (*fileattr_set)(const struct mnt_idmap *idmap, struct dentry *dentry, struct file_kattr *fa); int (*fileattr_get)(struct dentry *dentry, struct file_kattr *fa); struct offset_ctx *(*get_offset_ctx)(struct inode *inode); @@ -507,8 +507,8 @@ otherwise noted. dentry before the first mkdir returns. If there is any chance this could happen, then the new inode - should be d_drop()ed and attached with d_splice_alias(). The - returned dentry (if any) should be returned by ->mkdir(). + should be attached with d_splice_alias(). The returned + dentry (if any) should be returned by ->mkdir(). ``rmdir`` called by the rmdir(2) system call. Only required if you want diff --git a/Documentation/filesystems/xfs/xfs-online-fsck-design.rst b/Documentation/filesystems/xfs/xfs-online-fsck-design.rst index 3d9233f403db..14767ce9fad4 100644 --- a/Documentation/filesystems/xfs/xfs-online-fsck-design.rst +++ b/Documentation/filesystems/xfs/xfs-online-fsck-design.rst @@ -1973,8 +1973,7 @@ provide loading and storing of array elements at arbitrary array indices. Gaps are defined to be null records, and null records are defined to be a sequence of all zero bytes. Null records are detected by calling ``xfarray_element_is_null``. -They are created either by calling ``xfarray_unset`` to null out an existing -record or by never storing anything to an array index. +They are created by never storing anything to an array index. The second type of caller handles records that are not indexed by position and do not require multiple updates to a record. @@ -1991,9 +1990,7 @@ The typical use case here is constructing space extent reference counts from reverse mapping information. Records can be put in the bag in any order, they can be removed from the bag at any time, and uniqueness of records is left to callers. -The ``xfarray_store_anywhere`` function is used to insert a record in any -null record slot in the bag; and the ``xfarray_unset`` function removes a -record from the bag. +Note: Bags are now implemented with in-memory btrees for faster access. Iterating Array Elements ^^^^^^^^^^^^^^^^^^^^^^^^ @@ -2643,11 +2640,7 @@ generate refcount information from reverse mapping records. refcount record associating the block number range that we just walked to the size of the bag. -The bag-like structure in this case is a type 2 xfarray as discussed in the -:ref:`xfarray access patterns<xfarray_access_patterns>` section. -Reverse mappings are added to the bag using ``xfarray_store_anywhere`` and -removed via ``xfarray_unset``. -Bag members are examined through ``xfarray_iter`` loops. +The bag-like structure in this case is an in-memory btree. Case Study: Rebuilding File Fork Mapping Indices ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^ diff --git a/MAINTAINERS b/MAINTAINERS index 4f14fd2ccff4..c53778bc674b 100644 --- a/MAINTAINERS +++ b/MAINTAINERS @@ -4642,13 +4642,6 @@ S: Odd Fixes F: Documentation/block/bfq-iosched.rst F: block/bfq-* -BFS FILE SYSTEM -M: "Tigran A. Aivazian" <aivazian.tigran@gmail.com> -S: Maintained -F: Documentation/filesystems/bfs.rst -F: fs/bfs/ -F: include/uapi/linux/bfs_fs.h - BITMAP API M: Yury Norov <yury.norov@gmail.com> R: Rasmus Villemoes <linux@rasmusvillemoes.dk> @@ -6620,6 +6613,7 @@ S: Supported F: fs/configfs/ F: include/linux/configfs.h F: samples/configfs/ +F: tools/testing/selftests/filesystems/configfs/ CONFIGFS [RUST] M: Andreas Hindborg <a.hindborg@kernel.org> @@ -9848,6 +9842,7 @@ R: Yuezhang Mo <yuezhang.mo@sony.com> L: exfat@lists.linux.dev S: Maintained T: git git://git.kernel.org/pub/scm/linux/kernel/git/linkinjeon/exfat.git +F: Documentation/filesystems/exfat.rst F: fs/exfat/ EXPRESSWIRE PROTOCOL LIBRARY @@ -9995,6 +9990,7 @@ F: net/core/failover.c FANOTIFY M: Jan Kara <jack@suse.cz> R: Amir Goldstein <amir73il@gmail.com> +R: Matt Bobrowski <matt@bobrowski.net> L: linux-fsdevel@vger.kernel.org S: Maintained F: fs/notify/fanotify/ @@ -10677,7 +10673,8 @@ M: Jaegeuk Kim <jaegeuk@kernel.org> L: linux-fscrypt@vger.kernel.org S: Supported Q: https://patchwork.kernel.org/project/linux-fscrypt/list/ -T: git https://git.kernel.org/pub/scm/fs/fscrypt/linux.git +T: git https://git.kernel.org/pub/scm/fs/fscrypt/linux.git for-next +T: git https://git.kernel.org/pub/scm/fs/fscrypt/linux.git for-current F: Documentation/filesystems/fscrypt.rst F: fs/crypto/ F: include/linux/fscrypt.h @@ -10724,7 +10721,8 @@ M: Theodore Y. Ts'o <tytso@mit.edu> L: fsverity@lists.linux.dev S: Supported Q: https://patchwork.kernel.org/project/fsverity/list/ -T: git https://git.kernel.org/pub/scm/fs/fsverity/linux.git +T: git https://git.kernel.org/pub/scm/fs/fsverity/linux.git for-next +T: git https://git.kernel.org/pub/scm/fs/fsverity/linux.git for-current F: Documentation/filesystems/fsverity.rst F: fs/verity/ F: include/linux/fsverity.h @@ -14209,7 +14207,8 @@ L: linux-nfs@vger.kernel.org S: Supported P: Documentation/filesystems/nfs/nfsd-maintainer-entry-profile.rst B: https://bugzilla.kernel.org -T: git git://git.kernel.org/pub/scm/linux/kernel/git/cel/linux.git +T: git git://git.kernel.org/pub/scm/linux/kernel/git/cel/linux.git nfsd-testing +T: git git://git.kernel.org/pub/scm/linux/kernel/git/cel/linux.git nfsd-next F: Documentation/filesystems/nfs/ F: fs/lockd/ F: fs/nfs_common/ @@ -14435,6 +14434,7 @@ S: Supported T: git git://git.kernel.org/pub/scm/linux/kernel/git/driver-core/driver-core.git F: fs/kernfs/ F: include/linux/kernfs.h +F: tools/testing/selftests/filesystems/kernfs_test.c KEXEC M: Andrew Morton <akpm@linux-foundation.org> diff --git a/arch/mips/configs/malta_defconfig b/arch/mips/configs/malta_defconfig index 56a8f76dab41..4fb4d45f3cdc 100644 --- a/arch/mips/configs/malta_defconfig +++ b/arch/mips/configs/malta_defconfig @@ -333,7 +333,6 @@ CONFIG_AFFS_FS=m CONFIG_HFS_FS=m CONFIG_HFSPLUS_FS=m CONFIG_BEFS_FS=m -CONFIG_BFS_FS=m CONFIG_EFS_FS=m CONFIG_JFFS2_FS=m CONFIG_JFFS2_FS_XATTR=y diff --git a/arch/mips/configs/malta_kvm_defconfig b/arch/mips/configs/malta_kvm_defconfig index 85e95c0ae410..3d1c32d3475e 100644 --- a/arch/mips/configs/malta_kvm_defconfig +++ b/arch/mips/configs/malta_kvm_defconfig @@ -340,7 +340,6 @@ CONFIG_AFFS_FS=m CONFIG_HFS_FS=m CONFIG_HFSPLUS_FS=m CONFIG_BEFS_FS=m -CONFIG_BFS_FS=m CONFIG_EFS_FS=m CONFIG_JFFS2_FS=m CONFIG_JFFS2_FS_XATTR=y diff --git a/arch/mips/configs/maltaup_xpa_defconfig b/arch/mips/configs/maltaup_xpa_defconfig index 0498a0115349..2545168b75bb 100644 --- a/arch/mips/configs/maltaup_xpa_defconfig +++ b/arch/mips/configs/maltaup_xpa_defconfig @@ -339,7 +339,6 @@ CONFIG_AFFS_FS=m CONFIG_HFS_FS=m CONFIG_HFSPLUS_FS=m CONFIG_BEFS_FS=m -CONFIG_BFS_FS=m CONFIG_EFS_FS=m CONFIG_JFFS2_FS=m CONFIG_JFFS2_FS_XATTR=y diff --git a/arch/mips/configs/rm200_defconfig b/arch/mips/configs/rm200_defconfig index 09e006c7cd7c..cd2a831614c8 100644 --- a/arch/mips/configs/rm200_defconfig +++ b/arch/mips/configs/rm200_defconfig @@ -318,7 +318,6 @@ CONFIG_ADFS_FS=m CONFIG_AFFS_FS=m CONFIG_HFS_FS=m CONFIG_BEFS_FS=m -CONFIG_BFS_FS=m CONFIG_EFS_FS=m CONFIG_CRAMFS=m CONFIG_VXFS_FS=m diff --git a/arch/powerpc/Kconfig b/arch/powerpc/Kconfig index 877393207e3b..3f59b201b62f 100644 --- a/arch/powerpc/Kconfig +++ b/arch/powerpc/Kconfig @@ -160,7 +160,6 @@ config PPC select ARCH_HAS_UBSAN select ARCH_HAS_VDSO_ARCH_DATA select ARCH_HAVE_NMI_SAFE_CMPXCHG - select ARCH_HAVE_EXTRA_ELF_NOTES if SPU_BASE select ARCH_KEEP_MEMBLOCK select ARCH_MHP_MEMMAP_ON_MEMORY_ENABLE if PPC_RADIX_MMU select ARCH_MIGHT_HAVE_PC_PARPORT diff --git a/arch/powerpc/configs/fsl-emb-nonhw.config b/arch/powerpc/configs/fsl-emb-nonhw.config index 391c99117ee0..688051b11ce1 100644 --- a/arch/powerpc/configs/fsl-emb-nonhw.config +++ b/arch/powerpc/configs/fsl-emb-nonhw.config @@ -2,7 +2,6 @@ CONFIG_ADFS_FS=m CONFIG_AFFS_FS=m CONFIG_AUDIT=y CONFIG_BEFS_FS=m -CONFIG_BFS_FS=m CONFIG_BINFMT_MISC=m # CONFIG_BLK_DEV_BSG is not set CONFIG_BLK_DEV_INITRD=y diff --git a/arch/powerpc/configs/ppc6xx_defconfig b/arch/powerpc/configs/ppc6xx_defconfig index acffc5c17f92..8aab229fb909 100644 --- a/arch/powerpc/configs/ppc6xx_defconfig +++ b/arch/powerpc/configs/ppc6xx_defconfig @@ -945,7 +945,6 @@ CONFIG_ECRYPT_FS=m CONFIG_HFS_FS=m CONFIG_HFSPLUS_FS=m CONFIG_BEFS_FS=m -CONFIG_BFS_FS=m CONFIG_EFS_FS=m CONFIG_CRAMFS=m CONFIG_VXFS_FS=m diff --git a/arch/powerpc/include/asm/elf.h b/arch/powerpc/include/asm/elf.h index bb4b94444d3e..5dc8c4923eb1 100644 --- a/arch/powerpc/include/asm/elf.h +++ b/arch/powerpc/include/asm/elf.h @@ -123,12 +123,6 @@ extern int arch_setup_additional_pages(struct linux_binprm *bprm, (0x7ff >> (PAGE_SHIFT - 12)) : \ (0x3ffff >> (PAGE_SHIFT - 12))) -#ifdef CONFIG_SPU_BASE -/* Notes used in ET_CORE. Note name is "SPU/<fd>/<filename>". */ -#define NT_SPU 1 - -#endif /* CONFIG_SPU_BASE */ - #ifdef CONFIG_PPC64 #define get_cache_geometry(level) \ diff --git a/arch/powerpc/include/asm/spu.h b/arch/powerpc/include/asm/spu.h index 96ad4510c895..7152285b6268 100644 --- a/arch/powerpc/include/asm/spu.h +++ b/arch/powerpc/include/asm/spu.h @@ -210,15 +210,12 @@ extern long spu_sys_callback(struct spu_syscall_block *s); /* syscalls implemented in spufs */ struct file; -struct coredump_params; struct spufs_calls { long (*create_thread)(const char __user *name, unsigned int flags, umode_t mode, struct file *neighbor); long (*spu_run)(struct file *filp, __u32 __user *unpc, __u32 __user *ustatus); - int (*coredump_extra_notes_size)(void); - int (*coredump_extra_notes_write)(struct coredump_params *cprm); void (*notify_spus_active)(void); struct module *owner; }; diff --git a/arch/powerpc/platforms/cell/Kconfig b/arch/powerpc/platforms/cell/Kconfig index db65bfcd1e74..6bd26815c331 100644 --- a/arch/powerpc/platforms/cell/Kconfig +++ b/arch/powerpc/platforms/cell/Kconfig @@ -10,7 +10,6 @@ config SPU_FS tristate "SPU file system" default m depends on PPC_CELL - depends on COREDUMP select SPU_BASE help The SPU file system is used to access Synergistic Processing diff --git a/arch/powerpc/platforms/cell/spu_syscalls.c b/arch/powerpc/platforms/cell/spu_syscalls.c index 000894e07b02..8be81207e886 100644 --- a/arch/powerpc/platforms/cell/spu_syscalls.c +++ b/arch/powerpc/platforms/cell/spu_syscalls.c @@ -88,26 +88,6 @@ SYSCALL_DEFINE3(spu_run,int, fd, __u32 __user *, unpc, __u32 __user *, ustatus) return calls->spu_run(fd_file(arg), unpc, ustatus); } -#ifdef CONFIG_COREDUMP -int elf_coredump_extra_notes_size(void) -{ - CLASS(spufs_calls, calls)(); - if (!calls) - return 0; - - return calls->coredump_extra_notes_size(); -} - -int elf_coredump_extra_notes_write(struct coredump_params *cprm) -{ - CLASS(spufs_calls, calls)(); - if (!calls) - return 0; - - return calls->coredump_extra_notes_write(cprm); -} -#endif - void notify_spus_active(void) { struct spufs_calls *calls; diff --git a/arch/powerpc/platforms/cell/spufs/Makefile b/arch/powerpc/platforms/cell/spufs/Makefile index 52e4c80ec8d0..60319d4ff25a 100644 --- a/arch/powerpc/platforms/cell/spufs/Makefile +++ b/arch/powerpc/platforms/cell/spufs/Makefile @@ -4,7 +4,6 @@ obj-$(CONFIG_SPU_FS) += spufs.o spufs-y += inode.o file.o context.o syscalls.o spufs-y += sched.o backing_ops.o hw_ops.o run.o gang.o spufs-y += switch.o fault.o lscsa_alloc.o -spufs-$(CONFIG_COREDUMP) += coredump.o # magic for the trace events CFLAGS_sched.o := -I$(src) diff --git a/arch/powerpc/platforms/cell/spufs/coredump.c b/arch/powerpc/platforms/cell/spufs/coredump.c deleted file mode 100644 index 301ee7d8b7df..000000000000 --- a/arch/powerpc/platforms/cell/spufs/coredump.c +++ /dev/null @@ -1,183 +0,0 @@ -// SPDX-License-Identifier: GPL-2.0-or-later -/* - * SPU core dump code - * - * (C) Copyright 2006 IBM Corp. - * - * Author: Dwayne Grant McConnell <decimal@us.ibm.com> - */ - -#include <linux/elf.h> -#include <linux/file.h> -#include <linux/fdtable.h> -#include <linux/fs.h> -#include <linux/gfp.h> -#include <linux/list.h> -#include <linux/syscalls.h> -#include <linux/coredump.h> -#include <linux/binfmts.h> - -#include <linux/uaccess.h> - -#include "spufs.h" - -static int spufs_ctx_note_size(struct spu_context *ctx, int dfd) -{ - int i, sz, total = 0; - char *name; - char fullname[80]; - - for (i = 0; spufs_coredump_read[i].name != NULL; i++) { - name = spufs_coredump_read[i].name; - sz = spufs_coredump_read[i].size; - - sprintf(fullname, "SPU/%d/%s", dfd, name); - - total += sizeof(struct elf_note); - total += roundup(strlen(fullname) + 1, 4); - total += roundup(sz, 4); - } - - return total; -} - -static int match_context(const void *v, struct file *file, unsigned fd) -{ - struct spu_context *ctx; - if (file->f_op != &spufs_context_fops) - return 0; - ctx = SPUFS_I(file_inode(file))->i_ctx; - if (ctx->flags & SPU_CREATE_NOSCHED) - return 0; - return fd + 1; -} - -/* - * The additional architecture-specific notes for Cell are various - * context files in the spu context. - * - * This function iterates over all open file descriptors and sees - * if they are a directory in spufs. In that case we use spufs - * internal functionality to dump them without needing to actually - * open the files. - */ -/* - * descriptor table is not shared, so files can't change or go away. - */ -static struct spu_context *coredump_next_context(int *fd) -{ - struct spu_context *ctx = NULL; - struct file *file; - int n = iterate_fd(current->files, *fd, match_context, NULL); - if (!n) - return NULL; - *fd = n - 1; - - file = fget_raw(*fd); - if (file) { - ctx = SPUFS_I(file_inode(file))->i_ctx; - get_spu_context(ctx); - fput(file); - } - - return ctx; -} - -int spufs_coredump_extra_notes_size(void) -{ - struct spu_context *ctx; - int size = 0, rc, fd; - - fd = 0; - while ((ctx = coredump_next_context(&fd)) != NULL) { - rc = spu_acquire_saved(ctx); - if (rc) { - put_spu_context(ctx); - break; - } - - rc = spufs_ctx_note_size(ctx, fd); - spu_release_saved(ctx); - if (rc < 0) { - put_spu_context(ctx); - break; - } - - size += rc; - - /* start searching the next fd next time */ - fd++; - put_spu_context(ctx); - } - - return size; -} - -static int spufs_arch_write_note(struct spu_context *ctx, int i, - struct coredump_params *cprm, int dfd) -{ - size_t sz = spufs_coredump_read[i].size; - char fullname[80]; - struct elf_note en; - int ret; - - sprintf(fullname, "SPU/%d/%s", dfd, spufs_coredump_read[i].name); - en.n_namesz = strlen(fullname) + 1; - en.n_descsz = sz; - en.n_type = NT_SPU; - - if (!dump_emit(cprm, &en, sizeof(en))) - return -EIO; - if (!dump_emit(cprm, fullname, en.n_namesz)) - return -EIO; - if (!dump_align(cprm, 4)) - return -EIO; - - if (spufs_coredump_read[i].dump) { - ret = spufs_coredump_read[i].dump(ctx, cprm); - if (ret < 0) - return ret; - } else { - char buf[32]; - - ret = snprintf(buf, sizeof(buf), "0x%.16llx", - spufs_coredump_read[i].get(ctx)); - if (ret >= sizeof(buf)) - return sizeof(buf); - - /* count trailing the NULL: */ - if (!dump_emit(cprm, buf, ret + 1)) - return -EIO; - } - - dump_skip_to(cprm, roundup(cprm->pos - ret + sz, 4)); - return 0; -} - -int spufs_coredump_extra_notes_write(struct coredump_params *cprm) -{ - struct spu_context *ctx; - int fd, j, rc; - - fd = 0; - while ((ctx = coredump_next_context(&fd)) != NULL) { - rc = spu_acquire_saved(ctx); - if (rc) - return rc; - - for (j = 0; spufs_coredump_read[j].name != NULL; j++) { - rc = spufs_arch_write_note(ctx, j, cprm, fd); - if (rc) { - spu_release_saved(ctx); - return rc; - } - } - - spu_release_saved(ctx); - - /* start searching the next fd next time */ - fd++; - } - - return 0; -} diff --git a/arch/powerpc/platforms/cell/spufs/file.c b/arch/powerpc/platforms/cell/spufs/file.c index de7494748fec..98c47bafaf67 100644 --- a/arch/powerpc/platforms/cell/spufs/file.c +++ b/arch/powerpc/platforms/cell/spufs/file.c @@ -9,7 +9,6 @@ #undef DEBUG -#include <linux/coredump.h> #include <linux/fs.h> #include <linux/ioctl.h> #include <linux/export.h> @@ -130,14 +129,6 @@ out: return ret; } -static ssize_t spufs_dump_emit(struct coredump_params *cprm, void *buf, - size_t size) -{ - if (!dump_emit(cprm, buf, size)) - return -EIO; - return size; -} - #define DEFINE_SPUFS_SIMPLE_ATTRIBUTE(__fops, __get, __set, __fmt) \ static int __fops ## _open(struct inode *inode, struct file *file) \ { \ @@ -181,12 +172,6 @@ spufs_mem_release(struct inode *inode, struct file *file) } static ssize_t -spufs_mem_dump(struct spu_context *ctx, struct coredump_params *cprm) -{ - return spufs_dump_emit(cprm, ctx->ops->get_ls(ctx), LS_SIZE); -} - -static ssize_t spufs_mem_read(struct file *file, char __user *buffer, size_t size, loff_t *pos) { @@ -467,13 +452,6 @@ spufs_regs_open(struct inode *inode, struct file *file) } static ssize_t -spufs_regs_dump(struct spu_context *ctx, struct coredump_params *cprm) -{ - return spufs_dump_emit(cprm, ctx->csa.lscsa->gprs, - sizeof(ctx->csa.lscsa->gprs)); -} - -static ssize_t spufs_regs_read(struct file *file, char __user *buffer, size_t size, loff_t *pos) { @@ -524,13 +502,6 @@ static const struct file_operations spufs_regs_fops = { }; static ssize_t -spufs_fpcr_dump(struct spu_context *ctx, struct coredump_params *cprm) -{ - return spufs_dump_emit(cprm, &ctx->csa.lscsa->fpcr, - sizeof(ctx->csa.lscsa->fpcr)); -} - -static ssize_t spufs_fpcr_read(struct file *file, char __user * buffer, size_t size, loff_t * pos) { @@ -953,15 +924,6 @@ spufs_signal1_release(struct inode *inode, struct file *file) return 0; } -static ssize_t spufs_signal1_dump(struct spu_context *ctx, - struct coredump_params *cprm) -{ - if (!ctx->csa.spu_chnlcnt_RW[3]) - return 0; - return spufs_dump_emit(cprm, &ctx->csa.spu_chnldata_RW[3], - sizeof(ctx->csa.spu_chnldata_RW[3])); -} - static ssize_t __spufs_signal1_read(struct spu_context *ctx, char __user *buf, size_t len) { @@ -1086,15 +1048,6 @@ spufs_signal2_release(struct inode *inode, struct file *file) return 0; } -static ssize_t spufs_signal2_dump(struct spu_context *ctx, - struct coredump_params *cprm) -{ - if (!ctx->csa.spu_chnlcnt_RW[4]) - return 0; - return spufs_dump_emit(cprm, &ctx->csa.spu_chnldata_RW[4], - sizeof(ctx->csa.spu_chnldata_RW[4])); -} - static ssize_t __spufs_signal2_read(struct spu_context *ctx, char __user *buf, size_t len) { @@ -1924,15 +1877,6 @@ static const struct file_operations spufs_caps_fops = { .release = single_release, }; -static ssize_t spufs_mbox_info_dump(struct spu_context *ctx, - struct coredump_params *cprm) -{ - if (!(ctx->csa.prob.mb_stat_R & 0x0000ff)) - return 0; - return spufs_dump_emit(cprm, &ctx->csa.prob.pu_mb_R, - sizeof(ctx->csa.prob.pu_mb_R)); -} - static ssize_t spufs_mbox_info_read(struct file *file, char __user *buf, size_t len, loff_t *pos) { @@ -1962,15 +1906,6 @@ static const struct file_operations spufs_mbox_info_fops = { .llseek = generic_file_llseek, }; -static ssize_t spufs_ibox_info_dump(struct spu_context *ctx, - struct coredump_params *cprm) -{ - if (!(ctx->csa.prob.mb_stat_R & 0xff0000)) - return 0; - return spufs_dump_emit(cprm, &ctx->csa.priv2.puint_mb_R, - sizeof(ctx->csa.priv2.puint_mb_R)); -} - static ssize_t spufs_ibox_info_read(struct file *file, char __user *buf, size_t len, loff_t *pos) { @@ -2005,13 +1940,6 @@ static size_t spufs_wbox_info_cnt(struct spu_context *ctx) return (4 - ((ctx->csa.prob.mb_stat_R & 0x00ff00) >> 8)) * sizeof(u32); } -static ssize_t spufs_wbox_info_dump(struct spu_context *ctx, - struct coredump_params *cprm) -{ - return spufs_dump_emit(cprm, &ctx->csa.spu_mailbox_data, - spufs_wbox_info_cnt(ctx)); -} - static ssize_t spufs_wbox_info_read(struct file *file, char __user *buf, size_t len, loff_t *pos) { @@ -2059,15 +1987,6 @@ static void spufs_get_dma_info(struct spu_context *ctx, } } -static ssize_t spufs_dma_info_dump(struct spu_context *ctx, - struct coredump_params *cprm) -{ - struct spu_dma_info info; - - spufs_get_dma_info(ctx, &info); - return spufs_dump_emit(cprm, &info, sizeof(info)); -} - static ssize_t spufs_dma_info_read(struct file *file, char __user *buf, size_t len, loff_t *pos) { @@ -2112,15 +2031,6 @@ static void spufs_get_proxydma_info(struct spu_context *ctx, } } -static ssize_t spufs_proxydma_info_dump(struct spu_context *ctx, - struct coredump_params *cprm) -{ - struct spu_proxydma_info info; - - spufs_get_proxydma_info(ctx, &info); - return spufs_dump_emit(cprm, &info, sizeof(info)); -} - static ssize_t spufs_proxydma_info_read(struct file *file, char __user *buf, size_t len, loff_t *pos) { @@ -2580,27 +2490,3 @@ const struct spufs_tree_descr spufs_dir_debug_contents[] = { { ".ctx", &spufs_ctx_fops, 0444, }, {}, }; - -const struct spufs_coredump_reader spufs_coredump_read[] = { - { "regs", spufs_regs_dump, NULL, sizeof(struct spu_reg128[128])}, - { "fpcr", spufs_fpcr_dump, NULL, sizeof(struct spu_reg128) }, - { "lslr", NULL, spufs_lslr_get, 19 }, - { "decr", NULL, spufs_decr_get, 19 }, - { "decr_status", NULL, spufs_decr_status_get, 19 }, - { "mem", spufs_mem_dump, NULL, LS_SIZE, }, - { "signal1", spufs_signal1_dump, NULL, sizeof(u32) }, - { "signal1_type", NULL, spufs_signal1_type_get, 19 }, - { "signal2", spufs_signal2_dump, NULL, sizeof(u32) }, - { "signal2_type", NULL, spufs_signal2_type_get, 19 }, - { "event_mask", NULL, spufs_event_mask_get, 19 }, - { "event_status", NULL, spufs_event_status_get, 19 }, - { "mbox_info", spufs_mbox_info_dump, NULL, sizeof(u32) }, - { "ibox_info", spufs_ibox_info_dump, NULL, sizeof(u32) }, - { "wbox_info", spufs_wbox_info_dump, NULL, 4 * sizeof(u32)}, - { "dma_info", spufs_dma_info_dump, NULL, sizeof(struct spu_dma_info)}, - { "proxydma_info", spufs_proxydma_info_dump, - NULL, sizeof(struct spu_proxydma_info)}, - { "object-id", NULL, spufs_object_id_get, 19 }, - { "npc", NULL, spufs_npc_get, 19 }, - { NULL }, -}; diff --git a/arch/powerpc/platforms/cell/spufs/inode.c b/arch/powerpc/platforms/cell/spufs/inode.c index 2b54afb31529..f066d9c7344d 100644 --- a/arch/powerpc/platforms/cell/spufs/inode.c +++ b/arch/powerpc/platforms/cell/spufs/inode.c @@ -92,7 +92,7 @@ out: } static int -spufs_setattr(struct mnt_idmap *idmap, struct dentry *dentry, +spufs_setattr(const struct mnt_idmap *idmap, struct dentry *dentry, struct iattr *attr) { struct inode *inode = d_inode(dentry); @@ -266,9 +266,9 @@ spufs_mkdir(struct inode *dir, struct dentry *dentry, unsigned int flags, static int spufs_context_open(const struct path *path) { FD_PREPARE(fdf, 0, dentry_open(path, O_RDONLY, current_cred())); - if (fdf.err) - return fdf.err; - fd_prepare_file(fdf)->f_op = &spufs_context_fops; + if (fdf->fd < 0) + return fdf->fd; + fdf->file->f_op = &spufs_context_fops; return fd_publish(fdf); } @@ -499,9 +499,9 @@ static int spufs_gang_open(const struct path *path) * in error path of *_open(). */ FD_PREPARE(fdf, 0, dentry_open(path, O_RDONLY, current_cred())); - if (fdf.err) - return fdf.err; - fd_prepare_file(fdf)->f_op = &spufs_gang_fops; + if (fdf->fd < 0) + return fdf->fd; + fdf->file->f_op = &spufs_gang_fops; return fd_publish(fdf); } diff --git a/arch/powerpc/platforms/cell/spufs/spufs.h b/arch/powerpc/platforms/cell/spufs/spufs.h index d33787c57c39..612b5075d0ec 100644 --- a/arch/powerpc/platforms/cell/spufs/spufs.h +++ b/arch/powerpc/platforms/cell/spufs/spufs.h @@ -232,13 +232,9 @@ extern const struct spufs_tree_descr spufs_dir_debug_contents[]; /* system call implementation */ extern struct spufs_calls spufs_calls; -struct coredump_params; long spufs_run_spu(struct spu_context *ctx, u32 *npc, u32 *status); long spufs_create(const struct path *nd, struct dentry *dentry, unsigned int flags, umode_t mode, struct file *filp); -/* ELF coredump callbacks for writing SPU ELF notes */ -extern int spufs_coredump_extra_notes_size(void); -extern int spufs_coredump_extra_notes_write(struct coredump_params *cprm); extern const struct file_operations spufs_context_fops; @@ -335,14 +331,6 @@ void spufs_stop_callback(struct spu *spu, int irq); void spufs_mfc_callback(struct spu *spu); void spufs_dma_callback(struct spu *spu, int type); -struct spufs_coredump_reader { - char *name; - ssize_t (*dump)(struct spu_context *ctx, struct coredump_params *cprm); - u64 (*get)(struct spu_context *ctx); - size_t size; -}; -extern const struct spufs_coredump_reader spufs_coredump_read[]; - extern int spu_init_csa(struct spu_state *csa); extern void spu_fini_csa(struct spu_state *csa); extern int spu_save(struct spu_state *prev, struct spu *spu); diff --git a/arch/powerpc/platforms/cell/spufs/syscalls.c b/arch/powerpc/platforms/cell/spufs/syscalls.c index ea4ba1b6ce6a..b6de37150e73 100644 --- a/arch/powerpc/platforms/cell/spufs/syscalls.c +++ b/arch/powerpc/platforms/cell/spufs/syscalls.c @@ -82,8 +82,4 @@ struct spufs_calls spufs_calls = { .spu_run = do_spu_run, .notify_spus_active = do_notify_spus_active, .owner = THIS_MODULE, -#ifdef CONFIG_COREDUMP - .coredump_extra_notes_size = spufs_coredump_extra_notes_size, - .coredump_extra_notes_write = spufs_coredump_extra_notes_write, -#endif }; diff --git a/block/bio-integrity-fs.c b/block/bio-integrity-fs.c index 692403dfa047..c8e91ada8ca6 100644 --- a/block/bio-integrity-fs.c +++ b/block/bio-integrity-fs.c @@ -31,6 +31,7 @@ unsigned int fs_bio_integrity_alloc(struct bio *bio) bio_integrity_setup_default(bio); return action; } +EXPORT_SYMBOL_GPL(fs_bio_integrity_alloc); void fs_bio_integrity_free(struct bio *bio) { @@ -43,6 +44,7 @@ void fs_bio_integrity_free(struct bio *bio) bio->bi_integrity = NULL; bio->bi_opf &= ~REQ_INTEGRITY; } +EXPORT_SYMBOL_GPL(fs_bio_integrity_free); void fs_bio_integrity_generate(struct bio *bio) { @@ -52,14 +54,10 @@ void fs_bio_integrity_generate(struct bio *bio) } EXPORT_SYMBOL_GPL(fs_bio_integrity_generate); -int fs_bio_integrity_verify(struct bio *bio, sector_t sector, unsigned int size) +int fs_bio_integrity_verify(struct bio *bio, struct bvec_iter *data_iter) { struct blk_integrity *bi = blk_get_integrity(bio->bi_bdev->bd_disk); struct bio_integrity_payload *bip = bio_integrity(bio); - struct bvec_iter data_iter = { - .bi_sector = sector, - .bi_size = size, - }; if (!bip || !(bip->bip_flags & BIP_CHECK_FLAGS)) return 0; @@ -71,9 +69,10 @@ int fs_bio_integrity_verify(struct bio *bio, sector_t sector, unsigned int size) * bio. Requires the submitter to remember the sector and the size. */ memset(&bip->bip_iter, 0, sizeof(bip->bip_iter)); - bip->bip_iter.bi_sector = sector; - bip->bip_iter.bi_size = bio_integrity_bytes(bi, size >> SECTOR_SHIFT); - return blk_status_to_errno(bio_integrity_verify(bio, &data_iter)); + bip->bip_iter.bi_sector = data_iter->bi_sector; + bip->bip_iter.bi_size = + bio_integrity_bytes(bi, data_iter->bi_size >> SECTOR_SHIFT); + return blk_status_to_errno(bio_integrity_verify(bio, data_iter)); } static int __init fs_bio_integrity_init(void) diff --git a/block/bio-integrity.c b/block/bio-integrity.c index b23e2434d80c..d3df726e0f08 100644 --- a/block/bio-integrity.c +++ b/block/bio-integrity.c @@ -72,6 +72,7 @@ void bio_integrity_alloc_buf(struct bio *bio, gfp_t gfp, bool zero_buffer) unsigned int len = bio_integrity_bytes(bi, bio_sectors(bio)); void *buf; + WARN_ON_ONCE(len > BLK_INTEGRITY_MAX_SIZE); buf = kmalloc(len, gfp | __GFP_NOWARN | (zero_buffer ? __GFP_ZERO : 0)); if (unlikely(!buf)) { struct page *page; diff --git a/block/bio.c b/block/bio.c index f95b63c0604a..b48091c7663f 100644 --- a/block/bio.c +++ b/block/bio.c @@ -320,6 +320,26 @@ void bio_reuse(struct bio *bio, blk_opf_t opf) } EXPORT_SYMBOL_GPL(bio_reuse); +/** + * bio_prepare_reissue - prepare a bio for reuissing the original I/O + * @bio: bio to reuse + * @bdev: block device to use the bio for + * + * Prepare @bio to be resubmitted to retry the original operation. + * The caller must reset bio->bi_iter to the original state. + */ +void bio_prepare_reissue(struct bio *bio, struct block_device *bdev) +{ + bio->bi_bdev = bdev; + bio_associate_blkg(bio); + bio->bi_flags &= + (BIO_PAGE_PINNED | BIO_CLONED | BIO_QUIET | BIO_REFFED); + bio->bi_status = BLK_STS_OK; + bio->bi_bvec_gap_bit = 0; + atomic_set(&bio->__bi_remaining, 1); +} +EXPORT_SYMBOL_GPL(bio_prepare_reissue); + static struct bio *__bio_chain_endio(struct bio *bio) { struct bio *parent = bio->bi_private; @@ -1203,8 +1223,9 @@ bool bio_iov_iter_set(struct bio *bio, const struct iov_iter *iter) * for the next iteration. */ static int bio_iov_iter_align_down(struct bio *bio, struct iov_iter *iter, - struct bio_vec *bv, unsigned len_align_mask) + unsigned len_align_mask) { + struct bio_vec *bv = &bio->bi_io_vec[bio->bi_vcnt - 1]; size_t nbytes = bio->bi_iter.bi_size & len_align_mask; if (!nbytes) @@ -1262,6 +1283,7 @@ static inline bool bio_iov_bvec_aligned(const struct bio *bio, * bio_iov_iter_get_pages - add user or kernel pages to a bio * @bio: bio to add pages to * @iter: iov iterator describing the region to be added + * @maxlen: maximum size to consume from @iter * @mem_align_mask: the mask the source address and length must be aligned to, * 0 for no requirement * @len_align_mask: the mask to align the total size to, 0 for any length @@ -1282,7 +1304,8 @@ static inline bool bio_iov_bvec_aligned(const struct bio *bio, * is returned only if 0 pages could be pinned. */ int bio_iov_iter_get_pages(struct bio *bio, struct iov_iter *iter, - unsigned mem_align_mask, unsigned len_align_mask) + unsigned maxlen, unsigned mem_align_mask, + unsigned len_align_mask) { iov_iter_extraction_t flags = 0; @@ -1294,6 +1317,8 @@ int bio_iov_iter_get_pages(struct bio *bio, struct iov_iter *iter, !bio_iov_bvec_aligned(bio, mem_align_mask)) return -EINVAL; + /* Truncate to the maximum size that the caller can handle */ + bio->bi_iter.bi_size = min(bio->bi_iter.bi_size, maxlen); iov_iter_advance(iter, bio->bi_iter.bi_size); return 0; } @@ -1307,7 +1332,7 @@ int bio_iov_iter_get_pages(struct bio *bio, struct iov_iter *iter, ssize_t ret; ret = iov_iter_extract_bvecs(iter, bio->bi_io_vec, - BIO_MAX_SIZE - bio->bi_iter.bi_size, + maxlen - bio->bi_iter.bi_size, &bio->bi_vcnt, bio->bi_max_vecs, mem_align_mask, flags); if (ret <= 0) { @@ -1330,8 +1355,7 @@ int bio_iov_iter_get_pages(struct bio *bio, struct iov_iter *iter, if (is_pci_p2pdma_page(bio->bi_io_vec->bv_page)) bio->bi_opf |= REQ_NOMERGE; - return bio_iov_iter_align_down(bio, iter, - &bio->bi_io_vec[bio->bi_vcnt - 1], len_align_mask); + return bio_iov_iter_align_down(bio, iter, len_align_mask); } static struct folio *folio_alloc_greedy(gfp_t gfp, size_t *size, @@ -1350,7 +1374,7 @@ static struct folio *folio_alloc_greedy(gfp_t gfp, size_t *size, return folio_alloc(gfp, get_order(*size)); } -static void bio_free_folios(struct bio *bio) +void bio_free_folios(struct bio *bio) { struct bio_vec *bv; int i; @@ -1363,11 +1387,8 @@ static void bio_free_folios(struct bio *bio) } } -static int bio_iov_iter_bounce_write(struct bio *bio, struct iov_iter *iter, - size_t maxlen, size_t minsize) +int bio_alloc_bounce_folios(struct bio *bio, size_t total_len, size_t minsize) { - size_t total_len = min(maxlen, iov_iter_count(iter)); - if (WARN_ON_ONCE(bio_flagged(bio, BIO_CLONED))) return -EINVAL; if (WARN_ON_ONCE(bio->bi_iter.bi_size)) @@ -1377,7 +1398,6 @@ static int bio_iov_iter_bounce_write(struct bio *bio, struct iov_iter *iter, do { size_t this_len = min(total_len, SZ_1M); - size_t copied; struct folio *folio; if (this_len > minsize * 2) @@ -1389,164 +1409,69 @@ static int bio_iov_iter_bounce_write(struct bio *bio, struct iov_iter *iter, folio = folio_alloc_greedy(GFP_KERNEL, &this_len, minsize); if (!folio) break; - bio_add_folio_nofail(bio, folio, this_len, 0); - if (iter->nofault) - copied = copy_folio_from_iter_atomic(folio, 0, this_len, - iter); - else - copied = copy_folio_from_iter(folio, 0, this_len, iter); - if (copied < this_len) { - /* - * Need to revert the iov iter for all bytes we have - * copied. - * - * However the bio size differs from the real copied - * bytes as @this_len is queued but only advanced - * less than that. - * Need to compensate that for the revert. - */ - iov_iter_revert(iter, bio->bi_iter.bi_size - this_len + - copied); - bio_free_folios(bio); - return -EFAULT; - } + /* + * Align down the size to the minimum alignment. In practice + * this should not happen as minsize is expected to be a power + * of two, as is the allocation size, but it offers us a cheap + * extra safety belt. + */ + this_len &= ~(minsize - 1); + bio_add_folio_nofail(bio, folio, this_len, 0); total_len -= this_len; } while (total_len && bio->bi_vcnt < bio->bi_max_vecs); if (!bio->bi_iter.bi_size) return -ENOMEM; - return bio_iov_iter_align_down(bio, iter, - &bio->bi_io_vec[bio->bi_vcnt - 1], minsize - 1); -} - -static int bio_iov_iter_bounce_read(struct bio *bio, struct iov_iter *iter, - size_t maxlen, size_t minsize) -{ - size_t len = min3(iov_iter_count(iter), maxlen, SZ_1M); - struct folio *folio; - ssize_t ret; - - folio = folio_alloc_greedy(GFP_KERNEL, &len, minsize); - if (!folio) - return -ENOMEM; - - do { - ret = iov_iter_extract_bvecs(iter, bio->bi_io_vec + 1, len, - &bio->bi_vcnt, bio->bi_max_vecs - 1, 0, 0); - if (ret <= 0) { - if (!bio->bi_vcnt) - goto out_folio_put; - break; - } - len -= ret; - bio->bi_iter.bi_size += ret; - } while (len && bio->bi_vcnt < bio->bi_max_vecs - 1); - - /* - * Set the folio directly here. The above loop has already calculated - * the correct bi_size, and we use bi_vcnt for the user buffers. That - * is safe as bi_vcnt is only used by the submitter and not the actual - * I/O path. - */ - bvec_set_folio(&bio->bi_io_vec[0], folio, bio->bi_iter.bi_size, 0); - if (iov_iter_extract_will_pin(iter)) - bio_set_flag(bio, BIO_PAGE_PINNED); - - /* The first vec stores the bounce buffer, so do not subtract 1 here. */ - ret = bio_iov_iter_align_down(bio, iter, - &bio->bi_io_vec[bio->bi_vcnt], minsize - 1); - if (ret) - goto out_folio_put; - - /* Update the bounc buffer bv_len to the aligned down size. */ - bio->bi_io_vec[0].bv_len = bio->bi_iter.bi_size; return 0; - -out_folio_put: - folio_put(folio); - return ret; } /** - * bio_iov_iter_bounce - bounce buffer data from an iter into a bio + * bio_iov_iter_bounce_write - bounce buffer data from an iter into a bio * @bio: bio to send - * @iter: iter to read from / write into + * @iter: iter to read from * @maxlen: maximum size to bounce * @minsize: minimum folio allocation size * - * Helper for direct I/O implementations that need to bounce buffer because - * we need to checksum the data or perform other operations that require - * consistency. Allocates folios to back the bounce buffer, and for writes - * copies the data into it. Needs to be paired with bio_iov_iter_unbounce() - * called on completion. + * Helper for direct I/O write implementations that need to bounce buffer + * because they need need to checksum the data or perform other operations that + * require consistency. Allocates folios to back the bounce buffer, and copies + * the data into it. Needs to be paired with bio_free_folios() called on + * completion. */ -int bio_iov_iter_bounce(struct bio *bio, struct iov_iter *iter, size_t maxlen, - size_t minsize) -{ - if (op_is_write(bio_op(bio))) - return bio_iov_iter_bounce_write(bio, iter, maxlen, minsize); - return bio_iov_iter_bounce_read(bio, iter, maxlen, minsize); -} - -static void bvec_unpin(struct bio_vec *bv, bool mark_dirty) -{ - struct folio *folio = bvec_folio(bv); - size_t nr_pages = (bv->bv_offset + bv->bv_len - 1) / PAGE_SIZE - - bv->bv_offset / PAGE_SIZE + 1; - - if (mark_dirty) - folio_mark_dirty_lock(folio); - unpin_user_folio(folio, nr_pages); -} - -static void bio_iov_iter_unbounce_read(struct bio *bio, bool is_error, - bool mark_dirty) +int bio_iov_iter_bounce_write(struct bio *bio, struct iov_iter *iter, + size_t maxlen, size_t minsize) { - unsigned int len = bio->bi_io_vec[0].bv_len; - - if (likely(!is_error)) { - void *buf = bvec_virt(&bio->bi_io_vec[0]); - struct iov_iter to; + size_t total_len = min(maxlen, iov_iter_count(iter)); + size_t total_copied = 0; + struct bio_vec *bv; + int i, error; - iov_iter_bvec(&to, ITER_DEST, bio->bi_io_vec + 1, bio->bi_vcnt, - len); - /* copying to pinned pages should always work */ - WARN_ON_ONCE(copy_to_iter(buf, len, &to) != len); - } else { - /* No need to mark folios dirty if never copied to them */ - mark_dirty = false; - } + error = bio_alloc_bounce_folios(bio, total_len, minsize); + if (error) + return error; - if (bio_flagged(bio, BIO_PAGE_PINNED)) { - int i; + bio_for_each_bvec_all(bv, bio, i) { + struct folio *folio = page_folio(bv->bv_page); + size_t copied; - for (i = 0; i < bio->bi_vcnt; i++) - bvec_unpin(&bio->bi_io_vec[1 + i], mark_dirty); + if (iter->nofault) + copied = copy_folio_from_iter_atomic(folio, 0, + bv->bv_len, iter); + else + copied = copy_folio_from_iter(folio, 0, bv->bv_len, + iter); + total_copied += copied; + if (copied < bv->bv_len) { + iov_iter_revert(iter, total_copied); + bio_free_folios(bio); + return -EFAULT; + } } - folio_put(bvec_folio(&bio->bi_io_vec[0])); -} - -/** - * bio_iov_iter_unbounce - finish a bounce buffer operation - * @bio: completed bio - * @is_error: %true if an I/O error occurred and data should not be copied - * @mark_dirty: If %true, folios will be marked dirty. - * - * Helper for direct I/O implementations that need to bounce buffer because - * we need to checksum the data or perform other operations that require - * consistency. Called to complete a bio set up by bio_iov_iter_bounce(). - * Copies data back for reads, and marks the original folios dirty if - * requested and then frees the bounce buffer. - */ -void bio_iov_iter_unbounce(struct bio *bio, bool is_error, bool mark_dirty) -{ - if (op_is_write(bio_op(bio))) - bio_free_folios(bio); - else - bio_iov_iter_unbounce_read(bio, is_error, mark_dirty); + return 0; } +EXPORT_SYMBOL_GPL(bio_iov_iter_bounce_write); static void bio_wait_end_io(struct bio *bio) { diff --git a/block/blk-map.c b/block/blk-map.c index 9cb9605d1f62..81cba3af4e9c 100644 --- a/block/blk-map.c +++ b/block/blk-map.c @@ -274,7 +274,7 @@ static int bio_map_user_iov(struct request *rq, struct iov_iter *iter, * No alignment requirements on our part to support arbitrary * passthrough commands. */ - ret = bio_iov_iter_get_pages(bio, iter, 0, 0); + ret = bio_iov_iter_get_pages(bio, iter, BIO_MAX_SIZE, 0, 0); if (ret) goto out_put; ret = blk_rq_append_bio(rq, bio); diff --git a/block/blk-settings.c b/block/blk-settings.c index 8274631290db..e469baa1f08b 100644 --- a/block/blk-settings.c +++ b/block/blk-settings.c @@ -206,6 +206,12 @@ static int blk_validate_integrity_limits(struct queue_limits *lim) lim->max_sectors = min(lim->max_sectors, max_integrity_io_size(lim) >> SECTOR_SHIFT); + if (lim->features & BLK_FEAT_ATOMIC_WRITES) { + lim->atomic_write_max_sectors = + min(lim->atomic_write_max_sectors, + max_integrity_io_size(lim) >> SECTOR_SHIFT); + } + return 0; } diff --git a/block/fops.c b/block/fops.c index 2ce7c6c4714e..a83df69b175a 100644 --- a/block/fops.c +++ b/block/fops.c @@ -46,7 +46,8 @@ static bool blkdev_dio_invalid(struct block_device *bdev, struct kiocb *iocb, static inline int blkdev_iov_iter_get_pages(struct bio *bio, struct iov_iter *iter, struct block_device *bdev) { - return bio_iov_iter_get_pages(bio, iter, bdev_dma_alignment(bdev), + return bio_iov_iter_get_pages(bio, iter, BIO_MAX_SIZE, + bdev_dma_alignment(bdev), bdev_logical_block_size(bdev) - 1); } diff --git a/block/partitions/efi.h b/block/partitions/efi.h index 84b9f36b9e47..1f56f93b2804 100644 --- a/block/partitions/efi.h +++ b/block/partitions/efi.h @@ -75,18 +75,12 @@ typedef struct _gpt_header { */ } __packed gpt_header; -typedef struct _gpt_entry_attributes { - u64 required_to_function:1; - u64 reserved:47; - u64 type_guid_specific:16; -} __packed gpt_entry_attributes; - typedef struct _gpt_entry { efi_guid_t partition_type_guid; efi_guid_t unique_partition_guid; __le64 starting_lba; __le64 ending_lba; - gpt_entry_attributes attributes; + __le64 attributes; __le16 partition_name[72/sizeof(__le16)]; } __packed gpt_entry; diff --git a/drivers/acpi/acpi_configfs.c b/drivers/acpi/acpi_configfs.c index 12ffec795803..6071699c7165 100644 --- a/drivers/acpi/acpi_configfs.c +++ b/drivers/acpi/acpi_configfs.c @@ -91,7 +91,7 @@ static ssize_t acpi_table_aml_read(struct config_item *cfg, CONFIGFS_BIN_ATTR(acpi_table_, aml, NULL, MAX_ACPI_TABLE_SIZE); -static struct configfs_bin_attribute *acpi_table_bin_attrs[] = { +static const struct configfs_bin_attribute *const acpi_table_bin_attrs[] = { &acpi_table_attr_aml, NULL, }; diff --git a/drivers/android/binder/rust_binderfs.c b/drivers/android/binder/rust_binderfs.c index 46a37f1858dd..d48895f84596 100644 --- a/drivers/android/binder/rust_binderfs.c +++ b/drivers/android/binder/rust_binderfs.c @@ -340,7 +340,7 @@ static inline bool is_binderfs_control_device(const struct dentry *dentry) return info->control_dentry == dentry; } -static int binderfs_rename(struct mnt_idmap *idmap, +static int binderfs_rename(const struct mnt_idmap *idmap, struct inode *old_dir, struct dentry *old_dentry, struct inode *new_dir, struct dentry *new_dentry, unsigned int flags) diff --git a/drivers/android/binderfs.c b/drivers/android/binderfs.c index 1ca31b9a0583..1f0538767f95 100644 --- a/drivers/android/binderfs.c +++ b/drivers/android/binderfs.c @@ -343,7 +343,7 @@ static inline bool is_binderfs_control_device(const struct dentry *dentry) return info->control_dentry == dentry; } -static int binderfs_rename(struct mnt_idmap *idmap, +static int binderfs_rename(const struct mnt_idmap *idmap, struct inode *old_dir, struct dentry *old_dentry, struct inode *new_dir, struct dentry *new_dentry, unsigned int flags) diff --git a/drivers/base/devtmpfs.c b/drivers/base/devtmpfs.c index aef0fcc6aba1..11c70888f38b 100644 --- a/drivers/base/devtmpfs.c +++ b/drivers/base/devtmpfs.c @@ -72,39 +72,90 @@ static struct file_system_type internal_fs_type = { .kill_sb = kill_anon_super, }; -/* Simply take a ref on the existing mount */ +struct devtmpfs_context { + struct fs_context *fc; +}; + +static void devtmpfs_free(struct fs_context *fc) +{ + struct devtmpfs_context *ctx = fc->fs_private; + + if (ctx) { + put_fs_context(ctx->fc); + kfree(ctx); + } +} + +static int devtmpfs_parse_param(struct fs_context *fc, struct fs_parameter *param) +{ + struct devtmpfs_context *ctx = fc->fs_private; + + return ctx->fc->ops->parse_param(ctx->fc, param); +} + +static int devtmpfs_parse_monolithic(struct fs_context *fc, void *data) +{ + struct devtmpfs_context *ctx = fc->fs_private; + + if (ctx->fc->ops->parse_monolithic) + return ctx->fc->ops->parse_monolithic(ctx->fc, data); + return generic_parse_monolithic(ctx->fc, data); +} + static int devtmpfs_get_tree(struct fs_context *fc) { + struct devtmpfs_context *ctx = fc->fs_private; struct super_block *sb = mnt->mnt_sb; + int err; atomic_inc(&sb->s_active); down_write(&sb->s_umount); + + if (ctx->fc->ops->reconfigure) { + err = ctx->fc->ops->reconfigure(ctx->fc); + if (err) { + deactivate_locked_super(sb); + return err; + } + } + fc->root = dget(sb->s_root); return 0; } -/* Ops are filled in during init depending on underlying shmem or ramfs type */ -static struct fs_context_operations devtmpfs_context_ops = {}; +static const struct fs_context_operations devtmpfs_context_ops = { + .free = devtmpfs_free, + .parse_param = devtmpfs_parse_param, + .parse_monolithic = devtmpfs_parse_monolithic, + .get_tree = devtmpfs_get_tree, +}; -/* Call the underlying initialization and set to our ops */ static int devtmpfs_init_fs_context(struct fs_context *fc) { - int ret; -#ifdef CONFIG_TMPFS - ret = shmem_init_fs_context(fc); -#else - ret = ramfs_init_fs_context(fc); -#endif - if (ret < 0) - return ret; + struct devtmpfs_context *ctx; + int err; + + ctx = kzalloc_obj(struct devtmpfs_context); + if (!ctx) + return -ENOMEM; + + /* Each mount will reconfigure the shared superblock w/ new options */ + ctx->fc = fs_context_for_reconfigure(mnt->mnt_root, + mnt->mnt_sb->s_flags, MS_RMT_MASK); + if (IS_ERR(ctx->fc)) { + err = PTR_ERR(ctx->fc); + kfree(ctx); + return err; + } + fc->fs_private = ctx; fc->ops = &devtmpfs_context_ops; return 0; } static struct file_system_type dev_fs_type = { - .name = "devtmpfs", + .name = "devtmpfs", .init_fs_context = devtmpfs_init_fs_context, }; @@ -443,31 +494,6 @@ static int __ref devtmpfsd(void *p) } /* - * Get the underlying (shmem/ramfs) context ops to build ours - */ -static int devtmpfs_configure_context(void) -{ - struct fs_context *fc; - - fc = fs_context_for_reconfigure(mnt->mnt_root, mnt->mnt_sb->s_flags, - MS_RMT_MASK); - if (IS_ERR(fc)) - return PTR_ERR(fc); - - /* Set up devtmpfs_context_ops based on underlying type */ - devtmpfs_context_ops.free = fc->ops->free; - devtmpfs_context_ops.dup = fc->ops->dup; - devtmpfs_context_ops.parse_param = fc->ops->parse_param; - devtmpfs_context_ops.parse_monolithic = fc->ops->parse_monolithic; - devtmpfs_context_ops.get_tree = &devtmpfs_get_tree; - devtmpfs_context_ops.reconfigure = fc->ops->reconfigure; - - put_fs_context(fc); - - return 0; -} - -/* * Create devtmpfs instance, driver-core devices will add their device * nodes here. */ @@ -482,12 +508,6 @@ int __init devtmpfs_init(void) return PTR_ERR(mnt); } - err = devtmpfs_configure_context(); - if (err) { - pr_err("unable to configure devtmpfs type %d\n", err); - return err; - } - err = register_filesystem(&dev_fs_type); if (err) { pr_err("unable to register devtmpfs type %d\n", err); diff --git a/drivers/gpio/gpiolib-cdev.c b/drivers/gpio/gpiolib-cdev.c index 5d53bfcdf726..a127e6efd7a1 100644 --- a/drivers/gpio/gpiolib-cdev.c +++ b/drivers/gpio/gpiolib-cdev.c @@ -377,11 +377,11 @@ static int linehandle_create(struct gpio_device *gdev, void __user *ip) FD_PREPARE(fdf, O_RDONLY | O_CLOEXEC, anon_inode_getfile("gpio-linehandle", &linehandle_fileops, lh, O_RDONLY | O_CLOEXEC)); - if (fdf.err) - return fdf.err; + if (fdf->fd < 0) + return fdf->fd; retain_and_null_ptr(lh); - handlereq.fd = fd_prepare_fd(fdf); + handlereq.fd = fdf->fd; if (copy_to_user(ip, &handlereq, sizeof(handlereq))) return -EFAULT; @@ -1715,11 +1715,11 @@ static int linereq_create(struct gpio_device *gdev, void __user *ip) FD_PREPARE(fdf, O_RDONLY | O_CLOEXEC, anon_inode_getfile("gpio-line", &line_fileops, lr, O_RDONLY | O_CLOEXEC)); - if (fdf.err) - return fdf.err; + if (fdf->fd < 0) + return fdf->fd; retain_and_null_ptr(lr); - ulr.fd = fd_prepare_fd(fdf); + ulr.fd = fdf->fd; if (copy_to_user(ip, &ulr, sizeof(ulr))) return -EFAULT; @@ -2115,11 +2115,11 @@ static int lineevent_create(struct gpio_device *gdev, void __user *ip) FD_PREPARE(fdf, O_RDONLY | O_CLOEXEC, anon_inode_getfile("gpio-event", &lineevent_fileops, le, O_RDONLY | O_CLOEXEC)); - if (fdf.err) - return fdf.err; + if (fdf->fd < 0) + return fdf->fd; retain_and_null_ptr(le); - eventreq.fd = fd_prepare_fd(fdf); + eventreq.fd = fdf->fd; if (copy_to_user(ip, &eventreq, sizeof(eventreq))) return -EFAULT; diff --git a/drivers/gpu/drm/msm/msm_perfcntr.c b/drivers/gpu/drm/msm/msm_perfcntr.c index ce65b1160955..7fa2e858bd08 100644 --- a/drivers/gpu/drm/msm/msm_perfcntr.c +++ b/drivers/gpu/drm/msm/msm_perfcntr.c @@ -543,8 +543,8 @@ msm_ioctl_perfcntr_config(struct drm_device *dev, void *data, struct drm_file *f FD_PREPARE(fdf, O_CLOEXEC, anon_inode_getfile("[msm_perfcntrs]", &stream_fops, stream, 0)); - if (fdf.err) - return fdf.err; + if (fdf->fd < 0) + return fdf->fd; INIT_WORK(&stream->sel_work, sel_worker); kthread_init_work(&stream->sample_work, sample_worker); diff --git a/drivers/gpu/drm/xe/xe_configfs.c b/drivers/gpu/drm/xe/xe_configfs.c index 052cce962161..842c180b2e85 100644 --- a/drivers/gpu/drm/xe/xe_configfs.c +++ b/drivers/gpu/drm/xe/xe_configfs.c @@ -886,7 +886,7 @@ static struct configfs_item_operations xe_config_device_ops = { }; static bool xe_config_device_is_visible(struct config_item *item, - struct configfs_attribute *attr, int n) + const struct configfs_attribute *attr, int n) { struct xe_config_group_device *dev = to_xe_config_group_device(item); @@ -981,7 +981,7 @@ static struct configfs_attribute *xe_config_sriov_attrs[] = { }; static bool xe_config_sriov_is_visible(struct config_item *item, - struct configfs_attribute *attr, int n) + const struct configfs_attribute *attr, int n) { struct xe_config_group_device *dev = to_xe_config_group_device(item->ci_parent); diff --git a/drivers/media/mc/mc-request.c b/drivers/media/mc/mc-request.c index 13e77648807c..e1387f039780 100644 --- a/drivers/media/mc/mc-request.c +++ b/drivers/media/mc/mc-request.c @@ -316,15 +316,15 @@ int media_request_alloc(struct media_device *mdev, int *alloc_fd) FD_PREPARE(fdf, O_CLOEXEC, anon_inode_getfile("request", &request_fops, NULL, O_CLOEXEC)); - if (fdf.err) { - ret = fdf.err; + if (fdf->fd < 0) { + ret = fdf->fd; goto err_free_req; } - fd_prepare_file(fdf)->private_data = req; + fdf->file->private_data = req; snprintf(req->debug_str, sizeof(req->debug_str), "%u:%d", - atomic_inc_return(&mdev->request_id), fd_prepare_fd(fdf)); + atomic_inc_return(&mdev->request_id), fdf->fd); atomic_inc(&mdev->num_requests); dev_dbg(mdev->dev, "request: allocated %s\n", req->debug_str); diff --git a/drivers/misc/ntsync.c b/drivers/misc/ntsync.c index 4a805919bb0c..2857ae37d3c8 100644 --- a/drivers/misc/ntsync.c +++ b/drivers/misc/ntsync.c @@ -724,9 +724,9 @@ static int ntsync_obj_get_fd(struct ntsync_obj *obj) { FD_PREPARE(fdf, O_CLOEXEC, anon_inode_getfile("ntsync", &ntsync_obj_fops, obj, O_RDWR)); - if (fdf.err) - return fdf.err; - obj->file = fd_prepare_file(fdf); + if (fdf->fd < 0) + return fdf->fd; + obj->file = fdf->file; return fd_publish(fdf); } diff --git a/drivers/virt/coco/guest/report.c b/drivers/virt/coco/guest/report.c index b254a1416286..ad400fbe53f0 100644 --- a/drivers/virt/coco/guest/report.c +++ b/drivers/virt/coco/guest/report.c @@ -356,7 +356,7 @@ static struct configfs_attribute *tsm_report_attrs[] = { NULL, }; -static struct configfs_bin_attribute *tsm_report_bin_attrs[] = { +static const struct configfs_bin_attribute *const tsm_report_bin_attrs[] = { [TSM_REPORT_INBLOB] = &tsm_report_attr_inblob, [TSM_REPORT_OUTBLOB] = &tsm_report_attr_outblob, [TSM_REPORT_AUXBLOB] = &tsm_report_attr_auxblob, @@ -381,7 +381,7 @@ static struct configfs_item_operations tsm_report_item_ops = { }; static bool tsm_report_is_visible(struct config_item *item, - struct configfs_attribute *attr, int n) + const struct configfs_attribute *attr, int n) { guard(rwsem_read)(&tsm_rwsem); if (!provider.ops) @@ -394,7 +394,7 @@ static bool tsm_report_is_visible(struct config_item *item, } static bool tsm_report_is_bin_visible(struct config_item *item, - struct configfs_bin_attribute *attr, int n) + const struct configfs_bin_attribute *attr, int n) { guard(rwsem_read)(&tsm_rwsem); if (!provider.ops) diff --git a/fs/9p/acl.c b/fs/9p/acl.c index ae7e7cf7523a..c6c7c47d32b9 100644 --- a/fs/9p/acl.c +++ b/fs/9p/acl.c @@ -140,7 +140,7 @@ struct posix_acl *v9fs_iop_get_inode_acl(struct inode *inode, int type, bool rcu } -struct posix_acl *v9fs_iop_get_acl(struct mnt_idmap *idmap, +struct posix_acl *v9fs_iop_get_acl(const struct mnt_idmap *idmap, struct dentry *dentry, int type) { struct v9fs_session_info *v9ses; @@ -152,7 +152,7 @@ struct posix_acl *v9fs_iop_get_acl(struct mnt_idmap *idmap, return v9fs_get_cached_acl(d_inode(dentry), type); } -int v9fs_iop_set_acl(struct mnt_idmap *idmap, struct dentry *dentry, +int v9fs_iop_set_acl(const struct mnt_idmap *idmap, struct dentry *dentry, struct posix_acl *acl, int type) { int retval; diff --git a/fs/9p/acl.h b/fs/9p/acl.h index 333cfcc281da..2d1b24abcd3f 100644 --- a/fs/9p/acl.h +++ b/fs/9p/acl.h @@ -10,9 +10,9 @@ int v9fs_get_acl(struct inode *inode, struct p9_fid *fid); struct posix_acl *v9fs_iop_get_inode_acl(struct inode *inode, int type, bool rcu); -struct posix_acl *v9fs_iop_get_acl(struct mnt_idmap *idmap, +struct posix_acl *v9fs_iop_get_acl(const struct mnt_idmap *idmap, struct dentry *dentry, int type); -int v9fs_iop_set_acl(struct mnt_idmap *idmap, struct dentry *dentry, +int v9fs_iop_set_acl(const struct mnt_idmap *idmap, struct dentry *dentry, struct posix_acl *acl, int type); int v9fs_acl_chmod(struct inode *inode, struct p9_fid *fid); int v9fs_set_create_acl(struct inode *inode, struct p9_fid *fid, diff --git a/fs/9p/v9fs.h b/fs/9p/v9fs.h index a462bcbfc7da..54a4a4ec5c15 100644 --- a/fs/9p/v9fs.h +++ b/fs/9p/v9fs.h @@ -188,7 +188,7 @@ extern struct dentry *v9fs_vfs_lookup(struct inode *dir, struct dentry *dentry, unsigned int flags); extern int v9fs_vfs_unlink(struct inode *i, struct dentry *d); extern int v9fs_vfs_rmdir(struct inode *i, struct dentry *d); -extern int v9fs_vfs_rename(struct mnt_idmap *idmap, +extern int v9fs_vfs_rename(const struct mnt_idmap *idmap, struct inode *old_dir, struct dentry *old_dentry, struct inode *new_dir, struct dentry *new_dentry, unsigned int flags); diff --git a/fs/9p/v9fs_vfs.h b/fs/9p/v9fs_vfs.h index 1856d91f8703..5e7b60042498 100644 --- a/fs/9p/v9fs_vfs.h +++ b/fs/9p/v9fs_vfs.h @@ -74,7 +74,7 @@ int v9fs_file_open(struct inode *inode, struct file *file); int v9fs_uflags2omode(int uflags, int extended); void v9fs_blank_wstat(struct p9_wstat *wstat); -int v9fs_vfs_setattr_dotl(struct mnt_idmap *idmap, +int v9fs_vfs_setattr_dotl(const struct mnt_idmap *idmap, struct dentry *dentry, struct iattr *iattr); int v9fs_file_fsync_dotl(struct file *filp, loff_t start, loff_t end, int datasync); diff --git a/fs/9p/vfs_addr.c b/fs/9p/vfs_addr.c index 13cf87a5f90c..170a2b91c5f0 100644 --- a/fs/9p/vfs_addr.c +++ b/fs/9p/vfs_addr.c @@ -150,7 +150,6 @@ static int v9fs_init_request(struct netfs_io_request *rreq, struct file *file) struct p9_fid *fid; struct dentry *dentry; bool writing = (rreq->origin == NETFS_READ_FOR_WRITE || - rreq->origin == NETFS_WRITETHROUGH || rreq->origin == NETFS_UNBUFFERED_WRITE || rreq->origin == NETFS_DIO_WRITE); diff --git a/fs/9p/vfs_inode.c b/fs/9p/vfs_inode.c index 3829554ca369..c95e653344f6 100644 --- a/fs/9p/vfs_inode.c +++ b/fs/9p/vfs_inode.c @@ -652,7 +652,7 @@ error: */ static int -v9fs_vfs_create(struct mnt_idmap *idmap, struct inode *dir, +v9fs_vfs_create(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *dentry, umode_t mode) { struct v9fs_session_info *v9ses = v9fs_inode2v9ses(dir); @@ -679,7 +679,7 @@ v9fs_vfs_create(struct mnt_idmap *idmap, struct inode *dir, * */ -static struct dentry *v9fs_vfs_mkdir(struct mnt_idmap *idmap, struct inode *dir, +static struct dentry *v9fs_vfs_mkdir(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *dentry, umode_t mode) { u32 perm; @@ -858,7 +858,7 @@ int v9fs_vfs_rmdir(struct inode *i, struct dentry *d) */ int -v9fs_vfs_rename(struct mnt_idmap *idmap, struct inode *old_dir, +v9fs_vfs_rename(const struct mnt_idmap *idmap, struct inode *old_dir, struct dentry *old_dentry, struct inode *new_dir, struct dentry *new_dentry, unsigned int flags) { @@ -966,7 +966,7 @@ error: */ static int -v9fs_vfs_getattr(struct mnt_idmap *idmap, const struct path *path, +v9fs_vfs_getattr(const struct mnt_idmap *idmap, const struct path *path, struct kstat *stat, u32 request_mask, unsigned int flags) { struct dentry *dentry = path->dentry; @@ -1014,7 +1014,7 @@ v9fs_vfs_getattr(struct mnt_idmap *idmap, const struct path *path, * */ -static int v9fs_vfs_setattr(struct mnt_idmap *idmap, +static int v9fs_vfs_setattr(const struct mnt_idmap *idmap, struct dentry *dentry, struct iattr *iattr) { int retval, use_dentry = 0; @@ -1249,7 +1249,7 @@ static int v9fs_vfs_mkspecial(struct inode *dir, struct dentry *dentry, */ static int -v9fs_vfs_symlink(struct mnt_idmap *idmap, struct inode *dir, +v9fs_vfs_symlink(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *dentry, const char *symname) { p9_debug(P9_DEBUG_VFS, " %llu,%pd,%s\n", @@ -1304,7 +1304,7 @@ v9fs_vfs_link(struct dentry *old_dentry, struct inode *dir, */ static int -v9fs_vfs_mknod(struct mnt_idmap *idmap, struct inode *dir, +v9fs_vfs_mknod(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *dentry, umode_t mode, dev_t rdev) { struct v9fs_session_info *v9ses = v9fs_inode2v9ses(dir); diff --git a/fs/9p/vfs_inode_dotl.c b/fs/9p/vfs_inode_dotl.c index 116b29e95f21..2cd2898580a9 100644 --- a/fs/9p/vfs_inode_dotl.c +++ b/fs/9p/vfs_inode_dotl.c @@ -29,7 +29,7 @@ #include "acl.h" static int -v9fs_vfs_mknod_dotl(struct mnt_idmap *idmap, struct inode *dir, +v9fs_vfs_mknod_dotl(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *dentry, umode_t omode, dev_t rdev); /** @@ -216,7 +216,7 @@ int v9fs_open_to_dotl_flags(int flags) * */ static int -v9fs_vfs_create_dotl(struct mnt_idmap *idmap, struct inode *dir, +v9fs_vfs_create_dotl(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *dentry, umode_t omode) { return v9fs_vfs_mknod_dotl(idmap, dir, dentry, omode, 0); @@ -344,7 +344,7 @@ out: * */ -static struct dentry *v9fs_vfs_mkdir_dotl(struct mnt_idmap *idmap, +static struct dentry *v9fs_vfs_mkdir_dotl(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *dentry, umode_t omode) { @@ -414,7 +414,7 @@ error: } static int -v9fs_vfs_getattr_dotl(struct mnt_idmap *idmap, +v9fs_vfs_getattr_dotl(const struct mnt_idmap *idmap, const struct path *path, struct kstat *stat, u32 request_mask, unsigned int flags) { @@ -508,7 +508,7 @@ static int v9fs_mapped_iattr_valid(int iattr_valid) * */ -int v9fs_vfs_setattr_dotl(struct mnt_idmap *idmap, +int v9fs_vfs_setattr_dotl(const struct mnt_idmap *idmap, struct dentry *dentry, struct iattr *iattr) { int retval, use_dentry = 0; @@ -682,7 +682,7 @@ v9fs_stat2inode_dotl(struct p9_stat_dotl *stat, struct inode *inode, } static int -v9fs_vfs_symlink_dotl(struct mnt_idmap *idmap, struct inode *dir, +v9fs_vfs_symlink_dotl(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *dentry, const char *symname) { int err; @@ -809,7 +809,7 @@ v9fs_vfs_link_dotl(struct dentry *old_dentry, struct inode *dir, * */ static int -v9fs_vfs_mknod_dotl(struct mnt_idmap *idmap, struct inode *dir, +v9fs_vfs_mknod_dotl(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *dentry, umode_t omode, dev_t rdev) { int err; diff --git a/fs/9p/xattr.c b/fs/9p/xattr.c index 8604e3377ee7..dac06587f67a 100644 --- a/fs/9p/xattr.c +++ b/fs/9p/xattr.c @@ -153,7 +153,7 @@ static int v9fs_xattr_handler_get(const struct xattr_handler *handler, } static int v9fs_xattr_handler_set(const struct xattr_handler *handler, - struct mnt_idmap *idmap, + const struct mnt_idmap *idmap, struct dentry *dentry, struct inode *inode, const char *name, const void *value, size_t size, int flags) diff --git a/fs/Kconfig b/fs/Kconfig index 1454b7fe9641..70a1d0760d28 100644 --- a/fs/Kconfig +++ b/fs/Kconfig @@ -314,7 +314,6 @@ source "fs/ecryptfs/Kconfig" source "fs/hfs/Kconfig" source "fs/hfsplus/Kconfig" source "fs/befs/Kconfig" -source "fs/bfs/Kconfig" source "fs/jffs2/Kconfig" # UBIFS File system configuration source "fs/ubifs/Kconfig" @@ -421,4 +420,12 @@ source "fs/unicode/Kconfig" config IO_WQ bool +config FDTABLE_KUNIT_TEST + bool "KUnit test for fdtable" if !KUNIT_ALL_TESTS + depends on KUNIT=y + default KUNIT_ALL_TESTS + help + This builds the fdtable KUnit tests, which tests various aspects + of the fdtable structure and allocation. + endmenu diff --git a/fs/Makefile b/fs/Makefile index 055dfc23d82b..16f1108b64e1 100644 --- a/fs/Makefile +++ b/fs/Makefile @@ -76,7 +76,6 @@ obj-$(CONFIG_CODA_FS) += coda/ obj-$(CONFIG_MINIX_FS) += minix/ obj-$(CONFIG_FAT_FS) += fat/ obj-$(CONFIG_EXFAT_FS) += exfat/ -obj-$(CONFIG_BFS_FS) += bfs/ obj-$(CONFIG_ISO9660_FS) += isofs/ obj-$(CONFIG_HFSPLUS_FS) += hfsplus/ # Before hfs to find wrapped HFS+ obj-$(CONFIG_HFS_FS) += hfs/ diff --git a/fs/adfs/adfs.h b/fs/adfs/adfs.h index 0d32b7cd99b4..6003832277f8 100644 --- a/fs/adfs/adfs.h +++ b/fs/adfs/adfs.h @@ -144,7 +144,7 @@ struct adfs_discmap { /* Inode stuff */ struct inode *adfs_iget(struct super_block *sb, struct object_info *obj); int adfs_write_inode(struct inode *inode, struct writeback_control *wbc); -int adfs_setattr(struct mnt_idmap *idmap, struct dentry *dentry, +int adfs_setattr(const struct mnt_idmap *idmap, struct dentry *dentry, struct iattr *attr); /* map.c */ diff --git a/fs/adfs/dir.c b/fs/adfs/dir.c index 11afa9e157aa..b8cc6a697a05 100644 --- a/fs/adfs/dir.c +++ b/fs/adfs/dir.c @@ -191,7 +191,7 @@ static int adfs_dir_sync(struct adfs_dir *dir) for (i = dir->nr_buffers - 1; i >= 0; i--) { struct buffer_head *bh = dir->bhs[i]; sync_dirty_buffer(bh); - if (buffer_req(bh) && !buffer_uptodate(bh)) + if (buffer_write_io_error(bh)) err = -EIO; } diff --git a/fs/adfs/inode.c b/fs/adfs/inode.c index 4ac442d0a8c0..34598b499372 100644 --- a/fs/adfs/inode.c +++ b/fs/adfs/inode.c @@ -299,7 +299,7 @@ out: * later. */ int -adfs_setattr(struct mnt_idmap *idmap, struct dentry *dentry, struct iattr *attr) +adfs_setattr(const struct mnt_idmap *idmap, struct dentry *dentry, struct iattr *attr) { struct inode *inode = d_inode(dentry); struct super_block *sb = inode->i_sb; diff --git a/fs/affs/affs.h b/fs/affs/affs.h index d1c506c1f310..2518de96a0c8 100644 --- a/fs/affs/affs.h +++ b/fs/affs/affs.h @@ -166,17 +166,17 @@ extern const struct export_operations affs_export_ops; extern int affs_hash_name(struct super_block *sb, const u8 *name, unsigned int len); extern struct dentry *affs_lookup(struct inode *dir, struct dentry *dentry, unsigned int); extern int affs_unlink(struct inode *dir, struct dentry *dentry); -extern int affs_create(struct mnt_idmap *idmap, struct inode *dir, +extern int affs_create(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *dentry, umode_t mode); -extern struct dentry *affs_mkdir(struct mnt_idmap *idmap, struct inode *dir, +extern struct dentry *affs_mkdir(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *dentry, umode_t mode); extern int affs_rmdir(struct inode *dir, struct dentry *dentry); extern int affs_link(struct dentry *olddentry, struct inode *dir, struct dentry *dentry); -extern int affs_symlink(struct mnt_idmap *idmap, +extern int affs_symlink(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *dentry, const char *symname); -extern int affs_rename2(struct mnt_idmap *idmap, +extern int affs_rename2(const struct mnt_idmap *idmap, struct inode *old_dir, struct dentry *old_dentry, struct inode *new_dir, struct dentry *new_dentry, unsigned int flags); @@ -184,7 +184,7 @@ extern int affs_rename2(struct mnt_idmap *idmap, /* inode.c */ extern struct inode *affs_new_inode(struct inode *dir); -extern int affs_setattr(struct mnt_idmap *idmap, +extern int affs_setattr(const struct mnt_idmap *idmap, struct dentry *dentry, struct iattr *attr); extern void affs_evict_inode(struct inode *inode); extern struct inode *affs_iget(struct super_block *sb, diff --git a/fs/affs/inode.c b/fs/affs/inode.c index d4a3f381c4bc..2a48d7422091 100644 --- a/fs/affs/inode.c +++ b/fs/affs/inode.c @@ -213,7 +213,7 @@ affs_write_inode(struct inode *inode, struct writeback_control *wbc) } int -affs_setattr(struct mnt_idmap *idmap, struct dentry *dentry, struct iattr *attr) +affs_setattr(const struct mnt_idmap *idmap, struct dentry *dentry, struct iattr *attr) { struct inode *inode = d_inode(dentry); int error; diff --git a/fs/affs/namei.c b/fs/affs/namei.c index 6cb52efafe5f..2e32899a32d5 100644 --- a/fs/affs/namei.c +++ b/fs/affs/namei.c @@ -242,7 +242,7 @@ affs_unlink(struct inode *dir, struct dentry *dentry) } int -affs_create(struct mnt_idmap *idmap, struct inode *dir, +affs_create(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *dentry, umode_t mode) { struct super_block *sb = dir->i_sb; @@ -274,7 +274,7 @@ affs_create(struct mnt_idmap *idmap, struct inode *dir, } struct dentry * -affs_mkdir(struct mnt_idmap *idmap, struct inode *dir, +affs_mkdir(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *dentry, umode_t mode) { struct inode *inode; @@ -313,7 +313,7 @@ affs_rmdir(struct inode *dir, struct dentry *dentry) } int -affs_symlink(struct mnt_idmap *idmap, struct inode *dir, +affs_symlink(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *dentry, const char *symname) { struct super_block *sb = dir->i_sb; @@ -503,7 +503,7 @@ done: return retval; } -int affs_rename2(struct mnt_idmap *idmap, struct inode *old_dir, +int affs_rename2(const struct mnt_idmap *idmap, struct inode *old_dir, struct dentry *old_dentry, struct inode *new_dir, struct dentry *new_dentry, unsigned int flags) { diff --git a/fs/afs/dir.c b/fs/afs/dir.c index 2db534a2c7cc..75f8f70cfeff 100644 --- a/fs/afs/dir.c +++ b/fs/afs/dir.c @@ -33,17 +33,17 @@ static bool afs_lookup_one_filldir(struct dir_context *ctx, const char *name, in static bool afs_lookup_filldir(struct dir_context *ctx, const char *name, int nlen, u64 ino, u32 uniquifier); #define AFS_LOOKUP ((filldir_t)0x137UL) -static int afs_create(struct mnt_idmap *idmap, struct inode *dir, +static int afs_create(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *dentry, umode_t mode); -static struct dentry *afs_mkdir(struct mnt_idmap *idmap, struct inode *dir, +static struct dentry *afs_mkdir(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *dentry, umode_t mode); static int afs_rmdir(struct inode *dir, struct dentry *dentry); static int afs_unlink(struct inode *dir, struct dentry *dentry); static int afs_link(struct dentry *from, struct inode *dir, struct dentry *dentry); -static int afs_symlink(struct mnt_idmap *idmap, struct inode *dir, +static int afs_symlink(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *dentry, const char *content); -static int afs_rename(struct mnt_idmap *idmap, struct inode *old_dir, +static int afs_rename(const struct mnt_idmap *idmap, struct inode *old_dir, struct dentry *old_dentry, struct inode *new_dir, struct dentry *new_dentry, unsigned int flags); static int afs_dir_writepages(struct address_space *mapping, @@ -1310,7 +1310,7 @@ static const struct afs_operation_ops afs_mkdir_operation = { /* * create a directory on an AFS filesystem */ -static struct dentry *afs_mkdir(struct mnt_idmap *idmap, struct inode *dir, +static struct dentry *afs_mkdir(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *dentry, umode_t mode) { struct afs_operation *op; @@ -1632,7 +1632,7 @@ static const struct afs_operation_ops afs_create_operation = { /* * create a regular file on an AFS filesystem */ -static int afs_create(struct mnt_idmap *idmap, struct inode *dir, +static int afs_create(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *dentry, umode_t mode) { struct afs_operation *op; @@ -1779,7 +1779,7 @@ static const struct afs_operation_ops afs_symlink_operation = { /* * create a symlink in an AFS filesystem */ -static int afs_symlink(struct mnt_idmap *idmap, struct inode *dir, +static int afs_symlink(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *dentry, const char *content) { struct afs_operation *op; @@ -2067,7 +2067,7 @@ static const struct afs_operation_ops afs_rename_exchange_operation = { /* * rename a file in an AFS filesystem and/or move it between directories */ -static int afs_rename(struct mnt_idmap *idmap, struct inode *old_dir, +static int afs_rename(const struct mnt_idmap *idmap, struct inode *old_dir, struct dentry *old_dentry, struct inode *new_dir, struct dentry *new_dentry, unsigned int flags) { diff --git a/fs/afs/file.c b/fs/afs/file.c index 0467742bfeee..4c78d3441785 100644 --- a/fs/afs/file.c +++ b/fs/afs/file.c @@ -400,7 +400,6 @@ static int afs_init_request(struct netfs_io_request *rreq, struct file *file) } break; case NETFS_WRITEBACK: - case NETFS_WRITETHROUGH: case NETFS_UNBUFFERED_WRITE: case NETFS_DIO_WRITE: if (S_ISREG(rreq->inode->i_mode)) @@ -413,7 +412,7 @@ static int afs_init_request(struct netfs_io_request *rreq, struct file *file) return 0; } -static int afs_check_write_begin(struct file *file, loff_t pos, unsigned len, +static int afs_check_write_begin(struct file *file, uoff_t pos, unsigned len, struct folio **foliop, void **_fsdata) { struct afs_vnode *vnode = AFS_FS_I(file_inode(file)); @@ -434,7 +433,7 @@ static void afs_free_request(struct netfs_io_request *rreq) * Also, estimate the number of 512 bytes blocks used, rounded up to nearest 1K * for consistency with other AFS clients. */ -void afs_set_i_size(struct afs_vnode *vnode, loff_t new_i_size) +void afs_set_i_size(struct afs_vnode *vnode, uoff_t new_i_size) { struct inode *inode = &vnode->netfs.inode; loff_t i_size; @@ -448,10 +447,9 @@ void afs_set_i_size(struct afs_vnode *vnode, loff_t new_i_size) } spin_unlock(&inode->i_lock); write_sequnlock(&vnode->cb_lock); - fscache_update_cookie(afs_vnode_cache(vnode), NULL, &new_i_size); } -static void afs_update_i_size(struct inode *inode, loff_t new_i_size) +static void afs_update_i_size(struct inode *inode, uoff_t new_i_size) { afs_set_i_size(AFS_FS_I(inode), new_i_size); } diff --git a/fs/afs/inode.c b/fs/afs/inode.c index 14f39a9bea6c..8e6ca6b45c6b 100644 --- a/fs/afs/inode.c +++ b/fs/afs/inode.c @@ -596,7 +596,7 @@ error: /* * read the attributes of an inode */ -int afs_getattr(struct mnt_idmap *idmap, const struct path *path, +int afs_getattr(const struct mnt_idmap *idmap, const struct path *path, struct kstat *stat, u32 request_mask, unsigned int query_flags) { struct inode *inode = d_inode(path->dentry); @@ -759,7 +759,7 @@ static const struct afs_operation_ops afs_setattr_operation = { /* * set the attributes of an inode */ -int afs_setattr(struct mnt_idmap *idmap, struct dentry *dentry, +int afs_setattr(const struct mnt_idmap *idmap, struct dentry *dentry, struct iattr *attr) { const unsigned int supported = diff --git a/fs/afs/internal.h b/fs/afs/internal.h index 330654ed16ec..5744a347ce2e 100644 --- a/fs/afs/internal.h +++ b/fs/afs/internal.h @@ -1171,7 +1171,7 @@ extern int afs_open(struct inode *, struct file *); extern int afs_release(struct inode *, struct file *); void afs_fetch_data_async_rx(struct work_struct *work); void afs_fetch_data_immediate_cancel(struct afs_call *call); -void afs_set_i_size(struct afs_vnode *vnode, loff_t new_i_size); +void afs_set_i_size(struct afs_vnode *vnode, uoff_t new_i_size); /* * flock.c @@ -1266,9 +1266,9 @@ extern int afs_fetch_status(struct afs_vnode *, struct key *, bool, afs_access_t extern int afs_ilookup5_test_by_fid(struct inode *, void *); extern struct inode *afs_iget(struct afs_operation *, struct afs_vnode_param *); extern struct inode *afs_root_iget(struct super_block *, struct key *); -extern int afs_getattr(struct mnt_idmap *idmap, const struct path *, +extern int afs_getattr(const struct mnt_idmap *idmap, const struct path *, struct kstat *, u32, unsigned int); -extern int afs_setattr(struct mnt_idmap *idmap, struct dentry *, struct iattr *); +extern int afs_setattr(const struct mnt_idmap *idmap, struct dentry *, struct iattr *); extern void afs_evict_inode(struct inode *); extern int afs_drop_inode(struct inode *); @@ -1538,7 +1538,7 @@ extern void afs_cache_permit(struct afs_vnode *, struct key *, unsigned int, extern struct key *afs_request_key(struct afs_cell *); extern struct key *afs_request_key_rcu(struct afs_cell *); extern int afs_check_permit(struct afs_vnode *, struct key *, afs_access_t *); -extern int afs_permission(struct mnt_idmap *, struct inode *, int); +extern int afs_permission(const struct mnt_idmap *, struct inode *, int); extern void __exit afs_clean_up_permit_cache(void); /* diff --git a/fs/afs/security.c b/fs/afs/security.c index 6d00d62a65ed..fd040f7c6766 100644 --- a/fs/afs/security.c +++ b/fs/afs/security.c @@ -428,7 +428,7 @@ int afs_check_permit(struct afs_vnode *vnode, struct key *key, * - AFS ACLs are attached to directories only, and a file is controlled by its * parent directory's ACL */ -int afs_permission(struct mnt_idmap *idmap, struct inode *inode, +int afs_permission(const struct mnt_idmap *idmap, struct inode *inode, int mask) { struct afs_vnode *vnode = AFS_FS_I(inode); diff --git a/fs/afs/xattr.c b/fs/afs/xattr.c index 3770ed236f67..bcffd7236cc9 100644 --- a/fs/afs/xattr.c +++ b/fs/afs/xattr.c @@ -97,7 +97,7 @@ static const struct afs_operation_ops afs_store_acl_operation = { * Set a file's AFS3 ACL. */ static int afs_xattr_set_acl(const struct xattr_handler *handler, - struct mnt_idmap *idmap, + const struct mnt_idmap *idmap, struct dentry *dentry, struct inode *inode, const char *name, const void *buffer, size_t size, int flags) @@ -228,7 +228,7 @@ static const struct afs_operation_ops yfs_store_opaque_acl2_operation = { * Set a file's YFS ACL. */ static int afs_xattr_set_yfs(const struct xattr_handler *handler, - struct mnt_idmap *idmap, + const struct mnt_idmap *idmap, struct dentry *dentry, struct inode *inode, const char *name, const void *buffer, size_t size, int flags) @@ -936,7 +936,7 @@ static int kill_ioctx(struct mm_struct *mm, struct kioctx *ctx, /* * exit_aio: called when the last user of mm goes away. At this point, there is - * no way for any new requests to be submited or any of the io_* syscalls to be + * no way for any new requests to be submitted or any of the io_* syscalls to be * called on the context. * * There may be outstanding kiocbs, but free_ioctx() will explicitly wait on @@ -1280,7 +1280,7 @@ static long aio_read_events_ring(struct kioctx *ctx, * The mutex can block and wake us up and that will cause * wait_event_interruptible_hrtimeout() to schedule without sleeping * and repeat. This should be rare enough that it doesn't cause - * peformance issues. See the comment in read_events() for more detail. + * performance issues. See the comment in read_events() for more detail. */ sched_annotate_sleep(); mutex_lock(&ctx->ring_lock); @@ -1869,7 +1869,12 @@ static int aio_poll_wake(struct wait_queue_entry *wait, unsigned mode, int sync, list_del_init(&req->wait.entry); list_del(&iocb->ki_list); iocb->ki_res.res = mangle_poll(mask); - if (iocb->ki_eventfd && !eventfd_signal_allowed()) { + /* + * We hold an arbitrary provider waitqueue lock here. Signaling a + * result eventfd can feed back through epoll and try to take the same + * lock again. Defer all eventfd-backed poll completions. + */ + if (iocb->ki_eventfd) { iocb = NULL; INIT_WORK(&req->work, aio_poll_put_work); schedule_work(&req->work); diff --git a/fs/anon_inodes.c b/fs/anon_inodes.c index a7b9b948e33d..8f07c8d9bda0 100644 --- a/fs/anon_inodes.c +++ b/fs/anon_inodes.c @@ -46,7 +46,7 @@ static struct inode *anon_inode_inode __ro_after_init; * Rather than mess with our internal sane inode data, just fix it * up here in getattr() by masking off the format bits. */ -int anon_inode_getattr(struct mnt_idmap *idmap, const struct path *path, +int anon_inode_getattr(const struct mnt_idmap *idmap, const struct path *path, struct kstat *stat, u32 request_mask, unsigned int query_flags) { @@ -57,7 +57,7 @@ int anon_inode_getattr(struct mnt_idmap *idmap, const struct path *path, return 0; } -int anon_inode_setattr(struct mnt_idmap *idmap, struct dentry *dentry, +int anon_inode_setattr(const struct mnt_idmap *idmap, struct dentry *dentry, struct iattr *attr) { return -EOPNOTSUPP; diff --git a/fs/attr.c b/fs/attr.c index 71888ac903c2..9706d370f34a 100644 --- a/fs/attr.c +++ b/fs/attr.c @@ -30,7 +30,7 @@ * * Return: ATTR_KILL_SGID if setgid bit needs to be removed, 0 otherwise. */ -int setattr_should_drop_sgid(struct mnt_idmap *idmap, +int setattr_should_drop_sgid(const struct mnt_idmap *idmap, const struct inode *inode) { umode_t mode = inode->i_mode; @@ -60,7 +60,7 @@ EXPORT_SYMBOL(setattr_should_drop_sgid); * Return: A mask of ATTR_KILL_S{G,U}ID indicating which - if any - setid bits * to remove, 0 otherwise. */ -int setattr_should_drop_suidgid(struct mnt_idmap *idmap, +int setattr_should_drop_suidgid(const struct mnt_idmap *idmap, struct inode *inode) { umode_t mode = inode->i_mode; @@ -91,7 +91,7 @@ EXPORT_SYMBOL(setattr_should_drop_suidgid); * permissions. On non-idmapped mounts or if permission checking is to be * performed on the raw inode simply pass @nop_mnt_idmap. */ -static bool chown_ok(struct mnt_idmap *idmap, +static bool chown_ok(const struct mnt_idmap *idmap, const struct inode *inode, vfsuid_t ia_vfsuid) { vfsuid_t vfsuid = i_uid_into_vfsuid(idmap, inode); @@ -118,7 +118,7 @@ static bool chown_ok(struct mnt_idmap *idmap, * permissions. On non-idmapped mounts or if permission checking is to be * performed on the raw inode simply pass @nop_mnt_idmap. */ -static bool chgrp_ok(struct mnt_idmap *idmap, +static bool chgrp_ok(const struct mnt_idmap *idmap, const struct inode *inode, vfsgid_t ia_vfsgid) { vfsgid_t vfsgid = i_gid_into_vfsgid(idmap, inode); @@ -158,7 +158,7 @@ static bool chgrp_ok(struct mnt_idmap *idmap, * Should be called as the first thing in ->setattr implementations, * possibly after taking additional locks. */ -int setattr_prepare(struct mnt_idmap *idmap, struct dentry *dentry, +int setattr_prepare(const struct mnt_idmap *idmap, struct dentry *dentry, struct iattr *attr) { struct inode *inode = d_inode(dentry); @@ -339,7 +339,7 @@ static void setattr_copy_mgtime(struct inode *inode, const struct iattr *attr) * that for "simple" filesystems, the struct inode is the inode storage. * The caller is free to mark the inode dirty afterwards if needed. */ -void setattr_copy(struct mnt_idmap *idmap, struct inode *inode, +void setattr_copy(const struct mnt_idmap *idmap, struct inode *inode, const struct iattr *attr) { unsigned int ia_valid = attr->ia_valid; @@ -369,7 +369,7 @@ void setattr_copy(struct mnt_idmap *idmap, struct inode *inode, } EXPORT_SYMBOL(setattr_copy); -int may_setattr(struct mnt_idmap *idmap, struct inode *inode, +int may_setattr(const struct mnt_idmap *idmap, struct inode *inode, unsigned int ia_valid) { int error; @@ -424,7 +424,7 @@ EXPORT_SYMBOL(may_setattr); * permissions. On non-idmapped mounts or if permission checking is to be * performed on the raw inode simply pass @nop_mnt_idmap. */ -int notify_change(struct mnt_idmap *idmap, struct dentry *dentry, +int notify_change(const struct mnt_idmap *idmap, struct dentry *dentry, struct iattr *attr, struct delegated_inode *delegated_inode) { struct inode *inode = dentry->d_inode; diff --git a/fs/autofs/root.c b/fs/autofs/root.c index b36439f4521e..28f38f5d0236 100644 --- a/fs/autofs/root.c +++ b/fs/autofs/root.c @@ -11,12 +11,12 @@ #include "autofs_i.h" -static int autofs_dir_permission(struct mnt_idmap *, struct inode *, int); -static int autofs_dir_symlink(struct mnt_idmap *, struct inode *, +static int autofs_dir_permission(const struct mnt_idmap *, struct inode *, int); +static int autofs_dir_symlink(const struct mnt_idmap *, struct inode *, struct dentry *, const char *); static int autofs_dir_unlink(struct inode *, struct dentry *); static int autofs_dir_rmdir(struct inode *, struct dentry *); -static struct dentry *autofs_dir_mkdir(struct mnt_idmap *, struct inode *, +static struct dentry *autofs_dir_mkdir(const struct mnt_idmap *, struct inode *, struct dentry *, umode_t); static long autofs_root_ioctl(struct file *, unsigned int, unsigned long); #ifdef CONFIG_COMPAT @@ -552,7 +552,7 @@ static struct dentry *autofs_lookup(struct inode *dir, return NULL; } -static int autofs_dir_permission(struct mnt_idmap *idmap, +static int autofs_dir_permission(const struct mnt_idmap *idmap, struct inode *inode, int mask) { if (mask & MAY_WRITE) { @@ -572,7 +572,7 @@ static int autofs_dir_permission(struct mnt_idmap *idmap, return generic_permission(idmap, inode, mask); } -static int autofs_dir_symlink(struct mnt_idmap *idmap, +static int autofs_dir_symlink(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *dentry, const char *symname) { @@ -724,7 +724,7 @@ static int autofs_dir_rmdir(struct inode *dir, struct dentry *dentry) return 0; } -static struct dentry *autofs_dir_mkdir(struct mnt_idmap *idmap, +static struct dentry *autofs_dir_mkdir(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *dentry, umode_t mode) { diff --git a/fs/backing-file.c b/fs/backing-file.c index cc101143f921..5614cb7801e1 100644 --- a/fs/backing-file.c +++ b/fs/backing-file.c @@ -59,7 +59,7 @@ struct file *backing_tmpfile_open(const struct file *user_file, int flags, const struct path *real_parentpath, umode_t mode, const struct cred *cred) { - struct mnt_idmap *real_idmap = mnt_idmap(real_parentpath->mnt); + const struct mnt_idmap *real_idmap = mnt_idmap(real_parentpath->mnt); const struct path *user_path = &user_file->f_path; struct file *f; int error; diff --git a/fs/bad_inode.c b/fs/bad_inode.c index 486c40f73e51..bea9f4876ee0 100644 --- a/fs/bad_inode.c +++ b/fs/bad_inode.c @@ -27,7 +27,7 @@ static const struct file_operations bad_file_ops = .open = bad_file_open, }; -static int bad_inode_create(struct mnt_idmap *idmap, +static int bad_inode_create(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *dentry, umode_t mode) { @@ -51,14 +51,14 @@ static int bad_inode_unlink(struct inode *dir, struct dentry *dentry) return -EIO; } -static int bad_inode_symlink(struct mnt_idmap *idmap, +static int bad_inode_symlink(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *dentry, const char *symname) { return -EIO; } -static struct dentry *bad_inode_mkdir(struct mnt_idmap *idmap, struct inode *dir, +static struct dentry *bad_inode_mkdir(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *dentry, umode_t mode) { return ERR_PTR(-EIO); @@ -69,13 +69,13 @@ static int bad_inode_rmdir (struct inode *dir, struct dentry *dentry) return -EIO; } -static int bad_inode_mknod(struct mnt_idmap *idmap, struct inode *dir, +static int bad_inode_mknod(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *dentry, umode_t mode, dev_t rdev) { return -EIO; } -static int bad_inode_rename2(struct mnt_idmap *idmap, +static int bad_inode_rename2(const struct mnt_idmap *idmap, struct inode *old_dir, struct dentry *old_dentry, struct inode *new_dir, struct dentry *new_dentry, unsigned int flags) @@ -89,20 +89,20 @@ static int bad_inode_readlink(struct dentry *dentry, char __user *buffer, return -EIO; } -static int bad_inode_permission(struct mnt_idmap *idmap, +static int bad_inode_permission(const struct mnt_idmap *idmap, struct inode *inode, int mask) { return -EIO; } -static int bad_inode_getattr(struct mnt_idmap *idmap, +static int bad_inode_getattr(const struct mnt_idmap *idmap, const struct path *path, struct kstat *stat, u32 request_mask, unsigned int query_flags) { return -EIO; } -static int bad_inode_setattr(struct mnt_idmap *idmap, +static int bad_inode_setattr(const struct mnt_idmap *idmap, struct dentry *direntry, struct iattr *attrs) { return -EIO; @@ -146,14 +146,14 @@ static int bad_inode_atomic_open(struct inode *inode, struct dentry *dentry, return -EIO; } -static int bad_inode_tmpfile(struct mnt_idmap *idmap, +static int bad_inode_tmpfile(const struct mnt_idmap *idmap, struct inode *inode, struct file *file, umode_t mode) { return -EIO; } -static int bad_inode_set_acl(struct mnt_idmap *idmap, +static int bad_inode_set_acl(const struct mnt_idmap *idmap, struct dentry *dentry, struct posix_acl *acl, int type) { diff --git a/fs/bfs/Kconfig b/fs/bfs/Kconfig deleted file mode 100644 index 8e7ef866b62a..000000000000 --- a/fs/bfs/Kconfig +++ /dev/null @@ -1,21 +0,0 @@ -# SPDX-License-Identifier: GPL-2.0-only -config BFS_FS - tristate "BFS file system support" - depends on BLOCK - select BUFFER_HEAD - help - Boot File System (BFS) is a file system used under SCO UnixWare to - allow the bootloader access to the kernel image and other important - files during the boot process. It is usually mounted under /stand - and corresponds to the slice marked as "STAND" in the UnixWare - partition. You should say Y if you want to read or write the files - on your /stand slice from within Linux. You then also need to say Y - to "UnixWare slices support", below. More information about the BFS - file system is contained in the file - <file:Documentation/filesystems/bfs.rst>. - - If you don't know what this is about, say N. - - To compile this as a module, choose M here: the module will be called - bfs. Note that the file system of your root partition (the one - containing the directory /) cannot be compiled as a module. diff --git a/fs/bfs/Makefile b/fs/bfs/Makefile deleted file mode 100644 index 2b6bc5eb4de9..000000000000 --- a/fs/bfs/Makefile +++ /dev/null @@ -1,8 +0,0 @@ -# SPDX-License-Identifier: GPL-2.0-only -# -# Makefile for BFS filesystem. -# - -obj-$(CONFIG_BFS_FS) += bfs.o - -bfs-objs := inode.o file.o dir.o diff --git a/fs/bfs/bfs.h b/fs/bfs/bfs.h deleted file mode 100644 index b08afe733e63..000000000000 --- a/fs/bfs/bfs.h +++ /dev/null @@ -1,69 +0,0 @@ -/* SPDX-License-Identifier: GPL-2.0 */ -/* - * fs/bfs/bfs.h - * Copyright (C) 1999-2018 Tigran Aivazian <aivazian.tigran@gmail.com> - */ -#ifndef _FS_BFS_BFS_H -#define _FS_BFS_BFS_H - -#include <linux/bfs_fs.h> - -/* In theory BFS supports up to 512 inodes, numbered from 2 (for /) up to 513 inclusive. - In actual fact, attempting to create the 512th inode (i.e. inode No. 513 or file No. 511) - will fail with ENOSPC in bfs_add_entry(): the root directory cannot contain so many entries, counting '..'. - So, mkfs.bfs(8) should really limit its -N option to 511 and not 512. For now, we just print a warning - if a filesystem is mounted with such "impossible to fill up" number of inodes */ -#define BFS_MAX_LASTI 513 - -/* - * BFS file system in-core superblock info - */ -struct bfs_sb_info { - unsigned long si_blocks; - unsigned long si_freeb; - unsigned long si_freei; - unsigned long si_lf_eblk; - unsigned long si_lasti; - DECLARE_BITMAP(si_imap, BFS_MAX_LASTI+1); - struct mutex bfs_lock; -}; - -/* - * BFS file system in-core inode info - */ -struct bfs_inode_info { - unsigned long i_dsk_ino; /* inode number from the disk, can be 0 */ - unsigned long i_sblock; - unsigned long i_eblock; - struct mapping_metadata_bhs i_metadata_bhs; - struct inode vfs_inode; -}; - -static inline struct bfs_sb_info *BFS_SB(struct super_block *sb) -{ - return sb->s_fs_info; -} - -static inline struct bfs_inode_info *BFS_I(struct inode *inode) -{ - return container_of(inode, struct bfs_inode_info, vfs_inode); -} - - -#define printf(format, args...) \ - printk(KERN_ERR "BFS-fs: %s(): " format, __func__, ## args) - -/* inode.c */ -extern struct inode *bfs_iget(struct super_block *sb, unsigned long ino); -extern void bfs_dump_imap(const char *, struct super_block *); - -/* file.c */ -extern const struct inode_operations bfs_file_inops; -extern const struct file_operations bfs_file_operations; -extern const struct address_space_operations bfs_aops; - -/* dir.c */ -extern const struct inode_operations bfs_dir_inops; -extern const struct file_operations bfs_dir_operations; - -#endif /* _FS_BFS_BFS_H */ diff --git a/fs/bfs/dir.c b/fs/bfs/dir.c index 91a4871fa051..b944bd62f5d0 100644 --- a/fs/bfs/dir.c +++ b/fs/bfs/dir.c @@ -75,7 +75,7 @@ const struct file_operations bfs_dir_operations = { .llseek = generic_file_llseek, }; -static int bfs_create(struct mnt_idmap *idmap, struct inode *dir, +static int bfs_create(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *dentry, umode_t mode) { int err; @@ -199,7 +199,7 @@ out_brelse: return error; } -static int bfs_rename(struct mnt_idmap *idmap, struct inode *old_dir, +static int bfs_rename(const struct mnt_idmap *idmap, struct inode *old_dir, struct dentry *old_dentry, struct inode *new_dir, struct dentry *new_dentry, unsigned int flags) { diff --git a/fs/bfs/file.c b/fs/bfs/file.c deleted file mode 100644 index d33d6bde992b..000000000000 --- a/fs/bfs/file.c +++ /dev/null @@ -1,203 +0,0 @@ -// SPDX-License-Identifier: GPL-2.0 -/* - * fs/bfs/file.c - * BFS file operations. - * Copyright (C) 1999-2018 Tigran Aivazian <aivazian.tigran@gmail.com> - * - * Make the file block allocation algorithm understand the size - * of the underlying block device. - * Copyright (C) 2007 Dmitri Vorobiev <dmitri.vorobiev@gmail.com> - * - */ - -#include <linux/fs.h> -#include <linux/mpage.h> -#include <linux/buffer_head.h> -#include "bfs.h" - -#undef DEBUG - -#ifdef DEBUG -#define dprintf(x...) printf(x) -#else -#define dprintf(x...) -#endif - -const struct file_operations bfs_file_operations = { - .llseek = generic_file_llseek, - .read_iter = generic_file_read_iter, - .write_iter = generic_file_write_iter, - .mmap_prepare = generic_file_mmap_prepare, - .splice_read = filemap_splice_read, -}; - -static int bfs_move_block(unsigned long from, unsigned long to, - struct super_block *sb) -{ - struct buffer_head *bh, *new; - - bh = sb_bread(sb, from); - if (!bh) - return -EIO; - new = sb_getblk(sb, to); - memcpy(new->b_data, bh->b_data, bh->b_size); - mark_buffer_dirty(new); - bforget(bh); - brelse(new); - return 0; -} - -static int bfs_move_blocks(struct super_block *sb, unsigned long start, - unsigned long end, unsigned long where) -{ - unsigned long i; - - dprintf("%08lx-%08lx->%08lx\n", start, end, where); - for (i = start; i <= end; i++) - if(bfs_move_block(i, where + i, sb)) { - dprintf("failed to move block %08lx -> %08lx\n", i, - where + i); - return -EIO; - } - return 0; -} - -static int bfs_get_block(struct inode *inode, sector_t block, - struct buffer_head *bh_result, int create) -{ - unsigned long phys; - int err; - struct super_block *sb = inode->i_sb; - struct bfs_sb_info *info = BFS_SB(sb); - struct bfs_inode_info *bi = BFS_I(inode); - - phys = bi->i_sblock + block; - if (!create) { - if (phys <= bi->i_eblock) { - dprintf("c=%d, b=%08lx, phys=%09lx (granted)\n", - create, (unsigned long)block, phys); - map_bh(bh_result, sb, phys); - } - return 0; - } - - /* - * If the file is not empty and the requested block is within the - * range of blocks allocated for this file, we can grant it. - */ - if (bi->i_sblock && (phys <= bi->i_eblock)) { - dprintf("c=%d, b=%08lx, phys=%08lx (interim block granted)\n", - create, (unsigned long)block, phys); - map_bh(bh_result, sb, phys); - return 0; - } - - /* The file will be extended, so let's see if there is enough space. */ - if (phys >= info->si_blocks) - return -ENOSPC; - - /* The rest has to be protected against itself. */ - mutex_lock(&info->bfs_lock); - - /* - * If the last data block for this file is the last allocated - * block, we can extend the file trivially, without moving it - * anywhere. - */ - if (bi->i_eblock == info->si_lf_eblk) { - dprintf("c=%d, b=%08lx, phys=%08lx (simple extension)\n", - create, (unsigned long)block, phys); - map_bh(bh_result, sb, phys); - info->si_freeb -= phys - bi->i_eblock; - info->si_lf_eblk = bi->i_eblock = phys; - mark_inode_dirty(inode); - err = 0; - goto out; - } - - /* Ok, we have to move this entire file to the next free block. */ - phys = info->si_lf_eblk + 1; - if (phys + block >= info->si_blocks) { - err = -ENOSPC; - goto out; - } - - if (bi->i_sblock) { - err = bfs_move_blocks(inode->i_sb, bi->i_sblock, - bi->i_eblock, phys); - if (err) { - dprintf("failed to move ino=%08lx -> fs corruption\n", - inode->i_ino); - goto out; - } - } else - err = 0; - - dprintf("c=%d, b=%08lx, phys=%08lx (moved)\n", - create, (unsigned long)block, phys); - bi->i_sblock = phys; - phys += block; - info->si_lf_eblk = bi->i_eblock = phys; - - /* - * This assumes nothing can write the inode back while we are here - * and thus update inode->i_blocks! (XXX) - */ - info->si_freeb -= bi->i_eblock - bi->i_sblock + 1 - inode->i_blocks; - mark_inode_dirty(inode); - map_bh(bh_result, sb, phys); -out: - mutex_unlock(&info->bfs_lock); - return err; -} - -static int bfs_writepages(struct address_space *mapping, - struct writeback_control *wbc) -{ - return mpage_writepages(mapping, wbc, bfs_get_block); -} - -static int bfs_read_folio(struct file *file, struct folio *folio) -{ - return block_read_full_folio(folio, bfs_get_block); -} - -static void bfs_write_failed(struct address_space *mapping, loff_t to) -{ - struct inode *inode = mapping->host; - - if (to > inode->i_size) - truncate_pagecache(inode, inode->i_size); -} - -static int bfs_write_begin(const struct kiocb *iocb, - struct address_space *mapping, - loff_t pos, unsigned len, - struct folio **foliop, void **fsdata) -{ - int ret; - - ret = block_write_begin(mapping, pos, len, foliop, bfs_get_block); - if (unlikely(ret)) - bfs_write_failed(mapping, pos + len); - - return ret; -} - -static sector_t bfs_bmap(struct address_space *mapping, sector_t block) -{ - return generic_block_bmap(mapping, block, bfs_get_block); -} - -const struct address_space_operations bfs_aops = { - .dirty_folio = block_dirty_folio, - .invalidate_folio = block_invalidate_folio, - .read_folio = bfs_read_folio, - .writepages = bfs_writepages, - .write_begin = bfs_write_begin, - .write_end = generic_write_end, - .migrate_folio = buffer_migrate_folio, - .bmap = bfs_bmap, -}; - -const struct inode_operations bfs_file_inops; diff --git a/fs/bfs/inode.c b/fs/bfs/inode.c deleted file mode 100644 index 06e3a848b4ef..000000000000 --- a/fs/bfs/inode.c +++ /dev/null @@ -1,538 +0,0 @@ -// SPDX-License-Identifier: GPL-2.0-only -/* - * fs/bfs/inode.c - * BFS superblock and inode operations. - * Copyright (C) 1999-2018 Tigran Aivazian <aivazian.tigran@gmail.com> - * From fs/minix, Copyright (C) 1991, 1992 Linus Torvalds. - * Made endianness-clean by Andrew Stribblehill <ads@wompom.org>, 2005. - */ - -#include <linux/module.h> -#include <linux/mm.h> -#include <linux/slab.h> -#include <linux/init.h> -#include <linux/fs.h> -#include <linux/buffer_head.h> -#include <linux/vfs.h> -#include <linux/writeback.h> -#include <linux/uio.h> -#include <linux/uaccess.h> -#include <linux/fs_context.h> -#include "bfs.h" - -MODULE_AUTHOR("Tigran Aivazian <aivazian.tigran@gmail.com>"); -MODULE_DESCRIPTION("SCO UnixWare BFS filesystem for Linux"); -MODULE_LICENSE("GPL"); - -#undef DEBUG - -#ifdef DEBUG -#define dprintf(x...) printf(x) -#else -#define dprintf(x...) -#endif - -struct inode *bfs_iget(struct super_block *sb, unsigned long ino) -{ - struct bfs_inode *di; - struct inode *inode; - struct buffer_head *bh; - int block, off; - - inode = iget_locked(sb, ino); - if (!inode) - return ERR_PTR(-ENOMEM); - if (!(inode_state_read_once(inode) & I_NEW)) - return inode; - - if ((ino < BFS_ROOT_INO) || (ino > BFS_SB(inode->i_sb)->si_lasti)) { - printf("Bad inode number %s:%08lx\n", inode->i_sb->s_id, ino); - goto error; - } - - block = (ino - BFS_ROOT_INO) / BFS_INODES_PER_BLOCK + 1; - bh = sb_bread(inode->i_sb, block); - if (!bh) { - printf("Unable to read inode %s:%08lx\n", inode->i_sb->s_id, - ino); - goto error; - } - - off = (ino - BFS_ROOT_INO) % BFS_INODES_PER_BLOCK; - di = (struct bfs_inode *)bh->b_data + off; - - /* - * https://martin.hinner.info/fs/bfs/bfs-structure.html explains that - * BFS in SCO UnixWare environment used only lower 9 bits of di->i_mode - * value. This means that, although bfs_write_inode() saves whole - * inode->i_mode bits (which include S_IFMT bits and S_IS{UID,GID,VTX} - * bits), middle 7 bits of di->i_mode value can be garbage when these - * bits were not saved by bfs_write_inode(). - * Since we can't tell whether middle 7 bits are garbage, use only - * lower 12 bits (i.e. tolerate S_IS{UID,GID,VTX} bits possibly being - * garbage) and reconstruct S_IFMT bits for Linux environment from - * di->i_vtype value. - */ - inode->i_mode = 0x00000FFF & le32_to_cpu(di->i_mode); - if (le32_to_cpu(di->i_vtype) == BFS_VDIR) { - inode->i_mode |= S_IFDIR; - inode->i_op = &bfs_dir_inops; - inode->i_fop = &bfs_dir_operations; - } else if (le32_to_cpu(di->i_vtype) == BFS_VREG) { - inode->i_mode |= S_IFREG; - inode->i_op = &bfs_file_inops; - inode->i_fop = &bfs_file_operations; - inode->i_mapping->a_ops = &bfs_aops; - } else { - brelse(bh); - printf("Unknown vtype=%u %s:%08lx\n", - le32_to_cpu(di->i_vtype), inode->i_sb->s_id, ino); - goto error; - } - - BFS_I(inode)->i_sblock = le32_to_cpu(di->i_sblock); - BFS_I(inode)->i_eblock = le32_to_cpu(di->i_eblock); - BFS_I(inode)->i_dsk_ino = le16_to_cpu(di->i_ino); - i_uid_write(inode, le32_to_cpu(di->i_uid)); - i_gid_write(inode, le32_to_cpu(di->i_gid)); - set_nlink(inode, le32_to_cpu(di->i_nlink)); - inode->i_size = BFS_FILESIZE(di); - inode->i_blocks = BFS_FILEBLOCKS(di); - inode_set_atime(inode, le32_to_cpu(di->i_atime), 0); - inode_set_mtime(inode, le32_to_cpu(di->i_mtime), 0); - inode_set_ctime(inode, le32_to_cpu(di->i_ctime), 0); - - brelse(bh); - unlock_new_inode(inode); - return inode; - -error: - iget_failed(inode); - return ERR_PTR(-EIO); -} - -static struct bfs_inode *find_inode(struct super_block *sb, u16 ino, struct buffer_head **p) -{ - if ((ino < BFS_ROOT_INO) || (ino > BFS_SB(sb)->si_lasti)) { - printf("Bad inode number %s:%08x\n", sb->s_id, ino); - return ERR_PTR(-EIO); - } - - ino -= BFS_ROOT_INO; - - *p = sb_bread(sb, 1 + ino / BFS_INODES_PER_BLOCK); - if (!*p) { - printf("Unable to read inode %s:%08x\n", sb->s_id, ino); - return ERR_PTR(-EIO); - } - - return (struct bfs_inode *)(*p)->b_data + ino % BFS_INODES_PER_BLOCK; -} - -static int bfs_write_inode(struct inode *inode, struct writeback_control *wbc) -{ - struct bfs_sb_info *info = BFS_SB(inode->i_sb); - unsigned int ino = (u16)inode->i_ino; - unsigned long i_sblock; - struct bfs_inode *di; - struct buffer_head *bh; - - dprintf("ino=%08x\n", ino); - - di = find_inode(inode->i_sb, ino, &bh); - if (IS_ERR(di)) - return PTR_ERR(di); - - mutex_lock(&info->bfs_lock); - - if (ino == BFS_ROOT_INO) - di->i_vtype = cpu_to_le32(BFS_VDIR); - else - di->i_vtype = cpu_to_le32(BFS_VREG); - - di->i_ino = cpu_to_le16(ino); - di->i_mode = cpu_to_le32(inode->i_mode); - di->i_uid = cpu_to_le32(i_uid_read(inode)); - di->i_gid = cpu_to_le32(i_gid_read(inode)); - di->i_nlink = cpu_to_le32(inode->i_nlink); - di->i_atime = cpu_to_le32(inode_get_atime_sec(inode)); - di->i_mtime = cpu_to_le32(inode_get_mtime_sec(inode)); - di->i_ctime = cpu_to_le32(inode_get_ctime_sec(inode)); - i_sblock = BFS_I(inode)->i_sblock; - di->i_sblock = cpu_to_le32(i_sblock); - di->i_eblock = cpu_to_le32(BFS_I(inode)->i_eblock); - di->i_eoffset = cpu_to_le32(i_sblock * BFS_BSIZE + inode->i_size - 1); - - mark_buffer_dirty(bh); - brelse(bh); - mutex_unlock(&info->bfs_lock); - set_inode_metadata_writeback(inode); - return 0; -} - -static int bfs_sync_inode_metadata(struct inode *inode, - struct writeback_control *wbc) -{ - int err = 0; - struct bfs_inode *di; - struct buffer_head *bh; - - di = find_inode(inode->i_sb, (u16)inode->i_ino, &bh); - if (IS_ERR(di)) - return PTR_ERR(di); - - sync_dirty_buffer(bh); - if (buffer_write_io_error(bh)) { - err = -EIO; - goto out; - } - err = mmb_sync(&BFS_I(inode)->i_metadata_bhs); -out: - brelse(bh); - return err; -} - -static void bfs_evict_inode(struct inode *inode) -{ - unsigned long ino = inode->i_ino; - struct bfs_inode *di; - struct buffer_head *bh; - struct super_block *s = inode->i_sb; - struct bfs_sb_info *info = BFS_SB(s); - struct bfs_inode_info *bi = BFS_I(inode); - - dprintf("ino=%08lx\n", ino); - - truncate_inode_pages_final(&inode->i_data); - if (inode->i_nlink) - mmb_sync(&BFS_I(inode)->i_metadata_bhs); - mmb_invalidate(&BFS_I(inode)->i_metadata_bhs); - clear_inode(inode); - - if (inode->i_nlink) - return; - - di = find_inode(s, inode->i_ino, &bh); - if (IS_ERR(di)) - return; - - mutex_lock(&info->bfs_lock); - /* clear on-disk inode */ - memset(di, 0, sizeof(struct bfs_inode)); - mark_buffer_dirty(bh); - brelse(bh); - - if (bi->i_dsk_ino) { - if (bi->i_sblock) - info->si_freeb += bi->i_eblock + 1 - bi->i_sblock; - info->si_freei++; - clear_bit(ino, info->si_imap); - bfs_dump_imap("evict_inode", s); - } - - /* - * If this was the last file, make the previous block - * "last block of the last file" even if there is no - * real file there, saves us 1 gap. - */ - if (info->si_lf_eblk == bi->i_eblock) - info->si_lf_eblk = bi->i_sblock - 1; - mutex_unlock(&info->bfs_lock); -} - -static void bfs_put_super(struct super_block *s) -{ - struct bfs_sb_info *info = BFS_SB(s); - - if (!info) - return; - - mutex_destroy(&info->bfs_lock); - kfree(info); - s->s_fs_info = NULL; -} - -static int bfs_statfs(struct dentry *dentry, struct kstatfs *buf) -{ - struct super_block *s = dentry->d_sb; - struct bfs_sb_info *info = BFS_SB(s); - u64 id = huge_encode_dev(s->s_bdev->bd_dev); - buf->f_type = BFS_MAGIC; - buf->f_bsize = s->s_blocksize; - buf->f_blocks = info->si_blocks; - buf->f_bfree = buf->f_bavail = info->si_freeb; - buf->f_files = info->si_lasti + 1 - BFS_ROOT_INO; - buf->f_ffree = info->si_freei; - buf->f_fsid = u64_to_fsid(id); - buf->f_namelen = BFS_NAMELEN; - return 0; -} - -static struct kmem_cache *bfs_inode_cachep; - -static struct inode *bfs_alloc_inode(struct super_block *sb) -{ - struct bfs_inode_info *bi; - bi = alloc_inode_sb(sb, bfs_inode_cachep, GFP_KERNEL); - if (!bi) - return NULL; - mmb_init(&bi->i_metadata_bhs, &bi->vfs_inode.i_data); - - return &bi->vfs_inode; -} - -static void bfs_free_inode(struct inode *inode) -{ - kmem_cache_free(bfs_inode_cachep, BFS_I(inode)); -} - -static void init_once(void *foo) -{ - struct bfs_inode_info *bi = foo; - - inode_init_once(&bi->vfs_inode); -} - -static int __init init_inodecache(void) -{ - bfs_inode_cachep = kmem_cache_create("bfs_inode_cache", - sizeof(struct bfs_inode_info), - 0, (SLAB_RECLAIM_ACCOUNT| - SLAB_ACCOUNT), - init_once); - if (bfs_inode_cachep == NULL) - return -ENOMEM; - return 0; -} - -static void destroy_inodecache(void) -{ - /* - * Make sure all delayed rcu free inodes are flushed before we - * destroy cache. - */ - rcu_barrier(); - kmem_cache_destroy(bfs_inode_cachep); -} - -static const struct super_operations bfs_sops = { - .alloc_inode = bfs_alloc_inode, - .free_inode = bfs_free_inode, - .write_inode = bfs_write_inode, - .sync_inode_metadata = bfs_sync_inode_metadata, - .evict_inode = bfs_evict_inode, - .put_super = bfs_put_super, - .statfs = bfs_statfs, -}; - -void bfs_dump_imap(const char *prefix, struct super_block *s) -{ -#ifdef DEBUG - int i; - char *tmpbuf = kzalloc(PAGE_SIZE, GFP_KERNEL); - - if (!tmpbuf) - return; - for (i = BFS_SB(s)->si_lasti; i >= 0; i--) { - if (i > PAGE_SIZE - 100) break; - if (test_bit(i, BFS_SB(s)->si_imap)) - strcat(tmpbuf, "1"); - else - strcat(tmpbuf, "0"); - } - printf("%s: lasti=%08lx <%s>\n", prefix, BFS_SB(s)->si_lasti, tmpbuf); - kfree(tmpbuf); -#endif -} - -static int bfs_fill_super(struct super_block *s, struct fs_context *fc) -{ - struct buffer_head *bh, *sbh; - struct bfs_super_block *bfs_sb; - struct inode *inode; - unsigned i; - struct bfs_sb_info *info; - int ret = -EINVAL; - unsigned long i_sblock, i_eblock, i_eoff, s_size; - int silent = fc->sb_flags & SB_SILENT; - - info = kzalloc_obj(*info); - if (!info) - return -ENOMEM; - mutex_init(&info->bfs_lock); - s->s_fs_info = info; - s->s_time_min = 0; - s->s_time_max = U32_MAX; - - if (!sb_set_blocksize(s, BFS_BSIZE)) - goto out; - - sbh = sb_bread(s, 0); - if (!sbh) - goto out; - bfs_sb = (struct bfs_super_block *)sbh->b_data; - if (le32_to_cpu(bfs_sb->s_magic) != BFS_MAGIC) { - if (!silent) - printf("No BFS filesystem on %s (magic=%08x)\n", s->s_id, le32_to_cpu(bfs_sb->s_magic)); - goto out1; - } - if (BFS_UNCLEAN(bfs_sb, s) && !silent) - printf("%s is unclean, continuing\n", s->s_id); - - s->s_magic = BFS_MAGIC; - - if (le32_to_cpu(bfs_sb->s_start) > le32_to_cpu(bfs_sb->s_end) || - le32_to_cpu(bfs_sb->s_start) < sizeof(struct bfs_super_block) + sizeof(struct bfs_dirent)) { - printf("Superblock is corrupted on %s\n", s->s_id); - goto out1; - } - - info->si_lasti = (le32_to_cpu(bfs_sb->s_start) - BFS_BSIZE) / sizeof(struct bfs_inode) + BFS_ROOT_INO - 1; - if (info->si_lasti == BFS_MAX_LASTI) - printf("NOTE: filesystem %s was created with 512 inodes, the real maximum is 511, mounting anyway\n", s->s_id); - else if (info->si_lasti > BFS_MAX_LASTI) { - printf("Impossible last inode number %lu > %d on %s\n", info->si_lasti, BFS_MAX_LASTI, s->s_id); - goto out1; - } - for (i = 0; i < BFS_ROOT_INO; i++) - set_bit(i, info->si_imap); - - s->s_op = &bfs_sops; - inode = bfs_iget(s, BFS_ROOT_INO); - if (IS_ERR(inode)) { - ret = PTR_ERR(inode); - goto out1; - } - s->s_root = d_make_root(inode); - if (!s->s_root) { - ret = -ENOMEM; - goto out1; - } - - info->si_blocks = (le32_to_cpu(bfs_sb->s_end) + 1) >> BFS_BSIZE_BITS; - info->si_freeb = (le32_to_cpu(bfs_sb->s_end) + 1 - le32_to_cpu(bfs_sb->s_start)) >> BFS_BSIZE_BITS; - info->si_freei = 0; - info->si_lf_eblk = 0; - - /* can we read the last block? */ - bh = sb_bread(s, info->si_blocks - 1); - if (!bh) { - printf("Last block not available on %s: %lu\n", s->s_id, info->si_blocks - 1); - ret = -EIO; - goto out2; - } - brelse(bh); - - bh = NULL; - for (i = BFS_ROOT_INO; i <= info->si_lasti; i++) { - struct bfs_inode *di; - int block = (i - BFS_ROOT_INO) / BFS_INODES_PER_BLOCK + 1; - int off = (i - BFS_ROOT_INO) % BFS_INODES_PER_BLOCK; - unsigned long eblock; - - if (!off) { - brelse(bh); - bh = sb_bread(s, block); - } - - if (!bh) - continue; - - di = (struct bfs_inode *)bh->b_data + off; - - /* test if filesystem is not corrupted */ - - i_eoff = le32_to_cpu(di->i_eoffset); - i_sblock = le32_to_cpu(di->i_sblock); - i_eblock = le32_to_cpu(di->i_eblock); - s_size = le32_to_cpu(bfs_sb->s_end); - - if (i_sblock > info->si_blocks || - i_eblock > info->si_blocks || - i_sblock > i_eblock || - (i_eoff != le32_to_cpu(-1) && i_eoff > s_size) || - i_sblock * BFS_BSIZE > i_eoff) { - - printf("Inode 0x%08x corrupted on %s\n", i, s->s_id); - - brelse(bh); - ret = -EIO; - goto out2; - } - - if (!di->i_ino) { - info->si_freei++; - continue; - } - set_bit(i, info->si_imap); - info->si_freeb -= BFS_FILEBLOCKS(di); - - eblock = le32_to_cpu(di->i_eblock); - if (eblock > info->si_lf_eblk) - info->si_lf_eblk = eblock; - } - brelse(bh); - brelse(sbh); - bfs_dump_imap("fill_super", s); - return 0; - -out2: - dput(s->s_root); - s->s_root = NULL; -out1: - brelse(sbh); -out: - mutex_destroy(&info->bfs_lock); - kfree(info); - s->s_fs_info = NULL; - return ret; -} - -static int bfs_get_tree(struct fs_context *fc) -{ - return get_tree_bdev(fc, bfs_fill_super); -} - -static const struct fs_context_operations bfs_context_ops = { - .get_tree = bfs_get_tree, -}; - -static int bfs_init_fs_context(struct fs_context *fc) -{ - fc->ops = &bfs_context_ops; - - return 0; -} - -static struct file_system_type bfs_fs_type = { - .owner = THIS_MODULE, - .name = "bfs", - .init_fs_context = bfs_init_fs_context, - .kill_sb = kill_block_super, - .fs_flags = FS_REQUIRES_DEV, -}; -MODULE_ALIAS_FS("bfs"); - -static int __init init_bfs_fs(void) -{ - int err = init_inodecache(); - if (err) - goto out1; - err = register_filesystem(&bfs_fs_type); - if (err) - goto out; - return 0; -out: - destroy_inodecache(); -out1: - return err; -} - -static void __exit exit_bfs_fs(void) -{ - unregister_filesystem(&bfs_fs_type); - destroy_inodecache(); -} - -module_init(init_bfs_fs) -module_exit(exit_bfs_fs) diff --git a/fs/binfmt_elf.c b/fs/binfmt_elf.c index 06d0df105382..bf7f8f47548d 100644 --- a/fs/binfmt_elf.c +++ b/fs/binfmt_elf.c @@ -74,7 +74,7 @@ static int load_elf_binary(struct linux_binprm *bprm); * don't even try. */ #ifdef CONFIG_ELF_CORE -static int elf_core_dump(struct coredump_params *cprm); +static bool elf_core_dump(struct coredump_params *cprm); #else #define elf_core_dump NULL #endif @@ -1875,7 +1875,7 @@ static int fill_note_info(struct elfhdr *elf, int phdrs, return 0; info->thread->task = dump_task; - for (ct = dump_task->signal->core_state->dumper.next; ct; ct = ct->next) { + for (ct = dump_task->signal->core_state->tasks; ct; ct = ct->next) { t = kzalloc_flex(*t, notes, info->thread_notes); if (unlikely(!t)) return 0; @@ -1987,9 +1987,9 @@ static void fill_extnum_info(struct elfhdr *elf, struct elf_shdr *shdr4extnum, * and then they are actually written out. If we run out of core limit * we just truncate. */ -static int elf_core_dump(struct coredump_params *cprm) +static bool elf_core_dump(struct coredump_params *cprm) { - int has_dumped = 0; + bool ret = false; int segs, i; struct elfhdr elf; loff_t offset = 0, dataoff; @@ -2020,7 +2020,7 @@ static int elf_core_dump(struct coredump_params *cprm) if (!fill_note_info(&elf, e_phnum, &info, cprm)) goto end_coredump; - has_dumped = 1; + cprm->state |= COREDUMP_STATE_STARTED; offset += sizeof(elf); /* ELF header */ offset += segs * sizeof(struct elf_phdr); /* Program headers */ @@ -2029,7 +2029,7 @@ static int elf_core_dump(struct coredump_params *cprm) { size_t sz = info.size; - /* For cell spufs and x86 xstate */ + /* For x86 xstate */ sz += elf_coredump_extra_notes_size(); phdr4note = kmalloc_obj(*phdr4note); @@ -2093,7 +2093,7 @@ static int elf_core_dump(struct coredump_params *cprm) if (!write_note_info(&info, cprm)) goto end_coredump; - /* For cell spufs and x86 xstate */ + /* For x86 xstate */ if (elf_coredump_extra_notes_write(cprm)) goto end_coredump; @@ -2115,11 +2115,13 @@ static int elf_core_dump(struct coredump_params *cprm) goto end_coredump; } + ret = true; + end_coredump: free_note_info(&info); kfree(shdr4extnum); kfree(phdr4note); - return has_dumped; + return ret; } #endif /* CONFIG_ELF_CORE */ diff --git a/fs/binfmt_elf_fdpic.c b/fs/binfmt_elf_fdpic.c index 068c46875c74..d3872169f55e 100644 --- a/fs/binfmt_elf_fdpic.c +++ b/fs/binfmt_elf_fdpic.c @@ -75,7 +75,7 @@ static int elf_fdpic_map_file_by_direct_mmap(struct elf_fdpic_params *, struct file *, struct mm_struct *); #ifdef CONFIG_ELF_CORE -static int elf_fdpic_core_dump(struct coredump_params *cprm); +static bool elf_fdpic_core_dump(struct coredump_params *cprm); #endif static struct linux_binfmt elf_fdpic_format = { @@ -1477,9 +1477,9 @@ static bool elf_fdpic_dump_segments(struct coredump_params *cprm, * and then they are actually written out. If we run out of core limit * we just truncate. */ -static int elf_fdpic_core_dump(struct coredump_params *cprm) +static bool elf_fdpic_core_dump(struct coredump_params *cprm) { - int has_dumped = 0; + bool ret = false; int segs; int i; struct elfhdr *elf = NULL; @@ -1504,7 +1504,7 @@ static int elf_fdpic_core_dump(struct coredump_params *cprm) if (!psinfo) goto end_coredump; - for (ct = current->signal->core_state->dumper.next; + for (ct = current->signal->core_state->tasks; ct; ct = ct->next) { tmp = elf_dump_thread_status(cprm->siginfo->si_signo, ct->task, &thread_status_size); @@ -1536,7 +1536,7 @@ static int elf_fdpic_core_dump(struct coredump_params *cprm) /* Set up header */ fill_elf_fdpic_header(elf, e_phnum); - has_dumped = 1; + cprm->state |= COREDUMP_STATE_STARTED; /* * Set up the notes in similar form to SVR4 core dumps made * with info from their /proc. @@ -1656,6 +1656,8 @@ static int elf_fdpic_core_dump(struct coredump_params *cprm) cprm->file->f_pos, offset); } + ret = true; + end_coredump: while (thread_list) { tmp = thread_list; @@ -1666,7 +1668,7 @@ end_coredump: kfree(elf); kfree(psinfo); kfree(shdr4extnum); - return has_dumped; + return ret; } #endif /* CONFIG_ELF_CORE */ diff --git a/fs/binfmt_misc.c b/fs/binfmt_misc.c index 620da85948b4..d945b4f6e158 100644 --- a/fs/binfmt_misc.c +++ b/fs/binfmt_misc.c @@ -100,6 +100,13 @@ static const struct binfmt_misc_flag *misc_flag_by_char(const char c) return NULL; } +static bool misc_valid_delim(const char c) +{ + if (!isascii(c) || !ispunct(c)) + return false; + return c != '\\'; +} + struct binfmt_misc_entry { struct hlist_node node; unsigned long flags; /* type, status, etc. */ @@ -871,10 +878,9 @@ static struct binfmt_misc_entry *create_entry(const char __user *buffer, del = *p++; /* delimiter */ - pr_debug("register: delim: %#x {%c}\n", del, del); + pr_debug("register: delim: %#x\n", del); - /* A flag-char delimiter runs the flag scan off the buffer. */ - if (misc_flag_by_char(del)) + if (!misc_valid_delim(del)) return ERR_PTR(-EINVAL); /* Pad the buffer with the delim to simplify parsing below. */ diff --git a/fs/bpf_fs_kfuncs.c b/fs/bpf_fs_kfuncs.c index 357a379ef92a..abdfbd83dc57 100644 --- a/fs/bpf_fs_kfuncs.c +++ b/fs/bpf_fs_kfuncs.c @@ -237,7 +237,7 @@ int bpf_set_dentry_xattr_locked(struct dentry *dentry, const char *name__str, * @dentry: dentry to get xattr from * @name__str: name of the xattr * - * Rmove xattr *name__str* of *dentry*. + * Remove xattr *name__str* of *dentry*. * * For security reasons, only *name__str* with prefix "security.bpf." * is allowed. @@ -305,7 +305,7 @@ __bpf_kfunc int bpf_set_dentry_xattr(struct dentry *dentry, const char *name__st * @dentry: dentry to get xattr from * @name__str: name of the xattr * - * Rmove xattr *name__str* of *dentry*. + * Remove xattr *name__str* of *dentry*. * * For security reasons, only *name__str* with prefix "security.bpf." * is allowed. diff --git a/fs/btrfs/Kconfig b/fs/btrfs/Kconfig index 4b10d78ed99b..0b8d8905e38e 100644 --- a/fs/btrfs/Kconfig +++ b/fs/btrfs/Kconfig @@ -1,4 +1,5 @@ # SPDX-License-Identifier: GPL-2.0 +# misc-next marker config BTRFS_FS tristate "Btrfs filesystem support" diff --git a/fs/btrfs/acl.c b/fs/btrfs/acl.c index 662cdd1cbdef..10a0d733bfd1 100644 --- a/fs/btrfs/acl.c +++ b/fs/btrfs/acl.c @@ -101,7 +101,7 @@ int __btrfs_set_acl(struct btrfs_trans_handle *trans, struct inode *inode, return 0; } -int btrfs_set_acl(struct mnt_idmap *idmap, struct dentry *dentry, +int btrfs_set_acl(const struct mnt_idmap *idmap, struct dentry *dentry, struct posix_acl *acl, int type) { int ret; diff --git a/fs/btrfs/acl.h b/fs/btrfs/acl.h index 0458cd51ed48..6eae2db3654d 100644 --- a/fs/btrfs/acl.h +++ b/fs/btrfs/acl.h @@ -15,7 +15,7 @@ struct mnt_idmap; struct dentry; struct posix_acl *btrfs_get_acl(struct inode *inode, int type, bool rcu); -int btrfs_set_acl(struct mnt_idmap *idmap, struct dentry *dentry, +int btrfs_set_acl(const struct mnt_idmap *idmap, struct dentry *dentry, struct posix_acl *acl, int type); int __btrfs_set_acl(struct btrfs_trans_handle *trans, struct inode *inode, struct posix_acl *acl, int type); diff --git a/fs/btrfs/bio.c b/fs/btrfs/bio.c index cc0bd03048ba..2c6234f6182f 100644 --- a/fs/btrfs/bio.c +++ b/fs/btrfs/bio.c @@ -27,15 +27,9 @@ struct btrfs_failed_bio { atomic_t repair_count; }; -/* Is this a data path I/O that needs storage layer checksum and repair? */ -static inline bool is_data_bbio(const struct btrfs_bio *bbio) -{ - return bbio->inode && is_data_inode(bbio->inode); -} - static bool bbio_has_ordered_extent(const struct btrfs_bio *bbio) { - return is_data_bbio(bbio) && btrfs_op(&bbio->bio) == BTRFS_MAP_WRITE; + return is_data_inode(bbio->inode) && btrfs_op(&bbio->bio) == BTRFS_MAP_WRITE; } /* @@ -103,7 +97,6 @@ static struct btrfs_bio *btrfs_split_bio(struct btrfs_fs_info *fs_info, bbio->can_use_append = orig_bbio->can_use_append; bbio->is_scrub = orig_bbio->is_scrub; bbio->is_remap = orig_bbio->is_remap; - bbio->async_csum = orig_bbio->async_csum; atomic_inc(&orig_bbio->pending_ios); return bbio; @@ -114,9 +107,6 @@ void btrfs_bio_end_io(struct btrfs_bio *bbio, blk_status_t status) /* Make sure we're already in task context. */ ASSERT(in_task()); - if (bbio->async_csum) - wait_for_completion(&bbio->csum_done); - bbio->bio.bi_status = status; if (bbio->bio.bi_pool == &btrfs_clone_bioset) { struct btrfs_bio *orig_bbio = bbio->private; @@ -180,30 +170,13 @@ static void btrfs_end_repair_bio(struct btrfs_bio *repair_bbio, struct btrfs_failed_bio *fbio = repair_bbio->private; struct btrfs_inode *inode = repair_bbio->inode; struct btrfs_fs_info *fs_info = inode->root->fs_info; - /* - * We can not move forward the saved_iter, as it will be later - * utilized by repair_bbio again. - */ - struct bvec_iter saved_iter = repair_bbio->saved_iter; - const u32 step = min(fs_info->sectorsize, PAGE_SIZE); - const u64 logical = repair_bbio->saved_iter.bi_sector << SECTOR_SHIFT; - const u32 nr_steps = repair_bbio->saved_iter.bi_size / step; int mirror = repair_bbio->mirror_num; - phys_addr_t paddrs[BTRFS_MAX_BLOCKSIZE / PAGE_SIZE]; - phys_addr_t paddr; - unsigned int slot = 0; - /* Repair bbio should be eaxctly one block sized. */ + /* Repair bbio should be exactly one block sized. */ ASSERT(repair_bbio->saved_iter.bi_size == fs_info->sectorsize); - btrfs_bio_for_each_block(paddr, &repair_bbio->bio, &saved_iter, step) { - ASSERT(slot < nr_steps); - paddrs[slot] = paddr; - slot++; - } - if (repair_bbio->bio.bi_status || - !btrfs_data_csum_ok(repair_bbio, dev, 0, paddrs)) { + !btrfs_bio_data_csum_ok(repair_bbio, &repair_bbio->saved_iter, dev)) { bio_reset(&repair_bbio->bio, NULL, REQ_OP_READ); repair_bbio->bio.bi_iter = repair_bbio->saved_iter; @@ -220,9 +193,8 @@ static void btrfs_end_repair_bio(struct btrfs_bio *repair_bbio, do { mirror = prev_repair_mirror(fbio, mirror); - btrfs_repair_io_failure(fs_info, btrfs_ino(inode), - repair_bbio->file_offset, fs_info->sectorsize, - logical, paddrs, step, mirror); + btrfs_repair_bbio_failure(repair_bbio, &repair_bbio->saved_iter, + fs_info->sectorsize, mirror); } while (mirror != fbio->bbio->mirror_num); done: @@ -238,25 +210,21 @@ done: * read succeeded to restore the redundancy. */ static struct btrfs_failed_bio *repair_one_sector(struct btrfs_bio *failed_bbio, - u32 bio_offset, - phys_addr_t paddrs[], + const struct bvec_iter *orig_iter, struct btrfs_failed_bio *fbio) { struct btrfs_inode *inode = failed_bbio->inode; struct btrfs_fs_info *fs_info = inode->root->fs_info; - const u32 sectorsize = fs_info->sectorsize; - const u32 step = min(fs_info->sectorsize, PAGE_SIZE); - const u32 nr_steps = sectorsize / step; - /* - * For bs > ps cases, the saved_iter can be partially moved forward. - * In that case we should round it down to the block boundary. - */ - const u64 logical = round_down(failed_bbio->saved_iter.bi_sector << SECTOR_SHIFT, - sectorsize); struct btrfs_bio *repair_bbio; struct bio *repair_bio; + struct bvec_iter iter = *orig_iter; + const u32 sectorsize = fs_info->sectorsize; + const u32 bio_offset = ((iter.bi_sector - failed_bbio->saved_iter.bi_sector) << + SECTOR_SHIFT); + const u64 logical = (iter.bi_sector << SECTOR_SHIFT); int num_copies; int mirror; + u32 cur = 0; btrfs_debug(fs_info, "repair read error: read error at %llu", failed_bbio->file_offset + bio_offset); @@ -277,17 +245,21 @@ static struct btrfs_failed_bio *repair_one_sector(struct btrfs_bio *failed_bbio, atomic_inc(&fbio->repair_count); - repair_bio = bio_alloc_bioset(NULL, nr_steps, REQ_OP_READ, GFP_NOFS, - &btrfs_repair_bioset); + repair_bio = bio_alloc_bioset(NULL, max(1, sectorsize >> PAGE_SHIFT), + REQ_OP_READ, GFP_NOFS, &btrfs_repair_bioset); repair_bio->bi_iter.bi_sector = logical >> SECTOR_SHIFT; - for (int i = 0; i < nr_steps; i++) { + while (cur < sectorsize) { + struct page *page = bio_iter_page(&failed_bbio->bio, iter); + const u32 pg_off = bio_iter_offset(&failed_bbio->bio, iter); + const u32 cur_len = min(bio_iter_len(&failed_bbio->bio, iter), + sectorsize - cur); int ret; - ASSERT(offset_in_page(paddrs[i]) + step <= PAGE_SIZE); + ret = bio_add_page(repair_bio, page, cur_len, pg_off); + ASSERT(ret == cur_len); - ret = bio_add_page(repair_bio, phys_to_page(paddrs[i]), step, - offset_in_page(paddrs[i])); - ASSERT(ret == step); + bio_advance_iter_single(&failed_bbio->bio, &iter, cur_len); + cur += cur_len; } repair_bbio = btrfs_bio(repair_bio); @@ -305,18 +277,16 @@ static void btrfs_check_read_bio(struct btrfs_bio *bbio, struct btrfs_device *de struct btrfs_inode *inode = bbio->inode; struct btrfs_fs_info *fs_info = inode->root->fs_info; const u32 sectorsize = fs_info->sectorsize; - const u32 step = min(sectorsize, PAGE_SIZE); - const u32 nr_steps = sectorsize / step; - struct bvec_iter *iter = &bbio->saved_iter; + struct bvec_iter iter; blk_status_t status = bbio->bio.bi_status; struct btrfs_failed_bio *fbio = NULL; - phys_addr_t paddrs[BTRFS_MAX_BLOCKSIZE / PAGE_SIZE]; - phys_addr_t paddr; - u32 offset = 0; /* Read-repair requires the inode field to be set by the submitter. */ ASSERT(inode); + /* The original bbio should be sectorsize aligned. */ + ASSERT(IS_ALIGNED(bbio->saved_iter.bi_size, sectorsize)); + /* * Hand off repair bios to the repair code as there is no upper level * submitter for them. @@ -329,16 +299,10 @@ static void btrfs_check_read_bio(struct btrfs_bio *bbio, struct btrfs_device *de /* Clear the I/O error. A failed repair will reset it. */ bbio->bio.bi_status = BLK_STS_OK; - btrfs_bio_for_each_block(paddr, &bbio->bio, iter, step) { - paddrs[(offset / step) % nr_steps] = paddr; - offset += step; - - if (IS_ALIGNED(offset, sectorsize)) { - if (status || - !btrfs_data_csum_ok(bbio, dev, offset - sectorsize, paddrs)) - fbio = repair_one_sector(bbio, offset - sectorsize, - paddrs, fbio); - } + for (iter = bbio->saved_iter; iter.bi_size; + bio_advance_iter(&bbio->bio, &iter, sectorsize)) { + if (status || !btrfs_bio_data_csum_ok(bbio, &iter, dev)) + fbio = repair_one_sector(bbio, &iter, fbio); } if (bbio->csum != bbio->csum_inline) kvfree(bbio->csum); @@ -386,7 +350,7 @@ static void simple_end_io_work(struct work_struct *work) if (bio_op(bio) == REQ_OP_READ) { /* Metadata reads are checked and repaired by the submitter. */ - if (is_data_bbio(bbio)) + if (is_data_inode(bbio->inode)) return btrfs_check_read_bio(bbio, bbio->bio.bi_private); return btrfs_bio_end_io(bbio, bbio->bio.bi_status); } @@ -420,7 +384,7 @@ static void btrfs_raid56_end_io(struct bio *bio) btrfs_bio_counter_dec(bioc->fs_info); bbio->mirror_num = bioc->mirror_num; - if (bio_op(bio) == REQ_OP_READ && is_data_bbio(bbio)) + if (bio_op(bio) == REQ_OP_READ && is_data_inode(bbio->inode)) btrfs_check_read_bio(bbio, NULL); else btrfs_bio_end_io(bbio, bbio->bio.bi_status); @@ -783,7 +747,7 @@ static bool btrfs_submit_chunk(struct btrfs_bio *bbio, int mirror_num) * our bio to the physical disk location, so we need to save the * original bytenr so we know what we're checksumming. */ - if (bio_op(bio) == REQ_OP_WRITE && is_data_bbio(bbio)) + if (bio_op(bio) == REQ_OP_WRITE && is_data_inode(bbio->inode)) bbio->orig_logical = logical; bbio->can_use_append = btrfs_use_zone_append(bbio); @@ -809,7 +773,7 @@ static bool btrfs_submit_chunk(struct btrfs_bio *bbio, int mirror_num) * Save the iter for the end_io handler and preload the checksums for * data reads. */ - if (bio_op(bio) == REQ_OP_READ && is_data_bbio(bbio)) { + if (bio_op(bio) == REQ_OP_READ && is_data_inode(bbio->inode)) { bbio->saved_iter = bio->bi_iter; ret = btrfs_lookup_bio_sums(bbio); status = errno_to_blk_status(ret); @@ -818,7 +782,7 @@ static bool btrfs_submit_chunk(struct btrfs_bio *bbio, int mirror_num) } if (btrfs_op(bio) == BTRFS_MAP_WRITE) { - if (is_data_bbio(bbio) && bioc && bioc->use_rst) { + if (is_data_inode(bbio->inode) && bioc && bioc->use_rst) { /* * No locking for the list update, as we only add to * the list in the I/O submission path, and list @@ -925,21 +889,23 @@ void btrfs_submit_bbio(struct btrfs_bio *bbio, int mirror_num) * The I/O is issued synchronously to block the repair read completion from * freeing the bio. * - * @ino: Offending inode number - * @fileoff: File offset inside the inode + * @bbio: Original bbio where the repair is needed + * @orig_iter: Points to where the repair starts * @length: Length of the repair write - * @logical: Logical address of the range - * @paddrs: Physical address array of the content - * @step: Length of for each paddrs * @mirror_num: Mirror number to write to. Must not be zero */ -int btrfs_repair_io_failure(struct btrfs_fs_info *fs_info, u64 ino, u64 fileoff, - u32 length, u64 logical, const phys_addr_t paddrs[], - unsigned int step, int mirror_num) +int btrfs_repair_bbio_failure(struct btrfs_bio *bbio, const struct bvec_iter *orig_iter, + u32 length, int mirror_num) { - const u32 nr_steps = DIV_ROUND_UP_POW2(length, step); + struct btrfs_inode *inode = bbio->inode; + struct btrfs_fs_info *fs_info = inode->root->fs_info; struct btrfs_io_stripe smap = { 0 }; - struct bio *bio = NULL; + struct bvec_iter iter = *orig_iter; + struct bio *repair_bio = NULL; + const u64 logical = iter.bi_sector << SECTOR_SHIFT; + const u64 fileoff = bbio->file_offset + + ((iter.bi_sector - bbio->saved_iter.bi_sector) << SECTOR_SHIFT); + u32 cur = 0; int ret = 0; BUG_ON(!mirror_num); @@ -950,8 +916,9 @@ int btrfs_repair_io_failure(struct btrfs_fs_info *fs_info, u64 ino, u64 fileoff, ASSERT(IS_ALIGNED(fileoff, fs_info->sectorsize)); /* Either it's a single data or metadata block. */ ASSERT(length <= BTRFS_MAX_BLOCKSIZE); - ASSERT(step <= length); - ASSERT(is_power_of_2(step)); + + /* Our current iter should not be before the original bbio saved_iter. */ + ASSERT(iter.bi_sector >= bbio->saved_iter.bi_sector); /* * The fs either mounted RO or hit critical errors, no need @@ -979,15 +946,22 @@ int btrfs_repair_io_failure(struct btrfs_fs_info *fs_info, u64 ino, u64 fileoff, goto out_counter_dec; } - bio = bio_alloc(smap.dev->bdev, nr_steps, REQ_OP_WRITE | REQ_SYNC, GFP_NOFS); - bio->bi_iter.bi_sector = smap.physical >> SECTOR_SHIFT; - for (int i = 0; i < nr_steps; i++) { - ret = bio_add_page(bio, phys_to_page(paddrs[i]), step, offset_in_page(paddrs[i])); - /* We should have allocated enough slots to contain all the different pages. */ - ASSERT(ret == step); + repair_bio = bio_alloc(smap.dev->bdev, max(1, length >> PAGE_SHIFT), + REQ_OP_WRITE | REQ_SYNC, GFP_NOFS); + repair_bio->bi_iter.bi_sector = smap.physical >> SECTOR_SHIFT; + while (cur < length) { + struct page *page = bio_iter_page(&bbio->bio, iter); + const u32 pg_off = bio_iter_offset(&bbio->bio, iter); + const u32 cur_len = min(bio_iter_len(&bbio->bio, iter), length - cur); + + ret = bio_add_page(repair_bio, page, cur_len, pg_off); + ASSERT(ret == cur_len); + bio_advance_iter_single(&bbio->bio, &iter, cur_len); + cur += cur_len; } - ret = submit_bio_wait(bio); - bio_put(bio); + + ret = submit_bio_wait(repair_bio); + bio_put(repair_bio); if (ret) { /* try to remap that extent elsewhere? */ btrfs_dev_stat_inc_and_print(smap.dev, BTRFS_DEV_STAT_WRITE_ERRS); @@ -995,8 +969,9 @@ int btrfs_repair_io_failure(struct btrfs_fs_info *fs_info, u64 ino, u64 fileoff, } btrfs_info_rl(fs_info, - "read error corrected: ino %llu off %llu (dev %s sector %llu)", - ino, fileoff, btrfs_dev_name(smap.dev), + "read error corrected: root %llu ino %llu off %llu (dev %s sector %llu)", + btrfs_root_id(inode->root), btrfs_ino(inode), fileoff, + btrfs_dev_name(smap.dev), smap.physical >> SECTOR_SHIFT); ret = 0; diff --git a/fs/btrfs/bio.h b/fs/btrfs/bio.h index 303ed6c7103d..bbf362b8668b 100644 --- a/fs/btrfs/bio.h +++ b/fs/btrfs/bio.h @@ -58,7 +58,6 @@ struct btrfs_bio { struct btrfs_ordered_extent *ordered; struct btrfs_ordered_sum *sums; struct work_struct csum_work; - struct completion csum_done; struct bvec_iter csum_saved_iter; u64 orig_physical; u64 orig_logical; @@ -93,9 +92,6 @@ struct btrfs_bio { /* Whether the bio is coming from copy_remapped_data_io(). */ bool is_remap:1; - /* Whether the csum generation for data write is async. */ - bool async_csum:1; - /* Whether the bio is written using zone append. */ bool can_use_append:1; @@ -126,8 +122,7 @@ void btrfs_bio_end_io(struct btrfs_bio *bbio, blk_status_t status); void btrfs_submit_bbio(struct btrfs_bio *bbio, int mirror_num); void btrfs_submit_repair_write(struct btrfs_bio *bbio, int mirror_num, bool dev_replace); -int btrfs_repair_io_failure(struct btrfs_fs_info *fs_info, u64 ino, u64 fileoff, - u32 length, u64 logical, const phys_addr_t paddrs[], - unsigned int step, int mirror_num); +int btrfs_repair_bbio_failure(struct btrfs_bio *bbio, const struct bvec_iter *orig_iter, + u32 length, int mirror_num); #endif diff --git a/fs/btrfs/block-group.c b/fs/btrfs/block-group.c index ee182369254c..e6080cb47c89 100644 --- a/fs/btrfs/block-group.c +++ b/fs/btrfs/block-group.c @@ -904,22 +904,6 @@ static noinline void caching_thread(struct btrfs_work *work) down_read(&fs_info->commit_root_sem); load_block_group_size_class(caching_ctl); - if (btrfs_test_opt(fs_info, SPACE_CACHE)) { - ret = load_free_space_cache(block_group); - if (ret == 1) { - ret = 0; - goto done; - } - - /* - * We failed to load the space cache, set ourselves to - * CACHE_STARTED and carry on. - */ - spin_lock(&block_group->lock); - block_group->cached = BTRFS_CACHE_STARTED; - spin_unlock(&block_group->lock); - wake_up(&caching_ctl->wait); - } /* * If we are in the transaction that populated the free space tree we @@ -933,7 +917,7 @@ static noinline void caching_thread(struct btrfs_work *work) ret = btrfs_load_free_space_tree(caching_ctl); else ret = load_extent_tree_free(caching_ctl); -done: + spin_lock(&block_group->lock); block_group->caching_ctl = NULL; block_group->cached = ret ? BTRFS_CACHE_ERROR : BTRFS_CACHE_FINISHED; @@ -1194,36 +1178,21 @@ int btrfs_remove_block_group(struct btrfs_trans_handle *trans, goto out; } - /* - * get the inode first so any iput calls done for the io_list - * aren't the final iput (no unlinks allowed now) - */ inode = lookup_free_space_inode(block_group, path); - mutex_lock(&trans->transaction->cache_write_mutex); /* - * Make sure our free space cache IO is done before removing the - * free space inode + * Do not delete the block group item while + * btrfs_start_dirty_block_groups() is updating it. */ + mutex_lock(&trans->transaction->dirty_bgs_update_mutex); spin_lock(&trans->transaction->dirty_bgs_lock); - if (!list_empty(&block_group->io_list)) { - list_del_init(&block_group->io_list); - - WARN_ON(!IS_ERR(inode) && inode != block_group->io_ctl.inode); - - spin_unlock(&trans->transaction->dirty_bgs_lock); - btrfs_wait_cache_io(trans, block_group, path); - btrfs_put_block_group(block_group); - spin_lock(&trans->transaction->dirty_bgs_lock); - } - if (!list_empty(&block_group->dirty_list)) { list_del_init(&block_group->dirty_list); remove_rsv = true; btrfs_put_block_group(block_group); } spin_unlock(&trans->transaction->dirty_bgs_lock); - mutex_unlock(&trans->transaction->cache_write_mutex); + mutex_unlock(&trans->transaction->dirty_bgs_update_mutex); ret = btrfs_remove_free_space_inode(trans, inode, block_group); if (unlikely(ret)) { @@ -1287,7 +1256,6 @@ int btrfs_remove_block_group(struct btrfs_trans_handle *trans, spin_lock(&trans->transaction->dirty_bgs_lock); WARN_ON(!list_empty(&block_group->dirty_list)); - WARN_ON(!list_empty(&block_group->io_list)); spin_unlock(&trans->transaction->dirty_bgs_lock); btrfs_remove_free_space_cache(block_group); @@ -1612,7 +1580,7 @@ void btrfs_delete_unused_bgs(struct btrfs_fs_info *fs_info) space_info = block_group->space_info; - if (ret || btrfs_mixed_space_info(space_info)) { + if (btrfs_mixed_space_info(space_info)) { btrfs_put_block_group(block_group); continue; } @@ -1727,6 +1695,7 @@ void btrfs_delete_unused_bgs(struct btrfs_fs_info *fs_info) ret = inc_block_group_ro(block_group, false); up_write(&space_info->groups_sem); if (ret < 0) { + btrfs_link_bg_list(block_group, &retry_list); ret = 0; goto next; } @@ -1749,6 +1718,7 @@ void btrfs_delete_unused_bgs(struct btrfs_fs_info *fs_info) block_group->start); if (IS_ERR(trans)) { btrfs_dec_block_group_ro(block_group); + btrfs_link_bg_list(block_group, &retry_list); ret = PTR_ERR(trans); goto next; } @@ -1759,6 +1729,7 @@ void btrfs_delete_unused_bgs(struct btrfs_fs_info *fs_info) */ if (!clean_pinned_extents(trans, block_group)) { btrfs_dec_block_group_ro(block_group); + btrfs_link_bg_list(block_group, &retry_list); goto end_trans; } @@ -1845,6 +1816,8 @@ end_trans: next: btrfs_put_block_group(block_group); spin_lock(&fs_info->unused_bgs_lock); + if (ret) + break; } list_splice_tail(&retry_list, &fs_info->unused_bgs); spin_unlock(&fs_info->unused_bgs_lock); @@ -2431,7 +2404,6 @@ static struct btrfs_block_group *btrfs_create_block_group( INIT_LIST_HEAD(&cache->ro_list); INIT_LIST_HEAD(&cache->discard_list); INIT_LIST_HEAD(&cache->dirty_list); - INIT_LIST_HEAD(&cache->io_list); INIT_LIST_HEAD(&cache->active_bg_list); btrfs_init_free_space_ctl(cache, cache->free_space_ctl); atomic_set(&cache->frozen, 0); @@ -2487,8 +2459,7 @@ static int check_chunk_block_group_mappings(struct btrfs_fs_info *fs_info) static int read_one_block_group(struct btrfs_fs_info *info, struct btrfs_block_group_item_v2 *bgi, - const struct btrfs_key *key, - bool need_clear) + const struct btrfs_key *key) { struct btrfs_block_group *cache; const bool mixed = btrfs_fs_incompat(info, MIXED_GROUPS); @@ -2514,20 +2485,6 @@ static int read_one_block_group(struct btrfs_fs_info *info, btrfs_set_free_space_tree_thresholds(cache); - if (need_clear) { - /* - * When we mount with old space cache, we need to - * set BTRFS_DC_CLEAR and set dirty flag. - * - * a) Setting 'BTRFS_DC_CLEAR' makes sure that we - * truncate the old free space cache inode and - * setup a new one. - * b) Setting 'dirty flag' makes sure that we flush - * the new space cache info onto disk. - */ - if (btrfs_test_opt(info, SPACE_CACHE)) - cache->disk_cache_state = BTRFS_DC_CLEAR; - } if (!mixed && ((cache->flags & BTRFS_BLOCK_GROUP_METADATA) && (cache->flags & BTRFS_BLOCK_GROUP_DATA))) { btrfs_err(info, @@ -2668,8 +2625,6 @@ int btrfs_read_block_groups(struct btrfs_fs_info *info) struct btrfs_block_group *cache; struct btrfs_space_info *space_info; struct btrfs_key key; - bool need_clear = false; - u64 cache_gen; /* * Either no extent root (with ibadroots rescue option) or we have @@ -2690,13 +2645,6 @@ int btrfs_read_block_groups(struct btrfs_fs_info *info) if (!path) return -ENOMEM; - cache_gen = btrfs_super_cache_generation(info->super_copy); - if (btrfs_test_opt(info, SPACE_CACHE) && - btrfs_super_generation(info->super_copy) != cache_gen) - need_clear = true; - if (btrfs_test_opt(info, CLEAR_CACHE)) - need_clear = true; - while (1) { struct btrfs_block_group_item_v2 bgi; struct extent_buffer *leaf; @@ -2725,7 +2673,7 @@ int btrfs_read_block_groups(struct btrfs_fs_info *info) btrfs_item_key_to_cpu(leaf, &key, slot); btrfs_release_path(path); - ret = read_one_block_group(info, &bgi, &key, need_clear); + ret = read_one_block_group(info, &bgi, &key); if (ret < 0) goto error; key.objectid += key.offset; @@ -3373,207 +3321,16 @@ fail: } -static void cache_save_setup(struct btrfs_block_group *block_group, - struct btrfs_trans_handle *trans, - struct btrfs_path *path) -{ - struct btrfs_fs_info *fs_info = block_group->fs_info; - struct inode *inode = NULL; - struct extent_changeset *data_reserved = NULL; - u64 alloc_hint = 0; - int dcs = BTRFS_DC_ERROR; - u64 cache_size = 0; - int retries = 0; - int ret = 0; - - if (!btrfs_test_opt(fs_info, SPACE_CACHE)) - return; - - /* - * If this block group is smaller than 100 megs don't bother caching the - * block group. - */ - if (block_group->length < (100 * SZ_1M)) { - spin_lock(&block_group->lock); - block_group->disk_cache_state = BTRFS_DC_WRITTEN; - spin_unlock(&block_group->lock); - return; - } - - if (TRANS_ABORTED(trans)) - return; -again: - inode = lookup_free_space_inode(block_group, path); - if (IS_ERR(inode) && PTR_ERR(inode) != -ENOENT) { - ret = PTR_ERR(inode); - btrfs_release_path(path); - goto out; - } - - if (IS_ERR(inode)) { - if (retries) { - ret = PTR_ERR(inode); - btrfs_err(fs_info, - "failed to lookup free space inode after creation for block group %llu: %d", - block_group->start, ret); - goto out_free; - } - retries++; - - if (block_group->ro) - goto out_free; - - ret = create_free_space_inode(trans, block_group, path); - if (ret) - goto out_free; - goto again; - } - - /* - * We want to set the generation to 0, that way if anything goes wrong - * from here on out we know not to trust this cache when we load up next - * time. - */ - BTRFS_I(inode)->generation = 0; - ret = btrfs_update_inode(trans, BTRFS_I(inode)); - if (unlikely(ret)) { - /* - * So theoretically we could recover from this, simply set the - * super cache generation to 0 so we know to invalidate the - * cache, but then we'd have to keep track of the block groups - * that fail this way so we know we _have_ to reset this cache - * before the next commit or risk reading stale cache. So to - * limit our exposure to horrible edge cases lets just abort the - * transaction, this only happens in really bad situations - * anyway. - */ - btrfs_abort_transaction(trans, ret); - goto out_put; - } - - /* We've already setup this transaction, go ahead and exit */ - if (block_group->cache_generation == trans->transid && - i_size_read(inode)) { - dcs = BTRFS_DC_SETUP; - goto out_put; - } - - if (i_size_read(inode) > 0) { - ret = btrfs_check_trunc_cache_free_space(fs_info, - &fs_info->global_block_rsv); - if (ret) - goto out_put; - - ret = btrfs_truncate_free_space_cache(trans, NULL, inode); - if (ret) - goto out_put; - } - - spin_lock(&block_group->lock); - if (block_group->cached != BTRFS_CACHE_FINISHED || - !btrfs_test_opt(fs_info, SPACE_CACHE)) { - /* - * don't bother trying to write stuff out _if_ - * a) we're not cached, - * b) we're with nospace_cache mount option, - * c) we're with v2 space_cache (FREE_SPACE_TREE). - */ - dcs = BTRFS_DC_WRITTEN; - spin_unlock(&block_group->lock); - goto out_put; - } - spin_unlock(&block_group->lock); - - /* - * We hit an ENOSPC when setting up the cache in this transaction, just - * skip doing the setup, we've already cleared the cache so we're safe. - */ - if (test_bit(BTRFS_TRANS_CACHE_ENOSPC, &trans->transaction->flags)) - goto out_put; - - /* - * Try to preallocate enough space based on how big the block group is. - * Keep in mind this has to include any pinned space which could end up - * taking up quite a bit since it's not folded into the other space - * cache. - */ - cache_size = div_u64(block_group->length, SZ_256M); - if (!cache_size) - cache_size = 1; - - cache_size *= 16; - cache_size *= fs_info->sectorsize; - - ret = btrfs_check_data_free_space(BTRFS_I(inode), &data_reserved, 0, - cache_size, false); - if (ret) - goto out_put; - - ret = btrfs_prealloc_file_range_trans(inode, trans, 0, 0, cache_size, - cache_size, cache_size, - &alloc_hint); - /* - * Our cache requires contiguous chunks so that we don't modify a bunch - * of metadata or split extents when writing the cache out, which means - * we can enospc if we are heavily fragmented in addition to just normal - * out of space conditions. So if we hit this just skip setting up any - * other block groups for this transaction, maybe we'll unpin enough - * space the next time around. - */ - if (!ret) - dcs = BTRFS_DC_SETUP; - else if (ret == -ENOSPC) - set_bit(BTRFS_TRANS_CACHE_ENOSPC, &trans->transaction->flags); - -out_put: - iput(inode); -out_free: - btrfs_release_path(path); -out: - spin_lock(&block_group->lock); - if (!ret && dcs == BTRFS_DC_SETUP) - block_group->cache_generation = trans->transid; - block_group->disk_cache_state = dcs; - spin_unlock(&block_group->lock); - - extent_changeset_free(data_reserved); -} - -int btrfs_setup_space_cache(struct btrfs_trans_handle *trans) -{ - struct btrfs_fs_info *fs_info = trans->fs_info; - struct btrfs_block_group *cache, *tmp; - struct btrfs_transaction *cur_trans = trans->transaction; - BTRFS_PATH_AUTO_FREE(path); - - if (list_empty(&cur_trans->dirty_bgs) || - !btrfs_test_opt(fs_info, SPACE_CACHE)) - return 0; - - path = btrfs_alloc_path(); - if (!path) - return -ENOMEM; - - /* Could add new block groups, use _safe just in case */ - list_for_each_entry_safe(cache, tmp, &cur_trans->dirty_bgs, - dirty_list) { - if (cache->disk_cache_state == BTRFS_DC_CLEAR) - cache_save_setup(cache, trans, path); - } - - return 0; -} - /* - * Transaction commit does final block group cache writeback during a critical + * Transaction commit does the final block group item updates during a critical * section where nothing is allowed to change the FS. This is required in - * order for the cache to actually match the block group, but can introduce a + * order for the items to actually match the block groups, but can introduce a * lot of latency into the commit. * - * So, btrfs_start_dirty_block_groups is here to kick off block group cache IO. - * There's a chance we'll have to redo some of it if the block group changes - * again during the commit, but it greatly reduces the commit latency by - * getting rid of the easy block groups while we're still allowing others to + * So, btrfs_start_dirty_block_groups is here to update the block group items + * early. There's a chance we'll have to redo some of it if the block group + * changes again during the commit, but it greatly reduces the commit latency + * by getting rid of the easy block groups while we're still allowing others to * join the commit. */ int btrfs_start_dirty_block_groups(struct btrfs_trans_handle *trans) @@ -3582,10 +3339,8 @@ int btrfs_start_dirty_block_groups(struct btrfs_trans_handle *trans) struct btrfs_block_group *cache; struct btrfs_transaction *cur_trans = trans->transaction; int ret = 0; - int should_put; BTRFS_PATH_AUTO_FREE(path); LIST_HEAD(dirty); - struct list_head *io = &cur_trans->io_bgs; int loops = 0; spin_lock(&cur_trans->dirty_bgs_lock); @@ -3609,33 +3364,18 @@ again: } /* - * cache_write_mutex is here only to save us from balance or automatic - * removal of empty block groups deleting this block group while we are - * writing out the cache + * dirty_bgs_update_mutex is here only to save us from balance or + * automatic removal of empty block groups deleting this block group + * while we are updating its item */ - mutex_lock(&trans->transaction->cache_write_mutex); + mutex_lock(&trans->transaction->dirty_bgs_update_mutex); while (!list_empty(&dirty)) { bool drop_reserve = true; cache = list_first_entry(&dirty, struct btrfs_block_group, dirty_list); - /* - * This can happen if something re-dirties a block group that - * is already under IO. Just wait for it to finish and then do - * it all again - */ - if (!list_empty(&cache->io_list)) { - list_del_init(&cache->io_list); - btrfs_wait_cache_io(trans, cache, path); - btrfs_put_block_group(cache); - } - /* - * btrfs_wait_cache_io uses the cache->dirty_list to decide if - * it should update the cache_state. Don't delete until after - * we wait. - * * Since we're not running in the commit critical section * we need the dirty_bgs_lock to protect from update_block_group */ @@ -3643,72 +3383,39 @@ again: list_del_init(&cache->dirty_list); spin_unlock(&cur_trans->dirty_bgs_lock); - should_put = 1; - - cache_save_setup(cache, trans, path); - - if (cache->disk_cache_state == BTRFS_DC_SETUP) { - cache->io_ctl.inode = NULL; - ret = btrfs_write_out_cache(trans, cache, path); - if (ret == 0 && cache->io_ctl.inode) { - should_put = 0; - - /* - * The cache_write_mutex is protecting the - * io_list, also refer to the definition of - * btrfs_transaction::io_bgs for more details - */ - list_add_tail(&cache->io_list, io); - } else { - /* - * If we failed to write the cache, the - * generation will be bad and life goes on - */ - ret = 0; - } - } - if (!ret) { - ret = update_block_group_item(trans, path, cache); - /* - * Our block group might still be attached to the list - * of new block groups in the transaction handle of some - * other task (struct btrfs_trans_handle->new_bgs). This - * means its block group item isn't yet in the extent - * tree. If this happens ignore the error, as we will - * try again later in the critical section of the - * transaction commit. - */ - if (ret == -ENOENT) { - ret = 0; - spin_lock(&cur_trans->dirty_bgs_lock); - if (list_empty(&cache->dirty_list)) { - list_add_tail(&cache->dirty_list, - &cur_trans->dirty_bgs); - btrfs_get_block_group(cache); - drop_reserve = false; - } - spin_unlock(&cur_trans->dirty_bgs_lock); - } else if (ret) { - btrfs_abort_transaction(trans, ret); + ret = update_block_group_item(trans, path, cache); + /* + * Our block group might still be attached to the list of new + * block groups in the transaction handle of some other task + * (struct btrfs_trans_handle->new_bgs). This means its block + * group item isn't yet in the extent tree. If this happens + * ignore the error, as we will try again later in the critical + * section of the transaction commit. + */ + if (ret == -ENOENT) { + ret = 0; + spin_lock(&cur_trans->dirty_bgs_lock); + if (list_empty(&cache->dirty_list)) { + list_add_tail(&cache->dirty_list, + &cur_trans->dirty_bgs); + btrfs_get_block_group(cache); + drop_reserve = false; } + spin_unlock(&cur_trans->dirty_bgs_lock); + } else if (ret) { + btrfs_abort_transaction(trans, ret); } - /* If it's not on the io list, we need to put the block group */ - if (should_put) - btrfs_put_block_group(cache); + btrfs_put_block_group(cache); if (drop_reserve) btrfs_dec_delayed_refs_rsv_bg_updates(fs_info); - /* - * Avoid blocking other tasks for too long. It might even save - * us from writing caches for block groups that are going to be - * removed. - */ - mutex_unlock(&trans->transaction->cache_write_mutex); + /* Avoid blocking other tasks for too long. */ + mutex_unlock(&trans->transaction->dirty_bgs_update_mutex); if (ret) goto out; - mutex_lock(&trans->transaction->cache_write_mutex); + mutex_lock(&trans->transaction->dirty_bgs_update_mutex); } - mutex_unlock(&trans->transaction->cache_write_mutex); + mutex_unlock(&trans->transaction->dirty_bgs_update_mutex); /* * Go through delayed refs for all the stuff we've just kicked off @@ -3722,7 +3429,7 @@ again: list_splice_init(&cur_trans->dirty_bgs, &dirty); /* * dirty_bgs_lock protects us from concurrent block group - * deletes too (not just cache_write_mutex). + * deletes too (not just dirty_bgs_update_mutex). */ if (!list_empty(&dirty)) { spin_unlock(&cur_trans->dirty_bgs_lock); @@ -3747,121 +3454,34 @@ int btrfs_write_dirty_block_groups(struct btrfs_trans_handle *trans) struct btrfs_block_group *cache; struct btrfs_transaction *cur_trans = trans->transaction; int ret = 0; - int should_put; BTRFS_PATH_AUTO_FREE(path); - struct list_head *io = &cur_trans->io_bgs; path = btrfs_alloc_path(); if (!path) return -ENOMEM; - /* - * Even though we are in the critical section of the transaction commit, - * we can still have concurrent tasks adding elements to this - * transaction's list of dirty block groups. These tasks correspond to - * endio free space workers started when writeback finishes for a - * space cache, which run inode.c:btrfs_finish_ordered_io(), and can - * allocate new block groups as a result of COWing nodes of the root - * tree when updating the free space inode. The writeback for the space - * caches is triggered by an earlier call to - * btrfs_start_dirty_block_groups() and iterations of the following - * loop. - * Also we want to do the cache_save_setup first and then run the - * delayed refs to make sure we have the best chance at doing this all - * in one shot. - */ spin_lock(&cur_trans->dirty_bgs_lock); while (!list_empty(&cur_trans->dirty_bgs)) { cache = list_first_entry(&cur_trans->dirty_bgs, struct btrfs_block_group, dirty_list); - - /* - * This can happen if cache_save_setup re-dirties a block group - * that is already under IO. Just wait for it to finish and - * then do it all again - */ - if (!list_empty(&cache->io_list)) { - spin_unlock(&cur_trans->dirty_bgs_lock); - list_del_init(&cache->io_list); - btrfs_wait_cache_io(trans, cache, path); - btrfs_put_block_group(cache); - spin_lock(&cur_trans->dirty_bgs_lock); - } - - /* - * Don't remove from the dirty list until after we've waited on - * any pending IO - */ list_del_init(&cache->dirty_list); spin_unlock(&cur_trans->dirty_bgs_lock); - should_put = 1; - - cache_save_setup(cache, trans, path); if (!ret) ret = btrfs_run_delayed_refs(trans, U64_MAX); - - if (!ret && cache->disk_cache_state == BTRFS_DC_SETUP) { - cache->io_ctl.inode = NULL; - ret = btrfs_write_out_cache(trans, cache, path); - if (ret == 0 && cache->io_ctl.inode) { - should_put = 0; - list_add_tail(&cache->io_list, io); - } else { - /* - * If we failed to write the cache, the - * generation will be bad and life goes on - */ - ret = 0; - } - } if (!ret) { ret = update_block_group_item(trans, path, cache); - /* - * One of the free space endio workers might have - * created a new block group while updating a free space - * cache's inode (at inode.c:btrfs_finish_ordered_io()) - * and hasn't released its transaction handle yet, in - * which case the new block group is still attached to - * its transaction handle and its creation has not - * finished yet (no block group item in the extent tree - * yet, etc). If this is the case, wait for all free - * space endio workers to finish and retry. This is a - * very rare case so no need for a more efficient and - * complex approach. - */ - if (ret == -ENOENT) { - wait_event(cur_trans->writer_wait, - atomic_read(&cur_trans->num_writers) == 1); - ret = update_block_group_item(trans, path, cache); - if (ret) - btrfs_abort_transaction(trans, ret); - } else if (ret) { + if (ret) btrfs_abort_transaction(trans, ret); - } } - /* If its not on the io list, we need to put the block group */ - if (should_put) - btrfs_put_block_group(cache); + btrfs_put_block_group(cache); btrfs_dec_delayed_refs_rsv_bg_updates(fs_info); spin_lock(&cur_trans->dirty_bgs_lock); } spin_unlock(&cur_trans->dirty_bgs_lock); - /* - * Refer to the definition of io_bgs member for details why it's safe - * to use it without any locking - */ - while (!list_empty(io)) { - cache = list_first_entry(io, struct btrfs_block_group, - io_list); - list_del_init(&cache->io_list); - btrfs_wait_cache_io(trans, cache, path); - btrfs_put_block_group(cache); - } - return ret; } @@ -3905,10 +3525,10 @@ int btrfs_update_block_group(struct btrfs_trans_handle *trans, factor = btrfs_bg_type_to_factor(cache->flags); /* - * If this block group has free space cache written out, we need to make - * sure to load it if we are removing space. This is because we need - * the unpinning stage to actually add the space back to the block group, - * otherwise we will leak space. + * Make sure the free space of this block group is loaded if we are + * removing space. This is because we need the unpinning stage to + * actually add the space back to the block group, otherwise we will + * leak space. */ if (!alloc && !btrfs_block_group_done(cache)) btrfs_cache_block_group(cache, true); @@ -3916,10 +3536,6 @@ int btrfs_update_block_group(struct btrfs_trans_handle *trans, spin_lock(&space_info->lock); spin_lock(&cache->lock); - if (btrfs_test_opt(info, SPACE_CACHE) && - cache->disk_cache_state < BTRFS_DC_CLEAR) - cache->disk_cache_state = BTRFS_DC_CLEAR; - old_val = cache->used; if (alloc) { old_val += num_bytes; @@ -3964,7 +3580,7 @@ int btrfs_update_block_group(struct btrfs_trans_handle *trans, /* * No longer have used bytes in this block group, queue it for deletion. * We do this after adding the block group to the dirty list to avoid - * races between cleaner kthread and space cache writeout. + * races between the cleaner kthread and the dirty block group writeout. */ if (!alloc && old_val == 0) { if (!btrfs_test_opt(info, DISCARD_ASYNC)) @@ -4640,7 +4256,6 @@ void btrfs_put_block_group_cache(struct btrfs_fs_info *info) block_group->inode = NULL; spin_unlock(&block_group->lock); - ASSERT(block_group->io_ctl.inode == NULL); iput(&inode->vfs_inode); } else { spin_unlock(&block_group->lock); @@ -4777,7 +4392,6 @@ int btrfs_free_block_groups(struct btrfs_fs_info *info) btrfs_remove_free_space_cache(block_group); ASSERT(block_group->cached != BTRFS_CACHE_STARTED); ASSERT(list_empty(&block_group->dirty_list)); - ASSERT(list_empty(&block_group->io_list)); ASSERT(list_empty(&block_group->bg_list)); ASSERT(refcount_read(&block_group->refs) == 1); ASSERT(block_group->swap_extents == 0); diff --git a/fs/btrfs/block-group.h b/fs/btrfs/block-group.h index 69d56864d4ba..d567ed822e55 100644 --- a/fs/btrfs/block-group.h +++ b/fs/btrfs/block-group.h @@ -20,13 +20,6 @@ struct btrfs_fs_info; struct btrfs_inode; struct btrfs_trans_handle; -enum btrfs_disk_cache_state { - BTRFS_DC_WRITTEN, - BTRFS_DC_ERROR, - BTRFS_DC_CLEAR, - BTRFS_DC_SETUP, -}; - enum btrfs_block_group_size_class { /* Unset */ BTRFS_BG_SZ_NONE, @@ -131,11 +124,10 @@ struct btrfs_block_group { u64 delalloc_bytes; u64 bytes_super; u64 flags; - u64 cache_generation; u64 global_root_id; u64 remap_bytes; u32 identity_remap_count; - /* The last commited identity_remap_count value of this block group. */ + /* The last committed identity_remap_count value of this block group. */ u32 last_identity_remap_count; /* * The last committed used bytes of this block group, if the above @used @@ -171,8 +163,6 @@ struct btrfs_block_group { unsigned long full_stripe_len; unsigned long runtime_flags; - enum btrfs_disk_cache_state disk_cache_state; - /* Cache tracking stuff */ enum btrfs_caching_type cached; struct btrfs_caching_control *caching_ctl; @@ -228,9 +218,6 @@ struct btrfs_block_group { /* For dirty block groups */ struct list_head dirty_list; - struct list_head io_list; - - struct btrfs_io_ctl io_ctl; /* * Incremented when doing extent allocations and holding a read lock @@ -368,7 +355,6 @@ int btrfs_inc_block_group_ro(struct btrfs_block_group *cache, void btrfs_dec_block_group_ro(struct btrfs_block_group *cache); int btrfs_start_dirty_block_groups(struct btrfs_trans_handle *trans); int btrfs_write_dirty_block_groups(struct btrfs_trans_handle *trans); -int btrfs_setup_space_cache(struct btrfs_trans_handle *trans); int btrfs_update_block_group(struct btrfs_trans_handle *trans, u64 bytenr, u64 num_bytes, bool alloc); int btrfs_add_reserved_bytes(struct btrfs_block_group *cache, diff --git a/fs/btrfs/btrfs_inode.h b/fs/btrfs/btrfs_inode.h index 1082fa92c145..b673851d8d2e 100644 --- a/fs/btrfs/btrfs_inode.h +++ b/fs/btrfs/btrfs_inode.h @@ -32,6 +32,7 @@ struct btrfs_trans_handle; struct btrfs_bio; struct btrfs_file_extent; struct btrfs_delayed_node; +struct btrfs_dir_index_prealloc; /* * Since we search a directory based on f_pos (struct dir_context::pos) we have @@ -402,8 +403,6 @@ static inline void btrfs_mod_outstanding_extents(struct btrfs_inode *inode, { lockdep_assert_held(&inode->lock); inode->outstanding_extents += mod; - if (btrfs_is_free_space_inode(inode)) - return; trace_btrfs_inode_mod_outstanding_extents(inode->root, btrfs_ino(inode), mod, inode->outstanding_extents); } @@ -507,14 +506,10 @@ static inline void btrfs_set_inode_mapping_order(struct btrfs_inode *inode) inode->root->fs_info->block_max_order); } -void btrfs_calculate_block_csum_folio(struct btrfs_fs_info *fs_info, - const phys_addr_t paddr, u8 *dest); -void btrfs_calculate_block_csum_pages(struct btrfs_fs_info *fs_info, - const phys_addr_t paddrs[], u8 *dest); -int btrfs_check_block_csum(struct btrfs_fs_info *fs_info, phys_addr_t paddr, u8 *csum, - const u8 * const csum_expected); -bool btrfs_data_csum_ok(struct btrfs_bio *bbio, struct btrfs_device *dev, - u32 bio_offset, const phys_addr_t paddrs[]); +bool btrfs_bio_data_csum_ok(struct btrfs_bio *bbio, const struct bvec_iter *orig_iter, + struct btrfs_device *dev); +void btrfs_csum_one_bio_block(struct btrfs_fs_info *fs_info, struct bio *bio, + const struct bvec_iter *orig_iter, u8 *csum); noinline int can_nocow_extent(struct btrfs_inode *inode, u64 offset, u64 *len, struct btrfs_file_extent *file_extent, bool nowait); @@ -527,7 +522,8 @@ int btrfs_unlink_inode(struct btrfs_trans_handle *trans, const struct fscrypt_str *name); int btrfs_add_link(struct btrfs_trans_handle *trans, struct btrfs_inode *parent_inode, struct btrfs_inode *inode, - const struct fscrypt_str *name, bool add_backref, u64 index); + const struct fscrypt_str *name, bool add_backref, u64 index, + struct btrfs_dir_index_prealloc *prealloc); int btrfs_delete_subvolume(struct btrfs_inode *dir, struct dentry *dentry); int btrfs_truncate_block(struct btrfs_inode *inode, u64 offset, u64 start, u64 end); @@ -559,7 +555,7 @@ int btrfs_new_inode_prepare(struct btrfs_new_inode_args *args, int btrfs_create_new_inode(struct btrfs_trans_handle *trans, struct btrfs_new_inode_args *args); void btrfs_new_inode_args_destroy(struct btrfs_new_inode_args *args); -struct inode *btrfs_new_subvol_inode(struct mnt_idmap *idmap, +struct inode *btrfs_new_subvol_inode(const struct mnt_idmap *idmap, struct inode *dir); void btrfs_set_delalloc_extent(struct btrfs_inode *inode, struct extent_state *state, u32 bits); @@ -594,10 +590,6 @@ int btrfs_wait_on_delayed_iputs(struct btrfs_fs_info *fs_info); int btrfs_prealloc_file_range(struct inode *inode, int mode, u64 start, u64 num_bytes, u64 min_size, loff_t actual_len, u64 *alloc_hint); -int btrfs_prealloc_file_range_trans(struct inode *inode, - struct btrfs_trans_handle *trans, int mode, - u64 start, u64 num_bytes, u64 min_size, - loff_t actual_len, u64 *alloc_hint); int btrfs_run_delalloc_range(struct btrfs_inode *inode, struct folio *locked_folio, u64 start, u64 end, struct writeback_control *wbc); void btrfs_queue_writepage_fixup(struct btrfs_inode *inode, struct folio *folio); diff --git a/fs/btrfs/compression.c b/fs/btrfs/compression.c index c62b5148d5ac..ad90e14032db 100644 --- a/fs/btrfs/compression.c +++ b/fs/btrfs/compression.c @@ -168,10 +168,10 @@ static unsigned long btrfs_compr_pool_scan(struct shrinker *sh, struct shrink_co spin_unlock(&compr_pool.lock); list_for_each_safe(tmp, next, &remove) { - struct page *page = list_entry(tmp, struct page, lru); + struct folio *folio = list_entry(tmp, struct folio, lru); - ASSERT(page_ref_count(page) == 1); - put_page(page); + ASSERT(folio_ref_count(folio) == 1); + folio_put(folio); } return freed; @@ -431,7 +431,7 @@ static noinline int add_ra_bio_folios(struct inode *inode, u64 compressed_end, } /* - * Since add_ra_bio_pages() is always speculative, suppress + * Since add_ra_bio_folios() is always speculative, suppress * allocation warnings. */ masked_constraint_gfp = mapping_gfp_constraint(mapping, constraint_gfp); @@ -960,7 +960,7 @@ bool btrfs_compress_level_valid(unsigned int type, int level) return levels->min_level <= level && level <= levels->max_level; } -/* Wrapper around find_get_page(), with extra error message. */ +/* Wrapper around filemap_get_folio(), with extra error message. */ int btrfs_compress_filemap_get_folio(struct address_space *mapping, u64 start, struct folio **in_folio_ret) { @@ -1488,10 +1488,11 @@ static bool sample_repeated_patterns(struct heuristic_ws *ws) static void heuristic_collect_sample(struct inode *inode, u64 start, u64 end, struct heuristic_ws *ws) { - struct page *page; - pgoff_t index, index_end; - u32 i, curr_sample_pos; - u8 *in_data; + const u32 blocksize = BTRFS_I(inode)->root->fs_info->sectorsize; + u64 cur = start; + u32 curr_sample_pos = 0; + + ASSERT(IS_ALIGNED(start, blocksize) && IS_ALIGNED(end + 1, blocksize)); /* * Compression handles the input data by chunks of 128KiB @@ -1502,38 +1503,30 @@ static void heuristic_collect_sample(struct inode *inode, u64 start, u64 end, * MAX_SAMPLE_SIZE - calculated under assumption that heuristic will * process no more than BTRFS_MAX_UNCOMPRESSED at a time. */ - if (end - start > BTRFS_MAX_UNCOMPRESSED) - end = start + BTRFS_MAX_UNCOMPRESSED; - - index = start >> PAGE_SHIFT; - index_end = end >> PAGE_SHIFT; - - /* Don't miss unaligned end */ - if (!PAGE_ALIGNED(end)) - index_end++; - - curr_sample_pos = 0; - while (index < index_end) { - page = find_get_page(inode->i_mapping, index); - in_data = kmap_local_page(page); - /* Handle case where the start is not aligned to PAGE_SIZE */ - i = start % PAGE_SIZE; - while (i < PAGE_SIZE - SAMPLING_READ_SIZE) { - /* Don't sample any garbage from the last page */ - if (start > end - SAMPLING_READ_SIZE) - break; - memcpy(&ws->sample[curr_sample_pos], &in_data[i], - SAMPLING_READ_SIZE); - i += SAMPLING_INTERVAL; - start += SAMPLING_INTERVAL; + if (end + 1 - start > BTRFS_MAX_UNCOMPRESSED) + end = start + BTRFS_MAX_UNCOMPRESSED - 1; + + while (cur < end) { + struct folio *folio; + void *in_data; + u64 next_pos; + + folio = filemap_get_folio(inode->i_mapping, cur >> PAGE_SHIFT); + /* All folios inside the range should exist and be locked. */ + ASSERT(!IS_ERR(folio)); + next_pos = min_t(u64, end + 1, folio_next_pos(folio)); + in_data = kmap_local_folio(folio, 0); + + for (; cur < next_pos; cur += SAMPLING_INTERVAL) { + memcpy(&ws->sample[curr_sample_pos], + in_data + offset_in_folio(folio, cur), + SAMPLING_READ_SIZE); curr_sample_pos += SAMPLING_READ_SIZE; } kunmap_local(in_data); - put_page(page); - - index++; + folio_put(folio); + cur = next_pos; } - ws->sample_size = curr_sample_pos; } diff --git a/fs/btrfs/ctree.c b/fs/btrfs/ctree.c index 8fe330d81b8f..fef0e49dd918 100644 --- a/fs/btrfs/ctree.c +++ b/fs/btrfs/ctree.c @@ -3943,7 +3943,7 @@ static noinline int split_item(struct btrfs_trans_handle *trans, orig_offset = btrfs_item_offset(leaf, path->slots[0]); item_size = btrfs_item_size(leaf, path->slots[0]); - buf = kmalloc(item_size, GFP_NOFS); + buf = kvmalloc(item_size, GFP_NOFS); if (!buf) return -ENOMEM; @@ -3981,7 +3981,7 @@ static noinline int split_item(struct btrfs_trans_handle *trans, btrfs_mark_buffer_dirty(trans, leaf); BUG_ON(btrfs_leaf_free_space(leaf) < 0); - kfree(buf); + kvfree(buf); return 0; } diff --git a/fs/btrfs/delalloc-space.c b/fs/btrfs/delalloc-space.c index d357ed7efd99..77781852e417 100644 --- a/fs/btrfs/delalloc-space.c +++ b/fs/btrfs/delalloc-space.c @@ -132,9 +132,7 @@ int btrfs_alloc_data_chunk_ondemand(const struct btrfs_inode *inode, u64 bytes) /* Make sure bytes are sectorsize aligned */ bytes = ALIGN(bytes, fs_info->sectorsize); - if (btrfs_is_free_space_inode(inode)) - flush = BTRFS_RESERVE_FLUSH_FREE_SPACE_INODE; - else if (btrfs_is_zoned(fs_info) && btrfs_is_data_reloc_root(root)) + if (btrfs_is_zoned(fs_info) && btrfs_is_data_reloc_root(root)) flush = BTRFS_RESERVE_FLUSH_ZONED_RELOCATION; return btrfs_reserve_data_bytes(data_sinfo_for_inode(inode), bytes, flush); @@ -155,8 +153,6 @@ int btrfs_check_data_free_space(struct btrfs_inode *inode, if (noflush) flush = BTRFS_RESERVE_NO_FLUSH; - else if (btrfs_is_free_space_inode(inode)) - flush = BTRFS_RESERVE_FLUSH_FREE_SPACE_INODE; ret = btrfs_reserve_data_bytes(data_sinfo_for_inode(inode), len, flush); if (ret < 0) @@ -326,15 +322,10 @@ int btrfs_delalloc_reserve_metadata(struct btrfs_inode *inode, u64 num_bytes, int ret = 0; /* - * If we are a free space inode we need to not flush since we will be in - * the middle of a transaction commit. We also don't need the delalloc - * mutex since we won't race with anybody. We need this mostly to make - * lockdep shut its filthy mouth. - * * If we have a transaction open (can happen if we call truncate_block * from truncate), then we need FLUSH_LIMIT so we don't deadlock. */ - if (noflush || btrfs_is_free_space_inode(inode)) { + if (noflush) { flush = BTRFS_RESERVE_NO_FLUSH; } else { if (current->journal_info) diff --git a/fs/btrfs/delayed-inode.c b/fs/btrfs/delayed-inode.c index db2ffab0941a..1f20148c1fd4 100644 --- a/fs/btrfs/delayed-inode.c +++ b/fs/btrfs/delayed-inode.c @@ -6,6 +6,7 @@ #include <linux/slab.h> #include <linux/iversion.h> +#include <linux/error-injection.h> #include "ctree.h" #include "fs.h" #include "messages.h" @@ -686,7 +687,7 @@ static int btrfs_insert_delayed_item(struct btrfs_trans_handle *trans, /* * For delayed items to insert, we track reserved metadata bytes based * on the number of leaves that we will use. - * See btrfs_insert_delayed_dir_index() and + * See btrfs_insert_delayed_dir_index_prealloc() and * btrfs_delayed_item_reserve_metadata()). */ ASSERT(first_item->bytes_reserved == 0); @@ -1469,35 +1470,93 @@ static void btrfs_release_dir_index_item_space(struct btrfs_trans_handle *trans) trans->bytes_reserved -= bytes; } -/* Will return 0, -ENOMEM or -EEXIST (index number collision, unexpected). */ -int btrfs_insert_delayed_dir_index(struct btrfs_trans_handle *trans, - const char *name, int name_len, - struct btrfs_inode *dir, - const struct btrfs_disk_key *disk_key, u8 flags, - u64 index) +/* + * Pre-allocate a delayed node and delayed item for a dir index insertion and + * copy the name into the item. Call this before modifying the btree so that + * ENOMEM can be returned before any on-disk state has changed. + * + * The returned prealloc is consumed by either + * btrfs_insert_delayed_dir_index_prealloc() or + * btrfs_free_delayed_dir_index_prealloc(); it must not be used afterwards. + * + * Returns a prealloc on success, ERR_PTR on allocation failure. + */ +struct btrfs_dir_index_prealloc *btrfs_prealloc_delayed_dir_index(struct btrfs_inode *dir, + const char *name, + int name_len) +{ + struct btrfs_dir_index_prealloc *prealloc; + struct btrfs_delayed_node *node; + struct btrfs_delayed_item *item; + + prealloc = kzalloc_obj(*prealloc, GFP_NOFS); + if (!prealloc) + return ERR_PTR(-ENOMEM); + + node = btrfs_get_or_create_delayed_node(dir, &prealloc->tracker); + if (IS_ERR(node)) { + kfree(prealloc); + return ERR_CAST(node); + } + + item = btrfs_alloc_delayed_item(sizeof(struct btrfs_dir_item) + name_len, + node, BTRFS_DELAYED_INSERTION_ITEM); + if (!item) { + btrfs_release_delayed_node(node, &prealloc->tracker); + kfree(prealloc); + return ERR_PTR(-ENOMEM); + } + + memcpy(item->data + sizeof(struct btrfs_dir_item), name, name_len); + + prealloc->node = node; + prealloc->item = item; + return prealloc; +} +ALLOW_ERROR_INJECTION(btrfs_prealloc_delayed_dir_index, ERRNO); + +/* + * Free resources from btrfs_prealloc_delayed_dir_index() when the btree + * insertion failed and we will not commit the delayed dir index. Does nothing + * if @prealloc is NULL. + */ +void btrfs_free_delayed_dir_index_prealloc(struct btrfs_trans_handle *trans, + struct btrfs_dir_index_prealloc *prealloc) { + if (!prealloc) + return; + + btrfs_release_delayed_item(prealloc->item); + btrfs_release_dir_index_item_space(trans); + btrfs_release_delayed_node(prealloc->node, &prealloc->tracker); + kfree(prealloc); +} + +/* + * Commit a pre-allocated delayed dir index item. @prealloc must have been + * returned by btrfs_prealloc_delayed_dir_index(). This populates the item, + * adds it to the delayed node's rb-tree, and reserves metadata space. It + * cannot fail with ENOMEM. @prealloc is freed here in all cases. + * + * Return 0 or -EEXIST (index number collision, unexpected). + */ +int btrfs_insert_delayed_dir_index_prealloc(struct btrfs_trans_handle *trans, + struct btrfs_inode *dir, + struct btrfs_dir_index_prealloc *prealloc, + const struct btrfs_disk_key *disk_key, + u8 flags, u64 index) +{ + struct btrfs_delayed_node *delayed_node = prealloc->node; + struct btrfs_ref_tracker *tracker = &prealloc->tracker; + struct btrfs_delayed_item *delayed_item = prealloc->item; struct btrfs_fs_info *fs_info = trans->fs_info; const unsigned int leaf_data_size = BTRFS_LEAF_DATA_SIZE(fs_info); - struct btrfs_delayed_node *delayed_node; - struct btrfs_ref_tracker delayed_node_tracker; - struct btrfs_delayed_item *delayed_item; + const int name_len = delayed_item->data_len - sizeof(struct btrfs_dir_item); struct btrfs_dir_item *dir_item; bool reserve_leaf_space; u32 data_len; int ret; - delayed_node = btrfs_get_or_create_delayed_node(dir, &delayed_node_tracker); - if (IS_ERR(delayed_node)) - return PTR_ERR(delayed_node); - - delayed_item = btrfs_alloc_delayed_item(sizeof(*dir_item) + name_len, - delayed_node, - BTRFS_DELAYED_INSERTION_ITEM); - if (!delayed_item) { - ret = -ENOMEM; - goto release_node; - } - delayed_item->index = index; dir_item = (struct btrfs_dir_item *)delayed_item->data; @@ -1506,7 +1565,7 @@ int btrfs_insert_delayed_dir_index(struct btrfs_trans_handle *trans, btrfs_set_stack_dir_data_len(dir_item, 0); btrfs_set_stack_dir_name_len(dir_item, name_len); btrfs_set_stack_dir_flags(dir_item, flags); - memcpy((char *)(dir_item + 1), name, name_len); + /* Name was already copied by btrfs_prealloc_delayed_dir_index(). */ data_len = delayed_item->data_len + sizeof(struct btrfs_item); @@ -1524,7 +1583,9 @@ int btrfs_insert_delayed_dir_index(struct btrfs_trans_handle *trans, if (unlikely(ret)) { btrfs_err(trans->fs_info, "error adding delayed dir index item, name: %.*s, index: %llu, root: %llu, dir: %llu, dir->index_cnt: %llu, delayed_node->index_cnt: %llu, error: %pe", - name_len, name, index, btrfs_root_id(delayed_node->root), + name_len, + (const char *)(dir_item + 1), + index, btrfs_root_id(delayed_node->root), delayed_node->inode_id, dir->index_cnt, delayed_node->index_cnt, ERR_PTR(ret)); btrfs_release_delayed_item(delayed_item); @@ -1562,7 +1623,9 @@ int btrfs_insert_delayed_dir_index(struct btrfs_trans_handle *trans, mutex_unlock(&delayed_node->mutex); release_node: - btrfs_release_delayed_node(delayed_node, &delayed_node_tracker); + /* Must release the node before freeing @tracker's containing struct. */ + btrfs_release_delayed_node(delayed_node, tracker); + kfree(prealloc); return ret; } @@ -1581,7 +1644,7 @@ static bool btrfs_delete_delayed_insertion_item(struct btrfs_delayed_node *node, /* * For delayed items to insert, we track reserved metadata bytes based * on the number of leaves that we will use. - * See btrfs_insert_delayed_dir_index() and + * See btrfs_insert_delayed_dir_index_prealloc() and * btrfs_delayed_item_reserve_metadata()). */ ASSERT(item->bytes_reserved == 0); diff --git a/fs/btrfs/delayed-inode.h b/fs/btrfs/delayed-inode.h index fc752863f89b..57ba96cfaf9c 100644 --- a/fs/btrfs/delayed-inode.h +++ b/fs/btrfs/delayed-inode.h @@ -115,11 +115,23 @@ struct btrfs_delayed_item { }; void btrfs_init_delayed_root(struct btrfs_delayed_root *delayed_root); -int btrfs_insert_delayed_dir_index(struct btrfs_trans_handle *trans, - const char *name, int name_len, - struct btrfs_inode *dir, - const struct btrfs_disk_key *disk_key, u8 flags, - u64 index); + +struct btrfs_dir_index_prealloc { + struct btrfs_delayed_node *node; + struct btrfs_ref_tracker tracker; + struct btrfs_delayed_item *item; +}; + +struct btrfs_dir_index_prealloc *btrfs_prealloc_delayed_dir_index(struct btrfs_inode *dir, + const char *name, + int name_len); +void btrfs_free_delayed_dir_index_prealloc(struct btrfs_trans_handle *trans, + struct btrfs_dir_index_prealloc *prealloc); +int btrfs_insert_delayed_dir_index_prealloc(struct btrfs_trans_handle *trans, + struct btrfs_inode *dir, + struct btrfs_dir_index_prealloc *prealloc, + const struct btrfs_disk_key *disk_key, + u8 flags, u64 index); int btrfs_delete_delayed_dir_index(struct btrfs_trans_handle *trans, struct btrfs_inode *dir, u64 index); diff --git a/fs/btrfs/dev-replace.c b/fs/btrfs/dev-replace.c index 0284be0e4e82..cdfe093e5c5c 100644 --- a/fs/btrfs/dev-replace.c +++ b/fs/btrfs/dev-replace.c @@ -235,7 +235,8 @@ static int btrfs_init_dev_replace_tgtdev(struct btrfs_fs_info *fs_info, struct btrfs_device **device_out) { struct btrfs_fs_devices *fs_devices = fs_info->fs_devices; - struct btrfs_device *device; + struct btrfs_device *device = NULL; + struct btrfs_device *tmp_device; struct file *bdev_file; struct block_device *bdev; u64 devid = BTRFS_DEV_REPLACE_DEVID; @@ -264,8 +265,8 @@ static int btrfs_init_dev_replace_tgtdev(struct btrfs_fs_info *fs_info, sync_blockdev(bdev); - list_for_each_entry(device, &fs_devices->devices, dev_list) { - if (device->bdev == bdev) { + list_for_each_entry(tmp_device, &fs_devices->devices, dev_list) { + if (tmp_device->bdev == bdev) { btrfs_err(fs_info, "target device is in the filesystem!"); ret = -EEXIST; @@ -285,6 +286,7 @@ static int btrfs_init_dev_replace_tgtdev(struct btrfs_fs_info *fs_info, device = btrfs_alloc_device(NULL, &devid, NULL, device_path); if (IS_ERR(device)) { ret = PTR_ERR(device); + device = NULL; goto error; } @@ -328,6 +330,8 @@ static int btrfs_init_dev_replace_tgtdev(struct btrfs_fs_info *fs_info, error: /* Undo the open-time freeze deny. */ + if (device) + btrfs_free_device(device); btrfs_release_device_allow_freeze(bdev_file); return ret; } diff --git a/fs/btrfs/dir-item.c b/fs/btrfs/dir-item.c index 84f1c64423d3..3a90736915af 100644 --- a/fs/btrfs/dir-item.c +++ b/fs/btrfs/dir-item.c @@ -106,8 +106,11 @@ int btrfs_insert_xattr_item(struct btrfs_trans_handle *trans, * Will return 0 or -ENOMEM */ int btrfs_insert_dir_item(struct btrfs_trans_handle *trans, - const struct fscrypt_str *name, struct btrfs_inode *dir, - const struct btrfs_key *location, u8 type, u64 index) + const struct fscrypt_str *name, + struct btrfs_inode *dir, + const struct btrfs_key *location, u8 type, + u64 index, + struct btrfs_dir_index_prealloc *prealloc) { int ret = 0; int ret2 = 0; @@ -119,17 +122,27 @@ int btrfs_insert_dir_item(struct btrfs_trans_handle *trans, struct btrfs_key key; struct btrfs_disk_key disk_key; u32 data_size; + const bool need_delayed_index = (root != root->fs_info->tree_root); key.objectid = btrfs_ino(dir); key.type = BTRFS_DIR_ITEM_KEY; key.offset = btrfs_name_hash(name->name, name->len); path = btrfs_alloc_path(); - if (!path) - return -ENOMEM; + if (!path) { + ret = -ENOMEM; + goto out_free_prealloc; + } btrfs_cpu_key_to_disk(&disk_key, location); + /* Pre-allocate the delayed dir index before modifying the btree. */ + if (need_delayed_index && !prealloc) { + prealloc = btrfs_prealloc_delayed_dir_index(dir, name->name, name->len); + if (IS_ERR(prealloc)) + return PTR_ERR(prealloc); + } + data_size = sizeof(*dir_item) + name->len; dir_item = insert_with_overflow(trans, root, path, &key, data_size, name->name, name->len); @@ -137,7 +150,7 @@ int btrfs_insert_dir_item(struct btrfs_trans_handle *trans, ret = PTR_ERR(dir_item); if (ret == -EEXIST) goto second_insert; - goto out_free; + goto out_free_prealloc; } if (IS_ENCRYPTED(&dir->vfs_inode)) @@ -154,21 +167,21 @@ int btrfs_insert_dir_item(struct btrfs_trans_handle *trans, write_extent_buffer(leaf, name->name, name_ptr, name->len); second_insert: - /* FIXME, use some real flag for selecting the extra index */ - if (root == root->fs_info->tree_root) { + if (!need_delayed_index) { ret = 0; - goto out_free; + goto out_free_prealloc; } btrfs_release_path(path); - ret2 = btrfs_insert_delayed_dir_index(trans, name->name, name->len, dir, - &disk_key, type, index); -out_free: + ret2 = btrfs_insert_delayed_dir_index_prealloc(trans, dir, prealloc, + &disk_key, type, index); if (ret) return ret; - if (ret2) - return ret2; - return 0; + return ret2; + +out_free_prealloc: + btrfs_free_delayed_dir_index_prealloc(trans, prealloc); + return ret; } static struct btrfs_dir_item *btrfs_lookup_match_dir( diff --git a/fs/btrfs/dir-item.h b/fs/btrfs/dir-item.h index e52174a8baf9..8d22976a0a77 100644 --- a/fs/btrfs/dir-item.h +++ b/fs/btrfs/dir-item.h @@ -13,12 +13,14 @@ struct btrfs_path; struct btrfs_inode; struct btrfs_root; struct btrfs_trans_handle; +struct btrfs_dir_index_prealloc; int btrfs_check_dir_item_collision(struct btrfs_root *root, u64 dir_ino, const struct fscrypt_str *name); int btrfs_insert_dir_item(struct btrfs_trans_handle *trans, const struct fscrypt_str *name, struct btrfs_inode *dir, - const struct btrfs_key *location, u8 type, u64 index); + const struct btrfs_key *location, u8 type, u64 index, + struct btrfs_dir_index_prealloc *prealloc); struct btrfs_dir_item *btrfs_lookup_dir_item(struct btrfs_trans_handle *trans, struct btrfs_root *root, struct btrfs_path *path, u64 dir, @@ -53,5 +55,4 @@ static inline u64 btrfs_name_hash(const char *name, int len) { return crc32c((u32)~1, name, len); } - #endif diff --git a/fs/btrfs/disk-io.c b/fs/btrfs/disk-io.c index dc7ad92876c0..7fd8babd74a6 100644 --- a/fs/btrfs/disk-io.c +++ b/fs/btrfs/disk-io.c @@ -125,15 +125,13 @@ int btrfs_buffer_uptodate(struct extent_buffer *eb, u64 parent_transid, return 1; } - if (btrfs_header_generation(eb) != parent_transid) { - btrfs_err_rl(eb->fs_info, + btrfs_err_rl(eb->fs_info, "parent transid verify failed on logical %llu mirror %u wanted %llu found %llu", - eb->start, eb->read_mirror, - parent_transid, btrfs_header_generation(eb)); - clear_extent_buffer_uptodate(eb); - return 0; - } - return 1; + eb->start, eb->read_mirror, + parent_transid, btrfs_header_generation(eb)); + clear_extent_buffer_uptodate(eb); + + return 0; } static bool btrfs_supported_super_csum(u16 csum_type) @@ -176,19 +174,24 @@ static int btrfs_repair_eb_io_failure(const struct extent_buffer *eb, int mirror_num) { struct btrfs_fs_info *fs_info = eb->fs_info; - const u32 step = min(fs_info->nodesize, PAGE_SIZE); - const u32 nr_steps = eb->len / step; - phys_addr_t paddrs[BTRFS_MAX_BLOCKSIZE / PAGE_SIZE]; + struct btrfs_bio *bbio; + int ret; if (sb_rdonly(fs_info->sb)) return -EROFS; + /* + * This bbio is only to queue all pages for btrfs_repair_bbio_failure(). + * Thus it will never get its endio called. + */ + bbio = btrfs_bio_alloc(max(1, fs_info->nodesize >> PAGE_SHIFT), REQ_OP_READ, + BTRFS_I(fs_info->btree_inode), eb->start, NULL, NULL); + bbio->bio.bi_iter.bi_sector = eb->start >> SECTOR_SHIFT; for (int i = 0; i < num_extent_pages(eb); i++) { struct folio *folio = eb->folios[i]; /* No large folio support yet. */ ASSERT(folio_order(folio) == 0); - ASSERT(i < nr_steps); /* * For nodesize < page size, there is just one paddr, with some @@ -197,11 +200,17 @@ static int btrfs_repair_eb_io_failure(const struct extent_buffer *eb, * For nodesize >= page size, it's one or more paddrs, and eb->start * must be aligned to page boundary. */ - paddrs[i] = page_to_phys(&folio->page) + offset_in_page(eb->start); + ret = bio_add_page(&bbio->bio, &folio->page, min(PAGE_SIZE, fs_info->nodesize), + offset_in_page(eb->start)); + ASSERT(ret == min(PAGE_SIZE, fs_info->nodesize)); } + /* Since the bbio is never submitted, we have to save the iter manually. */ + bbio->saved_iter = bbio->bio.bi_iter; - return btrfs_repair_io_failure(fs_info, 0, eb->start, eb->len, - eb->start, paddrs, step, mirror_num); + ret = btrfs_repair_bbio_failure(bbio, &bbio->saved_iter, fs_info->nodesize, + mirror_num); + bio_put(&bbio->bio); + return ret; } /* @@ -1485,7 +1494,9 @@ static int cleaner_kthread(void *arg) btrfs_run_delayed_iputs(fs_info); + set_bit(BTRFS_QGROUP_RUNTIME_BIT_REJECT_RESCAN, &fs_info->qgroup_flags); again = btrfs_clean_one_deleted_snapshot(fs_info); + clear_bit(BTRFS_QGROUP_RUNTIME_BIT_REJECT_RESCAN, &fs_info->qgroup_flags); mutex_unlock(&fs_info->cleaner_mutex); /* @@ -1769,7 +1780,6 @@ static void btrfs_stop_all_workers(struct btrfs_fs_info *fs_info) if (fs_info->rmw_workers) destroy_workqueue(fs_info->rmw_workers); btrfs_destroy_workqueue(fs_info->endio_write_workers); - btrfs_destroy_workqueue(fs_info->endio_freespace_worker); btrfs_destroy_workqueue(fs_info->delayed_workers); btrfs_destroy_workqueue(fs_info->caching_workers); btrfs_destroy_workqueue(fs_info->flush_workers); @@ -1980,9 +1990,6 @@ static int btrfs_init_workqueues(struct btrfs_fs_info *fs_info) fs_info->endio_write_workers = btrfs_alloc_workqueue(fs_info, "endio-write", flags, max_active, 2); - fs_info->endio_freespace_worker = - btrfs_alloc_workqueue(fs_info, "freespace-write", flags, - max_active, 0); fs_info->delayed_workers = btrfs_alloc_workqueue(fs_info, "delayed-meta", flags, max_active, 0); @@ -1995,8 +2002,7 @@ static int btrfs_init_workqueues(struct btrfs_fs_info *fs_info) if (!(fs_info->workers && fs_info->delalloc_workers && fs_info->flush_workers && fs_info->endio_workers && fs_info->endio_meta_workers && - fs_info->endio_write_workers && - fs_info->endio_freespace_worker && fs_info->rmw_workers && + fs_info->endio_write_workers && fs_info->rmw_workers && fs_info->caching_workers && fs_info->fixup_workers && fs_info->delayed_workers && fs_info->qgroup_rescan_workers && fs_info->discard_ctl.discard_workers)) { @@ -2374,7 +2380,7 @@ static int validate_sys_chunk_array(const struct btrfs_fs_info *fs_info, } ret = btrfs_check_chunk_valid(fs_info, NULL, chunk, key.offset, sectorsize); - if (ret < 0) + if (unlikely(ret < 0)) return ret; cur += btrfs_chunk_item_size(num_stripes); } @@ -3063,7 +3069,6 @@ static int btrfs_cleanup_fs_roots(struct btrfs_fs_info *fs_info) int btrfs_start_pre_rw_mount(struct btrfs_fs_info *fs_info) { int ret; - const bool cache_opt = btrfs_test_opt(fs_info, SPACE_CACHE); bool rebuild_free_space_tree = false; if (btrfs_test_opt(fs_info, CLEAR_CACHE) && @@ -3158,8 +3163,8 @@ int btrfs_start_pre_rw_mount(struct btrfs_fs_info *fs_info) } } - if (cache_opt != btrfs_free_space_cache_v1_active(fs_info)) { - ret = btrfs_set_free_space_cache_v1_active(fs_info, cache_opt); + if (btrfs_free_space_cache_v1_active(fs_info)) { + ret = btrfs_cleanup_free_space_cache_v1(fs_info); if (ret) return ret; } @@ -3275,20 +3280,6 @@ int btrfs_check_features(struct btrfs_fs_info *fs_info, bool is_rw_mount) return -EINVAL; } - /* - * Subpage/bs > ps runtime limitation on v1 cache. - * - * V1 space cache still has some hard coded PAGE_SIZE usage, while - * we're already defaulting to v2 cache, no need to bother v1 as it's - * going to be deprecated anyway. - */ - if (fs_info->sectorsize != PAGE_SIZE && btrfs_test_opt(fs_info, SPACE_CACHE)) { - btrfs_warn(fs_info, - "v1 space cache is not supported for page size %lu with sectorsize %u", - PAGE_SIZE, fs_info->sectorsize); - return -EINVAL; - } - /* This can be called by remount, we need to protect the super block. */ spin_lock(&fs_info->super_lock); btrfs_set_super_incompat_flags(disk_super, incompat); @@ -4217,8 +4208,6 @@ int write_all_supers(struct btrfs_trans_handle *trans) total_errors++; } if (unlikely(total_errors > max_errors)) { - btrfs_err(fs_info, "%d errors while writing supers", - total_errors); mutex_unlock(&fs_info->fs_devices->device_list_mutex); /* FUA is masked off if unsupported and can't be the reason */ @@ -4446,9 +4435,8 @@ void __cold close_ctree(struct btrfs_fs_info *fs_info) * to finish an ordered extent - end_bbio_compressed_write() * calls btrfs_finish_ordered_extent() which in turns does a call to * btrfs_queue_ordered_fn(), and that queues the ordered extent - * completion either in the endio_write_workers work queue or in the - * fs_info->endio_freespace_worker work queue. We flush those queues - * below, so before we flush them we must flush this queue for the + * completion in the endio_write_workers work queue. We flush that + * queue below, so before we flush it we must flush this queue for the * workers of compressed writes. */ flush_workqueue(fs_info->endio_workers); @@ -4474,8 +4462,6 @@ void __cold close_ctree(struct btrfs_fs_info *fs_info) * btrfs_finish_ordered_io() when we are unmounting). */ btrfs_flush_workqueue(fs_info->endio_write_workers); - /* Ordered extents for free space inodes. */ - btrfs_flush_workqueue(fs_info->endio_freespace_worker); /* * Run delayed iputs in case an async reclaim worker is waiting for them * to be run as mentioned above. @@ -4864,26 +4850,6 @@ static void btrfs_destroy_pinned_extent(struct btrfs_fs_info *fs_info, } } -static void btrfs_cleanup_bg_io(struct btrfs_block_group *cache) -{ - struct inode *inode; - - inode = cache->io_ctl.inode; - if (inode) { - unsigned int nofs_flag; - - nofs_flag = memalloc_nofs_save(); - invalidate_inode_pages2(inode->i_mapping); - memalloc_nofs_restore(nofs_flag); - - BTRFS_I(inode)->generation = 0; - cache->io_ctl.inode = NULL; - iput(inode); - } - ASSERT(cache->io_ctl.pages == NULL); - btrfs_put_block_group(cache); -} - void btrfs_cleanup_dirty_bgs(struct btrfs_transaction *cur_trans, struct btrfs_fs_info *fs_info) { @@ -4895,40 +4861,13 @@ void btrfs_cleanup_dirty_bgs(struct btrfs_transaction *cur_trans, struct btrfs_block_group, dirty_list); - if (!list_empty(&cache->io_list)) { - spin_unlock(&cur_trans->dirty_bgs_lock); - list_del_init(&cache->io_list); - btrfs_cleanup_bg_io(cache); - spin_lock(&cur_trans->dirty_bgs_lock); - } - list_del_init(&cache->dirty_list); - spin_lock(&cache->lock); - cache->disk_cache_state = BTRFS_DC_ERROR; - spin_unlock(&cache->lock); - spin_unlock(&cur_trans->dirty_bgs_lock); btrfs_put_block_group(cache); btrfs_dec_delayed_refs_rsv_bg_updates(fs_info); spin_lock(&cur_trans->dirty_bgs_lock); } spin_unlock(&cur_trans->dirty_bgs_lock); - - /* - * Refer to the definition of io_bgs member for details why it's safe - * to use it without any locking - */ - while (!list_empty(&cur_trans->io_bgs)) { - cache = list_first_entry(&cur_trans->io_bgs, - struct btrfs_block_group, - io_list); - - list_del_init(&cache->io_list); - spin_lock(&cache->lock); - cache->disk_cache_state = BTRFS_DC_ERROR; - spin_unlock(&cache->lock); - btrfs_cleanup_bg_io(cache); - } } static void btrfs_free_all_qgroup_pertrans(struct btrfs_fs_info *fs_info) @@ -4964,7 +4903,6 @@ void btrfs_cleanup_one_transaction(struct btrfs_transaction *cur_trans) btrfs_cleanup_dirty_bgs(cur_trans, fs_info); ASSERT(list_empty(&cur_trans->dirty_bgs)); - ASSERT(list_empty(&cur_trans->io_bgs)); list_for_each_entry_safe(dev, tmp, &cur_trans->dev_update_list, post_commit_list) { diff --git a/fs/btrfs/extent-io-tree.c b/fs/btrfs/extent-io-tree.c index d6df11f6088c..992b8b42bdb4 100644 --- a/fs/btrfs/extent-io-tree.c +++ b/fs/btrfs/extent-io-tree.c @@ -751,7 +751,7 @@ hit_next: btrfs_split_delalloc_extent(tree->inode, state, start); /* - * Temporarilly ajdust this state's range to match the + * Temporarily ajdust this state's range to match the * range for which we are clearing bits. */ state->start = start; diff --git a/fs/btrfs/extent-tree.c b/fs/btrfs/extent-tree.c index d6a4390ee34a..a0d5ab03aae2 100644 --- a/fs/btrfs/extent-tree.c +++ b/fs/btrfs/extent-tree.c @@ -6315,6 +6315,16 @@ int btrfs_drop_snapshot(struct btrfs_root *root, bool update_ref, bool for_reloc set_bit(BTRFS_ROOT_DELETING, &root->state); unfinished_drop = test_bit(BTRFS_ROOT_UNFINISHED_DROP, &root->state); + /* + * For subvolume dropping, check if the subvolume is large enough so + * that we need to mark qgroup inconsistent to avoid long qgroup stall. + * + * Even for a subvolume without any snapshot, there can still be + * a lot of qgroup records queued into one transaction. + */ + if (!for_reloc) + btrfs_qgroup_check_tree_drop(fs_info, rootid, + btrfs_header_level(root->node)); if (btrfs_disk_key_objectid(&root_item->drop_progress) == 0) { level = btrfs_header_level(root->node); path->nodes[level] = btrfs_lock_root_node(root); diff --git a/fs/btrfs/extent_io.c b/fs/btrfs/extent_io.c index d7600e5fa3d9..55e9144d4759 100644 --- a/fs/btrfs/extent_io.c +++ b/fs/btrfs/extent_io.c @@ -110,14 +110,14 @@ struct btrfs_bio_ctrl { * make the decision when submitting the bio. * * The pattern between do_readpage(), submit_one_bio() and - * submit_extent_folio() is quite subtle, so tracking this is tricky. + * submit_one_block() is quite subtle, so tracking this is tricky. * * As we process extent E, we might submit a bio with existing built up * extents before adding E to a new bio, or we might just add E to the * bio. As a result, E's generation could apply to the current bio or * to the next one, so we need to be careful to update the bio_ctrl's * generation with E's only when we are sure E is added to bio_ctrl->bbio - * in submit_extent_folio(). + * in submit_one_block(). * * See the comment in btrfs_lookup_bio_sums() for more detail on the * need for this optimization. @@ -797,118 +797,95 @@ static int alloc_new_bio(struct btrfs_inode *inode, } /* - * @disk_bytenr: logical bytenr where the write will be - * @page: page to add to the bio - * @size: portion of page that we want to write to - * @pg_offset: offset of the new bio or to check whether we are adding - * a contiguous page to the previous one + * @disk_bytenr: logical bytenr where the read/write will be + * @folio: the folio the block belongs to + * @pg_offset: the offset inside the folio * @read_em_generation: generation of the extent_map we are submitting * (only used for read) * - * The will either add the page into the existing @bio_ctrl->bbio, or allocate a + * This will either add the block into the existing @bio_ctrl->bbio, or allocate a * new one in @bio_ctrl->bbio. * The mirror number for this IO should already be initialized in * @bio_ctrl->mirror_num. * - * Return the number of bytes that are queued into a bio. - * If the returned bytes is smaller than @size, it means we hit a critical error - * for data write, where there is no ordered extent for the range. + * Return 0 if the block is queued or submitted. + * Return <0 for error. */ -static unsigned int submit_extent_folio(struct btrfs_bio_ctrl *bio_ctrl, - u64 disk_bytenr, struct folio *folio, - size_t size, unsigned long pg_offset, - u64 read_em_generation) +static int submit_one_block(struct btrfs_bio_ctrl *bio_ctrl, + u64 disk_bytenr, struct folio *folio, + unsigned long pg_offset, u64 read_em_generation) { struct btrfs_inode *inode = folio_to_inode(folio); + const struct btrfs_fs_info *fs_info = inode->root->fs_info; + const u32 blocksize = fs_info->sectorsize; loff_t file_offset = folio_pos(folio) + pg_offset; - unsigned int queued = 0; - ASSERT(pg_offset + size <= folio_size(folio)); + ASSERT(pg_offset + blocksize <= folio_size(folio)); ASSERT(bio_ctrl->end_io_func); if (bio_ctrl->bbio && !btrfs_bio_is_contig(bio_ctrl, disk_bytenr, file_offset)) submit_one_bio(bio_ctrl); - do { - u32 len = size; - - /* Allocate new bio if needed */ - if (!bio_ctrl->bbio) { - int ret; - - ret = alloc_new_bio(inode, bio_ctrl, disk_bytenr, file_offset); - if (ret < 0) - break; - } - - /* Cap to the current ordered extent boundary if there is one. */ - if (len > bio_ctrl->len_to_oe_boundary) { - ASSERT(bio_ctrl->compress_type == BTRFS_COMPRESS_NONE); - ASSERT(is_data_inode(inode)); - len = bio_ctrl->len_to_oe_boundary; - } +again: + /* Allocate new bio if needed */ + if (!bio_ctrl->bbio) { + int ret; - if (!bio_add_folio(&bio_ctrl->bbio->bio, folio, len, pg_offset)) { - /* bio full: move on to a new one */ - submit_one_bio(bio_ctrl); - continue; - } - /* - * Now that the folio is definitely added to the bio, include its - * generation in the max generation calculation. - */ - bio_ctrl->generation = max(bio_ctrl->generation, read_em_generation); - bio_ctrl->next_file_offset += len; + ret = alloc_new_bio(inode, bio_ctrl, disk_bytenr, file_offset); + if (ret < 0) + return ret; + } - if (bio_ctrl->wbc) - wbc_account_cgroup_owner(bio_ctrl->wbc, folio, len); + if (!bio_add_folio(&bio_ctrl->bbio->bio, folio, blocksize, pg_offset)) { + /* bio full: move on to a new one */ + submit_one_bio(bio_ctrl); + goto again; + } - size -= len; - pg_offset += len; - disk_bytenr += len; - file_offset += len; - queued += len; + /* + * Now that the folio is definitely added to the bio, include its + * generation in the max generation calculation. + */ + bio_ctrl->generation = max(bio_ctrl->generation, read_em_generation); + bio_ctrl->next_file_offset += blocksize; - /* - * len_to_oe_boundary defaults to U32_MAX, which isn't folio or - * sector aligned. alloc_new_bio() then sets it to the end of - * our ordered extent for writes into zoned devices. - * - * When len_to_oe_boundary is tracking an ordered extent, we - * trust the ordered extent code to align things properly, and - * the check above to cap our write to the ordered extent - * boundary is correct. - * - * When len_to_oe_boundary is U32_MAX, the cap above would - * result in a 4095 byte IO for the last folio right before - * we hit the bio limit of UINT_MAX. bio_add_folio() has all - * the checks required to make sure we don't overflow the bio, - * and we should just ignore len_to_oe_boundary completely - * unless we're using it to track an ordered extent. - * - * It's pretty hard to make a bio sized U32_MAX, but it can - * happen when the page cache is able to feed us contiguous - * folios for large extents. - */ - if (bio_ctrl->len_to_oe_boundary != U32_MAX) - bio_ctrl->len_to_oe_boundary -= len; + if (bio_ctrl->wbc) + wbc_account_cgroup_owner(bio_ctrl->wbc, folio, blocksize); - /* Ordered extent boundary: move on to a new bio. */ - if (bio_ctrl->len_to_oe_boundary == 0) - submit_one_bio(bio_ctrl); - /* - * If we have accumulated decent amount of IO, send it to the - * block layer so that IO can run while we are accumulating - * more folios to write. - */ - else if (bio_ctrl->wbc && - bio_ctrl->bbio->bio.bi_iter.bi_size >= - inode->root->fs_info->writeback_bio_size) - submit_one_bio(bio_ctrl); + /* + * len_to_oe_boundary defaults to U32_MAX, which isn't folio or sector + * aligned. alloc_new_bio() then sets it to the end of our ordered + * extent for writes into zoned devices. + * + * When len_to_oe_boundary is tracking an ordered extent, the + * len_to_oe_boundary should follow that OE and never go beyond the max + * extent size (128MiB). + * + * When len_to_oe_boundary is U32_MAX, decreasing the length by + * blocksize will never make it reach 0, thus skipping the later + * submit_one_bio() call. So if len_to_oe_boundary() is not tracking + * an OE, do not decrease it. + * + * It's pretty hard to make a bio sized U32_MAX, but it can happen when + * the page cache is able to feed us contiguous folios for large + * extents. + */ + if (bio_ctrl->len_to_oe_boundary != U32_MAX) + bio_ctrl->len_to_oe_boundary -= blocksize; - } while (size); - return queued; + /* Ordered extent boundary: move on to a new bio. */ + if (bio_ctrl->len_to_oe_boundary == 0) + submit_one_bio(bio_ctrl); + /* + * If we have accumulated decent amount of IO, send it to the block + * layer so that IO can run while we are accumulating more folios to + * write. + */ + else if (bio_ctrl->wbc && + bio_ctrl->bbio->bio.bi_iter.bi_size >= fs_info->writeback_bio_size) + submit_one_bio(bio_ctrl); + return 0; } static int attach_extent_buffer_folio(struct extent_buffer *eb, @@ -1092,7 +1069,6 @@ static int btrfs_do_readpage(struct folio *folio, struct extent_map **em_cached, u64 disk_bytenr; u64 block_start; u64 em_gen; - unsigned int queued; ASSERT(IS_ALIGNED(cur, fs_info->sectorsize)); if (cur >= last_byte) { @@ -1206,10 +1182,9 @@ static int btrfs_do_readpage(struct folio *folio, struct extent_map **em_cached, if (force_bio_submit) submit_one_bio(bio_ctrl); - queued = submit_extent_folio(bio_ctrl, disk_bytenr, folio, blocksize, - pg_offset, em_gen); + ret = submit_one_block(bio_ctrl, disk_bytenr, folio, pg_offset, em_gen); /* Read submission should not fail. */ - ASSERT(queued == blocksize); + ASSERT(ret == 0); } return 0; } @@ -1808,33 +1783,51 @@ out: return 0; } +static struct btrfs_ordered_extent *get_oe_from_bbio(const struct btrfs_bio *bbio, + u64 filepos) +{ + struct btrfs_ordered_extent *oe; + + if (!bbio || !bbio->ordered) + return NULL; + + oe = bbio->ordered; + if (!in_range(filepos, oe->file_offset, oe->num_bytes)) + return NULL; + + refcount_inc(&oe->refs); + return oe; +} + /* * Return 0 if we have submitted or queued the sector for submission. * Return <0 for critical errors, and the involved sector will be cleaned up. * * Caller should make sure filepos < i_size and handle filepos >= i_size case. */ -static int submit_one_sector(struct btrfs_inode *inode, - struct folio *folio, - u64 filepos, struct btrfs_bio_ctrl *bio_ctrl, - loff_t i_size) +static int submit_write_sector(struct btrfs_inode *inode, + struct folio *folio, + u64 filepos, struct btrfs_bio_ctrl *bio_ctrl, + loff_t i_size) { struct btrfs_fs_info *fs_info = inode->root->fs_info; - struct extent_map *em; + struct btrfs_ordered_extent *oe; u64 block_start; u64 disk_bytenr; u64 extent_offset; - u64 em_end; const u32 sectorsize = fs_info->sectorsize; - unsigned int queued; + int ret; ASSERT(IS_ALIGNED(filepos, sectorsize)); /* @filepos >= i_size case should be handled by the caller. */ ASSERT(filepos < i_size); - em = btrfs_get_extent(inode, NULL, filepos, sectorsize); - if (IS_ERR(em)) { + /* Try to reuse the existing OE from bbio first. */ + oe = get_oe_from_bbio(bio_ctrl->bbio, filepos); + if (!oe) + oe = btrfs_lookup_ordered_extent(inode, filepos); + if (unlikely(!oe)) { /* * bio_ctrl may contain a bio crossing several folios. * Submit it immediately so that the bio has a chance @@ -1857,31 +1850,25 @@ static int submit_one_sector(struct btrfs_inode *inode, */ btrfs_mark_ordered_io_finished(inode, filepos, fs_info->sectorsize, false); - return PTR_ERR(em); + btrfs_err_rl(fs_info, + "no ordered extent for root %lld ino %llu filepos %llu", + btrfs_root_id(inode->root), btrfs_ino(inode), + filepos); + return -EUCLEAN; } - extent_offset = filepos - em->start; - em_end = btrfs_extent_map_end(em); - ASSERT(filepos <= em_end); - ASSERT(IS_ALIGNED(em->start, sectorsize)); - ASSERT(IS_ALIGNED(em->len, sectorsize)); + extent_offset = filepos - oe->file_offset; + ASSERT(filepos < oe->file_offset + oe->num_bytes); + ASSERT(IS_ALIGNED(oe->file_offset, sectorsize)); + ASSERT(IS_ALIGNED(oe->num_bytes, sectorsize)); + ASSERT(oe->compress_type == BTRFS_COMPRESS_NONE); + ASSERT(!test_bit(BTRFS_ORDERED_COMPRESSED, &oe->flags)); - block_start = btrfs_extent_map_block_start(em); - disk_bytenr = btrfs_extent_map_block_start(em) + extent_offset; + block_start = oe->disk_bytenr + oe->offset; + disk_bytenr = block_start + extent_offset; - ASSERT(!btrfs_extent_map_is_compressed(em)); - ASSERT(block_start != EXTENT_MAP_HOLE); - ASSERT(block_start != EXTENT_MAP_INLINE); + btrfs_put_ordered_extent(oe); - btrfs_free_extent_map(em); - em = NULL; - - /* - * Although the PageDirty bit is cleared before entering this - * function, subpage dirty bit is not cleared. - * So clear subpage dirty bit here so next time we won't submit - * a folio for a range already written to disk. - */ btrfs_folio_clear_dirty(fs_info, folio, filepos, sectorsize); btrfs_folio_set_writeback(fs_info, folio, filepos, sectorsize); /* @@ -1892,13 +1879,17 @@ static int submit_one_sector(struct btrfs_inode *inode, */ ASSERT(folio_test_writeback(folio)); - queued = submit_extent_folio(bio_ctrl, disk_bytenr, folio, - sectorsize, filepos - folio_pos(folio), 0); - if (unlikely(queued < sectorsize)) { + ret = submit_one_block(bio_ctrl, disk_bytenr, folio, + offset_in_folio(folio, filepos), 0); + if (unlikely(ret < 0)) { btrfs_folio_clear_writeback(fs_info, folio, filepos, sectorsize); btrfs_mark_ordered_io_finished(inode, filepos, fs_info->sectorsize, false); - return -EUCLEAN; + btrfs_err_rl(fs_info, + "failed to queue sector for root %lld ino %llu filepos %llu: %pe", + btrfs_root_id(inode->root), + btrfs_ino(inode), filepos, ERR_PTR(ret)); + return ret; } return 0; } @@ -1983,7 +1974,7 @@ static noinline_for_stack int extent_writepage_io(struct btrfs_inode *inode, btrfs_folio_clear_dirty(fs_info, folio, cur, fs_info->sectorsize); continue; } - ret = submit_one_sector(inode, folio, cur, bio_ctrl, i_size); + ret = submit_write_sector(inode, folio, cur, bio_ctrl, i_size); if (unlikely(ret < 0)) { if (!found_error) found_error = ret; diff --git a/fs/btrfs/extent_map.c b/fs/btrfs/extent_map.c index 6ad7b39ae358..86d9c6f5ff4b 100644 --- a/fs/btrfs/extent_map.c +++ b/fs/btrfs/extent_map.c @@ -1220,6 +1220,14 @@ static struct btrfs_inode *find_first_inode_to_shrink(struct btrfs_root *root, tree = &inode->extent_tree; /* + * Most inodes have no extent maps, so check without the lock. + * The race is harmless: a false empty just defers the inode to + * a later scan, and a false non-empty is caught under the lock. + */ + if (data_race(RB_EMPTY_ROOT(&tree->root))) + goto next; + + /* * We want to be fast so if the lock is busy we don't want to * spend time waiting for it (some task is about to do IO for * the inode). diff --git a/fs/btrfs/file-item.c b/fs/btrfs/file-item.c index cf50fd623f41..ff8f8cad00fc 100644 --- a/fs/btrfs/file-item.c +++ b/fs/btrfs/file-item.c @@ -397,17 +397,6 @@ int btrfs_lookup_bio_sums(struct btrfs_bio *bbio) path->reada = READA_FORWARD; /* - * the free space stuff is only read when it hasn't been - * updated in the current transaction. So, we can safely - * read from the commit root and sidestep a nasty deadlock - * between reading the free space cache and updating the csum tree. - */ - if (btrfs_is_free_space_inode(inode)) { - path->search_commit_root = true; - path->skip_locking = true; - } - - /* * If we are searching for a csum of an extent from a past * transaction, we can search in the commit root and reduce * lock contention on the csum tree extent buffers. @@ -797,29 +786,20 @@ fail: return ret; } -static void csum_one_bio(struct btrfs_bio *bbio, struct bvec_iter *src) +static void csum_one_bio(struct btrfs_bio *bbio) { struct btrfs_inode *inode = bbio->inode; struct btrfs_fs_info *fs_info = inode->root->fs_info; - struct bio *bio = &bbio->bio; struct btrfs_ordered_sum *sums = bbio->sums; - struct bvec_iter iter = *src; - phys_addr_t paddr; const u32 blocksize = fs_info->sectorsize; - const u32 step = min(blocksize, PAGE_SIZE); - const u32 nr_steps = blocksize / step; - phys_addr_t paddrs[BTRFS_MAX_BLOCKSIZE / PAGE_SIZE]; - u32 offset = 0; int index = 0; - btrfs_bio_for_each_block(paddr, bio, &iter, step) { - paddrs[(offset / step) % nr_steps] = paddr; - offset += step; + for (struct bvec_iter *iter = &bbio->csum_saved_iter; + iter->bi_size; + bio_advance_iter(&bbio->bio, iter, blocksize)) { + btrfs_csum_one_bio_block(fs_info, &bbio->bio, iter, sums->sums + index); - if (IS_ALIGNED(offset, blocksize)) { - btrfs_calculate_block_csum_pages(fs_info, paddrs, sums->sums + index); - index += fs_info->csum_size; - } + index += fs_info->csum_size; } } @@ -828,9 +808,8 @@ static void csum_one_bio_work(struct work_struct *work) struct btrfs_bio *bbio = container_of(work, struct btrfs_bio, csum_work); ASSERT(btrfs_op(&bbio->bio) == BTRFS_MAP_WRITE); - ASSERT(bbio->async_csum == true); - csum_one_bio(bbio, &bbio->csum_saved_iter); - complete(&bbio->csum_done); + csum_one_bio(bbio); + bio_endio(&bbio->bio); } /* @@ -859,15 +838,14 @@ int btrfs_csum_one_bio(struct btrfs_bio *bbio, bool async) bbio->sums = sums; btrfs_add_ordered_sum(ordered, sums); + bbio->csum_saved_iter = bio->bi_iter; if (!async) { - csum_one_bio(bbio, &bbio->bio.bi_iter); + csum_one_bio(bbio); return 0; } - init_completion(&bbio->csum_done); - bbio->async_csum = true; - bbio->csum_saved_iter = bbio->bio.bi_iter; + bio_inc_remaining(bio); INIT_WORK(&bbio->csum_work, csum_one_bio_work); - schedule_work(&bbio->csum_work); + queue_work(fs_info->endio_workers, &bbio->csum_work); return 0; } diff --git a/fs/btrfs/free-space-cache.c b/fs/btrfs/free-space-cache.c index e2af75a205ea..2a40167c3fb6 100644 --- a/fs/btrfs/free-space-cache.c +++ b/fs/btrfs/free-space-cache.c @@ -9,7 +9,6 @@ #include <linux/slab.h> #include <linux/math64.h> #include <linux/ratelimit.h> -#include <linux/error-injection.h> #include <linux/sched/mm.h> #include <linux/string_choices.h> #include "extent-tree.h" @@ -23,7 +22,6 @@ #include "space-info.h" #include "block-group.h" #include "discard.h" -#include "subpage.h" #include "inode-item.h" #include "accessors.h" #include "file-item.h" @@ -38,12 +36,6 @@ static struct kmem_cache *btrfs_free_space_cachep; static struct kmem_cache *btrfs_free_space_bitmap_cachep; -struct btrfs_trim_range { - u64 start; - u64 bytes; - struct list_head list; -}; - static int link_free_space(struct btrfs_free_space_ctl *ctl, struct btrfs_free_space *info); static void unlink_free_space(struct btrfs_free_space_ctl *ctl, @@ -57,11 +49,6 @@ static void bitmap_clear_bits(struct btrfs_free_space_ctl *ctl, struct btrfs_free_space *info, u64 offset, u64 bytes, bool update_stats); -static void btrfs_crc32c_final(u32 crc, u8 *result) -{ - put_unaligned_le32(~crc, result); -} - static void __btrfs_remove_free_space_cache(struct btrfs_free_space_ctl *ctl) { struct btrfs_free_space *info; @@ -123,10 +110,6 @@ static struct inode *__lookup_free_space_inode(struct btrfs_root *root, if (IS_ERR(inode)) return ERR_CAST(inode); - mapping_set_gfp_mask(inode->vfs_inode.i_mapping, - mapping_gfp_constraint(inode->vfs_inode.i_mapping, - ~(__GFP_FS | __GFP_HIGHMEM))); - return &inode->vfs_inode; } @@ -135,7 +118,6 @@ struct inode *lookup_free_space_inode(struct btrfs_block_group *block_group, { struct btrfs_fs_info *fs_info = block_group->fs_info; struct inode *inode = NULL; - u32 flags = BTRFS_INODE_NODATASUM | BTRFS_INODE_NODATACOW; spin_lock(&block_group->lock); if (block_group->inode) @@ -150,13 +132,6 @@ struct inode *lookup_free_space_inode(struct btrfs_block_group *block_group, return inode; spin_lock(&block_group->lock); - if (!((BTRFS_I(inode)->flags & flags) == flags)) { - btrfs_info(fs_info, "Old style space inode found, converting."); - BTRFS_I(inode)->flags |= BTRFS_INODE_NODATASUM | - BTRFS_INODE_NODATACOW; - block_group->disk_cache_state = BTRFS_DC_CLEAR; - } - if (!test_and_set_bit(BLOCK_GROUP_FLAG_IREF, &block_group->runtime_flags)) block_group->inode = BTRFS_I(igrab(inode)); spin_unlock(&block_group->lock); @@ -164,78 +139,6 @@ struct inode *lookup_free_space_inode(struct btrfs_block_group *block_group, return inode; } -static int __create_free_space_inode(struct btrfs_root *root, - struct btrfs_trans_handle *trans, - struct btrfs_path *path, - u64 ino, u64 offset) -{ - struct btrfs_key key; - struct btrfs_disk_key disk_key; - struct btrfs_free_space_header *header; - struct btrfs_inode_item *inode_item; - struct extent_buffer *leaf; - /* We inline CRCs for the free disk space cache */ - const u64 flags = BTRFS_INODE_NOCOMPRESS | BTRFS_INODE_PREALLOC | - BTRFS_INODE_NODATASUM | BTRFS_INODE_NODATACOW; - int ret; - - ret = btrfs_insert_empty_inode(trans, root, path, ino); - if (ret) - return ret; - - leaf = path->nodes[0]; - inode_item = btrfs_item_ptr(leaf, path->slots[0], - struct btrfs_inode_item); - btrfs_item_key(leaf, &disk_key, path->slots[0]); - memzero_extent_buffer(leaf, (unsigned long)inode_item, - sizeof(*inode_item)); - btrfs_set_inode_generation(leaf, inode_item, trans->transid); - btrfs_set_inode_size(leaf, inode_item, 0); - btrfs_set_inode_nbytes(leaf, inode_item, 0); - btrfs_set_inode_uid(leaf, inode_item, 0); - btrfs_set_inode_gid(leaf, inode_item, 0); - btrfs_set_inode_mode(leaf, inode_item, S_IFREG | 0600); - btrfs_set_inode_flags(leaf, inode_item, flags); - btrfs_set_inode_nlink(leaf, inode_item, 1); - btrfs_set_inode_transid(leaf, inode_item, trans->transid); - btrfs_set_inode_block_group(leaf, inode_item, offset); - btrfs_release_path(path); - - key.objectid = BTRFS_FREE_SPACE_OBJECTID; - key.type = 0; - key.offset = offset; - ret = btrfs_insert_empty_item(trans, root, path, &key, - sizeof(struct btrfs_free_space_header)); - if (ret < 0) { - btrfs_release_path(path); - return ret; - } - - leaf = path->nodes[0]; - header = btrfs_item_ptr(leaf, path->slots[0], - struct btrfs_free_space_header); - memzero_extent_buffer(leaf, (unsigned long)header, sizeof(*header)); - btrfs_set_free_space_key(leaf, header, &disk_key); - btrfs_release_path(path); - - return 0; -} - -int create_free_space_inode(struct btrfs_trans_handle *trans, - struct btrfs_block_group *block_group, - struct btrfs_path *path) -{ - int ret; - u64 ino; - - ret = btrfs_get_free_objectid(trans->fs_info->tree_root, &ino); - if (ret < 0) - return ret; - - return __create_free_space_inode(trans->fs_info->tree_root, trans, path, - ino, block_group->start); -} - /* * inode is an optional sink: if it is NULL, btrfs_remove_free_space_inode * handles lookup, otherwise it takes ownership and iputs the inode. @@ -292,7 +195,6 @@ int btrfs_remove_free_space_inode(struct btrfs_trans_handle *trans, } int btrfs_truncate_free_space_cache(struct btrfs_trans_handle *trans, - struct btrfs_block_group *block_group, struct inode *vfs_inode) { struct btrfs_truncate_control control = { @@ -306,33 +208,6 @@ int btrfs_truncate_free_space_cache(struct btrfs_trans_handle *trans, struct btrfs_root *root = inode->root; struct extent_state *cached_state = NULL; int ret = 0; - bool locked = false; - - if (block_group) { - BTRFS_PATH_AUTO_FREE(path); - - path = btrfs_alloc_path(); - if (!path) { - ret = -ENOMEM; - goto fail; - } - locked = true; - mutex_lock(&trans->transaction->cache_write_mutex); - if (!list_empty(&block_group->io_list)) { - list_del_init(&block_group->io_list); - - btrfs_wait_cache_io(trans, block_group, path); - btrfs_put_block_group(block_group); - } - - /* - * now that we've truncated the cache away, its no longer - * setup or written - */ - spin_lock(&block_group->lock); - block_group->disk_cache_state = BTRFS_DC_CLEAR; - spin_unlock(&block_group->lock); - } btrfs_i_size_write(inode, 0); truncate_pagecache(vfs_inode, 0); @@ -356,336 +231,12 @@ int btrfs_truncate_free_space_cache(struct btrfs_trans_handle *trans, ret = btrfs_update_inode(trans, inode); fail: - if (locked) - mutex_unlock(&trans->transaction->cache_write_mutex); if (ret) btrfs_abort_transaction(trans, ret); return ret; } -static void readahead_cache(struct inode *inode) -{ - struct file_ra_state ra; - pgoff_t last_index; - - file_ra_state_init(&ra, inode->i_mapping); - last_index = (i_size_read(inode) - 1) >> PAGE_SHIFT; - - page_cache_sync_readahead(inode->i_mapping, &ra, NULL, 0, last_index); -} - -static int io_ctl_init(struct btrfs_io_ctl *io_ctl, struct inode *inode, - int write) -{ - int num_pages; - - num_pages = DIV_ROUND_UP(i_size_read(inode), PAGE_SIZE); - - /* Make sure we can fit our crcs and generation into the first page */ - if (write && (num_pages * sizeof(u32) + sizeof(u64)) > PAGE_SIZE) - return -ENOSPC; - - memset(io_ctl, 0, sizeof(struct btrfs_io_ctl)); - - io_ctl->pages = kzalloc_objs(struct page *, num_pages, GFP_NOFS); - if (!io_ctl->pages) - return -ENOMEM; - - io_ctl->num_pages = num_pages; - io_ctl->fs_info = inode_to_fs_info(inode); - io_ctl->inode = inode; - - return 0; -} -ALLOW_ERROR_INJECTION(io_ctl_init, ERRNO); - -static void io_ctl_free(struct btrfs_io_ctl *io_ctl) -{ - kfree(io_ctl->pages); - io_ctl->pages = NULL; -} - -static void io_ctl_unmap_page(struct btrfs_io_ctl *io_ctl) -{ - if (io_ctl->cur) { - io_ctl->cur = NULL; - io_ctl->orig = NULL; - } -} - -static void io_ctl_map_page(struct btrfs_io_ctl *io_ctl, int clear) -{ - ASSERT(io_ctl->index < io_ctl->num_pages); - io_ctl->page = io_ctl->pages[io_ctl->index++]; - io_ctl->cur = page_address(io_ctl->page); - io_ctl->orig = io_ctl->cur; - io_ctl->size = PAGE_SIZE; - if (clear) - clear_page(io_ctl->cur); -} - -static void io_ctl_drop_pages(struct btrfs_io_ctl *io_ctl) -{ - int i; - - io_ctl_unmap_page(io_ctl); - - for (i = 0; i < io_ctl->num_pages; i++) { - if (io_ctl->pages[i]) { - unlock_page(io_ctl->pages[i]); - put_page(io_ctl->pages[i]); - } - } -} - -static int io_ctl_prepare_pages(struct btrfs_io_ctl *io_ctl, bool uptodate) -{ - struct folio *folio; - struct inode *inode = io_ctl->inode; - gfp_t mask = btrfs_alloc_write_mask(inode->i_mapping); - int i; - - for (i = 0; i < io_ctl->num_pages; i++) { - int ret; - - folio = __filemap_get_folio(inode->i_mapping, i, - FGP_LOCK | FGP_ACCESSED | FGP_CREAT, - mask); - if (IS_ERR(folio)) { - io_ctl_drop_pages(io_ctl); - return PTR_ERR(folio); - } - - ret = set_folio_extent_mapped(folio); - if (ret < 0) { - folio_unlock(folio); - folio_put(folio); - io_ctl_drop_pages(io_ctl); - return ret; - } - - io_ctl->pages[i] = &folio->page; - if (uptodate && !folio_test_uptodate(folio)) { - btrfs_read_folio(NULL, folio); - folio_lock(folio); - if (folio->mapping != inode->i_mapping) { - btrfs_err(BTRFS_I(inode)->root->fs_info, - "free space cache page truncated"); - io_ctl_drop_pages(io_ctl); - return -EIO; - } - if (!folio_test_uptodate(folio)) { - btrfs_err(BTRFS_I(inode)->root->fs_info, - "error reading free space cache"); - io_ctl_drop_pages(io_ctl); - return -EIO; - } - } - } - - for (i = 0; i < io_ctl->num_pages; i++) - clear_page_dirty_for_io(io_ctl->pages[i]); - - return 0; -} - -static void io_ctl_set_generation(struct btrfs_io_ctl *io_ctl, u64 generation) -{ - io_ctl_map_page(io_ctl, 1); - - /* - * Skip the csum areas. If we don't check crcs then we just have a - * 64bit chunk at the front of the first page. - */ - io_ctl->cur += (sizeof(u32) * io_ctl->num_pages); - io_ctl->size -= sizeof(u64) + (sizeof(u32) * io_ctl->num_pages); - - put_unaligned_le64(generation, io_ctl->cur); - io_ctl->cur += sizeof(u64); -} - -static int io_ctl_check_generation(struct btrfs_io_ctl *io_ctl, u64 generation) -{ - u64 cache_gen; - - /* - * Skip the crc area. If we don't check crcs then we just have a 64bit - * chunk at the front of the first page. - */ - io_ctl->cur += sizeof(u32) * io_ctl->num_pages; - io_ctl->size -= sizeof(u64) + (sizeof(u32) * io_ctl->num_pages); - - cache_gen = get_unaligned_le64(io_ctl->cur); - if (cache_gen != generation) { - btrfs_err_rl(io_ctl->fs_info, - "space cache generation (%llu) does not match inode (%llu)", - cache_gen, generation); - io_ctl_unmap_page(io_ctl); - return -EIO; - } - io_ctl->cur += sizeof(u64); - return 0; -} - -static void io_ctl_set_crc(struct btrfs_io_ctl *io_ctl, int index) -{ - u32 *tmp; - u32 crc = ~(u32)0; - unsigned offset = 0; - - if (index == 0) - offset = sizeof(u32) * io_ctl->num_pages; - - crc = crc32c(crc, io_ctl->orig + offset, PAGE_SIZE - offset); - btrfs_crc32c_final(crc, (u8 *)&crc); - io_ctl_unmap_page(io_ctl); - tmp = page_address(io_ctl->pages[0]); - tmp += index; - *tmp = crc; -} - -static int io_ctl_check_crc(struct btrfs_io_ctl *io_ctl, int index) -{ - u32 *tmp, val; - u32 crc = ~(u32)0; - unsigned offset = 0; - - if (index >= io_ctl->num_pages) - return -EIO; - - if (index == 0) - offset = sizeof(u32) * io_ctl->num_pages; - - tmp = page_address(io_ctl->pages[0]); - tmp += index; - val = *tmp; - - io_ctl_map_page(io_ctl, 0); - crc = crc32c(crc, io_ctl->orig + offset, PAGE_SIZE - offset); - btrfs_crc32c_final(crc, (u8 *)&crc); - if (val != crc) { - btrfs_err_rl(io_ctl->fs_info, - "csum mismatch on free space cache"); - io_ctl_unmap_page(io_ctl); - return -EIO; - } - - return 0; -} - -static int io_ctl_add_entry(struct btrfs_io_ctl *io_ctl, u64 offset, u64 bytes, - void *bitmap) -{ - struct btrfs_free_space_entry *entry; - - if (!io_ctl->cur) - return -ENOSPC; - - entry = io_ctl->cur; - put_unaligned_le64(offset, &entry->offset); - put_unaligned_le64(bytes, &entry->bytes); - entry->type = (bitmap) ? BTRFS_FREE_SPACE_BITMAP : - BTRFS_FREE_SPACE_EXTENT; - io_ctl->cur += sizeof(struct btrfs_free_space_entry); - io_ctl->size -= sizeof(struct btrfs_free_space_entry); - - if (io_ctl->size >= sizeof(struct btrfs_free_space_entry)) - return 0; - - io_ctl_set_crc(io_ctl, io_ctl->index - 1); - - /* No more pages to map */ - if (io_ctl->index >= io_ctl->num_pages) - return 0; - - /* map the next page */ - io_ctl_map_page(io_ctl, 1); - return 0; -} - -static int io_ctl_add_bitmap(struct btrfs_io_ctl *io_ctl, void *bitmap) -{ - if (!io_ctl->cur) - return -ENOSPC; - - /* - * If we aren't at the start of the current page, unmap this one and - * map the next one if there is any left. - */ - if (io_ctl->cur != io_ctl->orig) { - io_ctl_set_crc(io_ctl, io_ctl->index - 1); - if (io_ctl->index >= io_ctl->num_pages) - return -ENOSPC; - io_ctl_map_page(io_ctl, 0); - } - - copy_page(io_ctl->cur, bitmap); - io_ctl_set_crc(io_ctl, io_ctl->index - 1); - if (io_ctl->index < io_ctl->num_pages) - io_ctl_map_page(io_ctl, 0); - return 0; -} - -static void io_ctl_zero_remaining_pages(struct btrfs_io_ctl *io_ctl) -{ - /* - * If we're not on the boundary we know we've modified the page and we - * need to crc the page. - */ - if (io_ctl->cur != io_ctl->orig) - io_ctl_set_crc(io_ctl, io_ctl->index - 1); - else - io_ctl_unmap_page(io_ctl); - - while (io_ctl->index < io_ctl->num_pages) { - io_ctl_map_page(io_ctl, 1); - io_ctl_set_crc(io_ctl, io_ctl->index - 1); - } -} - -static int io_ctl_read_entry(struct btrfs_io_ctl *io_ctl, - struct btrfs_free_space *entry, u8 *type) -{ - struct btrfs_free_space_entry *e; - int ret; - - if (!io_ctl->cur) { - ret = io_ctl_check_crc(io_ctl, io_ctl->index); - if (ret) - return ret; - } - - e = io_ctl->cur; - entry->offset = get_unaligned_le64(&e->offset); - entry->bytes = get_unaligned_le64(&e->bytes); - *type = e->type; - io_ctl->cur += sizeof(struct btrfs_free_space_entry); - io_ctl->size -= sizeof(struct btrfs_free_space_entry); - - if (io_ctl->size >= sizeof(struct btrfs_free_space_entry)) - return 0; - - io_ctl_unmap_page(io_ctl); - - return 0; -} - -static int io_ctl_read_bitmap(struct btrfs_io_ctl *io_ctl, - struct btrfs_free_space *entry) -{ - int ret; - - ret = io_ctl_check_crc(io_ctl, io_ctl->index); - if (ret) - return ret; - - copy_page(entry->bitmap, io_ctl->cur); - io_ctl_unmap_page(io_ctl); - - return 0; -} - static void recalculate_thresholds(struct btrfs_free_space_ctl *ctl) { struct btrfs_block_group *block_group = ctl->block_group; @@ -731,824 +282,6 @@ static void recalculate_thresholds(struct btrfs_free_space_ctl *ctl) div_u64(extent_bytes, sizeof(struct btrfs_free_space)); } -static int __load_free_space_cache(struct btrfs_root *root, struct inode *inode, - struct btrfs_free_space_ctl *ctl, - struct btrfs_path *path, u64 offset) -{ - struct btrfs_fs_info *fs_info = root->fs_info; - struct btrfs_free_space_header *header; - struct extent_buffer *leaf; - struct btrfs_io_ctl io_ctl; - struct btrfs_key key; - struct btrfs_free_space *e, *n; - LIST_HEAD(bitmaps); - u64 num_entries; - u64 num_bitmaps; - u64 generation; - u8 type; - int ret = 0; - - /* Nothing in the space cache, goodbye */ - if (!i_size_read(inode)) - return 0; - - key.objectid = BTRFS_FREE_SPACE_OBJECTID; - key.type = 0; - key.offset = offset; - - ret = btrfs_search_slot(NULL, root, &key, path, 0, 0); - if (ret < 0) - return 0; - else if (ret > 0) { - btrfs_release_path(path); - return 0; - } - - ret = -1; - - leaf = path->nodes[0]; - header = btrfs_item_ptr(leaf, path->slots[0], - struct btrfs_free_space_header); - num_entries = btrfs_free_space_entries(leaf, header); - num_bitmaps = btrfs_free_space_bitmaps(leaf, header); - generation = btrfs_free_space_generation(leaf, header); - btrfs_release_path(path); - - if (!BTRFS_I(inode)->generation) { - btrfs_info(fs_info, - "the free space cache file (%llu) is invalid, skip it", - offset); - return 0; - } - - if (BTRFS_I(inode)->generation != generation) { - btrfs_err(fs_info, - "free space inode generation (%llu) did not match free space cache generation (%llu)", - BTRFS_I(inode)->generation, generation); - return 0; - } - - if (!num_entries) - return 0; - - ret = io_ctl_init(&io_ctl, inode, 0); - if (ret) - return ret; - - readahead_cache(inode); - - ret = io_ctl_prepare_pages(&io_ctl, true); - if (ret) - goto out; - - ret = io_ctl_check_crc(&io_ctl, 0); - if (ret) - goto free_cache; - - ret = io_ctl_check_generation(&io_ctl, generation); - if (ret) - goto free_cache; - - while (num_entries) { - e = kmem_cache_zalloc(btrfs_free_space_cachep, - GFP_NOFS); - if (!e) { - ret = -ENOMEM; - goto free_cache; - } - - ret = io_ctl_read_entry(&io_ctl, e, &type); - if (ret) { - kmem_cache_free(btrfs_free_space_cachep, e); - goto free_cache; - } - - if (!e->bytes) { - ret = -1; - kmem_cache_free(btrfs_free_space_cachep, e); - goto free_cache; - } - - if (type == BTRFS_FREE_SPACE_EXTENT) { - spin_lock(&ctl->tree_lock); - ret = link_free_space(ctl, e); - spin_unlock(&ctl->tree_lock); - if (ret) { - btrfs_err(fs_info, - "Duplicate entries in free space cache, dumping"); - kmem_cache_free(btrfs_free_space_cachep, e); - goto free_cache; - } - } else { - ASSERT(num_bitmaps); - num_bitmaps--; - e->bitmap = kmem_cache_zalloc( - btrfs_free_space_bitmap_cachep, GFP_NOFS); - if (!e->bitmap) { - ret = -ENOMEM; - kmem_cache_free( - btrfs_free_space_cachep, e); - goto free_cache; - } - spin_lock(&ctl->tree_lock); - ret = link_free_space(ctl, e); - if (ret) { - spin_unlock(&ctl->tree_lock); - btrfs_err(fs_info, - "Duplicate entries in free space cache, dumping"); - kmem_cache_free(btrfs_free_space_bitmap_cachep, e->bitmap); - kmem_cache_free(btrfs_free_space_cachep, e); - goto free_cache; - } - ctl->total_bitmaps++; - recalculate_thresholds(ctl); - spin_unlock(&ctl->tree_lock); - list_add_tail(&e->list, &bitmaps); - } - - num_entries--; - } - - io_ctl_unmap_page(&io_ctl); - - /* - * We add the bitmaps at the end of the entries in order that - * the bitmap entries are added to the cache. - */ - list_for_each_entry_safe(e, n, &bitmaps, list) { - list_del_init(&e->list); - ret = io_ctl_read_bitmap(&io_ctl, e); - if (ret) - goto free_cache; - } - - io_ctl_drop_pages(&io_ctl); - ret = 1; -out: - io_ctl_free(&io_ctl); - return ret; -free_cache: - io_ctl_drop_pages(&io_ctl); - - spin_lock(&ctl->tree_lock); - __btrfs_remove_free_space_cache(ctl); - spin_unlock(&ctl->tree_lock); - goto out; -} - -static int copy_free_space_cache(struct btrfs_free_space_ctl *ctl) -{ - struct btrfs_free_space *info; - struct rb_node *n; - int ret = 0; - - while (!ret && (n = rb_first(&ctl->free_space_offset)) != NULL) { - info = rb_entry(n, struct btrfs_free_space, offset_index); - if (!info->bitmap) { - const u64 offset = info->offset; - const u64 bytes = info->bytes; - - unlink_free_space(ctl, info, true); - spin_unlock(&ctl->tree_lock); - kmem_cache_free(btrfs_free_space_cachep, info); - ret = btrfs_add_free_space(ctl->block_group, offset, bytes); - spin_lock(&ctl->tree_lock); - } else { - u64 offset = info->offset; - u64 bytes = ctl->block_group->fs_info->sectorsize; - - ret = search_bitmap(ctl, info, &offset, &bytes, false); - if (ret == 0) { - bitmap_clear_bits(ctl, info, offset, bytes, true); - spin_unlock(&ctl->tree_lock); - ret = btrfs_add_free_space(ctl->block_group, offset, - bytes); - spin_lock(&ctl->tree_lock); - } else { - free_bitmap(ctl, info); - ret = 0; - } - } - cond_resched_lock(&ctl->tree_lock); - } - return ret; -} - -static struct lock_class_key btrfs_free_space_inode_key; - -int load_free_space_cache(struct btrfs_block_group *block_group) -{ - struct btrfs_fs_info *fs_info = block_group->fs_info; - struct btrfs_free_space_ctl *ctl = block_group->free_space_ctl; - struct btrfs_free_space_ctl tmp_ctl = {}; - struct inode *inode; - struct btrfs_path *path; - int ret = 0; - bool matched; - u64 used = block_group->used; - - /* - * Because we could potentially discard our loaded free space, we want - * to load everything into a temporary structure first, and then if it's - * valid copy it all into the actual free space ctl. - */ - btrfs_init_free_space_ctl(block_group, &tmp_ctl); - - /* - * If this block group has been marked to be cleared for one reason or - * another then we can't trust the on disk cache, so just return. - */ - spin_lock(&block_group->lock); - if (block_group->disk_cache_state != BTRFS_DC_WRITTEN) { - spin_unlock(&block_group->lock); - return 0; - } - spin_unlock(&block_group->lock); - - path = btrfs_alloc_path(); - if (!path) - return 0; - path->search_commit_root = true; - path->skip_locking = true; - - /* - * We must pass a path with search_commit_root set to btrfs_iget in - * order to avoid a deadlock when allocating extents for the tree root. - * - * When we are COWing an extent buffer from the tree root, when looking - * for a free extent, at extent-tree.c:find_free_extent(), we can find - * block group without its free space cache loaded. When we find one - * we must load its space cache which requires reading its free space - * cache's inode item from the root tree. If this inode item is located - * in the same leaf that we started COWing before, then we end up in - * deadlock on the extent buffer (trying to read lock it when we - * previously write locked it). - * - * It's safe to read the inode item using the commit root because - * block groups, once loaded, stay in memory forever (until they are - * removed) as well as their space caches once loaded. New block groups - * once created get their ->cached field set to BTRFS_CACHE_FINISHED so - * we will never try to read their inode item while the fs is mounted. - */ - inode = lookup_free_space_inode(block_group, path); - if (IS_ERR(inode)) { - btrfs_free_path(path); - return 0; - } - - /* We may have converted the inode and made the cache invalid. */ - spin_lock(&block_group->lock); - if (block_group->disk_cache_state != BTRFS_DC_WRITTEN) { - spin_unlock(&block_group->lock); - btrfs_free_path(path); - goto out; - } - spin_unlock(&block_group->lock); - - /* - * Reinitialize the class of struct inode's mapping->invalidate_lock for - * free space inodes to prevent false positives related to locks for normal - * inodes. - */ - lockdep_set_class(&(&inode->i_data)->invalidate_lock, - &btrfs_free_space_inode_key); - - ret = __load_free_space_cache(fs_info->tree_root, inode, &tmp_ctl, - path, block_group->start); - btrfs_free_path(path); - if (ret <= 0) - goto out; - - matched = (tmp_ctl.free_space == (block_group->length - used - - block_group->bytes_super)); - - if (matched) { - spin_lock(&tmp_ctl.tree_lock); - ret = copy_free_space_cache(&tmp_ctl); - spin_unlock(&tmp_ctl.tree_lock); - /* - * ret == 1 means we successfully loaded the free space cache, - * so we need to re-set it here. - */ - if (ret == 0) - ret = 1; - } else { - /* - * We need to call the _locked variant so we don't try to update - * the discard counters. - */ - spin_lock(&tmp_ctl.tree_lock); - __btrfs_remove_free_space_cache(&tmp_ctl); - spin_unlock(&tmp_ctl.tree_lock); - btrfs_warn(fs_info, - "block group %llu has wrong amount of free space", - block_group->start); - ret = -1; - } -out: - if (ret < 0) { - /* This cache is bogus, make sure it gets cleared */ - spin_lock(&block_group->lock); - block_group->disk_cache_state = BTRFS_DC_CLEAR; - spin_unlock(&block_group->lock); - ret = 0; - - btrfs_warn(fs_info, - "failed to load free space cache for block group %llu, rebuilding it now", - block_group->start); - } - - spin_lock(&ctl->tree_lock); - btrfs_discard_update_discardable(block_group); - spin_unlock(&ctl->tree_lock); - iput(inode); - return ret; -} - -static noinline_for_stack -int write_cache_extent_entries(struct btrfs_io_ctl *io_ctl, - struct btrfs_block_group *block_group, - int *entries, int *bitmaps, - struct list_head *bitmap_list) -{ - int ret; - struct btrfs_free_space_ctl *ctl = block_group->free_space_ctl; - struct btrfs_free_cluster *cluster = NULL; - struct btrfs_free_cluster *cluster_locked = NULL; - struct rb_node *node = rb_first(&ctl->free_space_offset); - struct btrfs_trim_range *trim_entry; - - /* Get the cluster for this block_group if it exists */ - if (!list_empty(&block_group->cluster_list)) { - cluster = list_first_entry(&block_group->cluster_list, - struct btrfs_free_cluster, block_group_list); - } - - if (!node && cluster) { - cluster_locked = cluster; - spin_lock(&cluster_locked->lock); - node = rb_first(&cluster->root); - cluster = NULL; - } - - /* Write out the extent entries */ - while (node) { - struct btrfs_free_space *e; - - e = rb_entry(node, struct btrfs_free_space, offset_index); - *entries += 1; - - ret = io_ctl_add_entry(io_ctl, e->offset, e->bytes, - e->bitmap); - if (ret) - goto fail; - - if (e->bitmap) { - list_add_tail(&e->list, bitmap_list); - *bitmaps += 1; - } - node = rb_next(node); - if (!node && cluster) { - node = rb_first(&cluster->root); - cluster_locked = cluster; - spin_lock(&cluster_locked->lock); - cluster = NULL; - } - } - if (cluster_locked) { - spin_unlock(&cluster_locked->lock); - cluster_locked = NULL; - } - - /* - * Make sure we don't miss any range that was removed from our rbtree - * because trimming is running. Otherwise after a umount+mount (or crash - * after committing the transaction) we would leak free space and get - * an inconsistent free space cache report from fsck. - */ - list_for_each_entry(trim_entry, &ctl->trimming_ranges, list) { - ret = io_ctl_add_entry(io_ctl, trim_entry->start, - trim_entry->bytes, NULL); - if (ret) - goto fail; - *entries += 1; - } - - return 0; -fail: - if (cluster_locked) - spin_unlock(&cluster_locked->lock); - return -ENOSPC; -} - -static noinline_for_stack int -update_cache_item(struct btrfs_trans_handle *trans, - struct btrfs_root *root, - struct inode *inode, - struct btrfs_path *path, u64 offset, - int entries, int bitmaps) -{ - struct btrfs_key key; - struct btrfs_free_space_header *header; - struct extent_buffer *leaf; - int ret; - - key.objectid = BTRFS_FREE_SPACE_OBJECTID; - key.type = 0; - key.offset = offset; - - ret = btrfs_search_slot(trans, root, &key, path, 0, 1); - if (ret < 0) { - btrfs_clear_extent_bit(&BTRFS_I(inode)->io_tree, 0, inode->i_size - 1, - EXTENT_DELALLOC, NULL); - return ret; - } - leaf = path->nodes[0]; - if (ret > 0) { - struct btrfs_key found_key; - ASSERT(path->slots[0]); - path->slots[0]--; - btrfs_item_key_to_cpu(leaf, &found_key, path->slots[0]); - if (found_key.objectid != BTRFS_FREE_SPACE_OBJECTID || - found_key.offset != offset) { - btrfs_clear_extent_bit(&BTRFS_I(inode)->io_tree, 0, - inode->i_size - 1, EXTENT_DELALLOC, - NULL); - btrfs_release_path(path); - return -ENOENT; - } - } - - BTRFS_I(inode)->generation = trans->transid; - header = btrfs_item_ptr(leaf, path->slots[0], - struct btrfs_free_space_header); - btrfs_set_free_space_entries(leaf, header, entries); - btrfs_set_free_space_bitmaps(leaf, header, bitmaps); - btrfs_set_free_space_generation(leaf, header, trans->transid); - btrfs_release_path(path); - - return 0; -} - -static noinline_for_stack int write_pinned_extent_entries( - struct btrfs_trans_handle *trans, - struct btrfs_block_group *block_group, - struct btrfs_io_ctl *io_ctl, - int *entries) -{ - u64 start, extent_start, extent_end, len; - const u64 block_group_end = btrfs_block_group_end(block_group); - struct extent_io_tree *unpin = NULL; - int ret; - - /* - * We want to add any pinned extents to our free space cache - * so we don't leak the space - * - * We shouldn't have switched the pinned extents yet so this is the - * right one - */ - unpin = &trans->transaction->pinned_extents; - - start = block_group->start; - - while (start < block_group_end) { - if (!btrfs_find_first_extent_bit(unpin, start, - &extent_start, &extent_end, - EXTENT_DIRTY, NULL)) - return 0; - - /* This pinned extent is out of our range */ - if (extent_start >= block_group_end) - return 0; - - extent_start = max(extent_start, start); - extent_end = min(block_group_end, extent_end + 1); - len = extent_end - extent_start; - - *entries += 1; - ret = io_ctl_add_entry(io_ctl, extent_start, len, NULL); - if (ret) - return -ENOSPC; - - start = extent_end; - } - - return 0; -} - -static noinline_for_stack int -write_bitmap_entries(struct btrfs_io_ctl *io_ctl, struct list_head *bitmap_list) -{ - struct btrfs_free_space *entry, *next; - int ret; - - /* Write out the bitmaps */ - list_for_each_entry_safe(entry, next, bitmap_list, list) { - ret = io_ctl_add_bitmap(io_ctl, entry->bitmap); - if (ret) - return -ENOSPC; - list_del_init(&entry->list); - } - - return 0; -} - -static int flush_dirty_cache(struct inode *inode) -{ - int ret; - - ret = btrfs_wait_ordered_range(BTRFS_I(inode), 0, (u64)-1); - if (ret) - btrfs_clear_extent_bit(&BTRFS_I(inode)->io_tree, 0, inode->i_size - 1, - EXTENT_DELALLOC, NULL); - - return ret; -} - -static void noinline_for_stack -cleanup_bitmap_list(struct list_head *bitmap_list) -{ - struct btrfs_free_space *entry, *next; - - list_for_each_entry_safe(entry, next, bitmap_list, list) - list_del_init(&entry->list); -} - -static void noinline_for_stack -cleanup_write_cache_enospc(struct inode *inode, - struct btrfs_io_ctl *io_ctl, - struct extent_state **cached_state) -{ - io_ctl_drop_pages(io_ctl); - btrfs_unlock_extent(&BTRFS_I(inode)->io_tree, 0, i_size_read(inode) - 1, - cached_state); -} - -static int __btrfs_wait_cache_io(struct btrfs_root *root, - struct btrfs_trans_handle *trans, - struct btrfs_block_group *block_group, - struct btrfs_io_ctl *io_ctl, - struct btrfs_path *path, u64 offset) -{ - int ret; - struct inode *inode = io_ctl->inode; - - if (!inode) - return 0; - - /* Flush the dirty pages in the cache file. */ - ret = flush_dirty_cache(inode); - if (ret) - goto out; - - /* Update the cache item to tell everyone this cache file is valid. */ - ret = update_cache_item(trans, root, inode, path, offset, - io_ctl->entries, io_ctl->bitmaps); -out: - if (ret) { - invalidate_inode_pages2(inode->i_mapping); - BTRFS_I(inode)->generation = 0; - if (block_group) - btrfs_debug(root->fs_info, - "failed to write free space cache for block group %llu error %d", - block_group->start, ret); - } - btrfs_update_inode(trans, BTRFS_I(inode)); - - if (block_group) { - /* the dirty list is protected by the dirty_bgs_lock */ - spin_lock(&trans->transaction->dirty_bgs_lock); - - /* the disk_cache_state is protected by the block group lock */ - spin_lock(&block_group->lock); - - /* - * only mark this as written if we didn't get put back on - * the dirty list while waiting for IO. Otherwise our - * cache state won't be right, and we won't get written again - */ - if (!ret && list_empty(&block_group->dirty_list)) - block_group->disk_cache_state = BTRFS_DC_WRITTEN; - else if (ret) - block_group->disk_cache_state = BTRFS_DC_ERROR; - - spin_unlock(&block_group->lock); - spin_unlock(&trans->transaction->dirty_bgs_lock); - io_ctl->inode = NULL; - iput(inode); - } - - return ret; - -} - -int btrfs_wait_cache_io(struct btrfs_trans_handle *trans, - struct btrfs_block_group *block_group, - struct btrfs_path *path) -{ - return __btrfs_wait_cache_io(block_group->fs_info->tree_root, trans, - block_group, &block_group->io_ctl, - path, block_group->start); -} - -/* - * Write out cached info to an inode. - * - * @inode: freespace inode we are writing out - * @ctl: free space cache we are going to write out - * @block_group: block_group for this cache if it belongs to a block_group - * @io_ctl: holds context for the io - * @trans: the trans handle - * - * This function writes out a free space cache struct to disk for quick recovery - * on mount. This will return 0 if it was successful in writing the cache out, - * or an errno if it was not. - */ -static int __btrfs_write_out_cache(struct inode *inode, - struct btrfs_block_group *block_group, - struct btrfs_trans_handle *trans) -{ - struct btrfs_free_space_ctl *ctl = block_group->free_space_ctl; - struct btrfs_io_ctl *io_ctl = &block_group->io_ctl; - struct extent_state *cached_state = NULL; - LIST_HEAD(bitmap_list); - int entries = 0; - int bitmaps = 0; - int ret; - bool must_iput = false; - int i_size; - - if (!i_size_read(inode)) - return -EIO; - - WARN_ON(io_ctl->pages); - ret = io_ctl_init(io_ctl, inode, 1); - if (ret) - return ret; - - if (block_group->flags & BTRFS_BLOCK_GROUP_DATA) { - down_write(&block_group->data_rwsem); - spin_lock(&block_group->lock); - if (block_group->delalloc_bytes) { - block_group->disk_cache_state = BTRFS_DC_WRITTEN; - spin_unlock(&block_group->lock); - up_write(&block_group->data_rwsem); - BTRFS_I(inode)->generation = 0; - ret = 0; - must_iput = true; - goto out; - } - spin_unlock(&block_group->lock); - } - - /* Lock all pages first so we can lock the extent safely. */ - ret = io_ctl_prepare_pages(io_ctl, false); - if (ret) - goto out_unlock; - - btrfs_lock_extent(&BTRFS_I(inode)->io_tree, 0, i_size_read(inode) - 1, - &cached_state); - - io_ctl_set_generation(io_ctl, trans->transid); - - mutex_lock(&ctl->cache_writeout_mutex); - /* Write out the extent entries in the free space cache */ - spin_lock(&ctl->tree_lock); - ret = write_cache_extent_entries(io_ctl, block_group, &entries, &bitmaps, - &bitmap_list); - if (ret) - goto out_nospc_locked; - - /* - * Some spaces that are freed in the current transaction are pinned, - * they will be added into free space cache after the transaction is - * committed, we shouldn't lose them. - * - * If this changes while we are working we'll get added back to - * the dirty list and redo it. No locking needed - */ - ret = write_pinned_extent_entries(trans, block_group, io_ctl, &entries); - if (ret) - goto out_nospc_locked; - - /* - * At last, we write out all the bitmaps and keep cache_writeout_mutex - * locked while doing it because a concurrent trim can be manipulating - * or freeing the bitmap. - */ - ret = write_bitmap_entries(io_ctl, &bitmap_list); - spin_unlock(&ctl->tree_lock); - mutex_unlock(&ctl->cache_writeout_mutex); - if (ret) - goto out_nospc; - - /* Zero out the rest of the pages just to make sure */ - io_ctl_zero_remaining_pages(io_ctl); - - /* Everything is written out, now we dirty the pages in the file. */ - i_size = i_size_read(inode); - for (int i = 0; i < round_up(i_size, PAGE_SIZE) / PAGE_SIZE; i++) { - u64 dirty_start = i * PAGE_SIZE; - u64 dirty_len = min_t(u64, dirty_start + PAGE_SIZE, i_size) - dirty_start; - - ret = btrfs_dirty_folio(BTRFS_I(inode), page_folio(io_ctl->pages[i]), - dirty_start, dirty_len, &cached_state, false); - if (ret < 0) - goto out_nospc; - } - - if (block_group->flags & BTRFS_BLOCK_GROUP_DATA) - up_write(&block_group->data_rwsem); - /* - * Release the pages and unlock the extent, we will flush - * them out later - */ - io_ctl_drop_pages(io_ctl); - io_ctl_free(io_ctl); - - btrfs_unlock_extent(&BTRFS_I(inode)->io_tree, 0, i_size_read(inode) - 1, - &cached_state); - - /* - * at this point the pages are under IO and we're happy, - * The caller is responsible for waiting on them and updating - * the cache and the inode - */ - io_ctl->entries = entries; - io_ctl->bitmaps = bitmaps; - - ret = btrfs_fdatawrite_range(BTRFS_I(inode), 0, (u64)-1); - if (ret) - goto out; - - return 0; - -out_nospc_locked: - cleanup_bitmap_list(&bitmap_list); - spin_unlock(&ctl->tree_lock); - mutex_unlock(&ctl->cache_writeout_mutex); - -out_nospc: - cleanup_write_cache_enospc(inode, io_ctl, &cached_state); - -out_unlock: - if (block_group->flags & BTRFS_BLOCK_GROUP_DATA) - up_write(&block_group->data_rwsem); - -out: - io_ctl->inode = NULL; - io_ctl_free(io_ctl); - if (ret) { - invalidate_inode_pages2(inode->i_mapping); - BTRFS_I(inode)->generation = 0; - } - btrfs_update_inode(trans, BTRFS_I(inode)); - if (must_iput) - iput(inode); - return ret; -} - -int btrfs_write_out_cache(struct btrfs_trans_handle *trans, - struct btrfs_block_group *block_group, - struct btrfs_path *path) -{ - struct btrfs_fs_info *fs_info = trans->fs_info; - struct inode *inode; - int ret = 0; - - spin_lock(&block_group->lock); - if (block_group->disk_cache_state < BTRFS_DC_SETUP) { - spin_unlock(&block_group->lock); - return 0; - } - spin_unlock(&block_group->lock); - - inode = lookup_free_space_inode(block_group, path); - if (IS_ERR(inode)) - return 0; - - ret = __btrfs_write_out_cache(inode, block_group, trans); - if (ret) { - btrfs_debug(fs_info, - "failed to write free space cache for block group %llu error %d", - block_group->start, ret); - spin_lock(&block_group->lock); - block_group->disk_cache_state = BTRFS_DC_ERROR; - spin_unlock(&block_group->lock); - - block_group->io_ctl.inode = NULL; - iput(inode); - } - - /* - * if ret == 0 the caller is expected to call btrfs_wait_cache_io - * to wait for IO and put the inode - */ - - return ret; -} - static inline unsigned long offset_to_bit(u64 bitmap_start, u32 unit, u64 offset) { @@ -2953,8 +1686,6 @@ void btrfs_init_free_space_ctl(struct btrfs_block_group *block_group, spin_lock_init(&ctl->tree_lock); ctl->block_group = block_group; ctl->free_space_bytes = RB_ROOT_CACHED; - INIT_LIST_HEAD(&ctl->trimming_ranges); - mutex_init(&ctl->cache_writeout_mutex); /* * we only want to have 32k of ram per block group for keeping @@ -3650,12 +2381,10 @@ void btrfs_init_free_cluster(struct btrfs_free_cluster *cluster) static int do_trimming(struct btrfs_block_group *block_group, u64 *total_trimmed, u64 start, u64 bytes, u64 reserved_start, u64 reserved_bytes, - enum btrfs_trim_state reserved_trim_state, - struct btrfs_trim_range *trim_entry) + enum btrfs_trim_state reserved_trim_state) { struct btrfs_space_info *space_info = block_group->space_info; struct btrfs_fs_info *fs_info = block_group->fs_info; - struct btrfs_free_space_ctl *ctl = block_group->free_space_ctl; int ret; bool bg_ro; const u64 end = start + bytes; @@ -3681,7 +2410,6 @@ static int do_trimming(struct btrfs_block_group *block_group, trim_state = BTRFS_TRIM_STATE_TRIMMED; } - mutex_lock(&ctl->cache_writeout_mutex); if (reserved_start < start) __btrfs_add_free_space(block_group, reserved_start, start - reserved_start, @@ -3690,8 +2418,6 @@ static int do_trimming(struct btrfs_block_group *block_group, __btrfs_add_free_space(block_group, end, reserved_end - end, reserved_trim_state); __btrfs_add_free_space(block_group, start, bytes, trim_state); - list_del(&trim_entry->list); - mutex_unlock(&ctl->cache_writeout_mutex); if (!bg_ro) { spin_lock(&space_info->lock); @@ -3729,9 +2455,6 @@ static int trim_no_bitmap(struct btrfs_block_group *block_group, const u64 max_discard_size = READ_ONCE(discard_ctl->max_discard_size); while (start < end) { - struct btrfs_trim_range trim_entry; - - mutex_lock(&ctl->cache_writeout_mutex); spin_lock(&ctl->tree_lock); if (ctl->free_space < minlen) @@ -3762,7 +2485,6 @@ static int trim_no_bitmap(struct btrfs_block_group *block_group, bytes = entry->bytes; if (bytes < minlen) { spin_unlock(&ctl->tree_lock); - mutex_unlock(&ctl->cache_writeout_mutex); goto next; } unlink_free_space(ctl, entry, true); @@ -3787,7 +2509,6 @@ static int trim_no_bitmap(struct btrfs_block_group *block_group, bytes = min(extent_start + extent_bytes, end) - start; if (bytes < minlen) { spin_unlock(&ctl->tree_lock); - mutex_unlock(&ctl->cache_writeout_mutex); goto next; } @@ -3796,14 +2517,9 @@ static int trim_no_bitmap(struct btrfs_block_group *block_group, } spin_unlock(&ctl->tree_lock); - trim_entry.start = extent_start; - trim_entry.bytes = extent_bytes; - list_add_tail(&trim_entry.list, &ctl->trimming_ranges); - mutex_unlock(&ctl->cache_writeout_mutex); ret = do_trimming(block_group, total_trimmed, start, bytes, - extent_start, extent_bytes, extent_trim_state, - &trim_entry); + extent_start, extent_bytes, extent_trim_state); if (ret) { block_group->discard_cursor = start + bytes; break; @@ -3827,7 +2543,6 @@ next: out_unlock: block_group->discard_cursor = btrfs_block_group_end(block_group); spin_unlock(&ctl->tree_lock); - mutex_unlock(&ctl->cache_writeout_mutex); return ret; } @@ -3938,16 +2653,13 @@ static int trim_bitmaps(struct btrfs_block_group *block_group, while (offset < end) { bool next_bitmap = false; - struct btrfs_trim_range trim_entry; - mutex_lock(&ctl->cache_writeout_mutex); spin_lock(&ctl->tree_lock); if (ctl->free_space < minlen) { block_group->discard_cursor = btrfs_block_group_end(block_group); spin_unlock(&ctl->tree_lock); - mutex_unlock(&ctl->cache_writeout_mutex); break; } @@ -3963,7 +2675,6 @@ static int trim_bitmaps(struct btrfs_block_group *block_group, if (!entry || (async && minlen && start == offset && btrfs_free_space_trimmed(entry))) { spin_unlock(&ctl->tree_lock); - mutex_unlock(&ctl->cache_writeout_mutex); next_bitmap = true; goto next; } @@ -3989,7 +2700,6 @@ static int trim_bitmaps(struct btrfs_block_group *block_group, else entry->trim_state = BTRFS_TRIM_STATE_UNTRIMMED; spin_unlock(&ctl->tree_lock); - mutex_unlock(&ctl->cache_writeout_mutex); next_bitmap = true; goto next; } @@ -4000,14 +2710,12 @@ static int trim_bitmaps(struct btrfs_block_group *block_group, */ if (async && *total_trimmed) { spin_unlock(&ctl->tree_lock); - mutex_unlock(&ctl->cache_writeout_mutex); return ret; } bytes = min(bytes, end - start); if (bytes < minlen || (async && maxlen && bytes > maxlen)) { spin_unlock(&ctl->tree_lock); - mutex_unlock(&ctl->cache_writeout_mutex); goto next; } @@ -4027,13 +2735,9 @@ static int trim_bitmaps(struct btrfs_block_group *block_group, free_bitmap(ctl, entry); spin_unlock(&ctl->tree_lock); - trim_entry.start = start; - trim_entry.bytes = bytes; - list_add_tail(&trim_entry.list, &ctl->trimming_ranges); - mutex_unlock(&ctl->cache_writeout_mutex); ret = do_trimming(block_group, total_trimmed, start, bytes, - start, bytes, 0, &trim_entry); + start, bytes, 0); if (ret) { reset_trimming_bitmap(ctl, offset); block_group->discard_cursor = @@ -4152,47 +2856,29 @@ bool btrfs_free_space_cache_v1_active(struct btrfs_fs_info *fs_info) return btrfs_super_cache_generation(fs_info->super_copy); } -static int cleanup_free_space_cache_v1(struct btrfs_fs_info *fs_info, - struct btrfs_trans_handle *trans) +int btrfs_cleanup_free_space_cache_v1(struct btrfs_fs_info *fs_info) { - struct btrfs_block_group *block_group; + struct btrfs_trans_handle *trans; struct rb_node *node; + int ret; btrfs_info(fs_info, "cleaning free space cache v1"); - node = rb_first_cached(&fs_info->block_group_cache_tree); - while (node) { - int ret; - - block_group = rb_entry(node, struct btrfs_block_group, cache_node); - ret = btrfs_remove_free_space_inode(trans, NULL, block_group); - if (ret) - return ret; - node = rb_next(node); - } - return 0; -} - -int btrfs_set_free_space_cache_v1_active(struct btrfs_fs_info *fs_info, bool active) -{ - struct btrfs_trans_handle *trans; - int ret; - /* - * update_super_roots will appropriately set or unset - * super_copy->cache_generation based on SPACE_CACHE and - * BTRFS_FS_CLEANUP_SPACE_CACHE_V1. For this reason, we need a - * transaction commit whether we are enabling space cache v1 and don't - * have any other work to do, or are disabling it and removing free - * space inodes. + * update_super_roots() zeroes super_copy->cache_generation while + * BTRFS_FS_CLEANUP_SPACE_CACHE_V1 is set, so this needs a commit. */ trans = btrfs_start_transaction(fs_info->tree_root, 0); if (IS_ERR(trans)) return PTR_ERR(trans); - if (!active) { - set_bit(BTRFS_FS_CLEANUP_SPACE_CACHE_V1, &fs_info->flags); - ret = cleanup_free_space_cache_v1(fs_info, trans); + set_bit(BTRFS_FS_CLEANUP_SPACE_CACHE_V1, &fs_info->flags); + for (node = rb_first_cached(&fs_info->block_group_cache_tree); node; + node = rb_next(node)) { + struct btrfs_block_group *block_group; + + block_group = rb_entry(node, struct btrfs_block_group, cache_node); + ret = btrfs_remove_free_space_inode(trans, NULL, block_group); if (unlikely(ret)) { btrfs_abort_transaction(trans, ret); btrfs_end_transaction(trans); diff --git a/fs/btrfs/free-space-cache.h b/fs/btrfs/free-space-cache.h index 53fe8e293af1..e22443598b8e 100644 --- a/fs/btrfs/free-space-cache.h +++ b/fs/btrfs/free-space-cache.h @@ -14,7 +14,6 @@ #include "fs.h" struct inode; -struct page; struct btrfs_fs_info; struct btrfs_path; struct btrfs_trans_handle; @@ -84,44 +83,18 @@ struct btrfs_free_space_ctl { s32 discardable_extents[BTRFS_STAT_NR_ENTRIES]; s64 discardable_bytes[BTRFS_STAT_NR_ENTRIES]; struct btrfs_block_group *block_group; - struct mutex cache_writeout_mutex; - struct list_head trimming_ranges; -}; - -struct btrfs_io_ctl { - void *cur, *orig; - struct page *page; - struct page **pages; - struct btrfs_fs_info *fs_info; - struct inode *inode; - unsigned long size; - int index; - int num_pages; - int entries; - int bitmaps; }; int __init btrfs_free_space_init(void); void __cold btrfs_free_space_exit(void); struct inode *lookup_free_space_inode(struct btrfs_block_group *block_group, struct btrfs_path *path); -int create_free_space_inode(struct btrfs_trans_handle *trans, - struct btrfs_block_group *block_group, - struct btrfs_path *path); int btrfs_remove_free_space_inode(struct btrfs_trans_handle *trans, struct inode *inode, struct btrfs_block_group *block_group); int btrfs_truncate_free_space_cache(struct btrfs_trans_handle *trans, - struct btrfs_block_group *block_group, struct inode *inode); -int load_free_space_cache(struct btrfs_block_group *block_group); -int btrfs_wait_cache_io(struct btrfs_trans_handle *trans, - struct btrfs_block_group *block_group, - struct btrfs_path *path); -int btrfs_write_out_cache(struct btrfs_trans_handle *trans, - struct btrfs_block_group *block_group, - struct btrfs_path *path); void btrfs_init_free_space_ctl(struct btrfs_block_group *block_group, struct btrfs_free_space_ctl *ctl); @@ -161,7 +134,7 @@ int btrfs_trim_block_group_bitmaps(struct btrfs_block_group *block_group, void btrfs_trim_fully_remapped_block_group(struct btrfs_block_group *bg); bool btrfs_free_space_cache_v1_active(struct btrfs_fs_info *fs_info); -int btrfs_set_free_space_cache_v1_active(struct btrfs_fs_info *fs_info, bool active); +int btrfs_cleanup_free_space_cache_v1(struct btrfs_fs_info *fs_info); /* Support functions for running our sanity tests */ #ifdef CONFIG_BTRFS_FS_RUN_SANITY_TESTS bool btrfs_use_bitmap(struct btrfs_free_space_ctl *ctl, diff --git a/fs/btrfs/fs.c b/fs/btrfs/fs.c index de160d29dde8..75a1217727a7 100644 --- a/fs/btrfs/fs.c +++ b/fs/btrfs/fs.c @@ -79,7 +79,7 @@ void btrfs_csum_init(struct btrfs_csum_ctx *ctx, u16 csum_type) blake2b_init(&ctx->blake2b, 32); break; default: - /* Checksume type is validated at mount time. */ + /* Checksum type is validated at mount time. */ BUG(); } } diff --git a/fs/btrfs/fs.h b/fs/btrfs/fs.h index 10e15a319b93..79d0828c51c7 100644 --- a/fs/btrfs/fs.h +++ b/fs/btrfs/fs.h @@ -259,7 +259,6 @@ enum { BTRFS_MOUNT_NOSSD = (1ULL << 9), BTRFS_MOUNT_DISCARD_SYNC = (1ULL << 10), BTRFS_MOUNT_FORCE_COMPRESS = (1ULL << 11), - BTRFS_MOUNT_SPACE_CACHE = (1ULL << 12), BTRFS_MOUNT_CLEAR_CACHE = (1ULL << 13), BTRFS_MOUNT_USER_SUBVOL_RM_ALLOWED = (1ULL << 14), BTRFS_MOUNT_ENOSPC_DEBUG = (1ULL << 15), @@ -712,7 +711,6 @@ struct btrfs_fs_info { struct workqueue_struct *endio_meta_workers; struct workqueue_struct *rmw_workers; struct btrfs_workqueue *endio_write_workers; - struct btrfs_workqueue *endio_freespace_worker; struct btrfs_workqueue *caching_workers; struct workqueue_struct *fixup_workers; @@ -811,7 +809,7 @@ struct btrfs_fs_info { struct btrfs_discard_ctl discard_ctl; /* Is qgroup tracking in a consistent state? */ - u64 qgroup_flags; + unsigned long qgroup_flags; /* Holds configuration and tracking. Protected by qgroup_lock. */ struct rb_root qgroup_tree; diff --git a/fs/btrfs/inode.c b/fs/btrfs/inode.c index 558b4a3f9633..1d79f5263de3 100644 --- a/fs/btrfs/inode.c +++ b/fs/btrfs/inode.c @@ -730,6 +730,9 @@ static inline int inode_need_compress(struct btrfs_inode *inode, u64 start, u64 end, bool check_inline) { struct btrfs_fs_info *fs_info = inode->root->fs_info; + const u32 blocksize = fs_info->sectorsize; + + ASSERT(IS_ALIGNED(start, blocksize) && IS_ALIGNED(end + 1, blocksize)); if (unlikely(!btrfs_inode_can_compress(inode))) { DEBUG_WARN("BTRFS: unexpected compression for ino %llu", btrfs_ino(inode)); @@ -1366,11 +1369,6 @@ static noinline int cow_file_range(struct btrfs_inode *inode, goto out_unlock; } - if (btrfs_is_free_space_inode(inode)) { - ret = -EINVAL; - goto out_unlock; - } - num_bytes = ALIGN(end - start + 1, blocksize); num_bytes = max(blocksize, num_bytes); ASSERT(num_bytes <= btrfs_super_total_bytes(fs_info->super_copy)); @@ -1680,7 +1678,6 @@ static int fallback_to_cow(struct btrfs_inode *inode, struct folio *locked_folio, const u64 start, const u64 end) { - const bool is_space_ino = btrfs_is_free_space_inode(inode); const bool is_reloc_ino = btrfs_is_data_reloc_root(inode->root); const u64 range_bytes = end + 1 - start; struct extent_io_tree *io_tree = &inode->io_tree; @@ -1713,23 +1710,22 @@ static int fallback_to_cow(struct btrfs_inode *inode, * extent_clear_unlock_delalloc()) the bytes_may_use counter of the * data space info, which we incremented in the step above. * - * If we need to fallback to cow and the inode corresponds to a free - * space cache inode or an inode of the data relocation tree, we must - * also increment bytes_may_use of the data space_info for the same - * reason. Space caches and relocated data extents always get a prealloc - * extent for them, however scrub or balance may have set the block - * group that contains that extent to RO mode and therefore force COW - * when starting writeback. + * If we need to fallback to cow and the inode is in the data relocation + * tree, we must also increment bytes_may_use of the data space_info for + * the same reason. Relocated data extents always get a prealloc extent, + * however scrub or balance may have set the block group that contains + * that extent to RO mode and therefore force COW when starting + * writeback. */ btrfs_lock_extent(io_tree, start, end, &cached_state); count = btrfs_count_range_bits(io_tree, &range_start, end, range_bytes, EXTENT_NORESERVE, false, NULL); - if (count > 0 || is_space_ino || is_reloc_ino) { + if (count > 0 || is_reloc_ino) { u64 bytes = count; struct btrfs_fs_info *fs_info = inode->root->fs_info; struct btrfs_space_info *sinfo = fs_info->data_sinfo; - if (is_space_ino || is_reloc_ino) + if (is_reloc_ino) bytes = range_bytes; spin_lock(&sinfo->lock); @@ -1794,7 +1790,6 @@ static int can_nocow_file_extent(struct btrfs_path *path, struct btrfs_inode *inode, struct can_nocow_file_extent_args *args) { - const bool is_freespace_inode = btrfs_is_free_space_inode(inode); struct extent_buffer *leaf = path->nodes[0]; struct btrfs_root *root = inode->root; struct btrfs_file_extent_item *fi; @@ -1807,8 +1802,7 @@ static int can_nocow_file_extent(struct btrfs_path *path, bool nowait = path->nowait; /* If there are pending snapshots for this root, we must do COW. */ - if (args->writeback_path && !is_freespace_inode && - atomic_read(&root->snapshot_force_cow)) + if (args->writeback_path && atomic_read(&root->snapshot_force_cow)) goto out; fi = btrfs_item_ptr(leaf, path->slots[0], struct btrfs_file_extent_item); @@ -1857,7 +1851,6 @@ static int can_nocow_file_extent(struct btrfs_path *path, ret = btrfs_cross_ref_exist(inode, key->offset - args->file_extent.offset, args->file_extent.disk_bytenr, path); - WARN_ON_ONCE(ret > 0 && is_freespace_inode); if (ret != 0) goto out; @@ -1892,7 +1885,6 @@ static int can_nocow_file_extent(struct btrfs_path *path, ret = btrfs_lookup_csums_list(csum_root, io_start, io_start + args->file_extent.num_bytes - 1, NULL, nowait); - WARN_ON_ONCE(ret > 0 && is_freespace_inode); if (ret != 0) goto out; @@ -2331,7 +2323,7 @@ static int run_delalloc_inline(struct btrfs_inode *inode, struct folio *locked_f btrfs_check_folio_write_protected(locked_folio); if (btrfs_inode_can_compress(inode) && - inode_need_compress(inode, 0, blocksize, true)) { + inode_need_compress(inode, 0, blocksize - 1, true)) { if (inode->defrag_compress > 0 && inode->defrag_compress < BTRFS_NR_COMPRESS_TYPES) { compress_type = inode->defrag_compress; @@ -2640,7 +2632,7 @@ void btrfs_set_delalloc_extent(struct btrfs_inode *inode, struct extent_state *s * and are therefore protected against concurrent calls of this * function and btrfs_clear_delalloc_extent(). */ - if (!btrfs_is_free_space_inode(inode) && prev_delalloc_bytes == 0) + if (prev_delalloc_bytes == 0) btrfs_add_delalloc_inode(inode); } @@ -2698,7 +2690,6 @@ void btrfs_clear_delalloc_extent(struct btrfs_inode *inode, return; if (!btrfs_is_data_reloc_root(root) && - !btrfs_is_free_space_inode(inode) && !(state->state & EXTENT_NORESERVE) && (bits & EXTENT_CLEAR_DATA_RESV)) btrfs_free_reserved_data_space_noquota(inode, len); @@ -2716,7 +2707,7 @@ void btrfs_clear_delalloc_extent(struct btrfs_inode *inode, * and are therefore protected against concurrent calls of this * function and btrfs_set_delalloc_extent(). */ - if (!btrfs_is_free_space_inode(inode) && new_delalloc_bytes == 0) { + if (new_delalloc_bytes == 0) { spin_lock(&root->delalloc_lock); btrfs_del_delalloc_inode(inode); spin_unlock(&root->delalloc_lock); @@ -3218,7 +3209,7 @@ int btrfs_finish_one_ordered(struct btrfs_ordered_extent *ordered_extent) int compress_type = 0; int ret = 0; u64 logical_len = ordered_extent->num_bytes; - bool freespace_inode; + u64 unwritten_start; bool truncated = false; bool clear_reserved_extent = true; unsigned int clear_bits = 0; @@ -3235,9 +3226,7 @@ int btrfs_finish_one_ordered(struct btrfs_ordered_extent *ordered_extent) if (!test_bit(BTRFS_ORDERED_NOCOW, &ordered_extent->flags)) clear_bits |= EXTENT_DEFRAG; - freespace_inode = btrfs_is_free_space_inode(inode); - if (!freespace_inode) - btrfs_lockdep_acquire(fs_info, btrfs_ordered_extent); + btrfs_lockdep_acquire(fs_info, btrfs_ordered_extent); if (unlikely(test_bit(BTRFS_ORDERED_IOERR, &ordered_extent->flags))) { ret = -EIO; @@ -3272,10 +3261,7 @@ int btrfs_finish_one_ordered(struct btrfs_ordered_extent *ordered_extent) &cached_state); } - if (freespace_inode) - trans = btrfs_join_transaction_spacecache(root); - else - trans = btrfs_join_transaction(root); + trans = btrfs_join_transaction(root); if (IS_ERR(trans)) { ret = PTR_ERR(trans); trans = NULL; @@ -3385,29 +3371,11 @@ out: if (ret) btrfs_mark_ordered_extent_error(ordered_extent); - /* - * Drop extent maps for the part of the extent we didn't write. - * - * We have an exception here for the free_space_inode, this is - * because when we do btrfs_get_extent() on the free space inode - * we will search the commit root. If this is a new block group - * we won't find anything, and we will trip over the assert in - * writepage where we do ASSERT(em->block_start != - * EXTENT_MAP_HOLE). - * - * Theoretically we could also skip this for any NOCOW extent as - * we don't mess with the extent map tree in the NOCOW case, but - * for now simply skip this if we are the free space inode. - */ - if (!btrfs_is_free_space_inode(inode)) { - u64 unwritten_start = start; - - if (truncated) - unwritten_start += logical_len; - - btrfs_drop_extent_map_range(inode, unwritten_start, - end, false); - } + /* Drop extent maps for the part of the extent we didn't write. */ + unwritten_start = start; + if (truncated) + unwritten_start += logical_len; + btrfs_drop_extent_map_range(inode, unwritten_start, end, false); /* * If the ordered extent had an IOERR or something else went @@ -3471,79 +3439,30 @@ int btrfs_finish_ordered_io(struct btrfs_ordered_extent *ordered) return btrfs_finish_one_ordered(ordered); } -/* - * Calculate the checksum of an fs block at physical memory address @paddr, - * and save the result to @dest. - * - * The folio containing @paddr must be large enough to contain a full fs block. - */ -void btrfs_calculate_block_csum_folio(struct btrfs_fs_info *fs_info, - const phys_addr_t paddr, u8 *dest) +/* Generate data checksum for a single fs block, pointed to by @orig_iter. */ +void btrfs_csum_one_bio_block(struct btrfs_fs_info *fs_info, struct bio *bio, + const struct bvec_iter *orig_iter, u8 *csum) { - struct folio *folio = page_folio(phys_to_page(paddr)); + struct btrfs_csum_ctx cctx; + struct bvec_iter iter = *orig_iter; const u32 blocksize = fs_info->sectorsize; - const u32 step = min(blocksize, PAGE_SIZE); - const u32 nr_steps = blocksize / step; - phys_addr_t paddrs[BTRFS_MAX_BLOCKSIZE / PAGE_SIZE]; - - /* The full block must be inside the folio. */ - ASSERT(offset_in_folio(folio, paddr) + blocksize <= folio_size(folio)); + u32 cur = 0; - for (int i = 0; i < nr_steps; i++) { - u32 pindex = offset_in_folio(folio, paddr + i * step) >> PAGE_SHIFT; - - /* - * For bs <= ps cases, we will only run the loop once, so the offset - * inside the page will only added to paddrs[0]. - * - * For bs > ps cases, the block must be page aligned, thus offset - * inside the page will always be 0. - */ - paddrs[i] = page_to_phys(folio_page(folio, pindex)) + offset_in_page(paddr); - } - return btrfs_calculate_block_csum_pages(fs_info, paddrs, dest); -} - -/* - * Calculate the checksum of a fs block backed by multiple noncontiguous pages - * at @paddrs[] and save the result to @dest. - * - * The folio containing @paddr must be large enough to contain a full fs block. - */ -void btrfs_calculate_block_csum_pages(struct btrfs_fs_info *fs_info, - const phys_addr_t paddrs[], u8 *dest) -{ - const u32 blocksize = fs_info->sectorsize; - const u32 step = min(blocksize, PAGE_SIZE); - const u32 nr_steps = blocksize / step; - struct btrfs_csum_ctx csum; - - btrfs_csum_init(&csum, fs_info->csum_type); - for (int i = 0; i < nr_steps; i++) { - const phys_addr_t paddr = paddrs[i]; + btrfs_csum_init(&cctx, fs_info->csum_type); + while (cur < blocksize) { + struct page *page = bio_iter_page(bio, iter); + const u32 pg_off = bio_iter_offset(bio, iter); + const u32 cur_len = min(bio_iter_len(bio, iter), blocksize - cur); void *kaddr; - ASSERT(offset_in_page(paddr) + step <= PAGE_SIZE); - kaddr = kmap_local_page(phys_to_page(paddr)) + offset_in_page(paddr); - btrfs_csum_update(&csum, kaddr, step); + kaddr = kmap_local_page(page) + pg_off; + btrfs_csum_update(&cctx, kaddr, cur_len); kunmap_local(kaddr); - } - btrfs_csum_final(&csum, dest); -} -/* - * Verify the checksum for a single sector without any extra action that depend - * on the type of I/O. - * - * @kaddr must be a properly kmapped address. - */ -int btrfs_check_block_csum(struct btrfs_fs_info *fs_info, phys_addr_t paddr, u8 *csum, - const u8 * const csum_expected) -{ - btrfs_calculate_block_csum_folio(fs_info, paddr, csum); - if (unlikely(memcmp(csum, csum_expected, fs_info->csum_size) != 0)) - return -EIO; - return 0; + bio_advance_iter_single(bio, &iter, cur_len); + cur += cur_len; + } + btrfs_csum_final(&cctx, csum); } /* @@ -3551,27 +3470,30 @@ int btrfs_check_block_csum(struct btrfs_fs_info *fs_info, phys_addr_t paddr, u8 * different noncontiguous pages. * * @bbio: btrfs_io_bio which contains the csum - * @dev: device the sector is on - * @bio_offset: offset to the beginning of the bio (in bytes) - * @paddrs: physical addresses which back the fs block + * @orig_iter: bvec iter pointing to the start of the block + * @dev: device the sector is on (optional) * * Check if the checksum on a data block is valid. When a checksum mismatch is * detected, report the error and fill the corrupted range with zero. * * Return %true if the sector is ok or had no checksum to start with, else %false. */ -bool btrfs_data_csum_ok(struct btrfs_bio *bbio, struct btrfs_device *dev, - u32 bio_offset, const phys_addr_t paddrs[]) +bool btrfs_bio_data_csum_ok(struct btrfs_bio *bbio, + const struct bvec_iter *orig_iter, + struct btrfs_device *dev) { struct btrfs_inode *inode = bbio->inode; struct btrfs_fs_info *fs_info = inode->root->fs_info; + struct bvec_iter iter = *orig_iter; const u32 blocksize = fs_info->sectorsize; - const u32 step = min(blocksize, PAGE_SIZE); - const u32 nr_steps = blocksize / step; + const u32 bio_offset = (iter.bi_sector - bbio->saved_iter.bi_sector) << SECTOR_SHIFT; u64 file_offset = bbio->file_offset + bio_offset; u64 end = file_offset + blocksize - 1; u8 *csum_expected; u8 csum[BTRFS_CSUM_SIZE]; + u32 cur = 0; + + ASSERT(iter.bi_sector >= bbio->saved_iter.bi_sector); if (!bbio->csum) return true; @@ -3587,7 +3509,7 @@ bool btrfs_data_csum_ok(struct btrfs_bio *bbio, struct btrfs_device *dev, csum_expected = bbio->csum + (bio_offset >> fs_info->sectorsize_bits) * fs_info->csum_size; - btrfs_calculate_block_csum_pages(fs_info, paddrs, csum); + btrfs_csum_one_bio_block(fs_info, &bbio->bio, orig_iter, csum); if (unlikely(memcmp(csum, csum_expected, fs_info->csum_size) != 0)) goto zeroit; return true; @@ -3597,8 +3519,16 @@ zeroit: bbio->mirror_num); if (dev) btrfs_dev_stat_inc_and_print(dev, BTRFS_DEV_STAT_CORRUPTION_ERRS); - for (int i = 0; i < nr_steps; i++) - memzero_page(phys_to_page(paddrs[i]), offset_in_page(paddrs[i]), step); + while (cur < blocksize) { + struct page *page = bio_iter_page(&bbio->bio, iter); + const u32 pg_off = bio_iter_offset(&bbio->bio, iter); + const u32 cur_len = min(bio_iter_len(&bbio->bio, iter), blocksize - cur); + + memzero_page(page, pg_off, cur_len); + + bio_advance_iter_single(&bbio->bio, &iter, cur_len); + cur += cur_len; + } return false; } @@ -5498,7 +5428,7 @@ static int btrfs_setsize(struct inode *inode, struct iattr *attr) return ret; } -static int btrfs_setattr(struct mnt_idmap *idmap, struct dentry *dentry, +static int btrfs_setattr(const struct mnt_idmap *idmap, struct dentry *dentry, struct iattr *attr) { struct inode *inode = d_inode(dentry); @@ -6881,8 +6811,28 @@ int btrfs_create_new_inode(struct btrfs_trans_handle *trans, } } else { ret = btrfs_add_link(trans, BTRFS_I(dir), BTRFS_I(inode), name, - false, BTRFS_I(inode)->dir_index); - if (unlikely(ret)) { + false, BTRFS_I(inode)->dir_index, NULL); + if (ret == -ENOMEM) { + /* + * Orphan the new inode instead of aborting. The inode + * item was already written with nlink 1, and discard's + * eviction won't delete a bad inode, so nlink 0 must be + * persisted here or orphan cleanup would see nlink > 0, + * drop the orphan item, and leak the inode. + */ + clear_nlink(inode); + /* btrfs_orphan_add() aborts the transaction on failure. */ + ret = btrfs_orphan_add(trans, BTRFS_I(inode)); + if (ret) + goto discard; + ret = btrfs_update_inode(trans, BTRFS_I(inode)); + if (ret) { + btrfs_abort_transaction(trans, ret); + goto discard; + } + ret = -ENOMEM; + goto discard; + } else if (unlikely(ret)) { btrfs_abort_transaction(trans, ret); goto discard; } @@ -6913,7 +6863,8 @@ out: */ int btrfs_add_link(struct btrfs_trans_handle *trans, struct btrfs_inode *parent_inode, struct btrfs_inode *inode, - const struct fscrypt_str *name, bool add_backref, u64 index) + const struct fscrypt_str *name, bool add_backref, u64 index, + struct btrfs_dir_index_prealloc *prealloc) { int ret = 0; struct btrfs_key key; @@ -6939,12 +6890,14 @@ int btrfs_add_link(struct btrfs_trans_handle *trans, } /* Nothing to clean up yet */ - if (ret) + if (ret) { + btrfs_free_delayed_dir_index_prealloc(trans, prealloc); return ret; + } ret = btrfs_insert_dir_item(trans, name, parent_inode, &key, - btrfs_inode_type(inode), index); - if (ret == -EEXIST || ret == -EOVERFLOW) + btrfs_inode_type(inode), index, prealloc); + if (ret == -EEXIST || ret == -EOVERFLOW || ret == -ENOMEM) goto fail_dir_item; else if (unlikely(ret)) { btrfs_abort_transaction(trans, ret); @@ -7023,7 +6976,7 @@ out_inode: return ret; } -static int btrfs_mknod(struct mnt_idmap *idmap, struct inode *dir, +static int btrfs_mknod(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *dentry, umode_t mode, dev_t rdev) { struct inode *inode; @@ -7037,7 +6990,7 @@ static int btrfs_mknod(struct mnt_idmap *idmap, struct inode *dir, return btrfs_create_common(dir, dentry, inode); } -static int btrfs_create(struct mnt_idmap *idmap, struct inode *dir, +static int btrfs_create(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *dentry, umode_t mode) { struct inode *inode; @@ -7097,7 +7050,7 @@ static int btrfs_link(struct dentry *old_dentry, struct inode *dir, inode_set_ctime_current(inode); ret = btrfs_add_link(trans, BTRFS_I(dir), BTRFS_I(inode), - &fname.disk_name, true, index); + &fname.disk_name, true, index, NULL); if (ret) goto fail; @@ -7134,7 +7087,7 @@ fail: return ret; } -static struct dentry *btrfs_mkdir(struct mnt_idmap *idmap, struct inode *dir, +static struct dentry *btrfs_mkdir(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *dentry, umode_t mode) { struct inode *inode; @@ -7164,7 +7117,7 @@ static noinline int uncompress_inline(struct btrfs_path *path, compress_type = btrfs_file_extent_compression(leaf, item); max_size = btrfs_file_extent_ram_bytes(leaf, item); inline_size = btrfs_file_extent_inline_item_len(leaf, path->slots[0]); - tmp = kmalloc(inline_size, GFP_NOFS); + tmp = kvmalloc(inline_size, GFP_NOFS); if (!tmp) return -ENOMEM; ptr = btrfs_file_extent_inline_start(item); @@ -7185,7 +7138,7 @@ static noinline int uncompress_inline(struct btrfs_path *path, if (max_size < blocksize) folio_zero_range(folio, max_size, blocksize - max_size); - kfree(tmp); + kvfree(tmp); return ret; } @@ -7281,16 +7234,6 @@ struct extent_map *btrfs_get_extent(struct btrfs_inode *inode, /* Chances are we'll be called again, so go ahead and do readahead */ path->reada = READA_FORWARD; - /* - * The same explanation in load_free_space_cache applies here as well, - * we only read when we're loading the free space cache, and at that - * point the commit_root has everything we need. - */ - if (btrfs_is_free_space_inode(inode)) { - path->search_commit_root = true; - path->skip_locking = true; - } - ret = btrfs_lookup_file_extent(NULL, root, path, objectid, start, 0); if (ret < 0) { goto out; @@ -8041,7 +7984,7 @@ out: return ret; } -struct inode *btrfs_new_subvol_inode(struct mnt_idmap *idmap, +struct inode *btrfs_new_subvol_inode(const struct mnt_idmap *idmap, struct inode *dir) { struct inode *inode; @@ -8147,7 +8090,6 @@ void btrfs_destroy_inode(struct inode *vfs_inode) struct btrfs_ordered_extent *ordered; struct btrfs_inode *inode = BTRFS_I(vfs_inode); struct btrfs_root *root = inode->root; - bool freespace_inode; WARN_ON(!hlist_empty(&vfs_inode->i_dentry)); WARN_ON(vfs_inode->i_data.nrpages); @@ -8170,12 +8112,6 @@ void btrfs_destroy_inode(struct inode *vfs_inode) if (!root) return; - /* - * If this is a free space inode do not take the ordered extents lockdep - * map. - */ - freespace_inode = btrfs_is_free_space_inode(inode); - while (1) { ordered = btrfs_lookup_first_ordered_extent(inode, (u64)-1); if (!ordered) @@ -8185,8 +8121,7 @@ void btrfs_destroy_inode(struct inode *vfs_inode) "found ordered extent %llu %llu on inode cleanup", ordered->file_offset, ordered->num_bytes); - if (!freespace_inode) - btrfs_lockdep_acquire(root->fs_info, btrfs_ordered_extent); + btrfs_lockdep_acquire(root->fs_info, btrfs_ordered_extent); btrfs_remove_ordered_extent(ordered); btrfs_put_ordered_extent(ordered); @@ -8243,7 +8178,7 @@ int __init btrfs_init_cachep(void) return 0; } -static int btrfs_getattr(struct mnt_idmap *idmap, +static int btrfs_getattr(const struct mnt_idmap *idmap, const struct path *path, struct kstat *stat, u32 request_mask, unsigned int flags) { @@ -8511,14 +8446,14 @@ static int btrfs_rename_exchange(struct inode *old_dir, } ret = btrfs_add_link(trans, BTRFS_I(new_dir), BTRFS_I(old_inode), - new_name, false, old_idx); + new_name, false, old_idx, NULL); if (unlikely(ret)) { btrfs_abort_transaction(trans, ret); goto out_fail; } ret = btrfs_add_link(trans, BTRFS_I(old_dir), BTRFS_I(new_inode), - old_name, false, new_idx); + old_name, false, new_idx, NULL); if (unlikely(ret)) { btrfs_abort_transaction(trans, ret); goto out_fail; @@ -8559,7 +8494,7 @@ out_notrans: return ret; } -static struct inode *new_whiteout_inode(struct mnt_idmap *idmap, +static struct inode *new_whiteout_inode(const struct mnt_idmap *idmap, struct inode *dir) { struct inode *inode; @@ -8574,7 +8509,7 @@ static struct inode *new_whiteout_inode(struct mnt_idmap *idmap, return inode; } -static int btrfs_rename(struct mnt_idmap *idmap, +static int btrfs_rename(const struct mnt_idmap *idmap, struct inode *old_dir, struct dentry *old_dentry, struct inode *new_dir, struct dentry *new_dentry, unsigned int flags) @@ -8591,6 +8526,7 @@ static int btrfs_rename(struct mnt_idmap *idmap, struct inode *new_inode = d_inode(new_dentry); struct inode *old_inode = d_inode(old_dentry); struct btrfs_rename_ctx rename_ctx; + struct btrfs_dir_index_prealloc *prealloc = NULL; u64 index = 0; int ret; int ret2; @@ -8714,6 +8650,23 @@ static int btrfs_rename(struct mnt_idmap *idmap, if (ret) goto out_fail; + /* + * When not overwriting an existing entry, pre-allocate the delayed dir + * index now so that ENOMEM is returned before any btree modifications. + * For the overwrite case, too many btree changes have already happened + * by the time btrfs_add_link() is called. + */ + if (!new_inode) { + prealloc = btrfs_prealloc_delayed_dir_index(BTRFS_I(new_dir), + new_fname.disk_name.name, + new_fname.disk_name.len); + if (IS_ERR(prealloc)) { + ret = PTR_ERR(prealloc); + prealloc = NULL; + goto out_fail; + } + } + BTRFS_I(old_inode)->dir_index = 0ULL; if (unlikely(old_ino == BTRFS_FIRST_FREE_OBJECTID)) { /* force full log commit if subvolume involved. */ @@ -8809,7 +8762,8 @@ static int btrfs_rename(struct mnt_idmap *idmap, } ret = btrfs_add_link(trans, BTRFS_I(new_dir), BTRFS_I(old_inode), - &new_fname.disk_name, false, index); + &new_fname.disk_name, false, index, prealloc); + prealloc = NULL; if (unlikely(ret)) { btrfs_abort_transaction(trans, ret); goto out_fail; @@ -8834,6 +8788,7 @@ static int btrfs_rename(struct mnt_idmap *idmap, } } out_fail: + btrfs_free_delayed_dir_index_prealloc(trans, prealloc); if (logs_pinned) { btrfs_end_log_trans(root); btrfs_end_log_trans(dest); @@ -8854,7 +8809,7 @@ out_fscrypt_names: return ret; } -static int btrfs_rename2(struct mnt_idmap *idmap, struct inode *old_dir, +static int btrfs_rename2(const struct mnt_idmap *idmap, struct inode *old_dir, struct dentry *old_dentry, struct inode *new_dir, struct dentry *new_dentry, unsigned int flags) { @@ -9042,7 +8997,7 @@ out: return ret; } -static int btrfs_symlink(struct mnt_idmap *idmap, struct inode *dir, +static int btrfs_symlink(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *dentry, const char *symname) { struct btrfs_fs_info *fs_info = inode_to_fs_info(dir); @@ -9148,14 +9103,13 @@ out_inode: } static struct btrfs_trans_handle *insert_prealloc_file_extent( - struct btrfs_trans_handle *trans_in, struct btrfs_inode *inode, struct btrfs_key *ins, u64 file_offset) { struct btrfs_file_extent_item stack_fi; struct btrfs_replace_extent_info extent_info; - struct btrfs_trans_handle *trans = trans_in; + struct btrfs_trans_handle *trans; struct btrfs_path *path; u64 start = ins->objectid; u64 len = ins->offset; @@ -9176,15 +9130,6 @@ static struct btrfs_trans_handle *insert_prealloc_file_extent( if (ret < 0) return ERR_PTR(ret); - if (trans) { - ret = insert_reserved_file_extent(trans, inode, - file_offset, &stack_fi, - true, qgroup_released); - if (ret) - goto free_qgroup; - return trans; - } - extent_info.disk_offset = start; extent_info.disk_len = len; extent_info.data_offset = 0; @@ -9224,12 +9169,12 @@ free_qgroup: return ERR_PTR(ret); } -static int __btrfs_prealloc_file_range(struct inode *inode, int mode, - u64 start, u64 num_bytes, u64 min_size, - loff_t actual_len, u64 *alloc_hint, - struct btrfs_trans_handle *trans) +int btrfs_prealloc_file_range(struct inode *inode, int mode, + u64 start, u64 num_bytes, u64 min_size, + loff_t actual_len, u64 *alloc_hint) { struct btrfs_fs_info *fs_info = inode_to_fs_info(inode); + struct btrfs_trans_handle *trans; struct extent_map *em; struct btrfs_root *root = BTRFS_I(inode)->root; struct btrfs_key ins; @@ -9239,11 +9184,8 @@ static int __btrfs_prealloc_file_range(struct inode *inode, int mode, u64 cur_bytes; u64 last_alloc = (u64)-1; int ret = 0; - bool own_trans = true; u64 end = start + num_bytes - 1; - if (trans) - own_trans = false; while (num_bytes > 0) { cur_bytes = min_t(u64, num_bytes, SZ_256M); cur_bytes = max(cur_bytes, min_size); @@ -9269,8 +9211,8 @@ static int __btrfs_prealloc_file_range(struct inode *inode, int mode, clear_offset += ins.offset; last_alloc = ins.offset; - trans = insert_prealloc_file_extent(trans, BTRFS_I(inode), - &ins, cur_offset); + trans = insert_prealloc_file_extent(BTRFS_I(inode), &ins, + cur_offset); /* * Now that we inserted the prealloc extent we can finally * decrement the number of reservations in the block group. @@ -9342,8 +9284,7 @@ next: range_start, range_end - range_start); if (ret) { btrfs_abort_transaction(trans, ret); - if (own_trans) - btrfs_end_transaction(trans); + btrfs_end_transaction(trans); break; } @@ -9355,15 +9296,11 @@ next: if (unlikely(ret)) { btrfs_abort_transaction(trans, ret); - if (own_trans) - btrfs_end_transaction(trans); + btrfs_end_transaction(trans); break; } - if (own_trans) { - btrfs_end_transaction(trans); - trans = NULL; - } + btrfs_end_transaction(trans); } if (clear_offset < end) btrfs_free_reserved_data_space(BTRFS_I(inode), NULL, clear_offset, @@ -9371,30 +9308,12 @@ next: return ret; } -int btrfs_prealloc_file_range(struct inode *inode, int mode, - u64 start, u64 num_bytes, u64 min_size, - loff_t actual_len, u64 *alloc_hint) -{ - return __btrfs_prealloc_file_range(inode, mode, start, num_bytes, - min_size, actual_len, alloc_hint, - NULL); -} - -int btrfs_prealloc_file_range_trans(struct inode *inode, - struct btrfs_trans_handle *trans, int mode, - u64 start, u64 num_bytes, u64 min_size, - loff_t actual_len, u64 *alloc_hint) -{ - return __btrfs_prealloc_file_range(inode, mode, start, num_bytes, - min_size, actual_len, alloc_hint, trans); -} - /* * NOTE: in case you are adding MAY_EXEC check for directories: * we are marking them with IOP_FASTPERM_MAY_EXEC, allowing path lookup to * elide calls here. */ -static int btrfs_permission(struct mnt_idmap *idmap, +static int btrfs_permission(const struct mnt_idmap *idmap, struct inode *inode, int mask) { struct btrfs_root *root = BTRFS_I(inode)->root; @@ -9410,7 +9329,7 @@ static int btrfs_permission(struct mnt_idmap *idmap, return generic_permission(idmap, inode, mask); } -static int btrfs_tmpfile(struct mnt_idmap *idmap, struct inode *dir, +static int btrfs_tmpfile(const struct mnt_idmap *idmap, struct inode *dir, struct file *file, umode_t mode) { struct btrfs_fs_info *fs_info = inode_to_fs_info(dir); @@ -9623,7 +9542,6 @@ int btrfs_encoded_read_regular_fill_pages(struct btrfs_inode *inode, struct completion sync_reads; unsigned long i = 0; struct btrfs_bio *bbio; - int ret; /* * Fast path for synchronous reads which completes in this call, io_uring @@ -9670,10 +9588,10 @@ int btrfs_encoded_read_regular_fill_pages(struct btrfs_inode *inode, if (uring_ctx) { if (refcount_dec_and_test(&priv->pending_refs)) { - ret = blk_status_to_errno(READ_ONCE(priv->status)); - btrfs_uring_read_extent_endio(uring_ctx, ret); + int err = blk_status_to_errno(READ_ONCE(priv->status)); + + btrfs_uring_read_extent_endio(uring_ctx, err); kfree(priv); - return ret; } return -EIOCBQUEUED; diff --git a/fs/btrfs/ioctl.c b/fs/btrfs/ioctl.c index e4b2da31a0d5..f3e2afe221be 100644 --- a/fs/btrfs/ioctl.c +++ b/fs/btrfs/ioctl.c @@ -278,7 +278,7 @@ int btrfs_fileattr_get(struct dentry *dentry, struct file_kattr *fa) return 0; } -int btrfs_fileattr_set(struct mnt_idmap *idmap, +int btrfs_fileattr_set(const struct mnt_idmap *idmap, struct dentry *dentry, struct file_kattr *fa) { struct btrfs_inode *inode = BTRFS_I(d_inode(dentry)); @@ -549,7 +549,7 @@ static unsigned int create_subvol_num_items(const struct btrfs_qgroup_inherit *i return num_items; } -static noinline int create_subvol(struct mnt_idmap *idmap, +static noinline int create_subvol(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *dentry, struct btrfs_qgroup_inherit *inherit) { @@ -879,7 +879,7 @@ free_pending: * inside this filesystem so it's quite a bit simpler. */ static noinline int btrfs_mksubvol(struct dentry *parent, - struct mnt_idmap *idmap, + const struct mnt_idmap *idmap, struct qstr *qname, struct btrfs_root *snap_src, bool readonly, struct btrfs_qgroup_inherit *inherit) @@ -926,7 +926,7 @@ out_dput: } static noinline int btrfs_mksnapshot(struct dentry *parent, - struct mnt_idmap *idmap, + const struct mnt_idmap *idmap, struct qstr *qname, struct btrfs_root *root, bool readonly, @@ -1164,7 +1164,7 @@ static noinline int __btrfs_ioctl_snap_create(struct file *file, { int ret; struct qstr qname = QSTR(name); - struct mnt_idmap *idmap = file_mnt_idmap(file); + const struct mnt_idmap *idmap = file_mnt_idmap(file); if (!S_ISDIR(file_inode(file)->i_mode)) return -ENOTDIR; @@ -1741,7 +1741,7 @@ static noinline int btrfs_search_path_in_tree(struct btrfs_root *root, u64 dirid return 0; } -static int btrfs_search_path_in_tree_user(struct mnt_idmap *idmap, +static int btrfs_search_path_in_tree_user(const struct mnt_idmap *idmap, struct inode *inode, struct btrfs_ioctl_ino_lookup_user_args *args) { @@ -2241,7 +2241,7 @@ static noinline int btrfs_ioctl_snap_destroy(struct file *file, struct btrfs_root *dest = NULL; struct btrfs_ioctl_vol_args AUTO_KFREE(vol_args); struct btrfs_ioctl_vol_args_v2 AUTO_KFREE(vol_args2); - struct mnt_idmap *idmap = file_mnt_idmap(file); + const struct mnt_idmap *idmap = file_mnt_idmap(file); char *subvol_name, *subvol_name_ptr = NULL; int ret = 0; bool destroy_parent = false; @@ -3881,7 +3881,7 @@ static long btrfs_ioctl_quota_rescan_status(struct btrfs_fs_info *fs_info, if (!capable(CAP_SYS_ADMIN)) return -EPERM; - if (fs_info->qgroup_flags & BTRFS_QGROUP_STATUS_FLAG_RESCAN) { + if (test_bit(BTRFS_QGROUP_STATUS_BIT_RESCAN, &fs_info->qgroup_flags)) { qsa.flags = 1; qsa.progress = fs_info->qgroup_rescan_progress.objectid; } @@ -3901,7 +3901,7 @@ static long btrfs_ioctl_quota_rescan_wait(struct btrfs_fs_info *fs_info) } static long _btrfs_ioctl_set_received_subvol(struct file *file, - struct mnt_idmap *idmap, + const struct mnt_idmap *idmap, struct btrfs_ioctl_received_subvol_args *sa) { struct inode *inode = file_inode(file); @@ -4601,7 +4601,7 @@ static void btrfs_uring_read_finished(struct io_tw_req tw_req, io_tw_token_t tw) size_t page_offset; ssize_t ret; - /* The inode lock has already been acquired in btrfs_uring_read_extent. */ + /* The inode lock has already been acquired in btrfs_encoded_read(). */ btrfs_lockdep_inode_acquire(inode, i_rwsem); if (priv->err) { @@ -4667,7 +4667,6 @@ static int btrfs_uring_read_extent(struct kiocb *iocb, struct iov_iter *iter, struct iovec *iov, struct io_uring_cmd *cmd) { struct btrfs_inode *inode = BTRFS_I(file_inode(iocb->ki_filp)); - struct extent_io_tree *io_tree = &inode->io_tree; struct page **pages = NULL; struct btrfs_uring_priv *priv = NULL; unsigned long nr_pages; @@ -4723,8 +4722,6 @@ static int btrfs_uring_read_extent(struct kiocb *iocb, struct iov_iter *iter, return -EIOCBQUEUED; out_fail: - btrfs_unlock_extent(io_tree, start, lockend, &cached_state); - btrfs_inode_unlock(inode, BTRFS_ILOCK_SHARED); kfree(priv); for (int i = 0; i < nr_pages; i++) { if (pages[i]) @@ -4752,9 +4749,6 @@ static int btrfs_uring_encoded_read(struct io_uring_cmd *cmd, unsigned int issue struct io_btrfs_cmd *bc = io_uring_cmd_to_pdu(cmd, struct io_btrfs_cmd); struct btrfs_uring_encoded_data *data = NULL; - if (cmd->flags & IORING_URING_CMD_REISSUE) - data = bc->data; - if (!capable(CAP_SYS_ADMIN)) { ret = -EPERM; goto out_acct; @@ -4836,8 +4830,6 @@ static int btrfs_uring_encoded_read(struct io_uring_cmd *cmd, unsigned int issue ret = btrfs_encoded_read(&kiocb, &data->iter, &data->args, &cached_state, &disk_bytenr, &disk_io_size); - if (ret == -EAGAIN) - goto out_acct; if (ret < 0 && ret != -EIOCBQUEUED) goto out_free; @@ -4865,8 +4857,10 @@ static int btrfs_uring_encoded_read(struct io_uring_cmd *cmd, unsigned int issue cached_state, disk_bytenr, disk_io_size, count, data->args.compression, data->iov, cmd); - - goto out_acct; + if (ret == -EIOCBQUEUED) + goto out_acct; + btrfs_unlock_extent(io_tree, start, lockend, &cached_state); + btrfs_inode_unlock(inode, BTRFS_ILOCK_SHARED); } out_free: @@ -4877,8 +4871,10 @@ out_acct: add_rchar(current, ret); inc_syscr(current); - if (ret != -EIOCBQUEUED && ret != -EAGAIN) + if (ret != -EIOCBQUEUED) { kfree(data); + bc->data = NULL; + } return ret; } @@ -4890,12 +4886,8 @@ static int btrfs_uring_encoded_write(struct io_uring_cmd *cmd, unsigned int issu struct kiocb kiocb; ssize_t ret; void __user *sqe_addr; - struct io_btrfs_cmd *bc = io_uring_cmd_to_pdu(cmd, struct io_btrfs_cmd); struct btrfs_uring_encoded_data *data = NULL; - if (cmd->flags & IORING_URING_CMD_REISSUE) - data = bc->data; - if (!capable(CAP_SYS_ADMIN)) { ret = -EPERM; goto out_acct; @@ -4907,6 +4899,11 @@ static int btrfs_uring_encoded_write(struct io_uring_cmd *cmd, unsigned int issu goto out_acct; } + if (issue_flags & IO_URING_F_NONBLOCK) { + ret = -EAGAIN; + goto out_acct; + } + if (!data) { data = kzalloc_obj(*data, GFP_NOFS); if (!data) { @@ -4914,8 +4911,6 @@ static int btrfs_uring_encoded_write(struct io_uring_cmd *cmd, unsigned int issu goto out_acct; } - bc->data = data; - if (issue_flags & IO_URING_F_COMPAT) { #if defined(CONFIG_64BIT) && defined(CONFIG_COMPAT) struct btrfs_ioctl_encoded_io_args_32 args32; @@ -4975,11 +4970,6 @@ static int btrfs_uring_encoded_write(struct io_uring_cmd *cmd, unsigned int issu } } - if (issue_flags & IO_URING_F_NONBLOCK) { - ret = -EAGAIN; - goto out_acct; - } - pos = data->args.offset; ret = rw_verify_area(WRITE, file, &pos, data->args.len); if (ret < 0) @@ -5005,8 +4995,7 @@ out_acct: add_wchar(current, ret); inc_syscw(current); - if (ret != -EAGAIN) - kfree(data); + kfree(data); return ret; } diff --git a/fs/btrfs/ioctl.h b/fs/btrfs/ioctl.h index ccf6bed9cc24..55f86aeb3500 100644 --- a/fs/btrfs/ioctl.h +++ b/fs/btrfs/ioctl.h @@ -17,7 +17,7 @@ struct btrfs_ioctl_balance_args; long btrfs_ioctl(struct file *file, unsigned int cmd, unsigned long arg); long btrfs_compat_ioctl(struct file *file, unsigned int cmd, unsigned long arg); int btrfs_fileattr_get(struct dentry *dentry, struct file_kattr *fa); -int btrfs_fileattr_set(struct mnt_idmap *idmap, +int btrfs_fileattr_set(const struct mnt_idmap *idmap, struct dentry *dentry, struct file_kattr *fa); int btrfs_ioctl_get_supported_features(void __user *arg); void btrfs_sync_inode_flags_to_i_flags(struct btrfs_inode *inode); diff --git a/fs/btrfs/ordered-data.c b/fs/btrfs/ordered-data.c index b32d4eabe0ab..df74c75d6c29 100644 --- a/fs/btrfs/ordered-data.c +++ b/fs/btrfs/ordered-data.c @@ -417,13 +417,10 @@ static bool can_finish_ordered_extent(struct btrfs_ordered_extent *ordered, static void btrfs_queue_ordered_fn(struct btrfs_ordered_extent *ordered) { - struct btrfs_inode *inode = ordered->inode; - struct btrfs_fs_info *fs_info = inode->root->fs_info; - struct btrfs_workqueue *wq = btrfs_is_free_space_inode(inode) ? - fs_info->endio_freespace_worker : fs_info->endio_write_workers; + struct btrfs_fs_info *fs_info = ordered->inode->root->fs_info; btrfs_init_work(&ordered->work, finish_ordered_fn, NULL); - btrfs_queue_work(wq, &ordered->work); + btrfs_queue_work(fs_info->endio_write_workers, &ordered->work); } void btrfs_finish_ordered_extent(struct btrfs_ordered_extent *ordered, @@ -657,13 +654,6 @@ void btrfs_remove_ordered_extent(struct btrfs_ordered_extent *entry) struct btrfs_fs_info *fs_info = root->fs_info; struct rb_node *node; bool pending; - bool freespace_inode; - - /* - * If this is a free space inode the thread has not acquired the ordered - * extents lockdep map. - */ - freespace_inode = btrfs_is_free_space_inode(btrfs_inode); btrfs_lockdep_acquire(fs_info, btrfs_trans_pending_ordered); /* This is paired with alloc_ordered_extent(). */ @@ -738,8 +728,7 @@ void btrfs_remove_ordered_extent(struct btrfs_ordered_extent *entry) } spin_unlock(&root->ordered_extent_lock); wake_up(&entry->wait); - if (!freespace_inode) - btrfs_lockdep_release(fs_info, btrfs_ordered_extent); + btrfs_lockdep_release(fs_info, btrfs_ordered_extent); } static void btrfs_run_ordered_extent_work(struct btrfs_work *work) @@ -870,17 +859,10 @@ void btrfs_start_ordered_extent_nowriteback(struct btrfs_ordered_extent *entry, u64 start = entry->file_offset; u64 end = start + entry->num_bytes - 1; struct btrfs_inode *inode = entry->inode; - bool freespace_inode; trace_btrfs_ordered_extent_start(inode, entry); /* - * If this is a free space inode do not take the ordered extents lockdep - * map. - */ - freespace_inode = btrfs_is_free_space_inode(inode); - - /* * pages in the range can be dirty, clean or writeback. We * start IO on any dirty ones so the wait doesn't stall waiting * for the flusher thread to find them @@ -899,8 +881,7 @@ void btrfs_start_ordered_extent_nowriteback(struct btrfs_ordered_extent *entry, } } - if (!freespace_inode) - btrfs_might_wait_for_event(inode->root->fs_info, btrfs_ordered_extent); + btrfs_might_wait_for_event(inode->root->fs_info, btrfs_ordered_extent); wait_event(entry->wait, test_bit(BTRFS_ORDERED_COMPLETE, &entry->flags)); } diff --git a/fs/btrfs/qgroup.c b/fs/btrfs/qgroup.c index f68b696b4bf7..05e35eb126dc 100644 --- a/fs/btrfs/qgroup.c +++ b/fs/btrfs/qgroup.c @@ -34,7 +34,7 @@ enum btrfs_qgroup_mode btrfs_qgroup_mode(const struct btrfs_fs_info *fs_info) { if (!test_bit(BTRFS_FS_QUOTA_ENABLED, &fs_info->flags)) return BTRFS_QGROUP_MODE_DISABLED; - if (fs_info->qgroup_flags & BTRFS_QGROUP_STATUS_FLAG_SIMPLE_MODE) + if (test_bit(BTRFS_QGROUP_STATUS_BIT_SIMPLE_MODE, &fs_info->qgroup_flags)) return BTRFS_QGROUP_MODE_SIMPLE; return BTRFS_QGROUP_MODE_FULL; } @@ -384,14 +384,14 @@ static bool squota_check_parent_usage(struct btrfs_fs_info *fs_info, struct btrf __printf(2, 3) static void qgroup_mark_inconsistent(struct btrfs_fs_info *fs_info, const char *fmt, ...) { - const u64 old_flags = fs_info->qgroup_flags; + const unsigned long old_flags = fs_info->qgroup_flags; if (btrfs_qgroup_mode(fs_info) == BTRFS_QGROUP_MODE_SIMPLE) return; - fs_info->qgroup_flags |= (BTRFS_QGROUP_STATUS_FLAG_INCONSISTENT | - BTRFS_QGROUP_RUNTIME_FLAG_CANCEL_RESCAN | - BTRFS_QGROUP_RUNTIME_FLAG_NO_ACCOUNTING); - if (!(old_flags & BTRFS_QGROUP_STATUS_FLAG_INCONSISTENT)) { + set_bit(BTRFS_QGROUP_STATUS_BIT_INCONSISTENT, &fs_info->qgroup_flags); + set_bit(BTRFS_QGROUP_RUNTIME_BIT_CANCEL_RESCAN, &fs_info->qgroup_flags); + set_bit(BTRFS_QGROUP_RUNTIME_BIT_NO_ACCOUNTING, &fs_info->qgroup_flags); + if (!test_bit(BTRFS_QGROUP_STATUS_BIT_INCONSISTENT, &old_flags)) { struct va_format vaf; va_list args; @@ -426,7 +426,6 @@ int btrfs_read_qgroup_config(struct btrfs_fs_info *fs_info) struct extent_buffer *l; int slot; int ret = 0; - u64 flags = 0; u64 rescan_progress = 0; if (!fs_info->quota_root) @@ -473,8 +472,12 @@ int btrfs_read_qgroup_config(struct btrfs_fs_info *fs_info) "old qgroup version, quota disabled"); goto out; } - fs_info->qgroup_flags = btrfs_qgroup_status_flags(l, ptr); - if (fs_info->qgroup_flags & BTRFS_QGROUP_STATUS_FLAG_SIMPLE_MODE) + if (btrfs_qgroup_status_flags(l, ptr) > ULONG_MAX) { + btrfs_err(fs_info, "invalid qgroup status flags, quota disabled"); + goto out; + } + fs_info->qgroup_flags = (unsigned long)btrfs_qgroup_status_flags(l, ptr); + if (test_bit(BTRFS_QGROUP_STATUS_BIT_SIMPLE_MODE, &fs_info->qgroup_flags)) qgroup_read_enable_gen(fs_info, l, slot, ptr); else if (btrfs_qgroup_status_generation(l, ptr) != fs_info->generation) qgroup_mark_inconsistent(fs_info, "qgroup generation mismatch"); @@ -609,14 +612,13 @@ next2: } out: btrfs_free_path(path); - fs_info->qgroup_flags |= flags; if (ret >= 0) { - if (fs_info->qgroup_flags & BTRFS_QGROUP_STATUS_FLAG_ON) + if (test_bit(BTRFS_QGROUP_STATUS_BIT_ON, &fs_info->qgroup_flags)) set_bit(BTRFS_FS_QUOTA_ENABLED, &fs_info->flags); - if (fs_info->qgroup_flags & BTRFS_QGROUP_STATUS_FLAG_RESCAN) + if (test_bit(BTRFS_QGROUP_STATUS_BIT_RESCAN, &fs_info->qgroup_flags)) ret = qgroup_rescan_init(fs_info, rescan_progress, 0); } else { - fs_info->qgroup_flags &= ~BTRFS_QGROUP_STATUS_FLAG_RESCAN; + clear_bit(BTRFS_QGROUP_STATUS_BIT_RESCAN, &fs_info->qgroup_flags); btrfs_sysfs_del_qgroups(fs_info); } @@ -1101,9 +1103,9 @@ int btrfs_quota_enable(struct btrfs_fs_info *fs_info, struct btrfs_qgroup_status_item); btrfs_set_qgroup_status_generation(leaf, ptr, trans->transid); btrfs_set_qgroup_status_version(leaf, ptr, BTRFS_QGROUP_STATUS_VERSION); - fs_info->qgroup_flags = BTRFS_QGROUP_STATUS_FLAG_ON; + set_bit(BTRFS_QGROUP_STATUS_BIT_ON, &fs_info->qgroup_flags); if (simple) { - fs_info->qgroup_flags |= BTRFS_QGROUP_STATUS_FLAG_SIMPLE_MODE; + set_bit(BTRFS_QGROUP_STATUS_BIT_SIMPLE_MODE, &fs_info->qgroup_flags); btrfs_set_fs_incompat(fs_info, SIMPLE_QUOTA); /* * Set the enable generation to the next transaction, as we cannot @@ -1113,7 +1115,7 @@ int btrfs_quota_enable(struct btrfs_fs_info *fs_info, */ btrfs_set_qgroup_status_enable_gen(leaf, ptr, trans->transid + 1); } else { - fs_info->qgroup_flags |= BTRFS_QGROUP_STATUS_FLAG_INCONSISTENT; + set_bit(BTRFS_QGROUP_STATUS_BIT_INCONSISTENT, &fs_info->qgroup_flags); } btrfs_set_qgroup_status_flags(leaf, ptr, fs_info->qgroup_flags & BTRFS_QGROUP_STATUS_FLAGS_MASK); @@ -1403,8 +1405,14 @@ int btrfs_quota_disable(struct btrfs_fs_info *fs_info) spin_lock(&fs_info->qgroup_lock); quota_root = fs_info->quota_root; fs_info->quota_root = NULL; - fs_info->qgroup_flags &= ~BTRFS_QGROUP_STATUS_FLAG_ON; - fs_info->qgroup_flags &= ~BTRFS_QGROUP_STATUS_FLAG_SIMPLE_MODE; + /* + * Clear all on-disk and runtime bits, except RESCAN related ones, that + * are either handled by rescan thread, or the caller who rejects rescan. + */ + clear_bit(BTRFS_QGROUP_STATUS_BIT_ON, &fs_info->qgroup_flags); + clear_bit(BTRFS_QGROUP_STATUS_BIT_SIMPLE_MODE, &fs_info->qgroup_flags); + clear_bit(BTRFS_QGROUP_STATUS_BIT_INCONSISTENT, &fs_info->qgroup_flags); + clear_bit(BTRFS_QGROUP_RUNTIME_BIT_NO_ACCOUNTING, &fs_info->qgroup_flags); fs_info->qgroup_drop_subtree_thres = BTRFS_QGROUP_DROP_SUBTREE_THRES_DEFAULT; spin_unlock(&fs_info->qgroup_lock); @@ -1554,7 +1562,7 @@ static int quick_update_accounting(struct btrfs_fs_info *fs_info, } out: if (ret) - fs_info->qgroup_flags |= BTRFS_QGROUP_STATUS_FLAG_INCONSISTENT; + set_bit(BTRFS_QGROUP_STATUS_BIT_INCONSISTENT, &fs_info->qgroup_flags); return ret; } @@ -1875,7 +1883,7 @@ int btrfs_remove_qgroup(struct btrfs_trans_handle *trans, u64 qgroupid) * very frequently. */ if (btrfs_qgroup_mode(fs_info) == BTRFS_QGROUP_MODE_FULL && - !(fs_info->qgroup_flags & BTRFS_QGROUP_STATUS_FLAG_INCONSISTENT)) { + !test_bit(BTRFS_QGROUP_STATUS_BIT_INCONSISTENT, &fs_info->qgroup_flags)) { if (unlikely(qgroup->rfer || qgroup->excl || qgroup->rfer_cmpr || qgroup->excl_cmpr)) { DEBUG_WARN(); @@ -2120,7 +2128,7 @@ int btrfs_qgroup_trace_extent_post(struct btrfs_trans_handle *trans, */ ASSERT(trans != NULL); - if (fs_info->qgroup_flags & BTRFS_QGROUP_RUNTIME_FLAG_NO_ACCOUNTING) + if (test_bit(BTRFS_QGROUP_RUNTIME_BIT_NO_ACCOUNTING, &fs_info->qgroup_flags)) return 0; ret = btrfs_find_all_roots(&ctx, true); @@ -2740,6 +2748,24 @@ walk_down: return 0; } +void btrfs_qgroup_check_tree_drop(struct btrfs_fs_info *fs_info, u64 rootid, u8 level) +{ + u8 drop_subtree_thres; + + if (btrfs_qgroup_mode(fs_info) != BTRFS_QGROUP_MODE_FULL) + return; + + if (!btrfs_is_fstree(rootid)) + return; + + spin_lock(&fs_info->qgroup_lock); + drop_subtree_thres = fs_info->qgroup_drop_subtree_thres; + spin_unlock(&fs_info->qgroup_lock); + + if (level >= drop_subtree_thres) + qgroup_mark_inconsistent(fs_info, "subtree level reached threshold"); +} + static void qgroup_iterator_nested_add(struct list_head *head, struct btrfs_qgroup *qgroup) { if (!list_empty(&qgroup->nested_iterator)) @@ -2961,7 +2987,7 @@ int btrfs_qgroup_account_extent(struct btrfs_trans_handle *trans, u64 bytenr, * we can't just exit here. */ if (!btrfs_qgroup_full_accounting(fs_info) || - fs_info->qgroup_flags & BTRFS_QGROUP_RUNTIME_FLAG_NO_ACCOUNTING) + test_bit(BTRFS_QGROUP_RUNTIME_BIT_NO_ACCOUNTING, &fs_info->qgroup_flags)) goto out_free; if (new_roots) { @@ -2983,7 +3009,7 @@ int btrfs_qgroup_account_extent(struct btrfs_trans_handle *trans, u64 bytenr, num_bytes, nr_old_roots, nr_new_roots); mutex_lock(&fs_info->qgroup_rescan_lock); - if (fs_info->qgroup_flags & BTRFS_QGROUP_STATUS_FLAG_RESCAN) { + if (test_bit(BTRFS_QGROUP_STATUS_BIT_RESCAN, &fs_info->qgroup_flags)) { if (fs_info->qgroup_rescan_progress.objectid <= bytenr) { mutex_unlock(&fs_info->qgroup_rescan_lock); ret = 0; @@ -3044,8 +3070,8 @@ int btrfs_qgroup_account_extents(struct btrfs_trans_handle *trans) num_dirty_extents++; trace_btrfs_qgroup_account_extents(fs_info, record, bytenr); - if (!ret && !(fs_info->qgroup_flags & - BTRFS_QGROUP_RUNTIME_FLAG_NO_ACCOUNTING)) { + if (!ret && !test_bit(BTRFS_QGROUP_RUNTIME_BIT_NO_ACCOUNTING, + &fs_info->qgroup_flags)) { struct btrfs_backref_walk_ctx ctx = { 0 }; ctx.bytenr = bytenr; @@ -3152,9 +3178,9 @@ int btrfs_run_qgroups(struct btrfs_trans_handle *trans) spin_lock(&fs_info->qgroup_lock); } if (btrfs_qgroup_enabled(fs_info)) - fs_info->qgroup_flags |= BTRFS_QGROUP_STATUS_FLAG_ON; + set_bit(BTRFS_QGROUP_STATUS_BIT_ON, &fs_info->qgroup_flags); else - fs_info->qgroup_flags &= ~BTRFS_QGROUP_STATUS_FLAG_ON; + clear_bit(BTRFS_QGROUP_STATUS_BIT_ON, &fs_info->qgroup_flags); spin_unlock(&fs_info->qgroup_lock); ret = update_qgroup_status_item(trans); @@ -3844,7 +3870,7 @@ static bool rescan_should_stop(struct btrfs_fs_info *fs_info) return true; if (!btrfs_qgroup_enabled(fs_info)) return true; - if (fs_info->qgroup_flags & BTRFS_QGROUP_RUNTIME_FLAG_CANCEL_RESCAN) + if (test_bit(BTRFS_QGROUP_RUNTIME_BIT_CANCEL_RESCAN, &fs_info->qgroup_flags)) return true; return false; } @@ -3894,12 +3920,10 @@ out: btrfs_free_path(path); mutex_lock(&fs_info->qgroup_rescan_lock); - if (ret > 0 && - fs_info->qgroup_flags & BTRFS_QGROUP_STATUS_FLAG_INCONSISTENT) { - fs_info->qgroup_flags &= ~BTRFS_QGROUP_STATUS_FLAG_INCONSISTENT; - } else if (ret < 0 || stopped) { - fs_info->qgroup_flags |= BTRFS_QGROUP_STATUS_FLAG_INCONSISTENT; - } + if (ret > 0) + clear_bit(BTRFS_QGROUP_STATUS_BIT_INCONSISTENT, &fs_info->qgroup_flags); + else if (ret < 0 || stopped) + set_bit(BTRFS_QGROUP_STATUS_BIT_INCONSISTENT, &fs_info->qgroup_flags); mutex_unlock(&fs_info->qgroup_rescan_lock); /* @@ -3923,9 +3947,9 @@ out: } mutex_lock(&fs_info->qgroup_rescan_lock); - if (!stopped || - fs_info->qgroup_flags & BTRFS_QGROUP_RUNTIME_FLAG_CANCEL_RESCAN) - fs_info->qgroup_flags &= ~BTRFS_QGROUP_STATUS_FLAG_RESCAN; + if (!stopped || test_bit(BTRFS_QGROUP_RUNTIME_BIT_CANCEL_RESCAN, + &fs_info->qgroup_flags)) + clear_bit(BTRFS_QGROUP_STATUS_BIT_RESCAN, &fs_info->qgroup_flags); if (trans) { int ret2 = update_qgroup_status_item(trans); @@ -3935,7 +3959,7 @@ out: } } fs_info->qgroup_rescan_running = false; - fs_info->qgroup_flags &= ~BTRFS_QGROUP_RUNTIME_FLAG_CANCEL_RESCAN; + clear_bit(BTRFS_QGROUP_RUNTIME_BIT_CANCEL_RESCAN, &fs_info->qgroup_flags); complete_all(&fs_info->qgroup_rescan_completion); mutex_unlock(&fs_info->qgroup_rescan_lock); @@ -3946,7 +3970,7 @@ out: if (stopped) { btrfs_info(fs_info, "qgroup scan paused"); - } else if (fs_info->qgroup_flags & BTRFS_QGROUP_RUNTIME_FLAG_CANCEL_RESCAN) { + } else if (test_bit(BTRFS_QGROUP_RUNTIME_BIT_CANCEL_RESCAN, &fs_info->qgroup_flags)) { btrfs_info(fs_info, "qgroup scan cancelled"); } else if (ret >= 0) { btrfs_info(fs_info, "qgroup scan completed%s", @@ -3973,13 +3997,11 @@ qgroup_rescan_init(struct btrfs_fs_info *fs_info, u64 progress_objectid, if (!init_flags) { /* we're resuming qgroup rescan at mount time */ - if (!(fs_info->qgroup_flags & - BTRFS_QGROUP_STATUS_FLAG_RESCAN)) { + if (!(test_bit(BTRFS_QGROUP_STATUS_BIT_RESCAN, &fs_info->qgroup_flags))) { btrfs_debug(fs_info, "qgroup rescan init failed, qgroup rescan is not queued"); ret = -EINVAL; - } else if (!(fs_info->qgroup_flags & - BTRFS_QGROUP_STATUS_FLAG_ON)) { + } else if (!(test_bit(BTRFS_QGROUP_STATUS_BIT_ON, &fs_info->qgroup_flags))) { btrfs_debug(fs_info, "qgroup rescan init failed, qgroup is not enabled"); ret = -ENOTCONN; @@ -3992,10 +4014,12 @@ qgroup_rescan_init(struct btrfs_fs_info *fs_info, u64 progress_objectid, mutex_lock(&fs_info->qgroup_rescan_lock); if (init_flags) { - if (fs_info->qgroup_flags & BTRFS_QGROUP_STATUS_FLAG_RESCAN) { + if (test_bit(BTRFS_QGROUP_STATUS_BIT_RESCAN, + &fs_info->qgroup_flags) || + test_bit(BTRFS_QGROUP_RUNTIME_BIT_REJECT_RESCAN, + &fs_info->qgroup_flags)) { ret = -EINPROGRESS; - } else if (!(fs_info->qgroup_flags & - BTRFS_QGROUP_STATUS_FLAG_ON)) { + } else if (!test_bit(BTRFS_QGROUP_STATUS_BIT_ON, &fs_info->qgroup_flags)) { btrfs_debug(fs_info, "qgroup rescan init failed, qgroup is not enabled"); ret = -ENOTCONN; @@ -4008,13 +4032,13 @@ qgroup_rescan_init(struct btrfs_fs_info *fs_info, u64 progress_objectid, mutex_unlock(&fs_info->qgroup_rescan_lock); return ret; } - fs_info->qgroup_flags |= BTRFS_QGROUP_STATUS_FLAG_RESCAN; + set_bit(BTRFS_QGROUP_STATUS_BIT_RESCAN, &fs_info->qgroup_flags); } memset(&fs_info->qgroup_rescan_progress, 0, sizeof(fs_info->qgroup_rescan_progress)); - fs_info->qgroup_flags &= ~(BTRFS_QGROUP_RUNTIME_FLAG_CANCEL_RESCAN | - BTRFS_QGROUP_RUNTIME_FLAG_NO_ACCOUNTING); + clear_bit(BTRFS_QGROUP_RUNTIME_BIT_CANCEL_RESCAN, &fs_info->qgroup_flags); + clear_bit(BTRFS_QGROUP_RUNTIME_BIT_NO_ACCOUNTING, &fs_info->qgroup_flags); fs_info->qgroup_rescan_progress.objectid = progress_objectid; init_completion(&fs_info->qgroup_rescan_completion); mutex_unlock(&fs_info->qgroup_rescan_lock); @@ -4065,7 +4089,7 @@ btrfs_qgroup_rescan(struct btrfs_fs_info *fs_info) ret = btrfs_commit_current_transaction(fs_info->fs_root); if (ret) { - fs_info->qgroup_flags &= ~BTRFS_QGROUP_STATUS_FLAG_RESCAN; + clear_bit(BTRFS_QGROUP_STATUS_BIT_RESCAN, &fs_info->qgroup_flags); return ret; } @@ -4118,7 +4142,7 @@ int btrfs_qgroup_wait_for_completion(struct btrfs_fs_info *fs_info, void btrfs_qgroup_rescan_resume(struct btrfs_fs_info *fs_info) { - if (fs_info->qgroup_flags & BTRFS_QGROUP_STATUS_FLAG_RESCAN) { + if (test_bit(BTRFS_QGROUP_STATUS_BIT_RESCAN, &fs_info->qgroup_flags)) { mutex_lock(&fs_info->qgroup_rescan_lock); fs_info->qgroup_rescan_running = true; btrfs_queue_work(fs_info->qgroup_rescan_workers, diff --git a/fs/btrfs/qgroup.h b/fs/btrfs/qgroup.h index 80dd2dacd56d..c64b26b09c22 100644 --- a/fs/btrfs/qgroup.h +++ b/fs/btrfs/qgroup.h @@ -121,8 +121,19 @@ struct btrfs_qgroup_swapped_blocks; * To minimize the chance of collision with new persisted status flags, these * count backwards from the MSB. */ -#define BTRFS_QGROUP_RUNTIME_FLAG_CANCEL_RESCAN (1ULL << 63) -#define BTRFS_QGROUP_RUNTIME_FLAG_NO_ACCOUNTING (1ULL << 62) +#define BTRFS_QGROUP_RUNTIME_BIT_CANCEL_RESCAN (BITS_PER_LONG - 1) +#define BTRFS_QGROUP_RUNTIME_BIT_NO_ACCOUNTING (BITS_PER_LONG - 2) + +/* + * No new rescan allowed when set. + * + * During huge subtree dropping, qgroup will be marked inconsistent, and skip + * all future accounting to avoid long stall. But, an immediate rescan will + * re-enable qgroup and still stall the system. + * + * This bit is to avoid such rescan during the duration of a subvolume dropping. + */ +#define BTRFS_QGROUP_RUNTIME_BIT_REJECT_RESCAN (BITS_PER_LONG - 3) #define BTRFS_QGROUP_DROP_SUBTREE_THRES_DEFAULT (3) @@ -365,6 +376,7 @@ int btrfs_qgroup_trace_leaf_items(struct btrfs_trans_handle *trans, int btrfs_qgroup_trace_subtree(struct btrfs_trans_handle *trans, struct extent_buffer *root_eb, u64 root_gen, int root_level); +void btrfs_qgroup_check_tree_drop(struct btrfs_fs_info *fs_info, u64 rootid, u8 level); int btrfs_qgroup_account_extent(struct btrfs_trans_handle *trans, u64 bytenr, u64 num_bytes, struct ulist *old_roots, struct ulist *new_roots); diff --git a/fs/btrfs/raid56.c b/fs/btrfs/raid56.c index 1ee52a9dcee3..8ec24dbb180f 100644 --- a/fs/btrfs/raid56.c +++ b/fs/btrfs/raid56.c @@ -953,7 +953,7 @@ static void rbio_orig_end_io(struct btrfs_raid_bio *rbio, blk_status_t status) /* * Clear the data bitmap, as the rbio may be cached for later usage. - * do this before before unlock_stripe() so there will be no new bio + * do this before unlock_stripe() so there will be no new bio * for this bio. */ bitmap_clear(&rbio->dbitmap, 0, rbio->stripe_nsectors); @@ -988,7 +988,7 @@ static void rbio_orig_end_io(struct btrfs_raid_bio *rbio, blk_status_t status) * as possible, and only use stripe_sectors as fallback. * * Return NULL if bio_list_only is set but the specified sector has no - * coresponding bio. + * corresponding bio. */ static phys_addr_t *sector_paddrs_in_rbio(struct btrfs_raid_bio *rbio, int stripe_nr, int sector_nr, @@ -1450,10 +1450,7 @@ static int rmw_assemble_write_bios(struct btrfs_raid_bio *rbio, /* We should have at least one data sector. */ ASSERT(bitmap_weight(&rbio->dbitmap, rbio->stripe_nsectors)); - /* - * Reset errors, as we may have errors inherited from from degraded - * write. - */ + /* Reset errors, as we may have errors inherited from degraded write. */ bitmap_clear(rbio->error_bitmap, 0, rbio->nr_sectors); /* @@ -1652,12 +1649,7 @@ static void verify_bio_data_sectors(struct btrfs_raid_bio *rbio, struct bio *bio) { struct btrfs_fs_info *fs_info = rbio->bioc->fs_info; - const u32 step = min(fs_info->sectorsize, PAGE_SIZE); - const u32 nr_steps = rbio->sector_nsteps; int total_sector_nr = get_bio_sector_nr(rbio, bio); - u32 offset = 0; - phys_addr_t paddrs[BTRFS_MAX_BLOCKSIZE / PAGE_SIZE]; - phys_addr_t paddr; /* No data csum for the whole stripe, no need to verify. */ if (!rbio->csum_bitmap || !rbio->csum_buf) @@ -1667,28 +1659,20 @@ static void verify_bio_data_sectors(struct btrfs_raid_bio *rbio, if (total_sector_nr >= rbio->nr_data * rbio->stripe_nsectors) return; - btrfs_bio_for_each_block_all(paddr, bio, step) { + for (struct bvec_iter iter = init_bvec_iter_for_bio(bio); + iter.bi_size; + bio_advance_iter(bio, &iter, fs_info->sectorsize), total_sector_nr++) { u8 csum_buf[BTRFS_CSUM_SIZE]; u8 *expected_csum; - paddrs[(offset / step) % nr_steps] = paddr; - offset += step; - - /* Not yet covering the full fs block, continue to the next step. */ - if (!IS_ALIGNED(offset, fs_info->sectorsize)) - continue; - /* No csum for this sector, skip to the next sector. */ - if (!test_bit(total_sector_nr, rbio->csum_bitmap)) { - total_sector_nr++; + if (!test_bit(total_sector_nr, rbio->csum_bitmap)) continue; - } expected_csum = rbio->csum_buf + total_sector_nr * fs_info->csum_size; - btrfs_calculate_block_csum_pages(fs_info, paddrs, csum_buf); + btrfs_csum_one_bio_block(fs_info, bio, &iter, csum_buf); if (unlikely(memcmp(csum_buf, expected_csum, fs_info->csum_size) != 0)) set_bit(total_sector_nr, rbio->error_bitmap); - total_sector_nr++; } } @@ -1879,6 +1863,27 @@ void raid56_parity_write(struct bio *bio, struct btrfs_io_context *bioc) start_async_work(rbio, rmw_rbio_work); } +static void calculate_block_csum_paddrs(struct btrfs_fs_info *fs_info, + const phys_addr_t paddrs[], u8 *dest) +{ + const u32 blocksize = fs_info->sectorsize; + const u32 step = min(blocksize, PAGE_SIZE); + const u32 nr_steps = blocksize / step; + struct btrfs_csum_ctx csum; + + btrfs_csum_init(&csum, fs_info->csum_type); + for (int i = 0; i < nr_steps; i++) { + const phys_addr_t paddr = paddrs[i]; + void *kaddr; + + ASSERT(offset_in_page(paddr) + step <= PAGE_SIZE); + kaddr = kmap_local_page(phys_to_page(paddr)) + offset_in_page(paddr); + btrfs_csum_update(&csum, kaddr, step); + kunmap_local(kaddr); + } + btrfs_csum_final(&csum, dest); +} + static int verify_one_sector(struct btrfs_raid_bio *rbio, int stripe_nr, int sector_nr) { @@ -1906,7 +1911,7 @@ static int verify_one_sector(struct btrfs_raid_bio *rbio, csum_expected = rbio->csum_buf + (stripe_nr * rbio->stripe_nsectors + sector_nr) * fs_info->csum_size; - btrfs_calculate_block_csum_pages(fs_info, paddrs, csum_buf); + calculate_block_csum_paddrs(fs_info, paddrs, csum_buf); if (unlikely(memcmp(csum_buf, csum_expected, fs_info->csum_size) != 0)) return -EIO; return 0; @@ -2624,7 +2629,7 @@ static int alloc_rbio_essential_pages(struct btrfs_raid_bio *rbio) return 0; } -/* Return true if the content of the step matches the caclulated one. */ +/* Return true if the content of the step matches the calculated one. */ static bool verify_one_parity_step(struct btrfs_raid_bio *rbio, void *pointers[], unsigned int sector_nr, unsigned int step_nr) diff --git a/fs/btrfs/relocation.c b/fs/btrfs/relocation.c index da54db75e7a9..630a7ad8f8e1 100644 --- a/fs/btrfs/relocation.c +++ b/fs/btrfs/relocation.c @@ -3357,7 +3357,7 @@ truncate: goto out; } - ret = btrfs_truncate_free_space_cache(trans, block_group, inode); + ret = btrfs_truncate_free_space_cache(trans, inode); btrfs_end_transaction(trans); btrfs_btree_balance_dirty(fs_info); diff --git a/fs/btrfs/send.c b/fs/btrfs/send.c index 5c59b9abedcd..c523bf950c89 100644 --- a/fs/btrfs/send.c +++ b/fs/btrfs/send.c @@ -7023,7 +7023,7 @@ static int changed_extent(struct send_ctx *sctx, * get modified or replaced with a new one). Note that deduplication * updates the inode item, but it only changes the iversion (sequence * field in the inode item) of the inode, so if a file is deduplicated - * the same amount of times in both the parent and send snapshots, its + * the same number of times in both the parent and send snapshots, its * iversion becomes the same in both snapshots, whence the inode item is * the same on both snapshots. */ diff --git a/fs/btrfs/space-info.c b/fs/btrfs/space-info.c index 39a28e1bec8a..01018152c054 100644 --- a/fs/btrfs/space-info.c +++ b/fs/btrfs/space-info.c @@ -1704,7 +1704,6 @@ static int handle_reserve_ticket(struct btrfs_space_info *space_info, evict_flush_states, ARRAY_SIZE(evict_flush_states)); break; - case BTRFS_RESERVE_FLUSH_FREE_SPACE_INODE: case BTRFS_RESERVE_FLUSH_ZONED_RELOCATION: priority_reclaim_data_space(space_info, ticket); break; @@ -1968,7 +1967,6 @@ int btrfs_reserve_data_bytes(struct btrfs_space_info *space_info, u64 bytes, int ret; ASSERT(flush == BTRFS_RESERVE_FLUSH_DATA || - flush == BTRFS_RESERVE_FLUSH_FREE_SPACE_INODE || flush == BTRFS_RESERVE_FLUSH_ZONED_RELOCATION || flush == BTRFS_RESERVE_NO_FLUSH, "flush=%d", flush); ASSERT(!current->journal_info || flush != BTRFS_RESERVE_FLUSH_DATA, diff --git a/fs/btrfs/space-info.h b/fs/btrfs/space-info.h index aa836e8a9d4a..d0130c8ba3dd 100644 --- a/fs/btrfs/space-info.h +++ b/fs/btrfs/space-info.h @@ -66,7 +66,6 @@ enum btrfs_reserve_flush_enum { * Can be interrupted by a fatal signal. */ BTRFS_RESERVE_FLUSH_DATA, - BTRFS_RESERVE_FLUSH_FREE_SPACE_INODE, BTRFS_RESERVE_FLUSH_ALL, /* @@ -82,9 +81,6 @@ enum btrfs_reserve_flush_enum { * priority flushing for this, because otherwise we can deadlock on * waiting for a ticket, that cannot be granted, because we cannot do * any allocations. - * - * Apart from being specific to zoned relocation, it is equal to - * BTRFS_FLUSH_FREE_SPACE_INODE. */ BTRFS_RESERVE_FLUSH_ZONED_RELOCATION, diff --git a/fs/btrfs/super.c b/fs/btrfs/super.c index ddb620ac241b..14ed0ed823a3 100644 --- a/fs/btrfs/super.c +++ b/fs/btrfs/super.c @@ -515,7 +515,6 @@ static int btrfs_parse_param(struct fs_context *fc, struct fs_parameter *param) btrfs_warn(NULL, "v1 space cache is deprecated, falling back to no space cache"); btrfs_set_opt(ctx->mount_opt, NOSPACECACHE); - btrfs_clear_opt(ctx->mount_opt, SPACE_CACHE); btrfs_clear_opt(ctx->mount_opt, FREE_SPACE_TREE); break; case Opt_space_cache_version: @@ -524,11 +523,9 @@ static int btrfs_parse_param(struct fs_context *fc, struct fs_parameter *param) btrfs_warn(NULL, "v1 space cache is deprecated, falling back to no space cache"); btrfs_set_opt(ctx->mount_opt, NOSPACECACHE); - btrfs_clear_opt(ctx->mount_opt, SPACE_CACHE); btrfs_clear_opt(ctx->mount_opt, FREE_SPACE_TREE); break; case Opt_space_cache_v2: - btrfs_clear_opt(ctx->mount_opt, SPACE_CACHE); btrfs_set_opt(ctx->mount_opt, FREE_SPACE_TREE); break; default: @@ -705,13 +702,6 @@ bool btrfs_check_options(const struct btrfs_fs_info *info, if (btrfs_check_mountopts_zoned(info, mount_opt)) ret = false; - if (!test_bit(BTRFS_FS_STATE_REMOUNTING, &info->fs_state)) { - if (btrfs_raw_test_opt(*mount_opt, SPACE_CACHE)) { - btrfs_warn(info, -"space cache v1 is being deprecated and will be removed in a future release, please use -o space_cache=v2"); - } - } - return ret; } @@ -729,14 +719,6 @@ bool btrfs_check_options(const struct btrfs_fs_info *info, */ void btrfs_set_free_space_cache_settings(struct btrfs_fs_info *fs_info) { - if (fs_info->sectorsize != PAGE_SIZE && btrfs_test_opt(fs_info, SPACE_CACHE)) { - btrfs_info(fs_info, - "forcing free space tree for sector size %u with page size %lu", - fs_info->sectorsize, PAGE_SIZE); - btrfs_clear_opt(fs_info->mount_opt, SPACE_CACHE); - btrfs_set_opt(fs_info->mount_opt, FREE_SPACE_TREE); - } - /* * At this point our mount options are populated, so we only mess with * these settings if we don't have any settings already. @@ -751,20 +733,17 @@ void btrfs_set_free_space_cache_settings(struct btrfs_fs_info *fs_info) return; } - if (btrfs_test_opt(fs_info, SPACE_CACHE)) - return; - if (btrfs_test_opt(fs_info, NOSPACECACHE)) return; /* * At this point we don't have explicit options set by the user, set - * them ourselves based on the state of the file system. + * them ourselves based on the state of the file system. An existing + * v1 space cache is no longer used and gets cleaned up once the + * filesystem is mounted read-write. */ if (btrfs_fs_compat_ro(fs_info, FREE_SPACE_TREE)) btrfs_set_opt(fs_info->mount_opt, FREE_SPACE_TREE); - else if (btrfs_free_space_cache_v1_active(fs_info)) - btrfs_set_opt(fs_info->mount_opt, SPACE_CACHE); } static void set_device_specific_options(struct btrfs_fs_info *fs_info) @@ -1107,9 +1086,7 @@ static int btrfs_show_options(struct seq_file *seq, struct dentry *dentry) seq_puts(seq, ",discard=async"); if (!(info->sb->s_flags & SB_POSIXACL)) seq_puts(seq, ",noacl"); - if (btrfs_free_space_cache_v1_active(info)) - seq_puts(seq, ",space_cache"); - else if (btrfs_fs_compat_ro(info, FREE_SPACE_TREE)) + if (btrfs_fs_compat_ro(info, FREE_SPACE_TREE)) seq_puts(seq, ",space_cache=v2"); else seq_puts(seq, ",nospace_cache"); @@ -1243,7 +1220,6 @@ static void btrfs_resize_thread_pool(struct btrfs_fs_info *fs_info, workqueue_set_max_active(fs_info->endio_workers, new_pool_size); workqueue_set_max_active(fs_info->endio_meta_workers, new_pool_size); btrfs_workqueue_set_max(fs_info->endio_write_workers, new_pool_size); - btrfs_workqueue_set_max(fs_info->endio_freespace_worker, new_pool_size); btrfs_workqueue_set_max(fs_info->delayed_workers, new_pool_size); } @@ -1264,8 +1240,6 @@ static inline void btrfs_remount_begin(struct btrfs_fs_info *fs_info, static inline void btrfs_remount_cleanup(struct btrfs_fs_info *fs_info, unsigned long long old_opts) { - const bool cache_opt = btrfs_test_opt(fs_info, SPACE_CACHE); - /* * We need to cleanup all defraggable inodes if the autodefragment is * close or the filesystem is read only. @@ -1282,10 +1256,6 @@ static inline void btrfs_remount_cleanup(struct btrfs_fs_info *fs_info, else if (btrfs_raw_test_opt(old_opts, DISCARD_ASYNC) && !btrfs_test_opt(fs_info, DISCARD_ASYNC)) btrfs_discard_cleanup(fs_info); - - /* If we toggled space cache */ - if (cache_opt != btrfs_free_space_cache_v1_active(fs_info)) - btrfs_set_free_space_cache_v1_active(fs_info, cache_opt); } static int btrfs_remount_rw(struct btrfs_fs_info *fs_info) @@ -1448,7 +1418,6 @@ static void btrfs_emit_options(struct btrfs_fs_info *info, btrfs_info_if_set(info, old, DISCARD_SYNC, "turning on sync discard"); btrfs_info_if_set(info, old, DISCARD_ASYNC, "turning on async discard"); btrfs_info_if_set(info, old, FREE_SPACE_TREE, "enabling free space tree"); - btrfs_info_if_set(info, old, SPACE_CACHE, "enabling disk space caching"); btrfs_info_if_set(info, old, CLEAR_CACHE, "force clearing of disk cache"); btrfs_info_if_set(info, old, AUTO_DEFRAG, "enabling auto defrag"); btrfs_info_if_set(info, old, FRAGMENT_DATA, "fragmenting data"); @@ -1466,7 +1435,6 @@ static void btrfs_emit_options(struct btrfs_fs_info *info, btrfs_info_if_unset(info, old, SSD_SPREAD, "not using spread ssd allocation scheme"); btrfs_info_if_unset(info, old, NOBARRIER, "turning on barriers"); btrfs_info_if_unset(info, old, NOTREELOG, "enabling tree log"); - btrfs_info_if_unset(info, old, SPACE_CACHE, "disabling disk space caching"); btrfs_info_if_unset(info, old, FREE_SPACE_TREE, "disabling free space tree"); btrfs_info_if_unset(info, old, AUTO_DEFRAG, "disabling auto defrag"); btrfs_info_if_unset(info, old, COMPRESS, "use no compression"); @@ -1531,14 +1499,8 @@ static int btrfs_reconfigure(struct fs_context *fc) btrfs_warn(fs_info, "remount supports changing free space tree only from RO to RW"); /* Make sure free space cache options match the state on disk. */ - if (btrfs_fs_compat_ro(fs_info, FREE_SPACE_TREE)) { + if (btrfs_fs_compat_ro(fs_info, FREE_SPACE_TREE)) btrfs_set_opt(fs_info->mount_opt, FREE_SPACE_TREE); - btrfs_clear_opt(fs_info->mount_opt, SPACE_CACHE); - } - if (btrfs_free_space_cache_v1_active(fs_info)) { - btrfs_clear_opt(fs_info->mount_opt, FREE_SPACE_TREE); - btrfs_set_opt(fs_info->mount_opt, SPACE_CACHE); - } } ret = 0; diff --git a/fs/btrfs/sysfs.c b/fs/btrfs/sysfs.c index 39cb01ee441a..c5bb1c7eac6a 100644 --- a/fs/btrfs/sysfs.c +++ b/fs/btrfs/sysfs.c @@ -83,8 +83,7 @@ struct raid_kobject { #define BTRFS_FEAT_ATTR(_name, _feature_set, _feature_prefix, _feature_bit) \ static struct btrfs_feature_attr btrfs_attr_features_##_name = { \ .kobj_attr = __INIT_KOBJ_ATTR(_name, S_IRUGO, \ - btrfs_feature_attr_show, \ - btrfs_feature_attr_store), \ + btrfs_feature_attr_show, NULL), \ .feature_set = _feature_set, \ .feature_bit = _feature_prefix ##_## _feature_bit, \ } @@ -130,130 +129,20 @@ static u64 get_features(struct btrfs_fs_info *fs_info, return btrfs_super_incompat_flags(disk_super); } -static void set_features(struct btrfs_fs_info *fs_info, - enum btrfs_feature_set set, u64 features) -{ - struct btrfs_super_block *disk_super = fs_info->super_copy; - if (set == FEAT_COMPAT) - btrfs_set_super_compat_flags(disk_super, features); - else if (set == FEAT_COMPAT_RO) - btrfs_set_super_compat_ro_flags(disk_super, features); - else - btrfs_set_super_incompat_flags(disk_super, features); -} - -static int can_modify_feature(struct btrfs_feature_attr *fa) -{ - int val = 0; - u64 set, clear; - switch (fa->feature_set) { - case FEAT_COMPAT: - set = BTRFS_FEATURE_COMPAT_SAFE_SET; - clear = BTRFS_FEATURE_COMPAT_SAFE_CLEAR; - break; - case FEAT_COMPAT_RO: - set = BTRFS_FEATURE_COMPAT_RO_SAFE_SET; - clear = BTRFS_FEATURE_COMPAT_RO_SAFE_CLEAR; - break; - case FEAT_INCOMPAT: - set = BTRFS_FEATURE_INCOMPAT_SAFE_SET; - clear = BTRFS_FEATURE_INCOMPAT_SAFE_CLEAR; - break; - default: - btrfs_warn(NULL, "sysfs: unknown feature set %d", fa->feature_set); - return 0; - } - - if (set & fa->feature_bit) - val |= 1; - if (clear & fa->feature_bit) - val |= 2; - - return val; -} - static ssize_t btrfs_feature_attr_show(struct kobject *kobj, struct kobj_attribute *a, char *buf) { int val = 0; struct btrfs_fs_info *fs_info = to_fs_info(kobj); struct btrfs_feature_attr *fa = to_btrfs_feature_attr(a); + if (fs_info) { u64 features = get_features(fs_info, fa->feature_set); if (features & fa->feature_bit) val = 1; - } else - val = can_modify_feature(fa); - - return sysfs_emit(buf, "%d\n", val); -} - -static ssize_t btrfs_feature_attr_store(struct kobject *kobj, - struct kobj_attribute *a, - const char *buf, size_t count) -{ - struct btrfs_fs_info *fs_info; - struct btrfs_feature_attr *fa = to_btrfs_feature_attr(a); - u64 features, set, clear; - unsigned long val; - int ret; - - fs_info = to_fs_info(kobj); - if (!fs_info) - return -EPERM; - - if (sb_rdonly(fs_info->sb)) - return -EROFS; - - ret = kstrtoul(skip_spaces(buf), 0, &val); - if (ret) - return ret; - - if (fa->feature_set == FEAT_COMPAT) { - set = BTRFS_FEATURE_COMPAT_SAFE_SET; - clear = BTRFS_FEATURE_COMPAT_SAFE_CLEAR; - } else if (fa->feature_set == FEAT_COMPAT_RO) { - set = BTRFS_FEATURE_COMPAT_RO_SAFE_SET; - clear = BTRFS_FEATURE_COMPAT_RO_SAFE_CLEAR; - } else { - set = BTRFS_FEATURE_INCOMPAT_SAFE_SET; - clear = BTRFS_FEATURE_INCOMPAT_SAFE_CLEAR; - } - - features = get_features(fs_info, fa->feature_set); - - /* Nothing to do */ - if ((val && (features & fa->feature_bit)) || - (!val && !(features & fa->feature_bit))) - return count; - - if ((val && !(set & fa->feature_bit)) || - (!val && !(clear & fa->feature_bit))) { - btrfs_info(fs_info, - "%sabling feature %s on mounted fs is not supported.", - val ? "En" : "Dis", fa->kobj_attr.attr.name); - return -EPERM; } - btrfs_info(fs_info, "%s %s feature flag", - val ? "Setting" : "Clearing", fa->kobj_attr.attr.name); - - spin_lock(&fs_info->super_lock); - features = get_features(fs_info, fa->feature_set); - if (val) - features |= fa->feature_bit; - else - features &= ~fa->feature_bit; - set_features(fs_info, fa->feature_set, features); - spin_unlock(&fs_info->super_lock); - - /* - * We don't want to do full transaction commit from inside sysfs - */ - set_bit(BTRFS_FS_NEED_TRANS_COMMIT, &fs_info->flags); - wake_up_process(fs_info->transaction_kthread); - - return count; + return sysfs_emit(buf, "%d\n", val); } static umode_t btrfs_feature_visible(struct kobject *kobj, @@ -269,9 +158,7 @@ static umode_t btrfs_feature_visible(struct kobject *kobj, fa = attr_to_btrfs_feature_attr(attr); features = get_features(fs_info, fa->feature_set); - if (can_modify_feature(fa)) - mode |= S_IWUSR; - else if (!(features & fa->feature_bit)) + if (!(features & fa->feature_bit)) mode = 0; } @@ -2359,9 +2246,7 @@ static ssize_t qgroup_enabled_show(struct kobject *qgroups_kobj, struct btrfs_fs_info *fs_info = to_fs_info(qgroups_kobj->parent); bool enabled; - spin_lock(&fs_info->qgroup_lock); - enabled = fs_info->qgroup_flags & BTRFS_QGROUP_STATUS_FLAG_ON; - spin_unlock(&fs_info->qgroup_lock); + enabled = test_bit(BTRFS_QGROUP_STATUS_BIT_ON, &fs_info->qgroup_flags); return sysfs_emit(buf, "%d\n", enabled); } @@ -2401,9 +2286,7 @@ static ssize_t qgroup_inconsistent_show(struct kobject *qgroups_kobj, struct btrfs_fs_info *fs_info = to_fs_info(qgroups_kobj->parent); bool inconsistent; - spin_lock(&fs_info->qgroup_lock); - inconsistent = (fs_info->qgroup_flags & BTRFS_QGROUP_STATUS_FLAG_INCONSISTENT); - spin_unlock(&fs_info->qgroup_lock); + inconsistent = test_bit(BTRFS_QGROUP_STATUS_BIT_INCONSISTENT, &fs_info->qgroup_flags); return sysfs_emit(buf, "%d\n", inconsistent); } diff --git a/fs/btrfs/tests/extent-io-tests.c b/fs/btrfs/tests/extent-io-tests.c index 23459cd4e503..cd045778400d 100644 --- a/fs/btrfs/tests/extent-io-tests.c +++ b/fs/btrfs/tests/extent-io-tests.c @@ -18,8 +18,8 @@ #define PROCESS_RELEASE (1U << 1) #define PROCESS_TEST_LOCKED (1U << 2) -static noinline int process_page_range(struct inode *inode, u64 start, u64 end, - unsigned long flags) +static noinline int process_folio_range(struct inode *inode, u64 start, u64 end, + unsigned long flags) { int ret; struct folio_batch fbatch; @@ -112,8 +112,8 @@ static int test_find_delalloc(u32 sectorsize, u32 nodesize) struct btrfs_root *root = NULL; struct inode *inode = NULL; struct extent_io_tree *tmp; - struct page *page; - struct page *locked_page = NULL; + struct folio *folio; + struct folio *locked_folio = NULL; /* In this test we need at least 2 file extents at its maximum size */ u64 max_bytes = BTRFS_MAX_EXTENT_SIZE; u64 total_dirty = 2 * max_bytes; @@ -152,23 +152,27 @@ static int test_find_delalloc(u32 sectorsize, u32 nodesize) btrfs_extent_io_tree_init(NULL, tmp, IO_TREE_SELFTEST); /* - * First go through and create and mark all of our pages dirty, we pin - * everything to make sure our pages don't get evicted and screw up our + * First go through and create and mark all of our folios dirty, we pin + * everything to make sure our folios don't get evicted and screw up our * test. */ for (pgoff_t index = 0; index < (total_dirty >> PAGE_SHIFT); index++) { - page = find_or_create_page(inode->i_mapping, index, GFP_KERNEL); - if (!page) { - test_err("failed to allocate test page"); - ret = -ENOMEM; + folio = __filemap_get_folio(inode->i_mapping, index, + FGP_LOCK | FGP_ACCESSED | FGP_CREAT, + GFP_KERNEL); + if (IS_ERR(folio)) { + test_err("failed to allocate test folio"); + ret = PTR_ERR(folio); goto out; } - SetPageDirty(page); + /* The ranges below assume page sized folios. */ + ASSERT(folio_order(folio) == 0); + folio_set_dirty(folio); if (index) { - unlock_page(page); + folio_unlock(folio); } else { - get_page(page); - locked_page = page; + folio_get(folio); + locked_folio = folio; } } @@ -179,8 +183,7 @@ static int test_find_delalloc(u32 sectorsize, u32 nodesize) btrfs_set_extent_bit(tmp, 0, sectorsize - 1, EXTENT_DELALLOC, NULL); start = 0; end = start + PAGE_SIZE - 1; - found = find_lock_delalloc_range(inode, page_folio(locked_page), &start, - &end); + found = find_lock_delalloc_range(inode, locked_folio, &start, &end); if (!found) { test_err("should have found at least one delalloc"); goto out_bits; @@ -191,8 +194,8 @@ static int test_find_delalloc(u32 sectorsize, u32 nodesize) goto out_bits; } btrfs_unlock_extent(tmp, start, end, NULL); - unlock_page(locked_page); - put_page(locked_page); + folio_unlock(locked_folio); + folio_put(locked_folio); /* * Test this scenario @@ -201,17 +204,17 @@ static int test_find_delalloc(u32 sectorsize, u32 nodesize) * |--- search ---| */ test_start = SZ_64M; - locked_page = find_lock_page(inode->i_mapping, - test_start >> PAGE_SHIFT); - if (!locked_page) { - test_err("couldn't find the locked page"); + locked_folio = filemap_lock_folio(inode->i_mapping, test_start >> PAGE_SHIFT); + if (IS_ERR(locked_folio)) { + test_err("couldn't find the locked folio"); + locked_folio = NULL; goto out_bits; } + ASSERT(folio_order(locked_folio) == 0); btrfs_set_extent_bit(tmp, sectorsize, max_bytes - 1, EXTENT_DELALLOC, NULL); start = test_start; end = start + PAGE_SIZE - 1; - found = find_lock_delalloc_range(inode, page_folio(locked_page), &start, - &end); + found = find_lock_delalloc_range(inode, locked_folio, &start, &end); if (!found) { test_err("couldn't find delalloc in our range"); goto out_bits; @@ -221,14 +224,14 @@ static int test_find_delalloc(u32 sectorsize, u32 nodesize) test_start, max_bytes - 1, start, end); goto out_bits; } - if (process_page_range(inode, start, end, - PROCESS_TEST_LOCKED | PROCESS_UNLOCK)) { - test_err("there were unlocked pages in the range"); + if (process_folio_range(inode, start, end, + PROCESS_TEST_LOCKED | PROCESS_UNLOCK)) { + test_err("there were unlocked folios in the range"); goto out_bits; } btrfs_unlock_extent(tmp, start, end, NULL); - /* locked_page was unlocked above */ - put_page(locked_page); + /* locked_folio was unlocked above */ + folio_put(locked_folio); /* * Test this scenario @@ -236,16 +239,16 @@ static int test_find_delalloc(u32 sectorsize, u32 nodesize) * |--- search ---| */ test_start = max_bytes + sectorsize; - locked_page = find_lock_page(inode->i_mapping, test_start >> - PAGE_SHIFT); - if (!locked_page) { - test_err("couldn't find the locked page"); + locked_folio = filemap_lock_folio(inode->i_mapping, test_start >> PAGE_SHIFT); + if (IS_ERR(locked_folio)) { + test_err("couldn't find the locked folio"); + locked_folio = NULL; goto out_bits; } + ASSERT(folio_order(locked_folio) == 0); start = test_start; end = start + PAGE_SIZE - 1; - found = find_lock_delalloc_range(inode, page_folio(locked_page), &start, - &end); + found = find_lock_delalloc_range(inode, locked_folio, &start, &end); if (found) { test_err("found range when we shouldn't have"); goto out_bits; @@ -265,8 +268,7 @@ static int test_find_delalloc(u32 sectorsize, u32 nodesize) btrfs_set_extent_bit(tmp, max_bytes, total_dirty - 1, EXTENT_DELALLOC, NULL); start = test_start; end = start + PAGE_SIZE - 1; - found = find_lock_delalloc_range(inode, page_folio(locked_page), &start, - &end); + found = find_lock_delalloc_range(inode, locked_folio, &start, &end); if (!found) { test_err("didn't find our range"); goto out_bits; @@ -276,38 +278,37 @@ static int test_find_delalloc(u32 sectorsize, u32 nodesize) test_start, total_dirty - 1, start, end); goto out_bits; } - if (process_page_range(inode, start, end, - PROCESS_TEST_LOCKED | PROCESS_UNLOCK)) { - test_err("pages in range were not all locked"); + if (process_folio_range(inode, start, end, + PROCESS_TEST_LOCKED | PROCESS_UNLOCK)) { + test_err("folios in range were not all locked"); goto out_bits; } btrfs_unlock_extent(tmp, start, end, NULL); /* - * Now to test where we run into a page that is no longer dirty in the + * Now to test where we run into a folio that is no longer dirty in the * range we want to find. */ - page = find_get_page(inode->i_mapping, - (max_bytes + SZ_1M) >> PAGE_SHIFT); - if (!page) { - test_err("couldn't find our page"); + folio = filemap_get_folio(inode->i_mapping, (max_bytes + SZ_1M) >> PAGE_SHIFT); + if (IS_ERR(folio)) { + test_err("couldn't find our folio"); goto out_bits; } - ClearPageDirty(page); - put_page(page); + ASSERT(folio_order(folio) == 0); + folio_clear_dirty(folio); + folio_put(folio); /* We unlocked it in the previous test */ - lock_page(locked_page); + folio_lock(locked_folio); start = test_start; end = start + PAGE_SIZE - 1; /* - * Currently if we fail to find dirty pages in the delalloc range we + * Currently if we fail to find dirty folios in the delalloc range we * will adjust max_bytes down to PAGE_SIZE and then re-search. If * this changes at any point in the future we will need to fix this * tests expected behavior. */ - found = find_lock_delalloc_range(inode, page_folio(locked_page), &start, - &end); + found = find_lock_delalloc_range(inode, locked_folio, &start, &end); if (!found) { test_err("didn't find our range"); goto out_bits; @@ -317,9 +318,9 @@ static int test_find_delalloc(u32 sectorsize, u32 nodesize) test_start, test_start + PAGE_SIZE - 1, start, end); goto out_bits; } - if (process_page_range(inode, start, end, PROCESS_TEST_LOCKED | - PROCESS_UNLOCK)) { - test_err("pages in range were not all locked"); + if (process_folio_range(inode, start, end, PROCESS_TEST_LOCKED | + PROCESS_UNLOCK)) { + test_err("folios in range were not all locked"); goto out_bits; } ret = 0; @@ -328,10 +329,10 @@ out_bits: dump_extent_io_tree(tmp); btrfs_clear_extent_bit(tmp, 0, total_dirty - 1, (unsigned)-1, NULL); out: - if (locked_page) - put_page(locked_page); - process_page_range(inode, 0, total_dirty - 1, - PROCESS_UNLOCK | PROCESS_RELEASE); + if (locked_folio) + folio_put(locked_folio); + process_folio_range(inode, 0, total_dirty - 1, + PROCESS_UNLOCK | PROCESS_RELEASE); iput(inode); out_root_info: btrfs_free_dummy_root(root); @@ -671,8 +672,9 @@ static void dump_eb_and_memory_contents(struct extent_buffer *eb, void *memory, const char *test_name) { for (int i = 0; i < eb->len; i++) { - struct page *page = folio_page(eb->folios[i >> PAGE_SHIFT], 0); - void *addr = page_address(page) + offset_in_page(i); + const unsigned long idx = get_eb_folio_index(eb, i); + void *addr = folio_address(eb->folios[idx]) + + get_eb_offset_in_folio(eb, i); if (memcmp(addr, memory + i, 1) != 0) { test_err("%s failed", test_name); @@ -687,9 +689,12 @@ static int verify_eb_and_memory(struct extent_buffer *eb, void *memory, const char *test_name) { for (int i = 0; i < (eb->len >> PAGE_SHIFT); i++) { - void *eb_addr = folio_address(eb->folios[i]); + const unsigned long offset = i << PAGE_SHIFT; + const unsigned long idx = get_eb_folio_index(eb, offset); + void *eb_addr = folio_address(eb->folios[idx]) + + get_eb_offset_in_folio(eb, offset); - if (memcmp(memory + (i << PAGE_SHIFT), eb_addr, PAGE_SIZE) != 0) { + if (memcmp(memory + offset, eb_addr, PAGE_SIZE) != 0) { dump_eb_and_memory_contents(eb, memory, test_name); return -EUCLEAN; } diff --git a/fs/btrfs/transaction.c b/fs/btrfs/transaction.c index 6802b94ed76f..84f012bfcffc 100644 --- a/fs/btrfs/transaction.c +++ b/fs/btrfs/transaction.c @@ -125,17 +125,14 @@ static const unsigned int btrfs_blocked_trans_types[TRANS_STATE_MAX] = { [TRANS_STATE_UNBLOCKED] = (__TRANS_START | __TRANS_ATTACH | __TRANS_JOIN | - __TRANS_JOIN_NOLOCK | __TRANS_JOIN_NOSTART), [TRANS_STATE_SUPER_COMMITTED] = (__TRANS_START | __TRANS_ATTACH | __TRANS_JOIN | - __TRANS_JOIN_NOLOCK | __TRANS_JOIN_NOSTART), [TRANS_STATE_COMPLETED] = (__TRANS_START | __TRANS_ATTACH | __TRANS_JOIN | - __TRANS_JOIN_NOLOCK | __TRANS_JOIN_NOSTART), }; @@ -310,12 +307,6 @@ loop: if (type == TRANS_ATTACH || type == TRANS_JOIN_NOSTART) return -ENOENT; - /* - * JOIN_NOLOCK only happens during the transaction commit, so - * it is impossible that ->running_transaction is NULL - */ - BUG_ON(type == TRANS_JOIN_NOLOCK); - cur_trans = kmalloc_obj(*cur_trans, GFP_NOFS); if (!cur_trans) return -ENOMEM; @@ -379,9 +370,8 @@ loop: INIT_LIST_HEAD(&cur_trans->dev_update_list); INIT_LIST_HEAD(&cur_trans->switch_commits); INIT_LIST_HEAD(&cur_trans->dirty_bgs); - INIT_LIST_HEAD(&cur_trans->io_bgs); INIT_LIST_HEAD(&cur_trans->dropped_roots); - mutex_init(&cur_trans->cache_write_mutex); + mutex_init(&cur_trans->dirty_bgs_update_mutex); spin_lock_init(&cur_trans->dirty_bgs_lock); INIT_LIST_HEAD(&cur_trans->deleted_bgs); spin_lock_init(&cur_trans->dropped_roots_lock); @@ -710,14 +700,8 @@ again: } /* - * If we are JOIN_NOLOCK we're already committing a transaction and - * waiting on this guy, so we don't need to do the sb_start_intwrite - * because we're already holding a ref. We need this because we could - * have raced in and did an fsync() on a file which can kick a commit - * and then we deadlock with somebody doing a freeze. - * * If we are ATTACH, it means we just want to catch the current - * transaction and commit it, so we needn't do sb_start_intwrite(). + * transaction and commit it, so we needn't do sb_start_intwrite(). */ if (type & __TRANS_FREEZABLE) sb_start_intwrite(fs_info->sb); @@ -855,12 +839,6 @@ struct btrfs_trans_handle *btrfs_join_transaction(struct btrfs_root *root) true); } -struct btrfs_trans_handle *btrfs_join_transaction_spacecache(struct btrfs_root *root) -{ - return start_transaction(root, 0, TRANS_JOIN_NOLOCK, - BTRFS_RESERVE_NO_FLUSH, true); -} - /* * Similar to regular join but it never starts a transaction when none is * running or when there's a running one at a state >= TRANS_STATE_UNBLOCKED. @@ -1363,7 +1341,6 @@ static noinline int commit_cowonly_roots(struct btrfs_trans_handle *trans) { struct btrfs_fs_info *fs_info = trans->fs_info; struct list_head *dirty_bgs = &trans->transaction->dirty_bgs; - struct list_head *io_bgs = &trans->transaction->io_bgs; struct extent_buffer *eb; int ret; @@ -1393,10 +1370,6 @@ static noinline int commit_cowonly_roots(struct btrfs_trans_handle *trans) if (ret) return ret; - ret = btrfs_setup_space_cache(trans); - if (ret) - return ret; - again: while (!list_empty(&fs_info->dirty_cowonly_roots)) { struct btrfs_root *root; @@ -1417,7 +1390,7 @@ again: if (ret) return ret; - while (!list_empty(dirty_bgs) || !list_empty(io_bgs)) { + while (!list_empty(dirty_bgs)) { ret = btrfs_write_dirty_block_groups(trans); if (ret) return ret; @@ -1890,8 +1863,7 @@ static noinline int create_pending_snapshot(struct btrfs_trans_handle *trans, goto fail; ret = btrfs_insert_dir_item(trans, &fname.disk_name, - parent_inode, &key, BTRFS_FT_DIR, - index); + parent_inode, &key, BTRFS_FT_DIR, index, NULL); if (unlikely(ret)) { btrfs_abort_transaction(trans, ret); goto fail; @@ -1990,9 +1962,7 @@ static void update_super_roots(struct btrfs_fs_info *fs_info) super->root = root_item->bytenr; super->generation = root_item->generation; super->root_level = root_item->level; - if (btrfs_test_opt(fs_info, SPACE_CACHE)) - super->cache_generation = root_item->generation; - else if (test_bit(BTRFS_FS_CLEANUP_SPACE_CACHE_V1, &fs_info->flags)) + if (test_bit(BTRFS_FS_CLEANUP_SPACE_CACHE_V1, &fs_info->flags)) super->cache_generation = 0; if (test_bit(BTRFS_FS_UPDATE_UUID_TREE_GEN, &fs_info->flags)) super->uuid_tree_generation = root_item->generation; @@ -2274,18 +2244,16 @@ int btrfs_commit_transaction(struct btrfs_trans_handle *trans) if (!test_bit(BTRFS_TRANS_DIRTY_BG_RUN, &cur_trans->flags)) { bool run_it = false; - /* this mutex is also taken before trying to set - * block groups readonly. We need to make sure - * that nobody has set a block group readonly - * after a extents from that block group have been - * allocated for cache files. btrfs_set_block_group_ro - * will wait for the transaction to commit if it - * finds BTRFS_TRANS_DIRTY_BG_RUN set. + /* + * This mutex is also taken before trying to set block groups + * readonly. btrfs_inc_block_group_ro() will wait for the + * transaction to commit if it finds BTRFS_TRANS_DIRTY_BG_RUN + * set. * * The BTRFS_TRANS_DIRTY_BG_RUN flag is also used to make sure - * only one process starts all the block group IO. It wouldn't - * hurt to have more than one go through, but there's no - * real advantage to it either. + * only one process starts all the block group item updates. It + * wouldn't hurt to have more than one go through, but there's + * no real advantage to it either. */ mutex_lock(&fs_info->ro_block_group_mutex); if (!test_and_set_bit(BTRFS_TRANS_DIRTY_BG_RUN, @@ -2519,10 +2487,7 @@ int btrfs_commit_transaction(struct btrfs_trans_handle *trans) if (unlikely(ret)) goto unlock_reloc; - /* - * The tasks which save the space cache and inode cache may also - * update ->aborted, check it. - */ + /* Other tasks may also have updated ->aborted, check it. */ if (TRANS_ABORTED(cur_trans)) { ret = cur_trans->aborted; goto unlock_reloc; @@ -2543,7 +2508,6 @@ int btrfs_commit_transaction(struct btrfs_trans_handle *trans) switch_commit_roots(trans); ASSERT(list_empty(&cur_trans->dirty_bgs)); - ASSERT(list_empty(&cur_trans->io_bgs)); update_super_roots(fs_info); btrfs_set_super_log_root(fs_info->super_copy, 0); diff --git a/fs/btrfs/transaction.h b/fs/btrfs/transaction.h index 3a57f227b5ed..68c724e70809 100644 --- a/fs/btrfs/transaction.h +++ b/fs/btrfs/transaction.h @@ -47,7 +47,6 @@ enum btrfs_trans_state { #define BTRFS_TRANS_HAVE_FREE_BGS 0 #define BTRFS_TRANS_DIRTY_BG_RUN 1 -#define BTRFS_TRANS_CACHE_ENOSPC 2 struct btrfs_transaction { u64 transid; @@ -78,32 +77,15 @@ struct btrfs_transaction { struct list_head dev_update_list; struct list_head switch_commits; struct list_head dirty_bgs; - - /* - * There is no explicit lock which protects io_bgs, rather its - * consistency is implied by the fact that all the sites which modify - * it do so under some form of transaction critical section, namely: - * - * - btrfs_start_dirty_block_groups - This function can only ever be - * run by one of the transaction committers. Refer to - * BTRFS_TRANS_DIRTY_BG_RUN usage in btrfs_commit_transaction - * - * - btrfs_write_dirty_blockgroups - this is called by - * commit_cowonly_roots from transaction critical section - * (TRANS_STATE_COMMIT_DOING) - * - * - btrfs_cleanup_dirty_bgs - called on transaction abort - */ - struct list_head io_bgs; struct list_head dropped_roots; struct extent_io_tree pinned_extents; /* - * we need to make sure block group deletion doesn't race with - * free space cache writeout. This mutex keeps them from stomping - * on each other + * We need to make sure block group deletion doesn't race with the + * dirty block group item updates done outside the commit critical + * section. This mutex keeps them from stomping on each other. */ - struct mutex cache_write_mutex; + struct mutex dirty_bgs_update_mutex; spinlock_t dirty_bgs_lock; /* Protected by spin lock fs_info->unused_bgs_lock. */ struct list_head deleted_bgs; @@ -124,7 +106,6 @@ enum { ENUM_BIT(__TRANS_START), ENUM_BIT(__TRANS_ATTACH), ENUM_BIT(__TRANS_JOIN), - ENUM_BIT(__TRANS_JOIN_NOLOCK), ENUM_BIT(__TRANS_DUMMY), ENUM_BIT(__TRANS_JOIN_NOSTART), }; @@ -132,7 +113,6 @@ enum { #define TRANS_START (__TRANS_START | __TRANS_FREEZABLE) #define TRANS_ATTACH (__TRANS_ATTACH) #define TRANS_JOIN (__TRANS_JOIN | __TRANS_FREEZABLE) -#define TRANS_JOIN_NOLOCK (__TRANS_JOIN_NOLOCK) #define TRANS_JOIN_NOSTART (__TRANS_JOIN_NOSTART) #define TRANS_EXTWRITERS (__TRANS_START | __TRANS_ATTACH) @@ -288,7 +268,7 @@ do { \ * Call btrfs_abort_transaction() as early as possible when an error condition * is detected, that way the exact stack trace is reported for some errors. * - * Error number must be negative as it encodes wheather it's the first abort. + * Error number must be negative as it encodes whether it's the first abort. */ #define btrfs_abort_transaction(trans, error) \ do { \ @@ -312,7 +292,6 @@ struct btrfs_trans_handle *btrfs_start_transaction_fallback_global_rsv( struct btrfs_root *root, unsigned int num_items); struct btrfs_trans_handle *btrfs_join_transaction(struct btrfs_root *root); -struct btrfs_trans_handle *btrfs_join_transaction_spacecache(struct btrfs_root *root); struct btrfs_trans_handle *btrfs_join_transaction_nostart(struct btrfs_root *root); struct btrfs_trans_handle *btrfs_attach_transaction(struct btrfs_root *root); struct btrfs_trans_handle *btrfs_attach_transaction_barrier( diff --git a/fs/btrfs/tree-checker.c b/fs/btrfs/tree-checker.c index 4b1e47173c63..9cd79d97b9b5 100644 --- a/fs/btrfs/tree-checker.c +++ b/fs/btrfs/tree-checker.c @@ -108,13 +108,14 @@ static void file_extent_err(const struct extent_buffer *eb, int slot, */ #define CHECK_FE_ALIGNED(leaf, slot, fi, name, alignment) \ ({ \ - if (unlikely(!IS_ALIGNED(btrfs_file_extent_##name((leaf), (fi)), \ - (alignment)))) \ + const u64 val = btrfs_file_extent_##name((leaf), (fi)); \ + const bool not_aligned = !IS_ALIGNED(val, (alignment)); \ + \ + if (unlikely(not_aligned)) \ file_extent_err((leaf), (slot), \ "invalid %s for file extent, have %llu, should be aligned to %u", \ - (#name), btrfs_file_extent_##name((leaf), (fi)), \ - (alignment)); \ - (!IS_ALIGNED(btrfs_file_extent_##name((leaf), (fi)), (alignment))); \ + (#name), val, (alignment)); \ + not_aligned; \ }) static u64 file_extent_end(struct extent_buffer *leaf, @@ -163,6 +164,12 @@ static void dir_item_err(const struct extent_buffer *eb, int slot, va_end(args); } +/* Record info for the last hit inode. */ +struct saved_inode_info { + u64 ino; + u32 mode; +}; + /* * This functions checks prev_key->objectid, to ensure current key and prev_key * share the same objectid as inode number. @@ -204,15 +211,41 @@ static bool check_prev_ino(struct extent_buffer *leaf, prev_key->objectid, key->objectid); return false; } + +static bool can_have_extent_data(struct extent_buffer *leaf, + struct btrfs_key *key, int slot, u8 fi_type, + const struct saved_inode_info *inode_info) +{ + /* No inode item in this leaf. */ + if (inode_info->ino != key->objectid) + return true; + if (S_ISREG(inode_info->mode)) + return true; + if (S_ISLNK(inode_info->mode)) { + /* For a symlink, the file extent item should always be inlined. */ + if (unlikely(fi_type != BTRFS_FILE_EXTENT_INLINE)) + return false; + return true; + } + + /* + * The rest are special files, e.g. block/FIFO files, which cannot + * have any file extent. + */ + return false; +} + static int check_extent_data_item(struct extent_buffer *leaf, struct btrfs_key *key, int slot, - struct btrfs_key *prev_key) + struct btrfs_key *prev_key, + const struct saved_inode_info *inode_info) { struct btrfs_fs_info *fs_info = leaf->fs_info; struct btrfs_file_extent_item *fi; u32 sectorsize = fs_info->sectorsize; u32 item_size = btrfs_item_size(leaf, slot); u64 extent_end; + u8 fi_type; if (unlikely(!IS_ALIGNED(key->offset, sectorsize))) { file_extent_err(leaf, slot, @@ -243,12 +276,18 @@ static int check_extent_data_item(struct extent_buffer *leaf, SZ_4K); return -EUCLEAN; } - if (unlikely(btrfs_file_extent_type(leaf, fi) >= - BTRFS_NR_FILE_EXTENT_TYPES)) { + fi_type = btrfs_file_extent_type(leaf, fi); + if (unlikely(fi_type >= BTRFS_NR_FILE_EXTENT_TYPES)) { file_extent_err(leaf, slot, "invalid type for file extent, have %u expect range [0, %u]", - btrfs_file_extent_type(leaf, fi), - BTRFS_NR_FILE_EXTENT_TYPES - 1); + fi_type, BTRFS_NR_FILE_EXTENT_TYPES - 1); + return -EUCLEAN; + } + + if (unlikely(!can_have_extent_data(leaf, key, slot, fi_type, inode_info))) { + file_extent_err(leaf, slot, + "unexpected file extent item type %u for inode mode 0%o", + fi_type, inode_info->mode); return -EUCLEAN; } @@ -270,7 +309,8 @@ static int check_extent_data_item(struct extent_buffer *leaf, btrfs_file_extent_encryption(leaf, fi)); return -EUCLEAN; } - if (btrfs_file_extent_type(leaf, fi) == BTRFS_FILE_EXTENT_INLINE) { + + if (fi_type == BTRFS_FILE_EXTENT_INLINE) { /* Inline extent must have 0 as key offset */ if (unlikely(key->offset)) { file_extent_err(leaf, slot, @@ -1081,9 +1121,9 @@ int btrfs_check_chunk_valid(const struct btrfs_fs_info *fs_info, return -EUCLEAN; } - if (!remapped && - !valid_stripe_count(type & BTRFS_BLOCK_GROUP_PROFILE_MASK, - num_stripes, sub_stripes)) { + if (unlikely(!remapped && + !valid_stripe_count(type & BTRFS_BLOCK_GROUP_PROFILE_MASK, + num_stripes, sub_stripes))) { chunk_err(fs_info, leaf, chunk, logical, "invalid num_stripes:sub_stripes %u:%u for profile %llu", num_stripes, sub_stripes, @@ -1206,7 +1246,8 @@ static int check_dev_item(struct extent_buffer *leaf, } static int check_inode_item(struct extent_buffer *leaf, - struct btrfs_key *key, int slot) + struct btrfs_key *key, int slot, + struct saved_inode_info *inode_info) { struct btrfs_fs_info *fs_info = leaf->fs_info; struct btrfs_inode_item *iitem; @@ -1291,6 +1332,8 @@ static int check_inode_item(struct extent_buffer *leaf, ro_flags); return -EUCLEAN; } + inode_info->ino = key->objectid; + inode_info->mode = mode; return 0; } @@ -2348,14 +2391,15 @@ static int check_free_space_bitmap(struct extent_buffer *leaf, static enum btrfs_tree_block_status check_leaf_item(struct extent_buffer *leaf, struct btrfs_key *key, int slot, - struct btrfs_key *prev_key) + struct btrfs_key *prev_key, + struct saved_inode_info *inode_info) { int ret = 0; struct btrfs_chunk *chunk; switch (key->type) { case BTRFS_EXTENT_DATA_KEY: - ret = check_extent_data_item(leaf, key, slot, prev_key); + ret = check_extent_data_item(leaf, key, slot, prev_key, inode_info); break; case BTRFS_EXTENT_CSUM_KEY: ret = check_csum_item(leaf, key, slot, prev_key); @@ -2385,7 +2429,7 @@ static enum btrfs_tree_block_status check_leaf_item(struct extent_buffer *leaf, ret = check_dev_extent_item(leaf, key, slot, prev_key); break; case BTRFS_INODE_ITEM_KEY: - ret = check_inode_item(leaf, key, slot); + ret = check_inode_item(leaf, key, slot, inode_info); break; case BTRFS_ROOT_ITEM_KEY: ret = check_root_item(leaf, key, slot); @@ -2433,6 +2477,7 @@ static enum btrfs_tree_block_status check_leaf_item(struct extent_buffer *leaf, enum btrfs_tree_block_status __btrfs_check_leaf(struct extent_buffer *leaf) { struct btrfs_fs_info *fs_info = leaf->fs_info; + struct saved_inode_info inode_info = { 0 }; /* No valid key type is 0, so all key should be larger than this key */ struct btrfs_key prev_key = {0, 0, 0}; struct btrfs_key key; @@ -2568,7 +2613,7 @@ enum btrfs_tree_block_status __btrfs_check_leaf(struct extent_buffer *leaf) } /* Check if the item size and content meet other criteria. */ - ret = check_leaf_item(leaf, &key, slot, &prev_key); + ret = check_leaf_item(leaf, &key, slot, &prev_key, &inode_info); if (unlikely(ret != BTRFS_TREE_BLOCK_CLEAN)) return ret; diff --git a/fs/btrfs/tree-log.c b/fs/btrfs/tree-log.c index a00094604e54..1d9dfb63fc5a 100644 --- a/fs/btrfs/tree-log.c +++ b/fs/btrfs/tree-log.c @@ -503,7 +503,7 @@ static int overwrite_item(struct walk_control *wc) btrfs_release_path(wc->subvol_path); return 0; } - src_copy = kmalloc(item_size, GFP_NOFS); + src_copy = kvmalloc(item_size, GFP_NOFS); if (!src_copy) { btrfs_abort_log_replay(wc, -ENOMEM, "failed to allocate memory for log leaf item"); @@ -514,7 +514,7 @@ static int overwrite_item(struct walk_control *wc) dst_ptr = btrfs_item_ptr_offset(dst_eb, dst_slot); ret = memcmp_extent_buffer(dst_eb, src_copy, dst_ptr, item_size); - kfree(src_copy); + kvfree(src_copy); /* * they have the same contents, just return, this saves * us from cowing blocks in the destination tree and doing @@ -1683,7 +1683,7 @@ static noinline int add_inode_ref(struct walk_control *wc) } /* insert our name */ - ret = btrfs_add_link(trans, dir, inode, &name, false, ref_index); + ret = btrfs_add_link(trans, dir, inode, &name, false, ref_index, NULL); if (ret) { btrfs_abort_log_replay(wc, ret, "failed to add link for inode %llu in dir %llu ref_index %llu name %.*s root %llu", @@ -2031,7 +2031,7 @@ static noinline int insert_one_name(struct btrfs_trans_handle *trans, return PTR_ERR(dir); } - ret = btrfs_add_link(trans, dir, inode, name, true, index); + ret = btrfs_add_link(trans, dir, inode, name, true, index, NULL); /* FIXME, put inode into FIXUP list */ @@ -5350,7 +5350,7 @@ static int btrfs_log_changed_extents(struct btrfs_trans_handle *trans, * have a bunch of extents we just want to commit since it will * be faster. */ - if (++num > 32768) { + if (++num > SZ_16K) { list_del_init(&tree->modified_extents); ret = -EFBIG; goto process; @@ -5368,7 +5368,6 @@ static int btrfs_log_changed_extents(struct btrfs_trans_handle *trans, refcount_inc(&em->refs); em->flags |= EXTENT_FLAG_LOGGING; list_add_tail(&em->list, &extents); - num++; } list_sort(NULL, &extents, extent_cmp); diff --git a/fs/btrfs/verity.c b/fs/btrfs/verity.c index 600337a84fbe..83dc7dee14cf 100644 --- a/fs/btrfs/verity.c +++ b/fs/btrfs/verity.c @@ -286,21 +286,17 @@ static int write_key_bytes(struct btrfs_inode *inode, u8 key_type, u64 offset, * @dest: Buffer to read into. This parameter has slightly tricky * semantics. If it is NULL, the function will not do any copying * and will just return the size of all the items up to len bytes. - * If dest_page is passed, then the function will kmap_local the - * page and ignore dest, but it must still be non-NULL to avoid the - * counting-only behavior. * @len: length in bytes to read - * @dest_folio: copy into this folio instead of the dest buffer * * Helper function to read items from the btree. This returns the number of * bytes read or < 0 for errors. We can return short reads if the items don't * exist on disk or aren't big enough to fill the desired length. Supports - * reading into a provided buffer (dest) or into the page cache + * reading into a provided buffer (dest). * * Returns number of bytes read or a negative error code on failure. */ static int read_key_bytes(struct btrfs_inode *inode, u8 key_type, u64 offset, - char *dest, u64 len, struct folio *dest_folio) + char *dest, u64 len) { BTRFS_PATH_AUTO_FREE(path); struct btrfs_root *root = inode->root; @@ -320,7 +316,11 @@ static int read_key_bytes(struct btrfs_inode *inode, u8 key_type, u64 offset, if (!path) return -ENOMEM; - if (dest_folio) + /* + * Merkle items can be large and split across multiple items, so enable + * readahead for such cases. + */ + if (key_type == BTRFS_VERITY_MERKLE_ITEM_KEY) path->reada = READA_FORWARD; key.objectid = btrfs_ino(inode); @@ -364,7 +364,7 @@ static int read_key_bytes(struct btrfs_inode *inode, u8 key_type, u64 offset, break; } - /* desc = NULL to just sum all the item lengths */ + /* dest == NULL to just sum all the item lengths */ if (!dest) copy_end = item_end; else @@ -377,16 +377,10 @@ static int read_key_bytes(struct btrfs_inode *inode, u8 key_type, u64 offset, copy_offset = offset - key.offset; if (dest) { - if (dest_folio) - kaddr = kmap_local_folio(dest_folio, 0); - data = btrfs_item_ptr(leaf, path->slots[0], void); read_extent_buffer(leaf, kaddr + dest_offset, (unsigned long)data + copy_offset, copy_bytes); - - if (dest_folio) - kunmap_local(kaddr); } offset += copy_bytes; @@ -677,7 +671,7 @@ int btrfs_get_verity_descriptor(struct inode *inode, void *buf, size_t buf_size) memset(&item, 0, sizeof(item)); ret = read_key_bytes(BTRFS_I(inode), BTRFS_VERITY_DESC_ITEM_KEY, 0, - (char *)&item, sizeof(item), NULL); + (char *)&item, sizeof(item)); if (ret < 0) return ret; @@ -694,7 +688,7 @@ int btrfs_get_verity_descriptor(struct inode *inode, void *buf, size_t buf_size) return -ERANGE; ret = read_key_bytes(BTRFS_I(inode), BTRFS_VERITY_DESC_ITEM_KEY, 1, - buf, buf_size, NULL); + buf, buf_size); if (ret < 0) return ret; if (ret != true_size) @@ -720,6 +714,7 @@ static struct page *btrfs_read_merkle_tree_page(struct inode *inode, struct folio *folio; u64 off = (u64)index << PAGE_SHIFT; loff_t merkle_pos = merkle_file_pos(inode); + void *kaddr; int ret; if (merkle_pos < 0) @@ -763,6 +758,7 @@ again: } read_folio: + kaddr = kmap_local_folio(folio, 0); /* * Merkle item keys are indexed from byte 0 in the merkle tree. * They have the form: @@ -770,7 +766,8 @@ read_folio: * [ inode objectid, BTRFS_MERKLE_ITEM_KEY, offset in bytes ] */ ret = read_key_bytes(BTRFS_I(inode), BTRFS_VERITY_MERKLE_ITEM_KEY, off, - folio_address(folio), PAGE_SIZE, folio); + kaddr, PAGE_SIZE); + kunmap_local(kaddr); if (ret < 0) { folio_unlock(folio); folio_put(folio); diff --git a/fs/btrfs/volumes.c b/fs/btrfs/volumes.c index 85ea9c5d4536..7c040f22dbc3 100644 --- a/fs/btrfs/volumes.c +++ b/fs/btrfs/volumes.c @@ -403,7 +403,7 @@ static struct btrfs_fs_devices *alloc_fs_devices(const u8 *fsid) return fs_devs; } -static void btrfs_free_device(struct btrfs_device *device) +void btrfs_free_device(struct btrfs_device *device) { WARN_ON(!list_empty(&device->post_commit_list)); /* @@ -1358,16 +1358,14 @@ int btrfs_open_devices(struct btrfs_fs_devices *fs_devices, void btrfs_release_disk_super(struct btrfs_super_block *super) { - struct page *page = virt_to_page(super); - - put_page(page); + folio_put(virt_to_folio(super)); } struct btrfs_super_block *btrfs_read_disk_super(struct block_device *bdev, int copy_num, bool drop_cache) { struct btrfs_super_block *super; - struct page *page; + struct folio *folio; u64 bytenr, bytenr_orig; struct address_space *mapping = bdev->bd_mapping; int ret; @@ -1388,7 +1386,7 @@ struct btrfs_super_block *btrfs_read_disk_super(struct block_device *bdev, ASSERT(copy_num == 0); /* - * Drop the page of the primary superblock, so later read will + * Drop the folio of the primary superblock, so later read will * always read from the device. */ invalidate_inode_pages2_range(mapping, bytenr >> PAGE_SHIFT, @@ -1396,12 +1394,12 @@ struct btrfs_super_block *btrfs_read_disk_super(struct block_device *bdev, } filemap_invalidate_lock_shared(mapping); - page = read_cache_page_gfp(mapping, bytenr >> PAGE_SHIFT, GFP_NOFS); + folio = mapping_read_folio_gfp(mapping, bytenr >> PAGE_SHIFT, GFP_NOFS); filemap_invalidate_unlock_shared(mapping); - if (IS_ERR(page)) - return ERR_CAST(page); + if (IS_ERR(folio)) + return ERR_CAST(folio); - super = page_address(page); + super = folio_address(folio) + offset_in_folio(folio, bytenr); if (btrfs_super_magic(super) != BTRFS_MAGIC || btrfs_super_bytenr(super) != bytenr_orig) { btrfs_release_disk_super(super); @@ -2814,6 +2812,41 @@ static void btrfs_setup_sprout(struct btrfs_fs_info *fs_info, btrfs_set_super_flags(disk_super, super_flags); } +static void btrfs_rollback_sprout(struct btrfs_fs_info *fs_info, + struct btrfs_fs_devices *seed_devices) +{ + struct btrfs_fs_devices *fs_devices = fs_info->fs_devices; + struct btrfs_super_block *disk_super = fs_info->super_copy; + struct btrfs_device *device; + u64 super_flags; + + lockdep_assert_held(&uuid_mutex); + lockdep_assert_held(&fs_devices->device_list_mutex); + + list_del_init(&seed_devices->seed_list); + list_splice_init_rcu(&seed_devices->devices, &fs_devices->devices, synchronize_rcu); + list_for_each_entry(device, &fs_devices->devices, dev_list) { + device->fs_devices = fs_devices; + } + + fs_devices->seeding = true; + fs_devices->num_devices = seed_devices->num_devices; + fs_devices->open_devices = seed_devices->open_devices; + fs_devices->missing_devices = seed_devices->missing_devices; + fs_devices->rotating = seed_devices->rotating; + fs_devices->latest_dev = seed_devices->latest_dev; + + memcpy(fs_devices->fsid, seed_devices->fsid, BTRFS_FSID_SIZE); + memcpy(fs_devices->metadata_uuid, seed_devices->metadata_uuid, BTRFS_FSID_SIZE); + memcpy(disk_super->fsid, seed_devices->fsid, BTRFS_FSID_SIZE); + + super_flags = (btrfs_super_flags(disk_super) | BTRFS_SUPER_FLAG_SEEDING); + btrfs_set_super_flags(disk_super, super_flags); + + seed_devices->opened = 0; + free_fs_devices(seed_devices); +} + /* * Store the expected generation for seed devices in device items. */ @@ -3165,6 +3198,8 @@ error_sysfs: orig_super_total_bytes); btrfs_set_super_num_devices(fs_info->super_copy, orig_super_num_devices); + if (seeding_dev) + btrfs_rollback_sprout(fs_info, seed_devices); btrfs_update_per_profile_avail(fs_info); mutex_unlock(&fs_info->chunk_mutex); mutex_unlock(&fs_info->fs_devices->device_list_mutex); diff --git a/fs/btrfs/volumes.h b/fs/btrfs/volumes.h index 0415d74cad9b..337d7007d9e2 100644 --- a/fs/btrfs/volumes.h +++ b/fs/btrfs/volumes.h @@ -799,6 +799,7 @@ void btrfs_rm_dev_replace_remove_srcdev(struct btrfs_device *srcdev); void btrfs_rm_dev_replace_free_srcdev(struct btrfs_device *srcdev); void btrfs_destroy_dev_replace_tgtdev(struct btrfs_device *tgtdev, bool allow_freeze); +void btrfs_free_device(struct btrfs_device *device); unsigned long btrfs_full_stripe_len(struct btrfs_fs_info *fs_info, u64 logical); u64 btrfs_calc_stripe_length(const struct btrfs_chunk_map *map); diff --git a/fs/btrfs/xattr.c b/fs/btrfs/xattr.c index ab55d10bd71f..a06420b9c662 100644 --- a/fs/btrfs/xattr.c +++ b/fs/btrfs/xattr.c @@ -353,7 +353,7 @@ static int btrfs_xattr_handler_get(const struct xattr_handler *handler, } static int btrfs_xattr_handler_set(const struct xattr_handler *handler, - struct mnt_idmap *idmap, + const struct mnt_idmap *idmap, struct dentry *unused, struct inode *inode, const char *name, const void *buffer, size_t size, int flags) @@ -395,7 +395,7 @@ static int btrfs_xattr_handler_get_security(const struct xattr_handler *handler, } static int btrfs_xattr_handler_set_security(const struct xattr_handler *handler, - struct mnt_idmap *idmap, + const struct mnt_idmap *idmap, struct dentry *unused, struct inode *inode, const char *name, @@ -413,7 +413,7 @@ static int btrfs_xattr_handler_set_security(const struct xattr_handler *handler, } static int btrfs_xattr_handler_set_prop(const struct xattr_handler *handler, - struct mnt_idmap *idmap, + const struct mnt_idmap *idmap, struct dentry *unused, struct inode *inode, const char *name, const void *value, size_t size, int flags) diff --git a/fs/btrfs/zlib.c b/fs/btrfs/zlib.c index 486b52db583e..e995d0b1452b 100644 --- a/fs/btrfs/zlib.c +++ b/fs/btrfs/zlib.c @@ -49,7 +49,7 @@ void zlib_free_workspace(struct list_head *ws) struct workspace *workspace = list_entry(ws, struct workspace, list); kvfree(workspace->strm.workspace); - kfree(workspace->buf); + kvfree(workspace->buf); kfree(workspace); } @@ -84,13 +84,13 @@ struct list_head *zlib_alloc_workspace(struct btrfs_fs_info *fs_info, unsigned i workspace->level = level; workspace->buf = NULL; if (need_special_buffer(fs_info)) { - workspace->buf = kmalloc(ZLIB_DFLTCC_BUF_SIZE, - __GFP_NOMEMALLOC | __GFP_NORETRY | - __GFP_NOWARN | GFP_NOIO); + workspace->buf = kvmalloc(ZLIB_DFLTCC_BUF_SIZE, + __GFP_NOMEMALLOC | __GFP_NORETRY | + __GFP_NOWARN | GFP_NOIO); workspace->buf_size = ZLIB_DFLTCC_BUF_SIZE; } if (!workspace->buf) { - workspace->buf = kmalloc(fs_info->sectorsize, GFP_KERNEL); + workspace->buf = kvmalloc(fs_info->sectorsize, GFP_KERNEL); workspace->buf_size = fs_info->sectorsize; } if (!workspace->strm.workspace || !workspace->buf) diff --git a/fs/btrfs/zoned.c b/fs/btrfs/zoned.c index 08a15465a087..c04a9955f2d3 100644 --- a/fs/btrfs/zoned.c +++ b/fs/btrfs/zoned.c @@ -123,24 +123,24 @@ static int sb_write_pointer(struct block_device *bdev, struct blk_zone *zones, } else if (full[0] && full[1]) { /* Compare two super blocks */ struct address_space *mapping = bdev->bd_mapping; - struct page *page[BTRFS_NR_SB_LOG_ZONES]; struct btrfs_super_block *super[BTRFS_NR_SB_LOG_ZONES]; for (int i = 0; i < BTRFS_NR_SB_LOG_ZONES; i++) { u64 zone_end = (zones[i].start + zones[i].capacity) << SECTOR_SHIFT; u64 bytenr = ALIGN_DOWN(zone_end, BTRFS_SUPER_INFO_SIZE) - BTRFS_SUPER_INFO_SIZE; + struct folio *folio; filemap_invalidate_lock_shared(mapping); - page[i] = read_cache_page_gfp(mapping, - bytenr >> PAGE_SHIFT, GFP_NOFS); + folio = mapping_read_folio_gfp(mapping, bytenr >> PAGE_SHIFT, + GFP_NOFS); filemap_invalidate_unlock_shared(mapping); - if (IS_ERR(page[i])) { + if (IS_ERR(folio)) { if (i == 1) btrfs_release_disk_super(super[0]); - return PTR_ERR(page[i]); + return PTR_ERR(folio); } - super[i] = page_address(page[i]); + super[i] = folio_address(folio) + offset_in_folio(folio, bytenr); } if (btrfs_super_generation(super[0]) > @@ -804,15 +804,6 @@ int btrfs_check_mountopts_zoned(const struct btrfs_fs_info *info, if (!btrfs_is_zoned(info)) return 0; - /* - * Space cache writing is not COWed. Disable that to avoid write errors - * in sequential zones. - */ - if (btrfs_raw_test_opt(*mount_opt, SPACE_CACHE)) { - btrfs_err(info, "zoned: space cache v1 is not supported"); - return -EINVAL; - } - if (btrfs_raw_test_opt(*mount_opt, NODATACOW)) { btrfs_err(info, "zoned: NODATACOW not supported"); return -EINVAL; diff --git a/fs/btrfs/zstd.c b/fs/btrfs/zstd.c index 58d9ff76fe07..cc92d0b1b948 100644 --- a/fs/btrfs/zstd.c +++ b/fs/btrfs/zstd.c @@ -373,7 +373,7 @@ void zstd_free_workspace(struct list_head *ws) struct workspace *workspace = list_entry(ws, struct workspace, list); kvfree(workspace->mem); - kfree(workspace->buf); + kvfree(workspace->buf); kfree(workspace); } @@ -391,7 +391,7 @@ struct list_head *zstd_alloc_workspace(struct btrfs_fs_info *fs_info, int level) workspace->req_level = level; workspace->last_used = jiffies; workspace->mem = kvmalloc(workspace->size, GFP_KERNEL | __GFP_NOWARN); - workspace->buf = kmalloc(fs_info->sectorsize, GFP_KERNEL); + workspace->buf = kvmalloc(fs_info->sectorsize, GFP_KERNEL); if (!workspace->mem || !workspace->buf) goto fail; @@ -589,10 +589,48 @@ out: return ret; } +/* + * Map the destination for the next chunk of output. + * + * @decompressed is the offset of the next output byte inside the fully + * decompressed extent. If that offset has reached the current destination + * segment, its page-bounded bio_vec is kmapped so that zstd can write into the + * page cache directly, and the number of bytes writable there is returned. + * Otherwise @kaddr_ret is set to NULL and the number of bytes to skip before + * that segment is returned. This covers both the initial prefix and gaps in + * the destination bio. + */ +static u32 zstd_map_dest(struct compressed_bio *cb, u32 decompressed, + void **kaddr_ret) +{ + struct bio *orig_bio = &cb->orig_bbio->bio; + struct bio_vec bvec; + u32 bvec_offset; + u32 off; + + bvec = bio_iter_iovec(orig_bio, orig_bio->bi_iter); + /* + * cb->start may underflow, but subtracting that value can still give us + * the correct offset inside the full decompressed extent. + */ + bvec_offset = page_offset(bvec.bv_page) + bvec.bv_offset - cb->start; + + if (decompressed < bvec_offset) { + *kaddr_ret = NULL; + return bvec_offset - decompressed; + } + + off = decompressed - bvec_offset; + ASSERT(off < bvec.bv_len); + *kaddr_ret = bvec_kmap_local(&bvec) + off; + return bvec.bv_len - off; +} + int zstd_decompress_bio(struct list_head *ws, struct compressed_bio *cb) { struct btrfs_fs_info *fs_info = cb_to_fs_info(cb); struct workspace *workspace = list_entry(ws, struct workspace, list); + struct bio *orig_bio = &cb->orig_bbio->bio; struct folio_iter fi; size_t srclen = bio_get_size(&cb->bbio.bio); zstd_dstream *stream; @@ -600,7 +638,6 @@ int zstd_decompress_bio(struct list_head *ws, struct compressed_bio *cb) const unsigned int min_folio_size = btrfs_min_folio_size(fs_info); unsigned long folio_in_index = 0; unsigned long total_folios_in = DIV_ROUND_UP(srclen, min_folio_size); - unsigned long buf_start; unsigned long total_out = 0; bio_first_folio(&fi, &cb->bbio.bio, 0); @@ -624,15 +661,26 @@ int zstd_decompress_bio(struct list_head *ws, struct compressed_bio *cb) workspace->in_buf.pos = 0; workspace->in_buf.size = min_t(size_t, srclen, min_folio_size); - workspace->out_buf.dst = workspace->buf; - workspace->out_buf.pos = 0; - workspace->out_buf.size = fs_info->sectorsize; - - while (1) { + while (orig_bio->bi_iter.bi_size) { size_t ret2; + void *kaddr; + u32 dstlen; + + dstlen = zstd_map_dest(cb, total_out, &kaddr); + if (kaddr) { + workspace->out_buf.dst = kaddr; + workspace->out_buf.size = dstlen; + } else { + workspace->out_buf.dst = workspace->buf; + workspace->out_buf.size = min_t(u32, dstlen, + fs_info->sectorsize); + } + workspace->out_buf.pos = 0; ret2 = zstd_decompress_stream(stream, &workspace->out_buf, &workspace->in_buf); + if (kaddr) + kunmap_local(kaddr); if (unlikely(zstd_is_error(ret2))) { struct btrfs_inode *inode = cb->bbio.inode; @@ -643,14 +691,9 @@ int zstd_decompress_bio(struct list_head *ws, struct compressed_bio *cb) ret = -EIO; goto done; } - buf_start = total_out; total_out += workspace->out_buf.pos; - workspace->out_buf.pos = 0; - - ret = btrfs_decompress_buf2page(workspace->out_buf.dst, - total_out - buf_start, cb, buf_start); - if (ret == 0) - break; + if (kaddr) + bio_advance(orig_bio, workspace->out_buf.pos); if (workspace->in_buf.pos >= srclen) break; diff --git a/fs/buffer.c b/fs/buffer.c index 020af5dbe2d0..5d3852b726d6 100644 --- a/fs/buffer.c +++ b/fs/buffer.c @@ -203,11 +203,10 @@ void bh_end_write(struct bio *bio) bool success = bio_endio_bh(bio, &bh); if (success) { - set_buffer_uptodate(bh); + clear_buffer_write_io_error(bh); } else { buffer_io_error(bh, ", lost sync page write"); mark_buffer_write_io_error(bh); - clear_buffer_uptodate(bh); } unlock_buffer(bh); } @@ -265,7 +264,7 @@ __find_get_block_slow(struct block_device *bdev, sector_t block, bool atomic) bh = bh->b_this_page; } while (bh != head); - /* we might be here because some of the buffers on this page are + /* we might be here because some of the buffers on this folio are * not mapped. This is due to various races between * file io on the block device and getblk. It gets dealt with * elsewhere, don't buffer_error if we had some unmapped buffers @@ -311,7 +310,7 @@ static void end_buffer_async_read(struct buffer_head *bh, int uptodate) /* * Be _very_ careful from here on. Bad things can happen if * two buffer heads end IO at almost the same time and both - * decide that the page is now completely done. + * decide that the folio is now completely done. */ first = folio_buffers(folio); spin_lock_irqsave(&first->b_uptodate_lock, flags); @@ -408,11 +407,10 @@ void bh_end_async_write(struct bio *bio) folio = bh->b_folio; if (success) { - set_buffer_uptodate(bh); + clear_buffer_write_io_error(bh); } else { buffer_io_error(bh, ", lost async page write"); mark_buffer_write_io_error(bh); - clear_buffer_uptodate(bh); } first = folio_buffers(folio); @@ -520,8 +518,8 @@ EXPORT_SYMBOL_GPL(mmb_has_buffers); * * Do this in two main stages: first we copy dirty buffers to a * temporary inode list, queueing the writes as we go. Then we clean - * up, waiting for those writes to complete. mark_buffer_dirty_inode() - * doesn't touch b_assoc_buffers list if b_mmb is not NULL so we are sure the + * up, waiting for those writes to complete. mmb_mark_buffer_dirty() + * doesn't touch b_assoc_buffers list if b_mmb is set so we are sure the * buffer stays on our list until IO completes (at which point it can be * reaped). */ @@ -542,7 +540,7 @@ int mmb_sync(struct mapping_metadata_bhs *mmb) bh = BH_ENTRY(mmb->list.next); WARN_ON_ONCE(bh->b_mmb != mmb); __remove_assoc_queue(mmb, bh); - /* Avoid race with mark_buffer_dirty_inode() which does + /* Avoid race with mmb_mark_buffer_dirty() which does * a lockless check and we rely on seeing the dirty bit */ smp_mb(); if (buffer_dirty(bh) || buffer_locked(bh)) { @@ -580,7 +578,7 @@ int mmb_sync(struct mapping_metadata_bhs *mmb) bh = BH_ENTRY(tmp.prev); get_bh(bh); __remove_assoc_queue(mmb, bh); - /* Avoid race with mark_buffer_dirty_inode() which does + /* Avoid race with mmb_mark_buffer_dirty() which does * a lockless check and we rely on seeing the dirty bit */ smp_mb(); if (buffer_dirty(bh)) { @@ -589,7 +587,7 @@ int mmb_sync(struct mapping_metadata_bhs *mmb) } spin_unlock(&mmb->lock); wait_on_buffer(bh); - if (!buffer_uptodate(bh)) + if (buffer_write_io_error(bh)) err = -EIO; brelse(bh); spin_lock(&mmb->lock); @@ -618,6 +616,14 @@ void write_boundary_block(struct block_device *bdev, } } +/** + * mmb_mark_buffer_dirty - Mark a metadata buffer dirty. + * @bh: The buffer to mark dirty. + * @mmb: The list of buffers to add the buffer to. + * + * Mark the buffer dirty and add it to the list if it is not already on + * a list. + */ void mmb_mark_buffer_dirty(struct buffer_head *bh, struct mapping_metadata_bhs *mmb) { @@ -686,7 +692,7 @@ bool block_dirty_folio(struct address_space *mapping, struct folio *folio) } while (bh != head); } /* - * Lock out page's memcg migration to keep PageDirty + * Lock out folio's memcg migration to keep folio dirty flag * synchronized with per-memcg dirty page counters. */ newly_dirty = !folio_test_set_dirty(folio); @@ -944,23 +950,23 @@ __getblk_slow(struct block_device *bdev, sector_t block, } /* - * The relationship between dirty buffers and dirty pages: + * The relationship between dirty buffers and dirty folios: * - * Whenever a page has any dirty buffers, the page's dirty bit is set, and - * the page is tagged dirty in the page cache. + * Whenever a folio has any dirty buffers, the folio's dirty flag is set, and + * the folio is tagged dirty in the page cache. * * At all times, the dirtiness of the buffers represents the dirtiness of - * subsections of the page. If the page has buffers, the page dirty bit is + * subsections of the folio. If the folio has buffers, the folio dirty flag is * merely a hint about the true dirty state. * - * When a page is set dirty in its entirety, all its buffers are marked dirty - * (if the page has buffers). + * When a folio is set dirty in its entirety, all its buffers are marked dirty + * (if the folio has buffers). * - * When a buffer is marked dirty, its page is dirtied, but the page's other + * When a buffer is marked dirty, its folio is dirtied, but the folio's other * buffers are not. * * Also. When blockdev buffers are explicitly read with bread(), they - * individually become uptodate. But their backing page remains not + * individually become uptodate. But their backing folio remains not * uptodate - even if all of its buffers are uptodate. A subsequent * block_read_full_folio() against that folio will discover all the uptodate * buffers, will set the folio uptodate and will perform no I/O. @@ -971,7 +977,7 @@ __getblk_slow(struct block_device *bdev, sector_t block, * @bh: the buffer_head to mark dirty * * mark_buffer_dirty() will set the dirty bit against the buffer, then set - * its backing page dirty, then tag the page as dirty in the page cache + * its backing folio dirty, then tag the folio as dirty in the page cache * and then attach the address_space's inode to its superblock's dirty * inode list. * @@ -1054,6 +1060,7 @@ EXPORT_SYMBOL(__brelse); void __bforget(struct buffer_head *bh) { clear_buffer_dirty(bh); + clear_buffer_write_io_error(bh); remove_assoc_queue(bh); __brelse(bh); } @@ -1062,12 +1069,16 @@ EXPORT_SYMBOL(__bforget); static void buffer_set_crypto_ctx(struct bio *bio, const struct buffer_head *bh, gfp_t gfp_mask) { - const struct address_space *mapping = folio_mapping(bh->b_folio); + const struct address_space *mapping; /* * The ext4 journal (jbd2) can submit a buffer_head it directly created - * for a non-pagecache page. fscrypt doesn't care about these. + * for memory that is not in the page cache at all. fscrypt doesn't + * care about these. */ + if (!bh->b_folio) + return; + mapping = bh->b_folio->mapping; if (!mapping) return; fscrypt_set_bio_crypt_ctx(bio, mapping->host, @@ -1078,7 +1089,6 @@ static void __bh_submit(struct buffer_head *bh, blk_opf_t opf, enum rw_hint write_hint, struct writeback_control *wbc, bio_end_io_t end_bio) { - const enum req_op op = opf & REQ_OP_MASK; struct bio *bio; BUG_ON(!buffer_locked(bh)); @@ -1086,11 +1096,7 @@ static void __bh_submit(struct buffer_head *bh, blk_opf_t opf, BUG_ON(buffer_delay(bh)); BUG_ON(buffer_unwritten(bh)); - /* - * Only clear out a write error when rewriting - */ - if (test_set_buffer_req(bh) && (op == REQ_OP_WRITE)) - clear_buffer_write_io_error(bh); + set_buffer_req(bh); if (buffer_meta(bh)) opf |= REQ_META; @@ -1099,7 +1105,8 @@ static void __bh_submit(struct buffer_head *bh, blk_opf_t opf, bio = bio_alloc(bh->b_bdev, 1, opf, GFP_NOIO); - if (folio_test_dropbehind(bh->b_folio) && op_is_write(opf)) + if (bh->b_folio && folio_test_dropbehind(bh->b_folio) && + op_is_write(opf)) bio_set_flag(bio, BIO_COMPLETE_IN_TASK); if (IS_ENABLED(CONFIG_FS_ENCRYPTION)) @@ -1108,7 +1115,11 @@ static void __bh_submit(struct buffer_head *bh, blk_opf_t opf, bio->bi_iter.bi_sector = bh->b_blocknr * (bh->b_size >> 9); bio->bi_write_hint = write_hint; - bio_add_folio_nofail(bio, bh->b_folio, bh->b_size, bh_offset(bh)); + if (bh->b_folio) + bio_add_folio_nofail(bio, bh->b_folio, bh->b_size, + bh_offset(bh)); + else + bio_add_virt_nofail(bio, bh->b_data, bh->b_size); bio->bi_end_io = end_bio; bio->bi_private = bh; @@ -1118,7 +1129,8 @@ static void __bh_submit(struct buffer_head *bh, blk_opf_t opf, if (wbc) { wbc_init_bio(wbc, bio); - wbc_account_cgroup_owner(wbc, bh->b_folio, bh->b_size); + if (bh->b_folio) + wbc_account_cgroup_owner(wbc, bh->b_folio, bh->b_size); } blk_crypto_submit_bio(bio); @@ -1208,7 +1220,7 @@ static void bh_lru_install(struct buffer_head *bh) /* * the refcount of buffer_head in bh_lru prevents dropping the - * attached page(i.e., try_to_free_buffers) so it could cause + * attached folio (i.e., try_to_free_buffers) so it could cause * failing page migration. * Skip putting upcoming bh into bh_lru until migration is done. */ @@ -1272,7 +1284,7 @@ lookup_bh_lru(struct block_device *bdev, sector_t block, unsigned size) * Perform a pagecache lookup for the matching buffer. If it's there, refresh * it in the LRU and mark it as accessed. If it is not present then return * NULL. Atomic context callers may also return NULL if the buffer is being - * migrated; similarly the page is not marked accessed either. + * migrated; similarly the folio is not marked accessed either. */ static struct buffer_head * find_get_block_common(struct block_device *bdev, sector_t block, @@ -1281,7 +1293,7 @@ find_get_block_common(struct block_device *bdev, sector_t block, struct buffer_head *bh = lookup_bh_lru(bdev, block, size); if (bh == NULL) { - /* __find_get_block_slow will mark the page accessed */ + /* __find_get_block_slow will mark the folio accessed */ bh = __find_get_block_slow(bdev, block, atomic); if (bh) bh_lru_install(bh); @@ -1467,15 +1479,14 @@ void folio_set_bh(struct buffer_head *bh, struct folio *folio, } EXPORT_SYMBOL(folio_set_bh); -/* - * Called when truncating a buffer on a page completely. - */ - /* Bits that are cleared during an invalidate */ #define BUFFER_FLAGS_DISCARD \ (1 << BH_Mapped | 1 << BH_New | 1 << BH_Req | \ - 1 << BH_Delay | 1 << BH_Unwritten) + 1 << BH_Delay | 1 << BH_Unwritten | 1 << BH_Write_EIO) +/* + * Called when truncating a buffer on a folio completely. + */ static void discard_buffer(struct buffer_head * bh) { unsigned long b_state; @@ -1603,9 +1614,7 @@ EXPORT_SYMBOL(create_empty_buffers); * moment when something will explicitly mark the buffer dirty (hopefully that * will not happen until we will free that block ;-) We don't even need to mark * it not-uptodate - nobody can expect anything from a newly allocated buffer - * anyway. We used to use unmap_buffer() for such invalidation, but that was - * wrong. We definitely don't want to mark the alias unmapped, for example - it - * would confuse anyone who might pick it with bread() afterwards... + * anyway. * * Also.. Note that bforget() doesn't lock the buffer. So there can be * writeout I/O going on against recently-freed buffers. We don't wait on that @@ -1641,7 +1650,7 @@ void clean_bdev_aliases(struct block_device *bdev, sector_t block, sector_t len) /* Recheck when the folio is locked which pins bhs */ head = folio_buffers(folio); if (!head) - goto unlock_page; + goto unlock_folio; bh = head; do { if (!buffer_mapped(bh) || (bh->b_blocknr < block)) @@ -1654,7 +1663,7 @@ void clean_bdev_aliases(struct block_device *bdev, sector_t block, sector_t len) next: bh = bh->b_this_page; } while (bh != head); -unlock_page: +unlock_folio: folio_unlock(folio); } folio_batch_release(&fbatch); @@ -1702,7 +1711,7 @@ static struct buffer_head *folio_create_buffers(struct folio *folio, * * If block_write_full_folio() is called for regular writeback * (wbc->sync_mode == WB_SYNC_NONE) then it will redirty a folio which - * has a locked buffer. This only can happen if someone has written + * has a locked buffer. This can only happen if someone has written * the buffer directly, with bh_submit(). At the address_space level * the folio writeback flag prevents this contention from occurring. * @@ -2205,9 +2214,9 @@ int generic_write_end(const struct kiocb *iocb, struct address_space *mapping, if (old_size < pos) pagecache_isize_extended(inode, old_size, pos); /* - * Don't mark the inode dirty under page lock. First, it unnecessarily - * makes the holding time of page lock longer. Second, it forces lock - * ordering of page lock and transaction start for journaling + * Don't mark the inode dirty under folio lock. First, it unnecessarily + * makes the holding time of folio lock longer. Second, it forces lock + * ordering of folio lock and transaction start for journaling * filesystems. */ if (i_size_changed) @@ -2333,7 +2342,7 @@ int block_read_full_folio(struct folio *folio, get_block_t *get_block) * BH_Async_Read tells end_buffer_async_read() that this * buffer is not under async I/O. * - * The folio comes unlocked when it has no locked + * The folio is unlocked when it has no locked * buffer_async buffers left. * * The folio lock prevents anyone starting new async @@ -2443,7 +2452,7 @@ static int cont_expand_zero(const struct kiocb *iocb, } } - /* page covers the boundary, find the boundary offset */ + /* folio crosses the boundary, find the boundary offset */ if (index == curidx) { zerofrom = curpos & ~PAGE_MASK; /* if we will expand the thing last block will be filled */ @@ -2501,18 +2510,18 @@ EXPORT_SYMBOL(cont_write_begin); /* * block_page_mkwrite() is not allowed to change the file size as it gets - * called from a page fault handler when a page is first dirtied. Hence we must - * be careful to check for EOF conditions here. We set the page up correctly - * for a written page which means we get ENOSPC checking when writing into + * called from a page fault handler when a folio is first dirtied. Hence we must + * be careful to check for EOF conditions here. We set the folio up correctly + * for a written folio which means we get ENOSPC checking when writing into * holes and correct delalloc and unwritten extent mapping on filesystems that * support these features. * * We are not allowed to take the i_rwsem here so we have to play games to - * protect against truncate races as the page could now be beyond EOF. Because - * truncate writes the inode size before removing pages, once we have the - * page lock we can determine safely if the page is beyond EOF. If it is not - * beyond EOF, then the page is guaranteed safe against truncation until we - * unlock the page. + * protect against truncate races as the folio could now be beyond EOF. Because + * truncate writes the inode size before removing folios, once we have the + * folio lock we can determine safely if the folio is beyond EOF. If it is not + * beyond EOF, then the folio is guaranteed safe against truncation until we + * unlock the folio. * * Direct callers of this function should protect against filesystem freezing * using sb_start_pagefault() - sb_end_pagefault() functions. @@ -2530,7 +2539,7 @@ int block_page_mkwrite(struct vm_area_struct *vma, struct vm_fault *vmf, size = i_size_read(inode); if ((folio->mapping != inode->i_mapping) || (folio_pos(folio) >= size)) { - /* We overload EFAULT to mean page got truncated */ + /* We overload EFAULT to mean folio got truncated */ ret = -EFAULT; goto out_unlock; } @@ -2702,7 +2711,7 @@ int __sync_dirty_buffer(struct buffer_head *bh, blk_opf_t op_flags) bh_submit(bh, REQ_OP_WRITE | op_flags, bh_end_write); wait_on_buffer(bh); - if (!buffer_uptodate(bh)) + if (buffer_write_io_error(bh)) return -EIO; } else { unlock_buffer(bh); diff --git a/fs/cachefiles/Kconfig b/fs/cachefiles/Kconfig index afb25b6af5aa..c9c168c7e072 100644 --- a/fs/cachefiles/Kconfig +++ b/fs/cachefiles/Kconfig @@ -17,7 +17,7 @@ config CACHEFILES_DEBUG help This permits debugging to be dynamically enabled in the filesystem caching on files module. If this is set, the debugging output may be - enabled by setting bits in /sys/modules/cachefiles/parameter/debug or + enabled by setting bits in /sys/module/cachefiles/parameters/debug or by including a debugging specifier in /etc/cachefilesd.conf. config CACHEFILES_ERROR_INJECTION diff --git a/fs/cachefiles/interface.c b/fs/cachefiles/interface.c index 50a000310a8c..789ff6abe926 100644 --- a/fs/cachefiles/interface.c +++ b/fs/cachefiles/interface.c @@ -100,73 +100,6 @@ void cachefiles_put_object(struct cachefiles_object *object, } /* - * Adjust the size of a cache file if necessary to match the DIO size. We keep - * the EOF marker a multiple of DIO blocks so that we don't fall back to doing - * non-DIO for a partial block straddling the EOF, but we also have to be - * careful of someone expanding the file and accidentally accreting the - * padding. - */ -static int cachefiles_adjust_size(struct cachefiles_object *object) -{ - struct iattr newattrs; - struct file *file = object->file; - uint64_t ni_size; - loff_t oi_size; - int ret; - - ni_size = object->cookie->object_size; - ni_size = round_up(ni_size, CACHEFILES_DIO_BLOCK_SIZE); - - _enter("{OBJ%x},[%llu]", - object->debug_id, (unsigned long long) ni_size); - - if (!file) - return -ENOBUFS; - - oi_size = i_size_read(file_inode(file)); - if (oi_size == ni_size) - return 0; - - inode_lock(file_inode(file)); - - /* if there's an extension to a partial page at the end of the backing - * file, we need to discard the partial page so that we pick up new - * data after it */ - if (oi_size & ~PAGE_MASK && ni_size > oi_size) { - _debug("discard tail %llx", oi_size); - newattrs.ia_valid = ATTR_SIZE; - newattrs.ia_size = oi_size & PAGE_MASK; - ret = cachefiles_inject_remove_error(); - if (ret == 0) - ret = notify_change(&nop_mnt_idmap, file->f_path.dentry, - &newattrs, NULL); - if (ret < 0) - goto truncate_failed; - } - - newattrs.ia_valid = ATTR_SIZE; - newattrs.ia_size = ni_size; - ret = cachefiles_inject_write_error(); - if (ret == 0) - ret = notify_change(&nop_mnt_idmap, file->f_path.dentry, - &newattrs, NULL); - -truncate_failed: - inode_unlock(file_inode(file)); - - if (ret < 0) - trace_cachefiles_io_error(NULL, file_inode(file), ret, - cachefiles_trace_notify_change_error); - if (ret == -EIO) { - cachefiles_io_error_obj(object, "Size set failed"); - ret = -ENOBUFS; - } - - _leave(" = %d", ret); - return ret; -} - -/* * Attempt to look up the nominated node in this cache */ static bool cachefiles_lookup_cookie(struct fscache_cookie *cookie) @@ -198,7 +131,6 @@ static bool cachefiles_lookup_cookie(struct fscache_cookie *cookie) spin_lock(&cache->object_list_lock); list_add(&object->cache_link, &cache->object_list); spin_unlock(&cache->object_list_lock); - cachefiles_adjust_size(object); cachefiles_end_secure(cache, saved_cred); _leave(" = t"); @@ -225,14 +157,14 @@ fail: * any unused granules. */ static bool cachefiles_shorten_object(struct cachefiles_object *object, - struct file *file, loff_t new_size) + struct file *file, uoff_t new_size) { struct cachefiles_cache *cache = object->volume->cache; struct inode *inode = file_inode(file); - loff_t i_size, dio_size; + uoff_t i_size, dio_size; int ret; - dio_size = round_up(new_size, CACHEFILES_DIO_BLOCK_SIZE); + dio_size = round_up(new_size, cache->bsize); i_size = i_size_read(inode); trace_cachefiles_trunc(object, inode, i_size, dio_size, @@ -264,6 +196,7 @@ static bool cachefiles_shorten_object(struct cachefiles_object *object, } } + object->object_size = new_size; return true; } @@ -271,29 +204,38 @@ static bool cachefiles_shorten_object(struct cachefiles_object *object, * Resize the backing object. */ static void cachefiles_resize_cookie(struct netfs_cache_resources *cres, - loff_t new_size) + uoff_t new_size) { struct cachefiles_object *object = cachefiles_cres_object(cres); struct cachefiles_cache *cache = object->volume->cache; struct fscache_cookie *cookie = object->cookie; const struct cred *saved_cred; struct file *file = cachefiles_cres_file(cres); - loff_t old_size = cookie->object_size; + uoff_t i_size = i_size_read(file_inode(file)); - _enter("%llu->%llu", old_size, new_size); + _enter("%llu->%llu", object->object_size, new_size); - if (new_size < old_size) { + /* If the file is being shrunk, we need to downsize the backing file + * and clear the end of the final block. + */ + if (new_size < object->object_size) { + if (new_size >= i_size) + goto out; cachefiles_begin_secure(cache, &saved_cred); cachefiles_shorten_object(object, file, new_size); cachefiles_end_secure(cache, saved_cred); object->cookie->object_size = new_size; + if (new_size == 0) + object->content_info = CACHEFILES_CONTENT_NO_DATA; return; } /* The file is being expanded. We don't need to do anything - * particularly. cookie->initial_size doesn't change and so the point - * at which we have to download before doesn't change. + * particularly. The tail of the last block should have been cleared + * both when it is written and when it is shrunk. */ +out: + object->object_size = new_size; cookie->object_size = new_size; } diff --git a/fs/cachefiles/internal.h b/fs/cachefiles/internal.h index c93324e0f98c..664be64ab538 100644 --- a/fs/cachefiles/internal.h +++ b/fs/cachefiles/internal.h @@ -16,8 +16,6 @@ #include <linux/cred.h> #include <linux/security.h> -#define CACHEFILES_DIO_BLOCK_SIZE 4096 - struct cachefiles_cache; struct cachefiles_object; @@ -51,12 +49,17 @@ struct cachefiles_object { struct list_head cache_link; /* Link in cache->*_list */ struct file *file; /* The file representing this object */ char *d_name; /* Backing file name */ + unsigned long flags; +#define CACHEFILES_OBJECT_USING_TMPFILE 0 /* Have an unlinked tmpfile */ + uoff_t object_size; /* Size of the object stored + * (independent of cookie->object_size for + * coherency reasons) + */ + atomic64_t read_limit; /* Point beyond which uncommitted writes */ int debug_id; spinlock_t lock; refcount_t ref; - enum cachefiles_content content_info:8; /* Info about content presence */ - unsigned long flags; -#define CACHEFILES_OBJECT_USING_TMPFILE 0 /* Have an unlinked tmpfile */ + enum cachefiles_content content_info; /* Info about content presence */ }; /* @@ -203,11 +206,11 @@ extern bool cachefiles_begin_operation(struct netfs_cache_resources *cres, enum fscache_want_state want_state); extern int __cachefiles_prepare_write(struct cachefiles_object *object, struct file *file, - loff_t *_start, size_t *_len, size_t upper_len, + uoff_t *_start, size_t *_len, size_t upper_len, bool no_space_allocated_yet); extern int __cachefiles_write(struct cachefiles_object *object, struct file *file, - loff_t start_pos, + uoff_t start_pos, struct iov_iter *iter, netfs_io_terminated_t term_func, void *term_func_priv); @@ -280,6 +283,7 @@ void cachefiles_withdraw_volume(struct cachefiles_volume *volume); /* * xattr.c */ +int cachefiles_preset_object_xattr(struct cachefiles_object *object, struct file *file); extern int cachefiles_set_object_xattr(struct cachefiles_object *object); extern int cachefiles_check_auxdata(struct cachefiles_object *object, struct file *file); diff --git a/fs/cachefiles/io.c b/fs/cachefiles/io.c index 9540ec25b3cb..4f547d97356e 100644 --- a/fs/cachefiles/io.c +++ b/fs/cachefiles/io.c @@ -19,7 +19,7 @@ struct cachefiles_kiocb { struct kiocb iocb; refcount_t ki_refcnt; - loff_t start; + uoff_t start; union { size_t skipped; size_t len; @@ -32,6 +32,8 @@ struct cachefiles_kiocb { u64 b_writing; }; +#define IS_ERR_VALUE_LL(x) unlikely((x) >= (unsigned long long)-MAX_ERRNO) + static inline void cachefiles_put_kiocb(struct cachefiles_kiocb *ki) { if (refcount_dec_and_test(&ki->ki_refcnt)) { @@ -73,7 +75,7 @@ static void cachefiles_read_complete(struct kiocb *iocb, long ret) * Initiate a read from the cache. */ static int cachefiles_read(struct netfs_cache_resources *cres, - loff_t start_pos, + uoff_t start_pos, struct iov_iter *iter, enum netfs_read_from_hole read_hole, netfs_io_terminated_t term_func, @@ -193,60 +195,81 @@ presubmission_error: } /* - * Query the occupancy of the cache in a region, returning where the next chunk - * of data starts and how long it is. + * Query the occupancy of the cache in a region, returning the extent of the + * next two chunks of cached data and the next hole. */ static int cachefiles_query_occupancy(struct netfs_cache_resources *cres, - loff_t start, size_t len, size_t granularity, - loff_t *_data_start, size_t *_data_len) + struct fscache_occupancy *occ) { struct cachefiles_object *object; + struct inode *inode; struct file *file; - loff_t off, off2; - - *_data_start = -1; - *_data_len = 0; + uoff_t read_limit; + loff_t ret; + int i; if (!fscache_wait_for_operation(cres, FSCACHE_WANT_READ)) return -ENOBUFS; object = cachefiles_cres_object(cres); file = cachefiles_cres_file(cres); - granularity = max_t(size_t, object->volume->cache->bsize, granularity); + inode = file_inode(file); + occ->granularity = object->volume->cache->bsize; + /* Read read_limit before content_info. */ + read_limit = atomic64_read_acquire(&object->read_limit); + + _enter("%pD,%llu,%llx-%llx/%llx", + file, inode->i_ino, occ->query_from, occ->query_to, read_limit); + + if (read_limit == 0) + goto done; + + switch (READ_ONCE(object->content_info)) { + case CACHEFILES_CONTENT_ALL: + case CACHEFILES_CONTENT_SINGLE: + if (read_limit > occ->query_from) { + occ->cached_from[0] = 0; + occ->cached_to[0] = read_limit; + occ->cached_type[0] = FSCACHE_EXTENT_DATA; + occ->query_from = ULLONG_MAX; + } + goto done; + default: + break; + } - _enter("%pD,%llu,%llx,%zx/%llx", - file, file_inode(file)->i_ino, start, len, - i_size_read(file_inode(file))); + for (i = 0; i < ARRAY_SIZE(occ->cached_from); i++) { + ret = cachefiles_inject_read_error(); + if (ret == 0) + ret = vfs_llseek(file, occ->query_from, SEEK_DATA); + if (IS_ERR_VALUE_LL(ret)) { + if (ret != -ENXIO) + return ret; + occ->query_from = ULLONG_MAX; + goto done; + } + occ->cached_type[i] = FSCACHE_EXTENT_DATA; + occ->cached_from[i] = ret; + occ->query_from = ret; + + ret = cachefiles_inject_read_error(); + if (ret == 0) + ret = vfs_llseek(file, occ->query_from, SEEK_HOLE); + if (IS_ERR_VALUE_LL(ret)) { + if (ret != -ENXIO) + return ret; + occ->query_from = ULLONG_MAX; + goto done; + } + occ->cached_to[i] = ret; + occ->query_from = ret; + if (occ->query_from >= occ->query_to) + break; + } - off = cachefiles_inject_read_error(); - if (off == 0) - off = vfs_llseek(file, start, SEEK_DATA); - if (off == -ENXIO) - return -ENODATA; /* Beyond EOF */ - if (off < 0 && off >= (loff_t)-MAX_ERRNO) - return -ENOBUFS; /* Error. */ - if (round_up(off, granularity) >= start + len) - return -ENODATA; /* No data in range */ - - off2 = cachefiles_inject_read_error(); - if (off2 == 0) - off2 = vfs_llseek(file, off, SEEK_HOLE); - if (off2 == -ENXIO) - return -ENODATA; /* Beyond EOF */ - if (off2 < 0 && off2 >= (loff_t)-MAX_ERRNO) - return -ENOBUFS; /* Error. */ - - /* Round away partial blocks */ - off = round_up(off, granularity); - off2 = round_down(off2, granularity); - if (off2 <= off) - return -ENODATA; - - *_data_start = off; - if (off2 > start + len) - *_data_len = len; - else - *_data_len = off2 - off; +done: + _debug("query[0] %llx-%llx", occ->cached_from[0], occ->cached_to[0]); + _debug("query[1] %llx-%llx", occ->cached_from[1], occ->cached_to[1]); return 0; } @@ -280,7 +303,7 @@ static void cachefiles_write_complete(struct kiocb *iocb, long ret) */ int __cachefiles_write(struct cachefiles_object *object, struct file *file, - loff_t start_pos, + uoff_t start_pos, struct iov_iter *iter, netfs_io_terminated_t term_func, void *term_func_priv) @@ -357,7 +380,7 @@ in_progress: } static int cachefiles_write(struct netfs_cache_resources *cres, - loff_t start_pos, + uoff_t start_pos, struct iov_iter *iter, netfs_io_terminated_t term_func, void *term_func_priv) @@ -375,127 +398,12 @@ static int cachefiles_write(struct netfs_cache_resources *cres, term_func, term_func_priv); } -static inline enum netfs_io_source -cachefiles_do_prepare_read(struct netfs_cache_resources *cres, - loff_t start, size_t *_len, loff_t i_size, - unsigned long *_flags, ino_t netfs_ino) -{ - enum cachefiles_prepare_read_trace why; - struct cachefiles_object *object = NULL; - struct cachefiles_cache *cache; - struct fscache_cookie *cookie = fscache_cres_cookie(cres); - const struct cred *saved_cred; - struct file *file = cachefiles_cres_file(cres); - enum netfs_io_source ret = NETFS_DOWNLOAD_FROM_SERVER; - size_t len = *_len; - loff_t off, to; - ino_t ino = file ? file_inode(file)->i_ino : 0; - - _enter("%zx @%llx/%llx", len, start, i_size); - - if (start >= i_size) { - ret = NETFS_FILL_WITH_ZEROES; - why = cachefiles_trace_read_after_eof; - goto out_no_object; - } - - if (test_bit(FSCACHE_COOKIE_NO_DATA_TO_READ, &cookie->flags)) { - __set_bit(NETFS_SREQ_COPY_TO_CACHE, _flags); - why = cachefiles_trace_read_no_data; - goto out_no_object; - } - - /* The object and the file may be being created in the background. */ - if (!file) { - why = cachefiles_trace_read_no_file; - if (!fscache_wait_for_operation(cres, FSCACHE_WANT_READ)) - goto out_no_object; - file = cachefiles_cres_file(cres); - if (!file) - goto out_no_object; - ino = file_inode(file)->i_ino; - } - - object = cachefiles_cres_object(cres); - cache = object->volume->cache; - cachefiles_begin_secure(cache, &saved_cred); - off = cachefiles_inject_read_error(); - if (off == 0) - off = vfs_llseek(file, start, SEEK_DATA); - if (off < 0 && off >= (loff_t)-MAX_ERRNO) { - if (off == (loff_t)-ENXIO) { - why = cachefiles_trace_read_seek_nxio; - goto download_and_store; - } - trace_cachefiles_io_error(object, file_inode(file), off, - cachefiles_trace_seek_error); - why = cachefiles_trace_read_seek_error; - goto out; - } - - if (off >= start + len) { - why = cachefiles_trace_read_found_hole; - goto download_and_store; - } - - if (off > start) { - off = round_up(off, cache->bsize); - len = off - start; - *_len = len; - why = cachefiles_trace_read_found_part; - goto download_and_store; - } - - to = cachefiles_inject_read_error(); - if (to == 0) - to = vfs_llseek(file, start, SEEK_HOLE); - if (to < 0 && to >= (loff_t)-MAX_ERRNO) { - trace_cachefiles_io_error(object, file_inode(file), to, - cachefiles_trace_seek_error); - why = cachefiles_trace_read_seek_error; - goto out; - } - - if (to < start + len) { - if (start + len >= i_size) - to = round_up(to, cache->bsize); - else - to = round_down(to, cache->bsize); - len = to - start; - *_len = len; - } - - why = cachefiles_trace_read_have_data; - ret = NETFS_READ_FROM_CACHE; - goto out; - -download_and_store: - __set_bit(NETFS_SREQ_COPY_TO_CACHE, _flags); -out: - cachefiles_end_secure(cache, saved_cred); -out_no_object: - trace_cachefiles_prep_read(object, start, len, *_flags, ret, why, ino, netfs_ino); - return ret; -} - -/* - * Prepare a read operation, shortening it to a cached/uncached - * boundary as appropriate. - */ -static enum netfs_io_source cachefiles_prepare_read(struct netfs_io_subrequest *subreq, - unsigned long long i_size) -{ - return cachefiles_do_prepare_read(&subreq->rreq->cache_resources, - subreq->start, &subreq->len, i_size, - &subreq->flags, subreq->rreq->inode->i_ino); -} - /* * Prepare for a write to occur. */ int __cachefiles_prepare_write(struct cachefiles_object *object, struct file *file, - loff_t *_start, size_t *_len, size_t upper_len, + uoff_t *_start, size_t *_len, size_t upper_len, bool no_space_allocated_yet) { struct cachefiles_cache *cache = object->volume->cache; @@ -504,7 +412,7 @@ int __cachefiles_prepare_write(struct cachefiles_object *object, int ret; /* Round to DIO size */ - start = round_down(*_start, PAGE_SIZE); + start = round_down(*_start, cache->bsize); if (start != *_start || *_len > upper_len) { /* Probably asked to cache a streaming write written into the * pagecache when the cookie was temporarily out of service to @@ -514,7 +422,7 @@ int __cachefiles_prepare_write(struct cachefiles_object *object, return -ENOBUFS; } - *_len = round_up(len, PAGE_SIZE); + *_len = round_up(len, cache->bsize); /* We need to work out whether there's sufficient disk space to perform * the write - but we can skip that check if we have space already @@ -540,10 +448,14 @@ int __cachefiles_prepare_write(struct cachefiles_object *object, * space, we need to see if it's fully allocated. If it's not, we may * want to cull it. */ - if (cachefiles_has_space(cache, 0, *_len / PAGE_SIZE, - cachefiles_has_space_check) == 0) + ret = cachefiles_has_space(cache, 0, *_len / cache->bsize, + cachefiles_has_space_check); + if (ret == 0) return 0; /* Enough space to simply overwrite the whole block */ + if (ret == -ENOBUFS) + trace_cachefiles_no_space(object, cachefiles_trace_write_nospace_2); + pos = cachefiles_inject_read_error(); if (pos == 0) pos = vfs_llseek(file, start, SEEK_HOLE); @@ -572,13 +484,16 @@ int __cachefiles_prepare_write(struct cachefiles_object *object, return ret; check_space: - return cachefiles_has_space(cache, 0, *_len / PAGE_SIZE, - cachefiles_has_space_for_write); + ret = cachefiles_has_space(cache, 0, *_len / cache->bsize, + cachefiles_has_space_for_write); + if (ret == -ENOBUFS) + trace_cachefiles_no_space(object, cachefiles_trace_write_nospace); + return ret; } static int cachefiles_prepare_write(struct netfs_cache_resources *cres, - loff_t *_start, size_t *_len, size_t upper_len, - loff_t i_size, bool no_space_allocated_yet) + uoff_t *_start, size_t *_len, size_t upper_len, + uoff_t i_size, bool no_space_allocated_yet) { struct cachefiles_object *object = cachefiles_cres_object(cres); struct cachefiles_cache *cache = object->volume->cache; @@ -612,10 +527,14 @@ static void cachefiles_prepare_write_subreq(struct netfs_io_subrequest *subreq) stream->sreq_max_segs = BIO_MAX_VECS; if (!cachefiles_cres_file(cres)) { - if (!fscache_wait_for_operation(cres, FSCACHE_WANT_WRITE)) + if (!fscache_wait_for_operation(cres, FSCACHE_WANT_WRITE)) { + trace_netfs_sreq(subreq, netfs_sreq_trace_cache_waitfail); return netfs_prepare_write_failed(subreq); - if (!cachefiles_cres_file(cres)) + } + if (!cachefiles_cres_file(cres)) { + trace_netfs_sreq(subreq, netfs_sreq_trace_cache_nofile); return netfs_prepare_write_failed(subreq); + } } } @@ -628,16 +547,16 @@ static void cachefiles_issue_write(struct netfs_io_subrequest *subreq) struct netfs_io_stream *stream = &wreq->io_streams[subreq->stream_nr]; const struct cred *saved_cred; size_t off, pre, post, len = subreq->len; - loff_t start = subreq->start; + uoff_t start = subreq->start; int ret; _enter("W=%x[%x] %llx-%llx", wreq->debug_id, subreq->debug_index, start, start + len - 1); /* We need to start on the cache granularity boundary */ - off = start & (CACHEFILES_DIO_BLOCK_SIZE - 1); + off = start & (cache->bsize - 1); if (off) { - pre = CACHEFILES_DIO_BLOCK_SIZE - off; + pre = cache->bsize - off; if (pre >= len) { fscache_count_dio_misfit(); netfs_write_subrequest_terminated(subreq, len); @@ -651,8 +570,8 @@ static void cachefiles_issue_write(struct netfs_io_subrequest *subreq) /* We also need to end on the cache granularity boundary */ if (start + len == wreq->i_size) { - size_t part = len % CACHEFILES_DIO_BLOCK_SIZE; - size_t need = CACHEFILES_DIO_BLOCK_SIZE - part; + size_t part = len & (cache->bsize - 1); + size_t need = cache->bsize - part; if (part && stream->submit_extendable_to >= need) { len += need; @@ -661,7 +580,7 @@ static void cachefiles_issue_write(struct netfs_io_subrequest *subreq) } } - post = len & (CACHEFILES_DIO_BLOCK_SIZE - 1); + post = len & (cache->bsize - 1); if (post) { len -= post; if (len == 0) { @@ -689,6 +608,198 @@ static void cachefiles_issue_write(struct netfs_io_subrequest *subreq) } /* + * Collect the result of buffered writeback to the cache. This includes + * copying a read to the cache. Netfslib collates the results, which might + * occur out of order, and delivers them to the cache so that it can update its + * content record. + * + * block_type is one of: + * - NETFS_CACHE_COLLECT_WRITE_DATA for a contiguous block of data + * - NETFS_CACHE_COLLECT_WRITE_GAP if a discontiguity was skipped + * - NETFS_CACHE_COLLECT_WRITE_CANCEL for a hole due to a failed/cancelled write + * + * The writes we made are all rounded out at both sides to the nearest DIO + * block boundary, so if the final block contains the EOF in the middle of it + * (rather than at the end), padding will have been written to the file. The + * backing file's filesize will have been updated if the write extended the + * file; the filesize may still change due to outstanding subreqs. + * + * The metadata in the cache file xattr records the size of the object we have + * stored, but the cache file EOF only goes up to where we've cached data to + * and, furthermore, is rounded up to the nearest DIO block boundary. + * + * Concurrent updates should be protected against by the caller. Netfslib + * holds NETFS_ICTX_WB_LOCK as a lock on writeback requests. DIO writes + * invalidate the cookie and caching is kept disabled until all users have + * unused the cookie. + */ +static void cachefiles_collect_write(struct netfs_io_request *wreq, + uoff_t start, size_t len, + enum netfs_cache_collect block_type) +{ + struct netfs_cache_resources *cres = &wreq->cache_resources; + struct cachefiles_object *object = cachefiles_cres_object(cres); + struct cachefiles_cache *cache = object->volume->cache; + struct inode *inode; + struct file *file = cachefiles_cres_file(cres); + uoff_t read_limit; + uoff_t old_size = cres->cache_i_size; + uoff_t new_size; + uoff_t data_to = object->object_size; + uoff_t end = start + len; + int ret; + + if (!file) + return; + + inode = file_inode(file); + new_size = i_size_read(inode); + + _enter("%llx,%zx,%x", start, len, cache->bsize); + + if (WARN_ON(old_size & (cache->bsize - 1)) || + WARN_ON(new_size & (cache->bsize - 1)) || + WARN_ON(start & (cache->bsize - 1)) || + WARN_ON(len & (cache->bsize - 1))) { + trace_cachefiles_io_error(object, inode, -EIO, + cachefiles_trace_alignment_error); + cachefiles_remove_object_xattr(cache, object, file->f_path.dentry); + return; + } + + /* If this is recording a gap, due to discontiguous writes or lack of + * cache space, then a hole may have been introduced into the backing + * file. Treat it as a zero-length data block. + */ + if (block_type == NETFS_CACHE_COLLECT_WRITE_GAP || + block_type == NETFS_CACHE_COLLECT_WRITE_CANCEL) { + start = end; + len = 0; + } + + /* Zeroth case: Single monolithic files are handled specially. + */ + if (wreq->origin == NETFS_WRITEBACK_SINGLE) { + if (block_type == NETFS_CACHE_COLLECT_WRITE_GAP || + block_type == NETFS_CACHE_COLLECT_WRITE_CANCEL) { + trace_cachefiles_trunc(object, inode, data_to, 0, + cachefiles_trunc_zap); + ret = cachefiles_inject_remove_error(); + if (ret == 0) + ret = vfs_truncate(&file->f_path, 0); + if (ret < 0) { + trace_cachefiles_io_error(object, inode, ret, + cachefiles_trace_trunc_error); + cachefiles_io_error_obj(object, "truncate failed %d", ret); + cachefiles_remove_object_xattr(cache, object, file->f_path.dentry); + return; + } + + object->content_info = CACHEFILES_CONTENT_NO_DATA; + read_limit = 0; + } else { + object->content_info = CACHEFILES_CONTENT_SINGLE; + read_limit = len; + } + goto update_sizes_2; + } + + /* First case: The backing file was empty. */ + if (old_size == 0) { + if (start == 0) + object->content_info = CACHEFILES_CONTENT_ALL; + else + object->content_info = CACHEFILES_CONTENT_BACKFS_MAP; + goto update_sizes; + } + + /* Second case: The backing file is entirely within the old object size + * and thus there can be no partial tail block to deal with in the + * cache file. + */ + if (old_size <= data_to) { + if (start > old_size) + goto discontiguous; + goto update_sizes; + } + + /* Third case: The write happened entirely within the bounds of the + * current cache file's size. + */ + if (end <= old_size) + goto update_sizes; + + /* Fourth case: The write overwrote the partial tail block and extended + * the file. We only need to update the object size because netfslib + * rounds out/pads cache writes to whole disk blocks. + */ + if (start < old_size) + goto update_sizes; + + /* Fifth case: The write started from the end of the whole tail block + * and extended the file. Just extend our notion of the filesize. + */ + if (start == old_size && old_size == data_to) + goto update_sizes; + + /* Sixth case: The write continued on from the partial tail block and + * extended the file. Need to clear the gap. + */ + if (start == old_size && old_size > data_to) + goto clear_gap; + +discontiguous: + /* Seventh case: The write was beyond the EOF on the cache file, so now + * there's a hole in the file and we can no longer say in the metadata + * that we can assume we have it all. We may also need to clear the + * end of the partial tail block. + */ + /* TODO: For the moment, we will have to use SEEK_HOLE/SEEK_DATA. */ + if (object->content_info != CACHEFILES_CONTENT_BACKFS_MAP) { + object->content_info = CACHEFILES_CONTENT_BACKFS_MAP; + trace_cachefiles_coherency(object, inode->i_ino, data_to, NULL, + CACHEFILES_CONTENT_BACKFS_MAP, + cachefiles_coherency_discontiguous); + } + +clear_gap: + /* We need to clear any partial padding that got jumped over. It + * *should* be all zeros, but shared-writable mmap exists... + */ + if (old_size > data_to) { + trace_cachefiles_trunc(object, inode, data_to, old_size, + cachefiles_trunc_clear_padding); + ret = cachefiles_inject_write_error(); + if (ret == 0) + ret = vfs_fallocate(file, FALLOC_FL_ZERO_RANGE, + data_to, old_size - data_to); + if (ret < 0) { + trace_cachefiles_io_error(object, inode, ret, + cachefiles_trace_fallocate_error); + cachefiles_io_error_obj(object, "fallocate zero pad failed %d", ret); + cachefiles_remove_object_xattr(cache, object, file->f_path.dentry); + return; + } + } + +update_sizes: + read_limit = umax(old_size, end); +update_sizes_2: + cres->cache_i_size = read_limit; + + /* We need to be careful setting the object_size: we may have written + * more to the cache than to the server (due to cache DIO rounding) and + * the i_size set on the netfs inode may include unwritten data that + * the server doesn't know about yet. + */ + object->object_size = umin(read_limit, wreq->i_size); + + /* Raise the limit at which reads can access the file. */ + /* Update read_limit after content_info */ + atomic64_set_release(&object->read_limit, read_limit); +} + +/* * Clean up an operation. */ static void cachefiles_end_operation(struct netfs_cache_resources *cres) @@ -705,10 +816,10 @@ static const struct netfs_cache_ops cachefiles_netfs_cache_ops = { .read = cachefiles_read, .write = cachefiles_write, .issue_write = cachefiles_issue_write, - .prepare_read = cachefiles_prepare_read, .prepare_write = cachefiles_prepare_write, .prepare_write_subreq = cachefiles_prepare_write_subreq, .query_occupancy = cachefiles_query_occupancy, + .collect_write = cachefiles_collect_write, }; /* @@ -718,13 +829,20 @@ bool cachefiles_begin_operation(struct netfs_cache_resources *cres, enum fscache_want_state want_state) { struct cachefiles_object *object = cachefiles_cres_object(cres); + struct file *file; + + cres->dio_size = object->volume->cache->bsize; if (!cachefiles_cres_file(cres)) { cres->ops = &cachefiles_netfs_cache_ops; + cres->object_id = object->debug_id; if (object->file) { spin_lock(&object->lock); - if (!cres->cache_priv2 && object->file) - cres->cache_priv2 = get_file(object->file); + file = object->file; + if (!cres->cache_priv2 && file) { + cres->cache_priv2 = get_file(file); + cres->cache_i_size = i_size_read(file_inode(file)); + } spin_unlock(&object->lock); } } diff --git a/fs/cachefiles/namei.c b/fs/cachefiles/namei.c index 88955249a1a6..ef656a319ede 100644 --- a/fs/cachefiles/namei.c +++ b/fs/cachefiles/namei.c @@ -117,8 +117,11 @@ retry: if (d_is_negative(subdir)) { ret = cachefiles_has_space(cache, 1, 0, cachefiles_has_space_for_create); - if (ret < 0) + if (ret < 0) { + if (ret == -ENOBUFS) + trace_cachefiles_no_space(NULL, cachefiles_trace_mkdir_nospace); goto mkdir_error; + } _debug("attempt mkdir"); @@ -414,7 +417,6 @@ struct file *cachefiles_create_tmpfile(struct cachefiles_object *object) struct dentry *fan = volume->fanout[(u8)object->cookie->key_hash]; struct file *file; const struct path parentpath = { .mnt = cache->mnt, .dentry = fan }; - uint64_t ni_size; long ret; @@ -442,31 +444,20 @@ struct file *cachefiles_create_tmpfile(struct cachefiles_object *object) if (!cachefiles_mark_inode_in_use(object, file_inode(file))) WARN_ON(1); - ni_size = object->cookie->object_size; - ni_size = round_up(ni_size, CACHEFILES_DIO_BLOCK_SIZE); - - if (ni_size > 0) { - trace_cachefiles_trunc(object, file_inode(file), 0, ni_size, - cachefiles_trunc_expand_tmpfile); - ret = cachefiles_inject_write_error(); - if (ret == 0) - ret = vfs_truncate(&file->f_path, ni_size); - if (ret < 0) { - trace_cachefiles_vfs_error( - object, file_inode(file), ret, - cachefiles_trace_trunc_error); - goto err_unuse; - } - } - ret = -EINVAL; if (unlikely(!file->f_op->read_iter) || unlikely(!file->f_op->write_iter)) { pr_notice("Cache does not support read_iter and write_iter\n"); goto err_unuse; } + + /* Preallocate space for the xattr. */ + ret = cachefiles_preset_object_xattr(object, file); + if (ret < 0) + goto err_unuse; out: cachefiles_end_secure(cache, saved_cred); + object->content_info = CACHEFILES_CONTENT_ALL; return file; err_unuse: @@ -487,8 +478,11 @@ static bool cachefiles_create_file(struct cachefiles_object *object) ret = cachefiles_has_space(object->volume->cache, 1, 0, cachefiles_has_space_for_create); - if (ret < 0) + if (ret < 0) { + if (ret == -ENOBUFS) + trace_cachefiles_no_space(object, cachefiles_trace_create_nospace); return false; + } file = cachefiles_create_tmpfile(object); if (IS_ERR(file)) diff --git a/fs/cachefiles/xattr.c b/fs/cachefiles/xattr.c index c70bf67e52b0..8ebb713482e3 100644 --- a/fs/cachefiles/xattr.c +++ b/fs/cachefiles/xattr.c @@ -35,6 +35,57 @@ struct cachefiles_vol_xattr { } __packed; /* + * Preset the state xattr on a cache file to allocate space for it. + */ +int cachefiles_preset_object_xattr(struct cachefiles_object *object, struct file *file) +{ + struct cachefiles_xattr *buf; + struct dentry *dentry = file->f_path.dentry; + unsigned int len = object->cookie->aux_len; + int ret; + + buf = kzalloc(sizeof(struct cachefiles_xattr) + min(len, sizeof(__be64)), GFP_KERNEL); + if (!buf) + return -ENOMEM; + + buf->type = CACHEFILES_COOKIE_TYPE_DATA; + buf->content = CACHEFILES_CONTENT_DIRTY; + + ret = cachefiles_inject_write_error(); + if (ret == 0) { + ret = mnt_want_write_file(file); + if (ret == 0) { + ret = vfs_setxattr(&nop_mnt_idmap, dentry, + cachefiles_xattr_cache, buf, + sizeof(struct cachefiles_xattr) + len, 0); + mnt_drop_write_file(file); + } + } + if (ret < 0) { + trace_cachefiles_vfs_error(object, file_inode(file), ret, + cachefiles_trace_setxattr_error); + trace_cachefiles_coherency(object, file_inode(file)->i_ino, + object->object_size, + buf->data, buf->content, + cachefiles_coherency_set_fail); + switch (ret) { + case -ENOMEM: + case -ENOSPC: + break; + default: + cachefiles_io_error_obj( + object, + "Failed to set xattr with error %d", ret); + break; + } + } + + kfree(buf); + _leave(" = %d", ret); + return ret; +} + +/* * set the state xattr on a cache file */ int cachefiles_set_object_xattr(struct cachefiles_object *object) @@ -43,6 +94,7 @@ int cachefiles_set_object_xattr(struct cachefiles_object *object) struct dentry *dentry; struct file *file = object->file; unsigned int len = object->cookie->aux_len; + uoff_t object_size = object->cookie->object_size; int ret; if (!file) @@ -55,7 +107,7 @@ int cachefiles_set_object_xattr(struct cachefiles_object *object) if (!buf) return -ENOMEM; - buf->object_size = cpu_to_be64(object->cookie->object_size); + buf->object_size = cpu_to_be64(object_size); buf->zero_point = 0; buf->type = CACHEFILES_COOKIE_TYPE_DATA; buf->content = object->content_info; @@ -79,15 +131,21 @@ int cachefiles_set_object_xattr(struct cachefiles_object *object) trace_cachefiles_vfs_error(object, file_inode(file), ret, cachefiles_trace_setxattr_error); trace_cachefiles_coherency(object, file_inode(file)->i_ino, - buf->data, buf->content, + object_size, buf->data, buf->content, cachefiles_coherency_set_fail); - if (ret != -ENOMEM) + switch (ret) { + case -ENOMEM: + break; + case -ENOSPC: + default: cachefiles_io_error_obj( object, "Failed to set xattr with error %d", ret); + break; + } } else { trace_cachefiles_coherency(object, file_inode(file)->i_ino, - buf->data, buf->content, + object_size, buf->data, buf->content, cachefiles_coherency_set_ok); } @@ -103,10 +161,12 @@ int cachefiles_check_auxdata(struct cachefiles_object *object, struct file *file { struct cachefiles_xattr *buf; struct dentry *dentry = file->f_path.dentry; + struct inode *inode = file_inode(file); unsigned int len = object->cookie->aux_len, tlen; const void *p = fscache_get_aux(object->cookie); enum cachefiles_coherency_trace why; ssize_t xlen; + uoff_t obj_size; int ret = -ESTALE; tlen = sizeof(struct cachefiles_xattr) + len; @@ -121,34 +181,39 @@ int cachefiles_check_auxdata(struct cachefiles_object *object, struct file *file if (xlen != tlen) { if (xlen < 0) { ret = xlen; - trace_cachefiles_vfs_error(object, file_inode(file), xlen, + trace_cachefiles_vfs_error(object, inode, xlen, cachefiles_trace_getxattr_error); } if (xlen == -EIO) cachefiles_io_error_obj( object, "Failed to read aux with error %zd", xlen); + obj_size = 0; why = cachefiles_coherency_check_xattr; goto out; } + obj_size = be64_to_cpu(buf->object_size); if (buf->type != CACHEFILES_COOKIE_TYPE_DATA) { why = cachefiles_coherency_check_type; } else if (memcmp(buf->data, p, len) != 0) { why = cachefiles_coherency_check_aux; - } else if (be64_to_cpu(buf->object_size) != object->cookie->object_size) { + } else if (obj_size != object->cookie->object_size) { why = cachefiles_coherency_check_objsize; } else if (buf->content == CACHEFILES_CONTENT_DIRTY) { // TODO: Begin conflict resolution pr_warn("Dirty object in cache\n"); why = cachefiles_coherency_check_dirty; } else { + object->content_info = buf->content; + object->object_size = obj_size; + atomic64_set(&object->read_limit, i_size_read(inode)); why = cachefiles_coherency_check_ok; ret = 0; } out: - trace_cachefiles_coherency(object, file_inode(file)->i_ino, + trace_cachefiles_coherency(object, inode->i_ino, obj_size, buf->data, buf->content, why); kfree(buf); return ret; @@ -163,6 +228,9 @@ int cachefiles_remove_object_xattr(struct cachefiles_cache *cache, { int ret; + trace_cachefiles_coherency(object, d_inode(dentry)->i_ino, 0, NULL, 0, + cachefiles_coherency_remove); + ret = cachefiles_inject_remove_error(); if (ret == 0) { ret = mnt_want_write(cache->mnt); diff --git a/fs/ceph/Kconfig b/fs/ceph/Kconfig index 3d64a316ca31..aa6ccd7794d2 100644 --- a/fs/ceph/Kconfig +++ b/fs/ceph/Kconfig @@ -4,6 +4,7 @@ config CEPH_FS depends on INET select CEPH_LIB select NETFS_SUPPORT + select NETFS_PGPRIV2 select FS_ENCRYPTION_ALGS if FS_ENCRYPTION default n help diff --git a/fs/ceph/acl.c b/fs/ceph/acl.c index 85d3dd48b167..124f07ae5b2d 100644 --- a/fs/ceph/acl.c +++ b/fs/ceph/acl.c @@ -87,7 +87,7 @@ retry: return acl; } -int ceph_set_acl(struct mnt_idmap *idmap, struct dentry *dentry, +int ceph_set_acl(const struct mnt_idmap *idmap, struct dentry *dentry, struct posix_acl *acl, int type) { int ret = 0; diff --git a/fs/ceph/addr.c b/fs/ceph/addr.c index 0c00e9636b51..21d52d849842 100644 --- a/fs/ceph/addr.c +++ b/fs/ceph/addr.c @@ -65,7 +65,7 @@ (CONGESTION_ON_THRESH(congestion_kb) - \ (CONGESTION_ON_THRESH(congestion_kb) >> 2)) -static int ceph_netfs_check_write_begin(struct file *file, loff_t pos, unsigned int len, +static int ceph_netfs_check_write_begin(struct file *file, uoff_t pos, unsigned int len, struct folio **foliop, void **_fsdata); static inline struct ceph_snap_context *page_snap_context(struct page *page) @@ -1866,7 +1866,7 @@ ceph_find_incompatible(struct folio *folio) return NULL; } -static int ceph_netfs_check_write_begin(struct file *file, loff_t pos, unsigned int len, +static int ceph_netfs_check_write_begin(struct file *file, uoff_t pos, unsigned int len, struct folio **foliop, void **_fsdata) { struct inode *inode = file_inode(file); diff --git a/fs/ceph/dir.c b/fs/ceph/dir.c index 2e5c0ccb1b34..d9615d67bf1c 100644 --- a/fs/ceph/dir.c +++ b/fs/ceph/dir.c @@ -921,7 +921,7 @@ int ceph_handle_notrace_create(struct inode *dir, struct dentry *dentry) return PTR_ERR(result); } -static int ceph_mknod(struct mnt_idmap *idmap, struct inode *dir, +static int ceph_mknod(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *dentry, umode_t mode, dev_t rdev) { struct ceph_mds_client *mdsc = ceph_sb_to_mdsc(dir->i_sb); @@ -988,7 +988,7 @@ out: return err; } -static int ceph_create(struct mnt_idmap *idmap, struct inode *dir, +static int ceph_create(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *dentry, umode_t mode) { return ceph_mknod(idmap, dir, dentry, mode, 0); @@ -1032,7 +1032,7 @@ static int prep_encrypted_symlink_target(struct ceph_mds_request *req, } #endif -static int ceph_symlink(struct mnt_idmap *idmap, struct inode *dir, +static int ceph_symlink(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *dentry, const char *dest) { struct ceph_mds_client *mdsc = ceph_sb_to_mdsc(dir->i_sb); @@ -1106,7 +1106,7 @@ out: return err; } -static struct dentry *ceph_mkdir(struct mnt_idmap *idmap, struct inode *dir, +static struct dentry *ceph_mkdir(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *dentry, umode_t mode) { struct ceph_mds_client *mdsc = ceph_sb_to_mdsc(dir->i_sb); @@ -1478,7 +1478,7 @@ out: return err; } -static int ceph_rename(struct mnt_idmap *idmap, struct inode *old_dir, +static int ceph_rename(const struct mnt_idmap *idmap, struct inode *old_dir, struct dentry *old_dentry, struct inode *new_dir, struct dentry *new_dentry, unsigned int flags) { diff --git a/fs/ceph/file.c b/fs/ceph/file.c index bd3e3f5c269e..2c994c08ed4b 100644 --- a/fs/ceph/file.c +++ b/fs/ceph/file.c @@ -795,7 +795,7 @@ static int ceph_finish_async_create(struct inode *dir, struct inode *inode, int ceph_atomic_open(struct inode *dir, struct dentry *dentry, struct file *file, unsigned flags, umode_t mode) { - struct mnt_idmap *idmap = file_mnt_idmap(file); + const struct mnt_idmap *idmap = file_mnt_idmap(file); struct ceph_fs_client *fsc = ceph_sb_to_fs_client(dir->i_sb); struct ceph_client *cl = fsc->client; struct ceph_mds_client *mdsc = fsc->mdsc; diff --git a/fs/ceph/inode.c b/fs/ceph/inode.c index d52e2b389e0b..a695dba82554 100644 --- a/fs/ceph/inode.c +++ b/fs/ceph/inode.c @@ -2398,7 +2398,7 @@ static const char *ceph_encrypted_get_link(struct dentry *dentry, done); } -static int ceph_encrypted_symlink_getattr(struct mnt_idmap *idmap, +static int ceph_encrypted_symlink_getattr(const struct mnt_idmap *idmap, const struct path *path, struct kstat *stat, u32 request_mask, unsigned int query_flags) @@ -2568,7 +2568,7 @@ out: return ret; } -int __ceph_setattr(struct mnt_idmap *idmap, struct inode *inode, +int __ceph_setattr(const struct mnt_idmap *idmap, struct inode *inode, struct iattr *attr, struct ceph_iattr *cia) { struct ceph_inode_info *ci = ceph_inode(inode); @@ -2921,7 +2921,7 @@ out: /* * setattr */ -int ceph_setattr(struct mnt_idmap *idmap, struct dentry *dentry, +int ceph_setattr(const struct mnt_idmap *idmap, struct dentry *dentry, struct iattr *attr) { struct inode *inode = d_inode(dentry); @@ -3098,7 +3098,7 @@ out: * Check inode permissions. We verify we have a valid value for * the AUTH cap, then call the generic handler. */ -int ceph_permission(struct mnt_idmap *idmap, struct inode *inode, +int ceph_permission(const struct mnt_idmap *idmap, struct inode *inode, int mask) { int err; @@ -3145,7 +3145,7 @@ static int statx_to_caps(u32 want, umode_t mode) * Get all the attributes. If we have sufficient caps for the requested attrs, * then we can avoid talking to the MDS at all. */ -int ceph_getattr(struct mnt_idmap *idmap, const struct path *path, +int ceph_getattr(const struct mnt_idmap *idmap, const struct path *path, struct kstat *stat, u32 request_mask, unsigned int flags) { struct inode *inode = d_inode(path->dentry); diff --git a/fs/ceph/mds_client.h b/fs/ceph/mds_client.h index e7a262c9c2ab..ea48ec5383ef 100644 --- a/fs/ceph/mds_client.h +++ b/fs/ceph/mds_client.h @@ -375,7 +375,7 @@ struct ceph_mds_request { int r_fmode; /* file mode, if expecting cap */ int r_request_release_offset; const struct cred *r_cred; - struct mnt_idmap *r_mnt_idmap; + const struct mnt_idmap *r_mnt_idmap; struct timespec64 r_stamp; /* for choosing which mds to send this request to */ diff --git a/fs/ceph/super.h b/fs/ceph/super.h index 72d4e30304dc..a033331bb151 100644 --- a/fs/ceph/super.h +++ b/fs/ceph/super.h @@ -1166,18 +1166,18 @@ static inline int ceph_do_getattr(struct inode *inode, int mask, bool force) { return __ceph_do_getattr(inode, NULL, mask, force); } -extern int ceph_permission(struct mnt_idmap *idmap, +extern int ceph_permission(const struct mnt_idmap *idmap, struct inode *inode, int mask); struct ceph_iattr { struct ceph_fscrypt_auth *fscrypt_auth; }; -extern int __ceph_setattr(struct mnt_idmap *idmap, struct inode *inode, +extern int __ceph_setattr(const struct mnt_idmap *idmap, struct inode *inode, struct iattr *attr, struct ceph_iattr *cia); -extern int ceph_setattr(struct mnt_idmap *idmap, +extern int ceph_setattr(const struct mnt_idmap *idmap, struct dentry *dentry, struct iattr *attr); -extern int ceph_getattr(struct mnt_idmap *idmap, +extern int ceph_getattr(const struct mnt_idmap *idmap, const struct path *path, struct kstat *stat, u32 request_mask, unsigned int flags); void ceph_inode_shutdown(struct inode *inode); @@ -1252,7 +1252,7 @@ void ceph_release_acl_sec_ctx(struct ceph_acl_sec_ctx *as_ctx); #ifdef CONFIG_CEPH_FS_POSIX_ACL struct posix_acl *ceph_get_acl(struct inode *, int, bool); -int ceph_set_acl(struct mnt_idmap *idmap, +int ceph_set_acl(const struct mnt_idmap *idmap, struct dentry *dentry, struct posix_acl *acl, int type); int ceph_pre_init_acls(struct inode *dir, umode_t *mode, struct ceph_acl_sec_ctx *as_ctx); diff --git a/fs/ceph/xattr.c b/fs/ceph/xattr.c index cc4ffbbcb719..7d77214c76c6 100644 --- a/fs/ceph/xattr.c +++ b/fs/ceph/xattr.c @@ -1352,7 +1352,7 @@ static int ceph_get_xattr_handler(const struct xattr_handler *handler, } static int ceph_set_xattr_handler(const struct xattr_handler *handler, - struct mnt_idmap *idmap, + const struct mnt_idmap *idmap, struct dentry *unused, struct inode *inode, const char *name, const void *value, size_t size, int flags) diff --git a/fs/char_dev.c b/fs/char_dev.c index 00229e25c10f..5ce5423f6c99 100644 --- a/fs/char_dev.c +++ b/fs/char_dev.c @@ -280,7 +280,9 @@ int __register_chrdev(unsigned int major, unsigned int baseminor, cdev->owner = fops->owner; cdev->ops = fops; - kobject_set_name(&cdev->kobj, "%s", name); + err = kobject_set_name(&cdev->kobj, "%s", name); + if (err) + goto out; err = cdev_add(cdev, MKDEV(cd->major, baseminor), count); if (err) diff --git a/fs/coda/coda_linux.h b/fs/coda/coda_linux.h index dd6277d87afb..0c0d5f81653c 100644 --- a/fs/coda/coda_linux.h +++ b/fs/coda/coda_linux.h @@ -46,12 +46,12 @@ extern const struct file_operations coda_ioctl_operations; /* operations shared over more than one file */ int coda_open(struct inode *i, struct file *f); int coda_release(struct inode *i, struct file *f); -int coda_permission(struct mnt_idmap *idmap, struct inode *inode, +int coda_permission(const struct mnt_idmap *idmap, struct inode *inode, int mask); int coda_revalidate_inode(struct inode *); -int coda_getattr(struct mnt_idmap *, const struct path *, struct kstat *, +int coda_getattr(const struct mnt_idmap *, const struct path *, struct kstat *, u32, unsigned int); -int coda_setattr(struct mnt_idmap *, struct dentry *, struct iattr *); +int coda_setattr(const struct mnt_idmap *, struct dentry *, struct iattr *); /* this file: helpers */ char *coda_f2s(struct CodaFid *f); diff --git a/fs/coda/dir.c b/fs/coda/dir.c index 67148edfadee..a85be5962e62 100644 --- a/fs/coda/dir.c +++ b/fs/coda/dir.c @@ -73,7 +73,7 @@ static struct dentry *coda_lookup(struct inode *dir, struct dentry *entry, unsig } -int coda_permission(struct mnt_idmap *idmap, struct inode *inode, +int coda_permission(const struct mnt_idmap *idmap, struct inode *inode, int mask) { int error; @@ -133,7 +133,7 @@ static inline void coda_dir_drop_nlink(struct inode *dir) } /* creation routines: create, mknod, mkdir, link, symlink */ -static int coda_create(struct mnt_idmap *idmap, struct inode *dir, +static int coda_create(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *de, umode_t mode) { int error; @@ -166,7 +166,7 @@ err_out: return error; } -static struct dentry *coda_mkdir(struct mnt_idmap *idmap, struct inode *dir, +static struct dentry *coda_mkdir(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *de, umode_t mode) { struct inode *inode; @@ -233,7 +233,7 @@ static int coda_link(struct dentry *source_de, struct inode *dir_inode, } -static int coda_symlink(struct mnt_idmap *idmap, +static int coda_symlink(const struct mnt_idmap *idmap, struct inode *dir_inode, struct dentry *de, const char *symname) { @@ -300,7 +300,7 @@ static int coda_rmdir(struct inode *dir, struct dentry *de) } /* rename */ -static int coda_rename(struct mnt_idmap *idmap, struct inode *old_dir, +static int coda_rename(const struct mnt_idmap *idmap, struct inode *old_dir, struct dentry *old_dentry, struct inode *new_dir, struct dentry *new_dentry, unsigned int flags) { diff --git a/fs/coda/inode.c b/fs/coda/inode.c index 40b43866e6a5..c449954e23c2 100644 --- a/fs/coda/inode.c +++ b/fs/coda/inode.c @@ -294,7 +294,7 @@ static void coda_evict_inode(struct inode *inode) coda_cache_clear_inode(inode); } -int coda_getattr(struct mnt_idmap *idmap, const struct path *path, +int coda_getattr(const struct mnt_idmap *idmap, const struct path *path, struct kstat *stat, u32 request_mask, unsigned int flags) { int err = coda_revalidate_inode(d_inode(path->dentry)); @@ -304,7 +304,7 @@ int coda_getattr(struct mnt_idmap *idmap, const struct path *path, return err; } -int coda_setattr(struct mnt_idmap *idmap, struct dentry *de, +int coda_setattr(const struct mnt_idmap *idmap, struct dentry *de, struct iattr *iattr) { struct inode *inode = d_inode(de); diff --git a/fs/coda/pioctl.c b/fs/coda/pioctl.c index 36e35c15561a..c457e9bab94b 100644 --- a/fs/coda/pioctl.c +++ b/fs/coda/pioctl.c @@ -24,7 +24,7 @@ #include "coda_linux.h" /* pioctl ops */ -static int coda_ioctl_permission(struct mnt_idmap *idmap, +static int coda_ioctl_permission(const struct mnt_idmap *idmap, struct inode *inode, int mask); static long coda_pioctl(struct file *filp, unsigned int cmd, unsigned long user_data); @@ -41,7 +41,7 @@ const struct file_operations coda_ioctl_operations = { }; /* the coda pioctl inode ops */ -static int coda_ioctl_permission(struct mnt_idmap *idmap, +static int coda_ioctl_permission(const struct mnt_idmap *idmap, struct inode *inode, int mask) { return (mask & MAY_EXEC) ? -EACCES : 0; diff --git a/fs/configfs/configfs_internal.h b/fs/configfs/configfs_internal.h index acdeea8e2d69..5f627e58f135 100644 --- a/fs/configfs/configfs_internal.h +++ b/fs/configfs/configfs_internal.h @@ -76,7 +76,7 @@ extern int configfs_make_dirent(struct configfs_dirent *, struct dentry *, extern int configfs_dirent_is_ready(struct configfs_dirent *); extern const unsigned char * configfs_get_name(struct configfs_dirent *sd); -extern int configfs_setattr(struct mnt_idmap *idmap, +extern int configfs_setattr(const struct mnt_idmap *idmap, struct dentry *dentry, struct iattr *iattr); extern struct dentry *configfs_pin_fs(void); @@ -90,7 +90,7 @@ extern const struct inode_operations configfs_root_inode_operations; extern const struct inode_operations configfs_symlink_inode_operations; extern const struct dentry_operations configfs_dentry_ops; -extern int configfs_symlink(struct mnt_idmap *idmap, +extern int configfs_symlink(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *dentry, const char *symname); extern int configfs_unlink(struct inode *dir, struct dentry *dentry); @@ -104,17 +104,17 @@ static inline struct config_item * to_item(struct dentry * dentry) return ((struct config_item *) sd->s_element); } -static inline struct configfs_attribute * to_attr(struct dentry * dentry) +static inline const struct configfs_attribute * to_attr(struct dentry * dentry) { struct configfs_dirent * sd = dentry->d_fsdata; - return ((struct configfs_attribute *) sd->s_element); + return ((const struct configfs_attribute *) sd->s_element); } -static inline struct configfs_bin_attribute *to_bin_attr(struct dentry *dentry) +static inline const struct configfs_bin_attribute *to_bin_attr(struct dentry *dentry) { - struct configfs_attribute *attr = to_attr(dentry); + const struct configfs_attribute *attr = to_attr(dentry); - return container_of(attr, struct configfs_bin_attribute, cb_attr); + return container_of_const(attr, struct configfs_bin_attribute, cb_attr); } static inline struct config_item *configfs_get_config_item(struct dentry *dentry) diff --git a/fs/configfs/dir.c b/fs/configfs/dir.c index eda80c2a2d38..0c80feec5926 100644 --- a/fs/configfs/dir.c +++ b/fs/configfs/dir.c @@ -470,7 +470,7 @@ static struct dentry * configfs_lookup(struct inode *dir, */ if ((sd->s_type & CONFIGFS_NOT_PINNED) && !strcmp(configfs_get_name(sd), dentry->d_name.name)) { - struct configfs_attribute *attr = sd->s_element; + const struct configfs_attribute *attr = sd->s_element; umode_t mode = (attr->ca_mode & S_IALLUGO) | S_IFREG; dentry->d_fsdata = configfs_get(sd); @@ -631,8 +631,8 @@ static int populate_attrs(struct config_item *item) { const struct config_item_type *t = item->ci_type; const struct configfs_group_operations *ops; - struct configfs_attribute *attr; - struct configfs_bin_attribute *bin_attr; + const struct configfs_attribute *attr; + const struct configfs_bin_attribute *bin_attr; int error = 0; int i; @@ -641,8 +641,8 @@ static int populate_attrs(struct config_item *item) ops = t->ct_group_ops; - if (t->ct_attrs) { - for (i = 0; (attr = t->ct_attrs[i]) != NULL; i++) { + if (t->ct_attrs_const) { + for (i = 0; (attr = t->ct_attrs_const[i]) != NULL; i++) { if (ops && ops->is_visible && !ops->is_visible(item, attr, i)) continue; @@ -1295,7 +1295,7 @@ out_root_unlock: } EXPORT_SYMBOL(configfs_depend_item_unlocked); -static struct dentry *configfs_mkdir(struct mnt_idmap *idmap, struct inode *dir, +static struct dentry *configfs_mkdir(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *dentry, umode_t mode) { int ret = 0; diff --git a/fs/configfs/file.c b/fs/configfs/file.c index a48cece775a3..6460b000c593 100644 --- a/fs/configfs/file.c +++ b/fs/configfs/file.c @@ -41,8 +41,8 @@ struct configfs_buffer { struct config_item *item; struct module *owner; union { - struct configfs_attribute *attr; - struct configfs_bin_attribute *bin_attr; + const struct configfs_attribute *attr; + const struct configfs_bin_attribute *bin_attr; }; }; @@ -291,7 +291,7 @@ static int __configfs_open_file(struct inode *inode, struct file *file, int type { struct dentry *dentry = file->f_path.dentry; struct configfs_fragment *frag = to_frag(file); - struct configfs_attribute *attr; + const struct configfs_attribute *attr; struct configfs_buffer *buffer; int error; diff --git a/fs/configfs/inode.c b/fs/configfs/inode.c index 68290fe0e374..c92a05251a47 100644 --- a/fs/configfs/inode.c +++ b/fs/configfs/inode.c @@ -32,7 +32,7 @@ static const struct inode_operations configfs_inode_operations ={ .setattr = configfs_setattr, }; -int configfs_setattr(struct mnt_idmap *idmap, struct dentry *dentry, +int configfs_setattr(const struct mnt_idmap *idmap, struct dentry *dentry, struct iattr *iattr) { struct inode * inode = d_inode(dentry); @@ -178,7 +178,7 @@ struct inode *configfs_create(struct dentry *dentry, umode_t mode) */ const unsigned char * configfs_get_name(struct configfs_dirent *sd) { - struct configfs_attribute *attr; + const struct configfs_attribute *attr; BUG_ON(!sd || !sd->s_element); diff --git a/fs/configfs/symlink.c b/fs/configfs/symlink.c index 3b31c714400f..89178f5d371a 100644 --- a/fs/configfs/symlink.c +++ b/fs/configfs/symlink.c @@ -146,7 +146,7 @@ static int get_target(const char *symname, struct config_item **target, } -int configfs_symlink(struct mnt_idmap *idmap, struct inode *dir, +int configfs_symlink(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *dentry, const char *symname) { int ret; diff --git a/fs/coredump.c b/fs/coredump.c index 4d03826dd249..4eac73eabe9b 100644 --- a/fs/coredump.c +++ b/fs/coredump.c @@ -39,6 +39,7 @@ #include <linux/oom.h> #include <linux/compat.h> #include <linux/fs.h> +#include <linux/wait_bit.h> #include <linux/path.h> #include <linux/timekeeping.h> #include <linux/sysctl.h> @@ -51,7 +52,6 @@ #include <net/sock.h> #include <uapi/linux/pidfd.h> #include <uapi/linux/un.h> -#include <uapi/linux/coredump.h> #include <linux/uaccess.h> #include <asm/mmu_context.h> @@ -68,6 +68,8 @@ static bool dump_vma_snapshot(struct coredump_params *cprm); static void free_vma_snapshot(struct coredump_params *cprm); +static void dump_end_record(struct coredump_params *cprm); +static bool dump_flush_skip(struct coredump_params *cprm); #define CORE_FILE_NOTE_SIZE_DEFAULT (4*1024*1024) /* Define a reasonable max cap */ @@ -83,6 +85,8 @@ static int core_uses_pid; static unsigned int core_pipe_limit; static unsigned int core_sort_vma; static char core_pattern[CORENAME_MAX_SIZE] = "core"; +/* Taken around every copy in and out of core_pattern. */ +static DEFINE_SPINLOCK(core_pattern_lock); static int core_name_size = CORENAME_MAX_SIZE; unsigned int core_file_note_size_limit = CORE_FILE_NOTE_SIZE_DEFAULT; static atomic_t core_pipe_count = ATOMIC_INIT(0); @@ -98,9 +102,7 @@ struct core_name { char *corename __counted_by_ptr(size); int used, size; unsigned int core_pipe_limit; - bool core_dumped; enum coredump_type_t core_type; - u64 mask; }; static int expand_corename(struct core_name *cn, int size) @@ -240,18 +242,22 @@ static bool coredump_parse(struct core_name *cn, struct coredump_params *cprm, size_t **argv, int *argc) { const struct cred *cred = current_cred(); - const char *pat_ptr = core_pattern; + char pattern[CORENAME_MAX_SIZE]; + const char *pat_ptr = pattern; bool was_space = false; int pid_in_pattern = 0; int err = 0; - cn->mask = COREDUMP_KERNEL; + /* The sysctl handler may be publishing a new pattern. */ + scoped_guard(spinlock, &core_pattern_lock) + strscpy(pattern, core_pattern); + + cprm->mask = COREDUMP_KERNEL; if (core_pipe_limit) - cn->mask |= COREDUMP_WAIT; + cprm->mask |= COREDUMP_WAIT; cn->used = 0; cn->corename = NULL; cn->core_pipe_limit = 0; - cn->core_dumped = false; if (*pat_ptr == '|') cn->core_type = COREDUMP_PIPE; else if (*pat_ptr == '@') @@ -508,60 +514,64 @@ static int zap_threads(struct task_struct *tsk, int nr = -EAGAIN; spin_lock_irq(&tsk->sighand->siglock); - if (!(signal->flags & SIGNAL_GROUP_EXIT) && !signal->group_exec_task) { + /* A freeze requested before the dump would be lost with TIF_SIGPENDING. */ + if (!(signal->flags & SIGNAL_GROUP_EXIT) && !signal->group_exec_task && + !freezing(tsk) && !(tsk->jobctl & JOBCTL_TRAP_FREEZE)) { /* Allow SIGKILL, see prepare_signal() */ signal->core_state = core_state; nr = zap_process(signal, exit_code); clear_tsk_thread_flag(tsk, TIF_SIGPENDING); tsk->flags |= PF_DUMPCORE; - atomic_set(&core_state->nr_threads, nr); + atomic_set(&core_state->threads_remaining, nr); } spin_unlock_irq(&tsk->sighand->siglock); return nr; } +static void coredump_wait_inactive(struct core_state *core_state) +{ + struct core_thread *ptr; + + wait_var_event_state(&core_state->threads_remaining, + !atomic_read_acquire(&core_state->threads_remaining), + TASK_UNINTERRUPTIBLE | TASK_FREEZABLE); + /* + * Wait for all the threads to become inactive, so that + * all the thread context (extended register state, like + * fpu etc) gets copied to the memory. + */ + for (ptr = core_state->tasks; ptr; ptr = ptr->next) + wait_task_inactive(ptr->task, TASK_ANY); +} + static int coredump_wait(int exit_code, struct core_state *core_state) { struct task_struct *tsk = current; int core_waiters = -EBUSY; - init_completion(&core_state->startup); - core_state->dumper.task = tsk; - core_state->dumper.next = NULL; + core_state->tasks = NULL; core_waiters = zap_threads(tsk, core_state, exit_code); - if (core_waiters > 0) { - struct core_thread *ptr; - - wait_for_completion_state(&core_state->startup, - TASK_UNINTERRUPTIBLE|TASK_FREEZABLE); - /* - * Wait for all the threads to become inactive, so that - * all the thread context (extended register state, like - * fpu etc) gets copied to the memory. - */ - ptr = core_state->dumper.next; - while (ptr != NULL) { - wait_task_inactive(ptr->task, TASK_ANY); - ptr = ptr->next; - } - } + if (core_waiters > 0) + coredump_wait_inactive(core_state); return core_waiters; } -static void coredump_finish(bool core_dumped) +static void coredump_finish(enum coredump_state state) { struct core_thread *curr, *next; struct task_struct *task; spin_lock_irq(¤t->sighand->siglock); - if (core_dumped && !__fatal_signal_pending(current)) + if ((state & COREDUMP_STATE_STARTED) && !__fatal_signal_pending(current)) current->signal->group_exit_code |= 0x80; - next = current->signal->core_state->dumper.next; + next = current->signal->core_state->tasks; current->signal->core_state = NULL; spin_unlock_irq(¤t->sighand->siglock); + /* A released thread may exit and be freed before it is woken. */ + guard(rcu)(); while ((curr = next) != NULL) { next = curr->next; task = curr->task; @@ -570,6 +580,7 @@ static void coredump_finish(bool core_dumped) * ->task == NULL before we read ->next. */ smp_mb(); + /* Any wakeup now lets the thread exit, rcu keeps it alive. */ curr->task = NULL; wake_up_process(task); } @@ -577,13 +588,8 @@ static void coredump_finish(bool core_dumped) static bool dump_interrupted(void) { - /* - * SIGKILL or freezing() interrupt the coredumping. Perhaps we - * can do try_to_freeze() and check __fatal_signal_pending(), - * but then we need to teach dump_write() to restart and clear - * TIF_SIGPENDING. - */ - return fatal_signal_pending(current) || freezing(current); + /* Only SIGKILL and the freezers set it after zap_threads(). */ + return task_sigpending(current); } static void wait_for_dump_helpers(struct file *file) @@ -664,7 +670,12 @@ static int umh_coredump_setup(struct subprocess_info *info, struct cred *new) return 0; } +static_assert(sizeof(struct coredump_record_header) == COREDUMP_RECORD_HEADER_SIZE_VER0); + #ifdef CONFIG_UNIX +/* af_unix halves the send buffer to size a single skb. */ +#define COREDUMP_SOCK_SNDBUF_MIN (3 * PAGE_SIZE) + static bool coredump_sock_connect(struct core_name *cn, struct coredump_params *cprm) { struct file *file __free(fput) = NULL; @@ -690,6 +701,10 @@ static bool coredump_sock_connect(struct core_name *cn, struct coredump_params * if (retval < 0) return false; + /* Don't let a page-sized write split into several skbs. */ + socket->sk->sk_sndbuf = max_t(int, socket->sk->sk_sndbuf, + COREDUMP_SOCK_SNDBUF_MIN); + file = sock_alloc_file(socket, 0, NULL); if (IS_ERR(file)) return false; @@ -752,8 +767,39 @@ static inline bool coredump_sock_send(struct file *file, struct coredump_req *re return ret == sizeof(*req); } +static_assert(sizeof(struct coredump_req) == COREDUMP_REQ_SIZE_VER1); +static_assert(sizeof(struct coredump_ack) == COREDUMP_ACK_SIZE_VER1); static_assert(sizeof(enum coredump_mark) == sizeof(__u32)); +/* Every memory type this kernel knows. */ +#define COREDUMP_MEMORY_ALL \ + (COREDUMP_MEMORY_ANON_PRIVATE | COREDUMP_MEMORY_ANON_SHARED | \ + COREDUMP_MEMORY_FILE_PRIVATE | COREDUMP_MEMORY_FILE_SHARED | \ + COREDUMP_MEMORY_ELF_HEADERS | \ + COREDUMP_MEMORY_HUGETLB_PRIVATE | COREDUMP_MEMORY_HUGETLB_SHARED | \ + COREDUMP_MEMORY_DAX_PRIVATE | COREDUMP_MEMORY_DAX_SHARED) + +#define COREDUMP_MEMORY_TYPE_BIT(mmf) BIT((mmf) - MMF_DUMP_FILTER_SHIFT) +static_assert(COREDUMP_MEMORY_ALL == (MMF_DUMP_FILTER_MASK >> MMF_DUMP_FILTER_SHIFT)); +static_assert(COREDUMP_MEMORY_ANON_PRIVATE == + COREDUMP_MEMORY_TYPE_BIT(MMF_DUMP_ANON_PRIVATE)); +static_assert(COREDUMP_MEMORY_ANON_SHARED == + COREDUMP_MEMORY_TYPE_BIT(MMF_DUMP_ANON_SHARED)); +static_assert(COREDUMP_MEMORY_FILE_PRIVATE == + COREDUMP_MEMORY_TYPE_BIT(MMF_DUMP_MAPPED_PRIVATE)); +static_assert(COREDUMP_MEMORY_FILE_SHARED == + COREDUMP_MEMORY_TYPE_BIT(MMF_DUMP_MAPPED_SHARED)); +static_assert(COREDUMP_MEMORY_ELF_HEADERS == + COREDUMP_MEMORY_TYPE_BIT(MMF_DUMP_ELF_HEADERS)); +static_assert(COREDUMP_MEMORY_HUGETLB_PRIVATE == + COREDUMP_MEMORY_TYPE_BIT(MMF_DUMP_HUGETLB_PRIVATE)); +static_assert(COREDUMP_MEMORY_HUGETLB_SHARED == + COREDUMP_MEMORY_TYPE_BIT(MMF_DUMP_HUGETLB_SHARED)); +static_assert(COREDUMP_MEMORY_DAX_PRIVATE == + COREDUMP_MEMORY_TYPE_BIT(MMF_DUMP_DAX_PRIVATE)); +static_assert(COREDUMP_MEMORY_DAX_SHARED == + COREDUMP_MEMORY_TYPE_BIT(MMF_DUMP_DAX_SHARED)); + static inline bool coredump_sock_mark(struct file *file, enum coredump_mark mark) { struct msghdr msg = { .msg_flags = MSG_NOSIGNAL }; @@ -795,10 +841,14 @@ static inline void coredump_sock_shutdown(struct file *file) static bool coredump_sock_request(struct core_name *cn, struct coredump_params *cprm) { struct coredump_req req = { - .size = sizeof(struct coredump_req), - .mask = COREDUMP_KERNEL | COREDUMP_USERSPACE | - COREDUMP_REJECT | COREDUMP_WAIT, - .size_ack = sizeof(struct coredump_ack), + .size = sizeof(struct coredump_req), + .mask = COREDUMP_KERNEL | COREDUMP_USERSPACE | + COREDUMP_REJECT | COREDUMP_WAIT | + COREDUMP_RECORDS | COREDUMP_SPARSE | + COREDUMP_MEMORY_TYPES, + .size_ack = sizeof(struct coredump_ack), + .memory_types = cprm->memory_types, + .memory_types_mask = COREDUMP_MEMORY_ALL, }; struct coredump_ack ack = {}; ssize_t usize; @@ -851,7 +901,54 @@ static bool coredump_sock_request(struct core_name *cn, struct coredump_params * return false; } - cn->mask = ack.mask; + /* Records only describe a coredump the kernel writes. */ + if ((ack.mask & COREDUMP_RECORDS) && !(ack.mask & COREDUMP_KERNEL)) { + coredump_sock_mark(cprm->file, COREDUMP_MARK_CONFLICTING); + return false; + } + + /* Zero records only exist inside a record stream. */ + if ((ack.mask & COREDUMP_SPARSE) && !(ack.mask & COREDUMP_RECORDS)) { + coredump_sock_mark(cprm->file, COREDUMP_MARK_CONFLICTING); + return false; + } + + if (ack.mask & COREDUMP_MEMORY_TYPES) { + /* The memory types need the whole field. */ + if (usize < COREDUMP_ACK_SIZE_VER1) { + coredump_sock_mark(cprm->file, COREDUMP_MARK_MINSIZE); + return false; + } + + /* The memory types only select what the kernel writes. */ + if (!(ack.mask & COREDUMP_KERNEL)) { + coredump_sock_mark(cprm->file, COREDUMP_MARK_CONFLICTING); + return false; + } + + /* Refuse unknown memory types. */ + if (ack.memory_types & ~req.memory_types_mask) { + coredump_sock_mark(cprm->file, COREDUMP_MARK_UNSUPPORTED); + return false; + } + } else if (ack.memory_types) { + /* Like @spare the field must be zero when it isn't used. */ + coredump_sock_mark(cprm->file, COREDUMP_MARK_UNSUPPORTED); + return false; + } + + /* Record header scratch; a bvec can't point at the stack. */ + if (ack.mask & COREDUMP_RECORDS) { + cprm->record_hdr = kmalloc_obj(*cprm->record_hdr); + if (!cprm->record_hdr) + return false; + } + + /* The server's selection replaces the task's entirely. */ + if (ack.mask & COREDUMP_MEMORY_TYPES) + cprm->memory_types = ack.memory_types; + + cprm->mask = ack.mask; return coredump_sock_mark(cprm->file, COREDUMP_MARK_REQACK); } @@ -878,7 +975,7 @@ static inline bool coredump_force_suid_safe(const struct coredump_params *cprm) static bool coredump_file(struct core_name *cn, struct coredump_params *cprm, const struct linux_binfmt *binfmt) { - struct mnt_idmap *idmap; + const struct mnt_idmap *idmap; struct inode *inode; struct file *file __free(fput) = NULL; int open_flags = O_CREAT | O_WRONLY | O_NOFOLLOW | O_LARGEFILE | O_EXCL; @@ -1032,29 +1129,41 @@ static bool coredump_pipe(struct core_name *cn, struct coredump_params *cprm, return true; } -static bool coredump_write(struct core_name *cn, - struct coredump_params *cprm, - const struct linux_binfmt *binfmt) +static bool coredump_write(struct coredump_params *cprm, + const struct linux_binfmt *binfmt) { - - if (dump_interrupted()) + if (dump_interrupted()) { + cprm->state |= COREDUMP_STATE_TRUNCATED; return true; + } - if (!dump_vma_snapshot(cprm)) + if (!dump_vma_snapshot(cprm)) { + cprm->state |= COREDUMP_STATE_TRUNCATED; return false; + } file_start_write(cprm->file); - cn->core_dumped = binfmt->core_dump(cprm); + if (!binfmt->core_dump(cprm)) + cprm->state |= COREDUMP_STATE_TRUNCATED; /* - * Ensures that file size is big enough to contain the current - * file postion. This prevents gdb from complaining about - * a truncated file if the last "write" to the file was - * dump_skip. + * A trailing hole still has to land in the coredump. Seeking over + * it doesn't grow the file, so the last byte of it is written + * instead and gdb doesn't see a truncated file. Everything else + * puts the hole on the wire as it flushes it. */ if (cprm->to_skip) { - cprm->to_skip--; - dump_emit(cprm, "", 1); + bool flushed; + + if (cprm->file->f_mode & FMODE_LSEEK) { + cprm->to_skip--; + flushed = dump_emit(cprm, "", 1); + } else { + flushed = dump_flush_skip(cprm); + } + if (!flushed) + cprm->state |= COREDUMP_STATE_TRUNCATED; } + dump_end_record(cprm); file_end_write(cprm->file); free_vma_snapshot(cprm); return true; @@ -1069,7 +1178,8 @@ static void coredump_cleanup(struct core_name *cn, struct coredump_params *cprm) atomic_dec(&core_pipe_count); } kfree(cn->corename); - coredump_finish(cn->core_dumped); + kfree(cprm->record_hdr); + coredump_finish(cprm->state); } static inline bool coredump_skip(const struct coredump_params *cprm, @@ -1115,29 +1225,24 @@ static void do_coredump(struct core_name *cn, struct coredump_params *cprm, } /* Don't even generate the coredump. */ - if (cn->mask & COREDUMP_REJECT) - return; - - /* get us an unshared descriptor table; almost always a no-op */ - /* The cell spufs coredump code reads the file descriptor tables */ - if (unshare_files()) + if (cprm->mask & COREDUMP_REJECT) return; - if ((cn->mask & COREDUMP_KERNEL) && !coredump_write(cn, cprm, binfmt)) + if ((cprm->mask & COREDUMP_KERNEL) && !coredump_write(cprm, binfmt)) return; coredump_sock_shutdown(cprm->file); /* Let the parent know that a coredump was generated. */ - if (cn->mask & COREDUMP_USERSPACE) - cn->core_dumped = true; + if (cprm->mask & COREDUMP_USERSPACE) + cprm->state |= COREDUMP_STATE_STARTED; /* * When core_pipe_limit is set we wait for the coredump server * or usermodehelper to finish before exiting so it can e.g., * inspect /proc/<pid>. */ - if (cn->mask & COREDUMP_WAIT) { + if (cprm->mask & COREDUMP_WAIT) { switch (cn->core_type) { case COREDUMP_PIPE: wait_for_dump_helpers(cprm->file); @@ -1153,6 +1258,10 @@ static void do_coredump(struct core_name *cn, struct coredump_params *cprm, } } +#define COREDUMP_TASK_MEMORY_TYPES(mm) \ + ((__mm_flags_get_word((mm)) & MMF_DUMP_FILTER_MASK) >> \ + MMF_DUMP_FILTER_SHIFT) + void vfs_coredump(const kernel_siginfo_t *siginfo) { size_t *argv __free(kfree) = NULL; @@ -1164,8 +1273,8 @@ void vfs_coredump(const kernel_siginfo_t *siginfo) struct coredump_params cprm = { .siginfo = siginfo, .limit = rlimit(RLIMIT_CORE), - /* Snapshot MMF_DUMP_FILTER_* (unlocked) and dumpable for the dump. */ - .mm_flags = __mm_flags_get_word(mm), + /* Snapshot the memory types (unlocked) and dumpable for the dump. */ + .memory_types = COREDUMP_TASK_MEMORY_TYPES(mm), .dumpable = task_exec_state_get_dumpable(current), .vma_meta = NULL, .cpu = raw_smp_processor_id(), @@ -1191,6 +1300,8 @@ void vfs_coredump(const kernel_siginfo_t *siginfo) if (coredump_wait(siginfo->si_signo, &core_state) < 0) return; + /* Task work must not cut the dump short, see signal_pending(). */ + guard(no_notify_signal)(); scoped_with_creds(cred) do_coredump(&cn, &cprm, &argv, &argc, binfmt); coredump_cleanup(&cn, &cprm); @@ -1202,60 +1313,181 @@ void vfs_coredump(const kernel_siginfo_t *siginfo) * do on a core-file: use only these functions to write out all the * necessary info. */ -static int __dump_emit(struct coredump_params *cprm, const void *addr, int nr) +static bool dump_records(const struct coredump_params *cprm) +{ + return cprm->mask & COREDUMP_RECORDS; +} + +static bool dump_sparse(const struct coredump_params *cprm) +{ + return cprm->mask & COREDUMP_SPARSE; +} + +/* Describe the next @len bytes of the coredump. Returns the header size. */ +static size_t dump_record_init(struct coredump_params *cprm, + enum coredump_record_type type, u64 flags, + u64 len) +{ + if (!dump_records(cprm)) + return 0; + + *cprm->record_hdr = (struct coredump_record_header) { + .size = sizeof(*cprm->record_hdr), + .type = type, + .flags = flags, + .offset = cprm->pos, + .len = len, + }; + + return sizeof(*cprm->record_hdr); +} + +/* Write @iter whole or fail. @len is what it advances the coredump by. */ +static bool dump_write_iter(struct coredump_params *cprm, struct iov_iter *iter, + size_t len) { struct file *file = cprm->file; + size_t count = iov_iter_count(iter); loff_t pos = file->f_pos; ssize_t n; - if (cprm->written + nr > cprm->limit) - return 0; - if (dump_interrupted()) - return 0; - n = __kernel_write(file, addr, nr, &pos); - if (n != nr) - return 0; + n = __kernel_write_iter(file, iter, &pos); + if (n != (ssize_t)count) + return false; file->f_pos = pos; - cprm->written += n; - cprm->pos += n; + cprm->written += count; + cprm->pos += len; + + return true; +} + +/* One record, never more than a page. See __dump_emit(). */ +static bool dump_emit_chunk(struct coredump_params *cprm, const void *addr, + int nr) +{ + struct kvec kvec[2]; + struct iov_iter iter; + unsigned int nseg = 0; + size_t hdrlen; - return 1; + if (dump_interrupted()) + return false; + + hdrlen = dump_record_init(cprm, COREDUMP_RECORD_DATA, 0, nr); + if (hdrlen) { + kvec[nseg].iov_base = cprm->record_hdr; + kvec[nseg].iov_len = hdrlen; + nseg++; + } + kvec[nseg].iov_base = (void *)addr; + kvec[nseg].iov_len = nr; + nseg++; + + iov_iter_kvec(&iter, ITER_SOURCE, kvec, nseg, hdrlen + nr); + + return dump_write_iter(cprm, &iter, nr); +} + +static bool __dump_emit(struct coredump_params *cprm, const void *addr, int nr) +{ + if (cprm->written + nr > cprm->limit) + return false; + + while (nr) { + int chunk = min_t(int, nr, PAGE_SIZE); + + if (!dump_emit_chunk(cprm, addr, chunk)) + return false; + + addr += chunk; + nr -= chunk; + } + + return true; } -static int __dump_skip(struct coredump_params *cprm, size_t nr) +/* Send a record that stands on its own: a header and nothing else. */ +static bool dump_emit_record(struct coredump_params *cprm, + enum coredump_record_type type, u64 flags, u64 len) +{ + struct kvec kvec; + struct iov_iter iter; + size_t hdrlen; + + hdrlen = dump_record_init(cprm, type, flags, len); + if (!hdrlen) + return false; + + kvec.iov_base = cprm->record_hdr; + kvec.iov_len = hdrlen; + iov_iter_kvec(&iter, ITER_SOURCE, &kvec, 1, hdrlen); + + return dump_write_iter(cprm, &iter, len); +} + +/* Close the record stream. Only a whole coredump gets an end record. */ +static void dump_end_record(struct coredump_params *cprm) +{ + if (cprm->state & COREDUMP_STATE_TRUNCATED) + return; + + dump_emit_record(cprm, COREDUMP_RECORD_END, 0, 0); +} + +static bool __dump_skip(struct coredump_params *cprm, size_t nr) { static char zeroes[PAGE_SIZE]; struct file *file = cprm->file; + if (dump_sparse(cprm)) { + /* Hand the server the length of the hole instead of the hole itself. */ + if (dump_interrupted()) + return false; + return dump_emit_record(cprm, COREDUMP_RECORD_ZERO, 0, nr); + } + if (file->f_mode & FMODE_LSEEK) { if (dump_interrupted() || vfs_llseek(file, nr, SEEK_CUR) < 0) - return 0; + return false; cprm->pos += nr; - return 1; + return true; } - while (nr > PAGE_SIZE) { - if (!__dump_emit(cprm, zeroes, PAGE_SIZE)) - return 0; - nr -= PAGE_SIZE; + while (nr) { + size_t chunk = min_t(size_t, nr, PAGE_SIZE); + + if (!__dump_emit(cprm, zeroes, chunk)) + return false; + + nr -= chunk; } - return __dump_emit(cprm, zeroes, nr); + return true; } -int dump_emit(struct coredump_params *cprm, const void *addr, int nr) +/* Flush the accumulated hole before writing data. */ +static bool dump_flush_skip(struct coredump_params *cprm) { if (cprm->to_skip) { if (!__dump_skip(cprm, cprm->to_skip)) - return 0; + return false; cprm->to_skip = 0; } + return true; +} + +bool dump_emit(struct coredump_params *cprm, const void *addr, int nr) +{ + if (!dump_flush_skip(cprm)) + return false; return __dump_emit(cprm, addr, nr); } EXPORT_SYMBOL(dump_emit); void dump_skip_to(struct coredump_params *cprm, unsigned long pos) { + if (WARN_ON_ONCE(pos < cprm->pos)) + return; cprm->to_skip = pos - cprm->pos; } EXPORT_SYMBOL(dump_skip_to); @@ -1267,37 +1499,32 @@ void dump_skip(struct coredump_params *cprm, size_t nr) EXPORT_SYMBOL(dump_skip); #ifdef CONFIG_ELF_CORE -static int dump_emit_page(struct coredump_params *cprm, struct page *page) +static bool dump_emit_page(struct coredump_params *cprm, struct page *page) { - struct bio_vec bvec; + struct bio_vec bvec[2]; struct iov_iter iter; - struct file *file = cprm->file; - loff_t pos; - ssize_t n; + unsigned int nseg = 0; + size_t hdrlen; if (!page) - return 0; + return false; - if (cprm->to_skip) { - if (!__dump_skip(cprm, cprm->to_skip)) - return 0; - cprm->to_skip = 0; - } + if (!dump_flush_skip(cprm)) + return false; if (cprm->written + PAGE_SIZE > cprm->limit) - return 0; + return false; if (dump_interrupted()) - return 0; - pos = file->f_pos; - bvec_set_page(&bvec, page, PAGE_SIZE, 0); - iov_iter_bvec(&iter, ITER_SOURCE, &bvec, 1, PAGE_SIZE); - n = __kernel_write_iter(cprm->file, &iter, &pos); - if (n != PAGE_SIZE) - return 0; - file->f_pos = pos; - cprm->written += PAGE_SIZE; - cprm->pos += PAGE_SIZE; + return false; - return 1; + /* Hand the record header to the same write as the page it describes. */ + hdrlen = dump_record_init(cprm, COREDUMP_RECORD_DATA, 0, PAGE_SIZE); + if (hdrlen) + bvec_set_virt(&bvec[nseg++], cprm->record_hdr, hdrlen); + bvec_set_page(&bvec[nseg++], page, PAGE_SIZE, 0); + + iov_iter_bvec(&iter, ITER_SOURCE, bvec, nseg, hdrlen + PAGE_SIZE); + + return dump_write_iter(cprm, &iter, PAGE_SIZE); } /* @@ -1329,18 +1556,19 @@ static inline struct page *dump_page_copy(struct page *src, struct page *dst) } #endif -int dump_user_range(struct coredump_params *cprm, unsigned long start, - unsigned long len) +bool dump_user_range(struct coredump_params *cprm, unsigned long start, + unsigned long len) { unsigned long addr; struct page *dump_page; - int locked, ret; + int locked; + bool ret; dump_page = dump_page_alloc(); if (!dump_page) - return 0; + return false; - ret = 0; + ret = false; locked = 0; for (addr = start; addr < start + len; addr += PAGE_SIZE) { struct page *page; @@ -1364,7 +1592,7 @@ int dump_user_range(struct coredump_params *cprm, unsigned long start, mmap_read_unlock(current->mm); locked = 0; } - int stop = !dump_emit_page(cprm, dump_page_copy(page, dump_page)); + bool stop = !dump_emit_page(cprm, dump_page_copy(page, dump_page)); put_page(page); if (stop) goto out; @@ -1383,7 +1611,7 @@ int dump_user_range(struct coredump_params *cprm, unsigned long start, } cond_resched(); } - ret = 1; + ret = true; out: if (locked) mmap_read_unlock(current->mm); @@ -1393,14 +1621,14 @@ out: } #endif -int dump_align(struct coredump_params *cprm, int align) +bool dump_align(struct coredump_params *cprm, int align) { unsigned mod = (cprm->pos + cprm->to_skip) & (align - 1); if (align & (align - 1)) - return 0; + return false; if (mod) cprm->to_skip += align - mod; - return 1; + return true; } EXPORT_SYMBOL(dump_align); @@ -1417,11 +1645,11 @@ void validate_coredump_safety(void) } } -static inline bool check_coredump_socket(void) +static inline bool check_coredump_socket(const char *pattern) { const char *p; - if (core_pattern[0] != '@') + if (pattern[0] != '@') return true; /* @@ -1433,16 +1661,16 @@ static inline bool check_coredump_socket(void) return false; /* Must be an absolute path... */ - if (core_pattern[1] != '/') { + if (pattern[1] != '/') { /* ... or the socket request protocol... */ - if (core_pattern[1] != '@') + if (pattern[1] != '@') return false; /* ... and if so must be an absolute path. */ - if (core_pattern[2] != '/') + if (pattern[2] != '/') return false; - p = &core_pattern[2]; + p = &pattern[2]; } else { - p = &core_pattern[1]; + p = &pattern[1]; } /* The path obviously cannot exceed UNIX_PATH_MAX. */ @@ -1450,7 +1678,7 @@ static inline bool check_coredump_socket(void) return false; /* Must not contain ".." in the path. */ - if (name_contains_dotdot(core_pattern)) + if (name_contains_dotdot(pattern)) return false; return true; @@ -1459,27 +1687,35 @@ static inline bool check_coredump_socket(void) static int proc_dostring_coredump(const struct ctl_table *table, int write, void *buffer, size_t *lenp, loff_t *ppos) { + char pattern[CORENAME_MAX_SIZE]; + const struct ctl_table tmp = { + .procname = table->procname, + .data = pattern, + .maxlen = sizeof(pattern), + }; + bool changed = false; int error; - ssize_t retval; - char old_core_pattern[CORENAME_MAX_SIZE]; - if (!write) - return proc_dostring(table, write, buffer, lenp, ppos); + /* Work on a copy, proc_dostring() appends at *ppos. */ + scoped_guard(spinlock, &core_pattern_lock) + strscpy(pattern, core_pattern); - retval = strscpy(old_core_pattern, core_pattern, CORENAME_MAX_SIZE); - - error = proc_dostring(table, write, buffer, lenp, ppos); - if (error) + error = proc_dostring(&tmp, write, buffer, lenp, ppos); + if (error || !write) return error; - if (!check_coredump_socket()) { - strscpy(core_pattern, old_core_pattern, retval + 1); + if (!check_coredump_socket(pattern)) return -EINVAL; - } - if (strncmp(old_core_pattern, core_pattern, CORENAME_MAX_SIZE)) + /* Publish the validated pattern whole. */ + scoped_guard(spinlock, &core_pattern_lock) { + changed = strncmp(pattern, core_pattern, CORENAME_MAX_SIZE); + if (changed) + strscpy(core_pattern, pattern); + } + if (changed) validate_coredump_safety(); - return error; + return 0; } static const unsigned int core_file_note_size_min = CORE_FILE_NOTE_SIZE_DEFAULT; @@ -1582,15 +1818,15 @@ static bool always_dump_vma(struct vm_area_struct *vma) } #define DUMP_SIZE_MAYBE_ELFHDR_PLACEHOLDER 1 +#define COREDUMP_MEMORY_TYPE_INCLUDE(types, type) \ + ((types) & COREDUMP_MEMORY_##type) /* * Decide how much of @vma's contents should be included in a core dump. */ static unsigned long vma_dump_size(struct vm_area_struct *vma, - unsigned long mm_flags) + u64 memory_types) { -#define FILTER(type) (mm_flags & (1UL << MMF_DUMP_##type)) - /* always dump the vdso and vsyscall sections */ if (always_dump_vma(vma)) goto whole; @@ -1600,18 +1836,22 @@ static unsigned long vma_dump_size(struct vm_area_struct *vma, /* support for DAX */ if (vma_is_dax(vma)) { - if ((vma->vm_flags & VM_SHARED) && FILTER(DAX_SHARED)) + if ((vma->vm_flags & VM_SHARED) && + COREDUMP_MEMORY_TYPE_INCLUDE(memory_types, DAX_SHARED)) goto whole; - if (!(vma->vm_flags & VM_SHARED) && FILTER(DAX_PRIVATE)) + if (!(vma->vm_flags & VM_SHARED) && + COREDUMP_MEMORY_TYPE_INCLUDE(memory_types, DAX_PRIVATE)) goto whole; return 0; } /* Hugetlb memory check */ if (vma_is_hugetlb(vma)) { - if ((vma->vm_flags & VM_SHARED) && FILTER(HUGETLB_SHARED)) + if ((vma->vm_flags & VM_SHARED) && + COREDUMP_MEMORY_TYPE_INCLUDE(memory_types, HUGETLB_SHARED)) goto whole; - if (!(vma->vm_flags & VM_SHARED) && FILTER(HUGETLB_PRIVATE)) + if (!(vma->vm_flags & VM_SHARED) && + COREDUMP_MEMORY_TYPE_INCLUDE(memory_types, HUGETLB_PRIVATE)) goto whole; return 0; } @@ -1623,25 +1863,27 @@ static unsigned long vma_dump_size(struct vm_area_struct *vma, /* By default, dump shared memory if mapped from an anonymous file. */ if (vma->vm_flags & VM_SHARED) { if (file_inode(vma->vm_file)->i_nlink == 0 ? - FILTER(ANON_SHARED) : FILTER(MAPPED_SHARED)) + COREDUMP_MEMORY_TYPE_INCLUDE(memory_types, ANON_SHARED) : + COREDUMP_MEMORY_TYPE_INCLUDE(memory_types, FILE_SHARED)) goto whole; return 0; } /* Dump segments that have been written to. */ - if ((!IS_ENABLED(CONFIG_MMU) || vma->anon_vma) && FILTER(ANON_PRIVATE)) + if ((!IS_ENABLED(CONFIG_MMU) || vma->anon_vma) && + COREDUMP_MEMORY_TYPE_INCLUDE(memory_types, ANON_PRIVATE)) goto whole; if (vma->vm_file == NULL) return 0; - if (FILTER(MAPPED_PRIVATE)) + if (COREDUMP_MEMORY_TYPE_INCLUDE(memory_types, FILE_PRIVATE)) goto whole; /* * If this is the beginning of an executable file mapping, * dump the first page to aid in determining what was mapped here. */ - if (FILTER(ELF_HEADERS) && + if (COREDUMP_MEMORY_TYPE_INCLUDE(memory_types, ELF_HEADERS) && vma->vm_pgoff == 0 && (vma->vm_flags & VM_READ)) { if ((READ_ONCE(file_inode(vma->vm_file)->i_mode) & 0111) != 0) return PAGE_SIZE; @@ -1657,8 +1899,6 @@ static unsigned long vma_dump_size(struct vm_area_struct *vma, return DUMP_SIZE_MAYBE_ELFHDR_PLACEHOLDER; } -#undef FILTER - return 0; whole: @@ -1743,7 +1983,7 @@ static bool dump_vma_snapshot(struct coredump_params *cprm) m->start = vma->vm_start; m->end = vma->vm_end; m->flags = vma->vm_flags; - m->dump_size = vma_dump_size(vma, cprm->mm_flags); + m->dump_size = vma_dump_size(vma, cprm->memory_types); m->pgoff = vma->vm_pgoff; m->file = vma->vm_file; if (m->file) diff --git a/fs/crypto/Kconfig b/fs/crypto/Kconfig index cd934e31dec4..4cfbf4b51be7 100644 --- a/fs/crypto/Kconfig +++ b/fs/crypto/Kconfig @@ -5,7 +5,7 @@ config FS_ENCRYPTION select BLK_INLINE_ENCRYPTION_FALLBACK if BLOCK select CRYPTO select CRYPTO_SKCIPHER - select CRYPTO_LIB_AES + select CRYPTO_LIB_AES_ECB # for deprecated v1 key derivation function select CRYPTO_LIB_SHA256 select CRYPTO_LIB_SHA512 select KEYS diff --git a/fs/crypto/crypto.c b/fs/crypto/crypto.c index aced5c50a460..b6a253217a5a 100644 --- a/fs/crypto/crypto.c +++ b/fs/crypto/crypto.c @@ -97,6 +97,11 @@ void fscrypt_generate_iv(union fscrypt_iv *iv, u64 index, iv->index = cpu_to_le64(index); } +typedef enum { + FS_DECRYPT = 0, + FS_ENCRYPT, +} fscrypt_direction_t; + /* Encrypt or decrypt a single "data unit" of file contents. */ static int fscrypt_crypt_data_unit(const struct fscrypt_inode_info *ci, fscrypt_direction_t rw, u64 index, diff --git a/fs/crypto/fscrypt_private.h b/fs/crypto/fscrypt_private.h index 74329e0953d1..e602c31b3844 100644 --- a/fs/crypto/fscrypt_private.h +++ b/fs/crypto/fscrypt_private.h @@ -219,15 +219,6 @@ fscrypt_policy_du_bits(const union fscrypt_policy *policy, BUG(); } -/* - * For encrypted symlinks, the ciphertext length is stored at the beginning - * of the string in little-endian format. - */ -struct fscrypt_symlink_data { - __le16 len; - char encrypted_path[]; -} __packed; - /** * struct fscrypt_prepared_key - a key prepared for actual encryption/decryption * @tfm: crypto API transform object @@ -321,11 +312,6 @@ struct fscrypt_inode_info { u8 ci_nonce[FSCRYPT_FILE_NONCE_SIZE]; }; -typedef enum { - FS_DECRYPT = 0, - FS_ENCRYPT, -} fscrypt_direction_t; - /* crypto.c */ extern struct kmem_cache *fscrypt_inode_info_cachep; int fscrypt_initialize(struct super_block *sb); @@ -474,19 +460,6 @@ fscrypt_is_key_prepared(const struct fscrypt_prepared_key *prep_key, /* keyring.c */ /* - * fscrypt_master_key_user - a user's claim to a master key - */ -struct fscrypt_master_key_user { - struct list_head link; - kuid_t uid; - /* - * This 'struct key' contains no secret. It exists solely to charge the - * appropriate user's key quota. - */ - struct key *quota_key; -}; - -/* * fscrypt_master_key_secret - secret key material of an in-use master key */ struct fscrypt_master_key_secret { diff --git a/fs/crypto/hooks.c b/fs/crypto/hooks.c index a7a8a3f581a0..149498bdab94 100644 --- a/fs/crypto/hooks.c +++ b/fs/crypto/hooks.c @@ -9,6 +9,15 @@ #include "fscrypt_private.h" +/* + * For encrypted symlinks, the ciphertext length is stored at the beginning + * of the string in little-endian format. + */ +struct fscrypt_symlink_data { + __le16 len; + char encrypted_path[]; +} __packed; + /** * fscrypt_file_open() - prepare to open a possibly-encrypted regular file * @inode: the inode being opened diff --git a/fs/crypto/keyring.c b/fs/crypto/keyring.c index 76e28d1e0064..874cc183ee62 100644 --- a/fs/crypto/keyring.c +++ b/fs/crypto/keyring.c @@ -145,6 +145,19 @@ static inline bool valid_key_spec(const struct fscrypt_key_specifier *spec) return master_key_spec_len(spec) != 0; } +/* + * fscrypt_master_key_user - a user's claim to a master key + */ +struct fscrypt_master_key_user { + struct list_head link; + kuid_t uid; + /* + * This 'struct key' contains no secret. It exists solely to charge the + * appropriate user's key quota. + */ + struct key *quota_key; +}; + static int fscrypt_user_key_instantiate(struct key *key, struct key_preparsed_payload *prep) { diff --git a/fs/crypto/keysetup_v1.c b/fs/crypto/keysetup_v1.c index 87fe13ccb253..d9ac7c4328f8 100644 --- a/fs/crypto/keysetup_v1.c +++ b/fs/crypto/keysetup_v1.c @@ -20,7 +20,7 @@ * managed alongside the master keys in the filesystem-level keyring) */ -#include <crypto/aes.h> +#include <crypto/aes-ecb.h> #include <crypto/utils.h> #include <keys/user-type.h> #include <linux/hashtable.h> @@ -246,8 +246,7 @@ static int setup_v1_file_key_derived(struct fscrypt_inode_info *ci, static_assert(FSCRYPT_FILE_NONCE_SIZE == AES_KEYSIZE_128); aes_prepareenckey(&aes, ci->ci_nonce, FSCRYPT_FILE_NONCE_SIZE); - for (unsigned int i = 0; i < derived_keysize; i += AES_BLOCK_SIZE) - aes_encrypt(&aes, &derived_key[i], &raw_master_key[i]); + aes_ecb_encrypt(derived_key, raw_master_key, derived_keysize, &aes); err = fscrypt_set_per_file_enc_key(ci, derived_key); @@ -469,8 +469,6 @@ static void dax_folio_init(void *entry) if (order > 0) { prep_compound_page(&folio->page, order); - if (order > 1) - INIT_LIST_HEAD(&folio->_deferred_list); WARN_ON_ONCE(folio_ref_count(folio)); } } @@ -775,24 +773,23 @@ fallback: /** * dax_layout_busy_page_range - find first pinned page in @mapping - * @mapping: address space to scan for a page with ref count > 1 + * @mapping: address space to scan for a pinned page * @start: Starting offset. Page containing 'start' is included. * @end: End offset. Page containing 'end' is included. If 'end' is LLONG_MAX, * pages from 'start' till the end of file are included. * - * DAX requires ZONE_DEVICE mapped pages. These pages are never - * 'onlined' to the page allocator so they are considered idle when - * page->count == 1. A filesystem uses this interface to determine if - * any page in the mapping is busy, i.e. for DMA, or other - * get_user_pages() usages. + * DAX requires ZONE_DEVICE mapped pages. A page is considered busy when + * folio_ref_count(folio) exceeds folio_mapcount(folio). This helper is + * used to determine if any page in the mapping is busy, i.e. for DMA, + * or other get_user_pages() usages. * * It is expected that the filesystem is holding locks to block the * establishment of new mappings in this address_space. I.e. it expects - * to be able to run unmap_mapping_range() and subsequently not race + * to be able to run unmap_mapping_pages() and subsequently not race * mapping_mapped() becoming true. */ -struct page *dax_layout_busy_page_range(struct address_space *mapping, - loff_t start, loff_t end) +static struct page *dax_layout_busy_page_range(struct address_space *mapping, + loff_t start, loff_t end) { void *entry; unsigned int scanned = 0; @@ -844,13 +841,6 @@ struct page *dax_layout_busy_page_range(struct address_space *mapping, xas_unlock_irq(&xas); return page; } -EXPORT_SYMBOL_GPL(dax_layout_busy_page_range); - -struct page *dax_layout_busy_page(struct address_space *mapping) -{ - return dax_layout_busy_page_range(mapping, 0, LLONG_MAX); -} -EXPORT_SYMBOL_GPL(dax_layout_busy_page); static int __dax_invalidate_entry(struct address_space *mapping, pgoff_t index, bool trunc) diff --git a/fs/dcache.c b/fs/dcache.c index a66be85f9d01..7a9346c4f2e4 100644 --- a/fs/dcache.c +++ b/fs/dcache.c @@ -32,6 +32,7 @@ #include <linux/bit_spinlock.h> #include <linux/rculist_bl.h> #include <linux/list_lru.h> +#include <linux/namei.h> #include "internal.h" #include "mount.h" @@ -451,6 +452,17 @@ static void dentry_free(struct dentry *dentry) } /* + * If inode is unlinked and doesn't have any aliases (i.e., all fds pointing to + * it are closed), it is pretty much dead. Except that file handle lookup could + * still revive it which causes issues to fsnotify. So once inode reaches this + * state we make sure to block creating any new aliases. + */ +static bool inode_notify_dead(struct inode *inode) +{ + return !inode->i_nlink && hlist_empty(&inode->i_dentry); +} + +/* * Release the dentry's inode, using the filesystem * d_iput() operation if defined. */ @@ -459,6 +471,7 @@ static void dentry_unlink_inode(struct dentry * dentry) __releases(dentry->d_inode->i_lock) { struct inode *inode = dentry->d_inode; + bool notify_dead; raw_write_seqcount_begin(&dentry->d_seq); __d_clear_type_and_inode(dentry); @@ -469,9 +482,10 @@ static void dentry_unlink_inode(struct dentry * dentry) */ dentry->waiters = NULL; raw_write_seqcount_end(&dentry->d_seq); + notify_dead = inode_notify_dead(inode); spin_unlock(&dentry->d_lock); spin_unlock(&inode->i_lock); - if (!inode->i_nlink) + if (notify_dead) fsnotify_inoderemove(inode); if (dentry->d_op && dentry->d_op->d_iput) dentry->d_op->d_iput(dentry, inode); @@ -830,7 +844,7 @@ static struct dentry *dentry_kill(struct dentry *dentry) if (dentry->d_op && dentry->d_op->d_release) dentry->d_op->d_release(dentry); - cond_resched(); + cond_resched_tasks_rcu_qs(); /* now that it's negative, ->d_parent is stable */ if (!IS_ROOT(dentry)) { parent = dentry->d_parent; @@ -1900,6 +1914,7 @@ EXPORT_SYMBOL(d_invalidate); static struct dentry *__d_alloc(struct super_block *sb, const struct qstr *name) { + static struct lock_class_key __lookup_key; struct dentry *dentry; char *dname; int err; @@ -1961,6 +1976,8 @@ static struct dentry *__d_alloc(struct super_block *sb, const struct qstr *name) dentry->waiters = NULL; INIT_HLIST_NODE(&dentry->d_sib); + lockdep_init_map(&dentry->lookup_map, "DCACHE_PAR_LOOKUP", &__lookup_key, 0); + if (dentry->d_op && dentry->d_op->d_init) { err = dentry->d_op->d_init(dentry); if (err) { @@ -2003,6 +2020,58 @@ struct dentry *d_alloc(struct dentry * parent, const struct qstr *name) } EXPORT_SYMBOL(d_alloc); +/** + * d_duplicate - duplicate a dentry for combined atomic operation + * @dentry: the dentry to duplicate + * + * Some rename operations need to be combined with another operation + * inside the filesystem. + * 1/ A cluster filesystem when renaming to an in-use file might need to + * first "silly-rename" that target out of the way before the main rename + * 2/ A filesystem that supports white-out might want to create a whiteout + * in place of the file being moved. + * + * For this they need two dentries which temporarily have the same name, + * before one is renamed. d_duplicate() provides for this. Given a + * positive hashed dentry, it creates a second in-lookup dentry. + * Because the original dentry exists, no other thread will try to + * create an in-lookup dentry, so there can be no race in this create. + * + * The caller should d_move() the original to a new name, often via a + * rename request, and should call d_lookup_done() on the newly created + * dentry. If the new is instantiated then the old MUST either be moved + * or dropped. + * + * Parent must be locked. + * + * Returns: an in-lookup dentry, or -ENOMEM. + */ +struct dentry *d_duplicate(struct dentry *dentry) +{ + unsigned int hash = dentry->d_name.hash; + struct dentry *parent = dentry->d_parent; + struct hlist_bl_head *b = in_lookup_hash(parent, hash); + struct dentry *new = __d_alloc(parent->d_sb, &dentry->d_name); + + if (unlikely(!new)) + return ERR_PTR(-ENOMEM); + + new->d_flags |= DCACHE_PAR_LOOKUP; + lock_map_acquire_try(&new->lookup_map); + spin_lock(&parent->d_lock); + new->d_parent = dget_dlock(parent); + hlist_add_head(&new->d_sib, &parent->d_children); + if (parent->d_flags & DCACHE_DISCONNECTED) + new->d_flags |= DCACHE_DISCONNECTED; + spin_unlock(&parent->d_lock); + + hlist_bl_lock(b); + hlist_bl_add_head(&new->d_in_lookup_hash, b); + hlist_bl_unlock(b); + return new; +} +EXPORT_SYMBOL(d_duplicate); + struct dentry *d_alloc_anon(struct super_block *sb) { return __d_alloc(sb, NULL); @@ -2172,7 +2241,6 @@ static void __d_instantiate(struct dentry *dentry, struct inode *inode) * (or otherwise set) by the caller to indicate that it is now * in use by the dcache. */ - void d_instantiate(struct dentry *entry, struct inode * inode) { BUG_ON(d_really_is_positive(entry)); @@ -2241,7 +2309,12 @@ static struct dentry *__d_obtain_alias(struct inode *inode, bool disconnected) sb = inode->i_sb; - res = d_find_any_alias(inode); /* existing alias? */ + spin_lock(&inode->i_lock); + if (!inode_notify_dead(inode)) + res = __d_find_any_alias(inode); /* existing alias? */ + else + res = ERR_PTR(-ESTALE); + spin_unlock(&inode->i_lock); if (res) goto out; @@ -2253,7 +2326,10 @@ static struct dentry *__d_obtain_alias(struct inode *inode, bool disconnected) security_d_instantiate(new, inode); spin_lock(&inode->i_lock); - res = __d_find_any_alias(inode); /* recheck under lock */ + if (!inode_notify_dead(inode)) + res = __d_find_any_alias(inode); /* recheck under lock */ + else + res = ERR_PTR(-ESTALE); if (likely(!res)) { /* still no alias, attach a disconnected dentry */ unsigned add_flags = d_flags_for_inode(inode); @@ -2754,6 +2830,15 @@ static inline void end_dir_add(struct inode *dir, unsigned int n) static void d_wait_lookup(struct dentry *dentry) { if (likely(d_in_lookup(dentry))) { + /* + * Tell lockdep we will wait for the lookup lock, after + * dropping ->d_lock, but won't actually take it. + */ + spin_release(&dentry->d_lock.dep_map, _THIS_IP_); + lock_map_acquire(&dentry->lookup_map); + lock_map_release(&dentry->lookup_map); + spin_acquire(&dentry->d_lock.dep_map, 0, 1, _THIS_IP_); + dentry->d_flags |= DCACHE_LOOKUP_WAITERS; wait_var_event_spinlock(&dentry->d_flags, !d_in_lookup(dentry), @@ -2761,8 +2846,16 @@ static void d_wait_lookup(struct dentry *dentry) } } -struct dentry *d_alloc_parallel(struct dentry *parent, - const struct qstr *name) +/* What to do when __d_alloc_parallel finds a d_in_lookup dentry */ +enum alloc_para { + ALLOC_PARA_WAIT, + ALLOC_PARA_FAIL, +}; + +static inline +struct dentry *__d_alloc_parallel(struct dentry *parent, + const struct qstr *name, + enum alloc_para how) { unsigned int hash = name->hash; struct hlist_bl_head *b = in_lookup_hash(parent, hash); @@ -2835,6 +2928,12 @@ retry: spin_unlock(&dentry->d_lock); goto retry; } + if (unlikely(how == ALLOC_PARA_FAIL)) { + /* mustn't wait for concurrent lookup to complete */ + spin_unlock(&dentry->d_lock); + dput(new); + return ERR_PTR(-EWOULDBLOCK); + } /* * somebody is likely to be still doing lookup for it; * pin it and wait for them to finish @@ -2862,14 +2961,77 @@ retry: } hlist_bl_add_head(&new->d_in_lookup_hash, b); hlist_bl_unlock(b); + lock_map_acquire_try(&new->lookup_map); return new; mismatch: spin_unlock(&dentry->d_lock); dput(dentry); goto retry; } + +/** + * d_alloc_parallel() - allocate a new dentry and ensure uniqueness + * @parent: dentry of the parent + * @name: name of the dentry within that parent. + * + * A new dentry is allocated and, providing it is unique, added to the + * relevant index. + * If an existing dentry is found with the same parent/name that is + * not d_in_lookup(), then that is returned instead. + * If the existing dentry is d_in_lookup(), d_alloc_parallel() waits for + * that lookup to complete before returning the dentry and then ensures the + * match is still valid. + * Thus if the returned dentry is d_in_lookup() then the caller has + * exclusive access until it completes the lookup. + * If the returned dentry is not d_in_lookup() then a lookup has + * already completed. + * + * The @name must already have ->hash set, as can be achieved + * by e.g. try_lookup_noperm(). + * + * Returns: the dentry, whether found or allocated, or an error %-ENOMEM. + */ +struct dentry *d_alloc_parallel(struct dentry *parent, + const struct qstr *name) +{ + return __d_alloc_parallel(parent, name, ALLOC_PARA_WAIT); +} EXPORT_SYMBOL(d_alloc_parallel); +/** + * d_alloc_trylock() - find or allocate a new dentry + * @parent: dentry of the parent + * @name: name of the dentry within that parent. + * + * A new dentry is allocated and, providing it is unique, added to the + * relevant index. + * If an existing dentry is found with the same parent/name that is + * not d_in_lookup() then that is returned instead. + * If the existing dentry is d_in_lookup(), d_alloc_trylock() + * returns with error %-EWOULDBLOCK. + * Thus if the returned dentry is d_in_lookup() then the caller has + * exclusive access until it completes the lookup. + * If the returned dentry is not d_in_lookup() then a lookup has + * already completed. + * + * The @name need not already have ->hash set. + * + * Returns: the dentry, whether found or allocated, or an error + * %-ENOMEM, %-EWOULDBLOCK, %-EACCES (for a bad name) or + * anything returned by ->d_hash(). + */ +struct dentry *d_alloc_trylock(struct dentry *parent, + struct qstr *name) +{ + struct dentry *de; + + de = try_lookup_noperm(name, parent); + if (!de) + de = __d_alloc_parallel(parent, name, ALLOC_PARA_FAIL); + return de; +} +EXPORT_SYMBOL(d_alloc_trylock); + /* * Move dentry from in-lookup state to busy-negative one. * @@ -2898,6 +3060,7 @@ static void __d_lookup_unhash(struct dentry *dentry) b = in_lookup_hash(dentry->d_parent, dentry->d_name.hash); hlist_bl_lock(b); dentry->d_flags &= ~DCACHE_PAR_LOOKUP; + lock_map_release(&dentry->lookup_map); __hlist_bl_del(&dentry->d_in_lookup_hash); hlist_bl_unlock(b); dentry->waiters = NULL; @@ -2935,15 +3098,10 @@ static inline void __d_add(struct dentry *dentry, struct inode *inode, } if (unlikely(ops)) d_set_d_op(dentry, ops); - if (inode) { - unsigned add_flags = d_flags_for_inode(inode); - hlist_add_head(&dentry->d_alias, &inode->i_dentry); - raw_write_seqcount_begin(&dentry->d_seq); - __d_set_inode_and_type(dentry, inode, add_flags); - raw_write_seqcount_end(&dentry->d_seq); - fsnotify_update_flags(dentry); - } - __d_rehash(dentry); + if (inode) + __d_instantiate(dentry, inode); + if (d_unhashed(dentry)) + __d_rehash(dentry); if (dir) { end_dir_add(dir, n); __d_wake_in_lookup_waiters(dentry); @@ -3245,7 +3403,7 @@ struct dentry *d_splice_alias_ops(struct inode *inode, struct dentry *dentry, if (IS_ERR(inode)) return ERR_CAST(inode); - BUG_ON(!d_unhashed(dentry)); + BUG_ON(d_really_is_positive(dentry)); if (!inode) goto out; @@ -3301,6 +3459,8 @@ out: * @inode: the inode which may have a disconnected dentry * @dentry: a negative dentry which we want to point to the inode. * + * @dentry must be negative and may be in-lookup or unhashed or hashed. + * * If inode is a directory and has an IS_ROOT alias, then d_move that in * place of the given dentry and return it, else simply d_add the inode * to the dentry and return NULL. @@ -3308,16 +3468,14 @@ out: * If a non-IS_ROOT directory is found, the filesystem is corrupt, and * we should error out: directories can't have multiple aliases. * - * This is needed in the lookup routine of any filesystem that is exportable - * (via knfsd) so that we can build dcache paths to directories effectively. + * This should be used to return the result of ->lookup() and to + * instantiate the result of ->mkdir(), is often useful for + * ->atomic_open, and may be used to instantiate other objects. * * If a dentry was found and moved, then it is returned. Otherwise NULL - * is returned. This matches the expected return value of ->lookup. + * is returned. This matches the expected return value of ->lookup and + * ->mkdir. * - * Cluster filesystems may call this function with a negative, hashed dentry. - * In that case, we know that the inode will be a regular file, and also this - * will only occur during atomic_open. So we need to check for the dentry - * being already hashed only in the final case. */ struct dentry *d_splice_alias(struct inode *inode, struct dentry *dentry) { diff --git a/fs/debugfs/inode.c b/fs/debugfs/inode.c index a4d08bd3743b..b4915551ad31 100644 --- a/fs/debugfs/inode.c +++ b/fs/debugfs/inode.c @@ -42,7 +42,7 @@ static bool debugfs_enabled __ro_after_init = IS_ENABLED(CONFIG_DEBUG_FS_ALLOW_A * so that we can use the file mode as part of a heuristic to determine whether * to lock down individual files. */ -static int debugfs_setattr(struct mnt_idmap *idmap, +static int debugfs_setattr(const struct mnt_idmap *idmap, struct dentry *dentry, struct iattr *ia) { int ret; diff --git a/fs/devpts/inode.c b/fs/devpts/inode.c index 9844dcf354ee..bd1e8eb26edb 100644 --- a/fs/devpts/inode.c +++ b/fs/devpts/inode.c @@ -249,6 +249,8 @@ static int devpts_parse_param(struct fs_context *fc, struct fs_parameter *param) case Opt_max: if (result.uint_32 > NR_UNIX98_PTY_MAX) return invalf(fc, "max out of range"); + if (result.uint_32 == 0) + return invalf(fc, "max must be greater than 0"); opts->max = result.uint_32; break; } diff --git a/fs/dlm/config.c b/fs/dlm/config.c index 53cd33293042..6c5c3f049b33 100644 --- a/fs/dlm/config.c +++ b/fs/dlm/config.c @@ -64,11 +64,37 @@ static void release_node(struct config_item *); static struct configfs_attribute *comm_attrs[]; static struct configfs_attribute *node_attrs[]; +static u32 rsb_hashfn(const void *data, u32 len, u32 seed) +{ + const struct dlm_rsb_key *key = data; + + return jhash(key->name, key->len, 0); +} + +static u32 rsb_obj_hashfn(const void *data, u32 len, u32 seed) +{ + const struct dlm_rsb *r = data; + + return r->res_hash; +} + +static int rsb_obj_cmpfn(struct rhashtable_compare_arg *arg, const void *obj) +{ + const struct dlm_rsb_key *key = arg->key; + const struct dlm_rsb *r = obj; + + if (key->len != r->res_length) + return -1; + + return memcmp(&r->res_name, key->name, key->len); +} + const struct rhashtable_params dlm_rhash_rsb_params = { .nelem_hint = 3, /* start small */ - .key_len = DLM_RESNAME_MAXLEN, - .key_offset = offsetof(struct dlm_rsb, res_name), .head_offset = offsetof(struct dlm_rsb, res_node), + .hashfn = rsb_hashfn, + .obj_hashfn = rsb_obj_hashfn, + .obj_cmpfn = rsb_obj_cmpfn, .automatic_shrinking = true, }; diff --git a/fs/dlm/dlm_internal.h b/fs/dlm/dlm_internal.h index 9df842421ae0..74c2e77d1b55 100644 --- a/fs/dlm/dlm_internal.h +++ b/fs/dlm/dlm_internal.h @@ -340,6 +340,11 @@ struct dlm_rsb { char res_name[DLM_RESNAME_MAXLEN+1]; }; +struct dlm_rsb_key { + char name[DLM_RESNAME_MAXLEN]; + size_t len; +}; + /* dlm_master_lookup() flags */ #define DLM_LU_RECOVER_DIR 1 diff --git a/fs/dlm/lock.c b/fs/dlm/lock.c index c381e1028446..2609e4fdeba8 100644 --- a/fs/dlm/lock.c +++ b/fs/dlm/lock.c @@ -622,13 +622,17 @@ static int get_rsb_struct(struct dlm_ls *ls, const void *name, int len, return 0; } -int dlm_search_rsb_tree(struct rhashtable *rhash, const void *name, int len, - struct dlm_rsb **r_ret) +int dlm_search_rsb_tree(struct rhashtable *rhash, const void *name, + unsigned int len, struct dlm_rsb **r_ret) { - char key[DLM_RESNAME_MAXLEN] = {}; + struct dlm_rsb_key key = { + .len = len, + }; + if (len > DLM_RESNAME_MAXLEN) return -EINVAL; - memcpy(key, name, len); + + memcpy(key.name, name, len); *r_ret = rhashtable_lookup_fast(rhash, &key, dlm_rhash_rsb_params); if (*r_ret) return 0; @@ -1421,7 +1425,7 @@ void dlm_dump_rsb_name(struct dlm_ls *ls, const char *name, int len) rcu_read_lock(); error = dlm_search_rsb_tree(&ls->ls_rsbtbl, name, len, &r); - if (!error) + if (error) goto out; dlm_dump_rsb(r); @@ -5529,6 +5533,10 @@ static int receive_rcom_lock_args(struct dlm_ls *ls, struct dlm_lkb *lkb, { struct rcom_lock *rl = (struct rcom_lock *) rc->rc_buf; + if (rl->rl_rqmode < DLM_LOCK_IV || rl->rl_rqmode > DLM_LOCK_EX || + rl->rl_grmode < DLM_LOCK_IV || rl->rl_grmode > DLM_LOCK_EX) + return -EINVAL; + lkb->lkb_nodeid = le32_to_cpu(rc->rc_header.h_nodeid); lkb->lkb_ownpid = le32_to_cpu(rl->rl_ownpid); lkb->lkb_remid = le32_to_cpu(rl->rl_lkid); diff --git a/fs/dlm/lock.h b/fs/dlm/lock.h index b23d7b854ed4..c75975937331 100644 --- a/fs/dlm/lock.h +++ b/fs/dlm/lock.h @@ -31,8 +31,8 @@ void resume_scan_timer(struct dlm_ls *ls); int dlm_master_lookup(struct dlm_ls *ls, int from_nodeid, const char *name, int len, unsigned int flags, int *r_nodeid, int *result); -int dlm_search_rsb_tree(struct rhashtable *rhash, const void *name, int len, - struct dlm_rsb **r_ret); +int dlm_search_rsb_tree(struct rhashtable *rhash, const void *name, + unsigned int len, struct dlm_rsb **r_ret); void dlm_recover_purge(struct dlm_ls *ls, const struct list_head *root_list); void dlm_purge_mstcpy_locks(struct dlm_rsb *r); diff --git a/fs/dlm/lowcomms.c b/fs/dlm/lowcomms.c index 2aff1c7c17de..ea8353c4638d 100644 --- a/fs/dlm/lowcomms.c +++ b/fs/dlm/lowcomms.c @@ -1984,4 +1984,5 @@ void dlm_lowcomms_exit(void) } } srcu_read_unlock(&connections_srcu, idx); + srcu_barrier(&connections_srcu); } diff --git a/fs/dlm/midcomms.c b/fs/dlm/midcomms.c index 8964164600d2..045431524494 100644 --- a/fs/dlm/midcomms.c +++ b/fs/dlm/midcomms.c @@ -1178,6 +1178,7 @@ void dlm_midcomms_exit(void) } } srcu_read_unlock(&nodes_srcu, idx); + srcu_barrier(&nodes_srcu); dlm_lowcomms_exit(); } diff --git a/fs/dlm/plock.c b/fs/dlm/plock.c index e9598b3fe5d0..711e8bc3a46a 100644 --- a/fs/dlm/plock.c +++ b/fs/dlm/plock.c @@ -4,6 +4,7 @@ */ #include <linux/fs.h> +#include <linux/capability.h> #include <linux/filelock.h> #include <linux/miscdevice.h> #include <linux/poll.h> @@ -477,6 +478,15 @@ out: } EXPORT_SYMBOL_GPL(dlm_posix_get); +static int dev_open(struct inode *inode, struct file *file) +{ + /* Userspace plock daemon is a privileged cluster component. */ + if (!capable(CAP_SYS_ADMIN)) + return -EPERM; + + return 0; +} + /* a read copies out one plock request from the send list */ static ssize_t dev_read(struct file *file, char __user *u, size_t count, loff_t *ppos) @@ -598,6 +608,7 @@ static __poll_t dev_poll(struct file *file, poll_table *wait) } static const struct file_operations dev_fops = { + .open = dev_open, .read = dev_read, .write = dev_write, .poll = dev_poll, @@ -608,7 +619,8 @@ static const struct file_operations dev_fops = { static struct miscdevice plock_dev_misc = { .minor = MISC_DYNAMIC_MINOR, .name = DLM_PLOCK_MISC_NAME, - .fops = &dev_fops + .fops = &dev_fops, + .mode = 0600, }; int dlm_plock_init(void) diff --git a/fs/dlm/user.c b/fs/dlm/user.c index a8ed4c8fdc5b..0b0a7e1bd109 100644 --- a/fs/dlm/user.c +++ b/fs/dlm/user.c @@ -4,6 +4,7 @@ */ #include <linux/miscdevice.h> +#include <linux/capability.h> #include <linux/init.h> #include <linux/wait.h> #include <linux/file.h> @@ -511,6 +512,7 @@ static ssize_t device_write(struct file *file, const char __user *buf, size_t count, loff_t *ppos) { struct dlm_user_proc *proc = file->private_data; + size_t name_payload = 0; struct dlm_write_request *kbuf; int error; @@ -544,6 +546,7 @@ static ssize_t device_write(struct file *file, const char __user *buf, if (count > sizeof(struct dlm_write_request32)) namelen = count - sizeof(struct dlm_write_request32); + name_payload = namelen; k32buf = (struct dlm_write_request32 *)kbuf; @@ -560,7 +563,13 @@ static ssize_t device_write(struct file *file, const char __user *buf, compat_input(kbuf, k32buf, namelen); kfree(k32buf); + } else { + if (count > sizeof(*kbuf)) + name_payload = count - sizeof(*kbuf); } +#else + if (count > sizeof(*kbuf)) + name_payload = count - sizeof(*kbuf); #endif /* do we really need this? can a write happen after a close? */ @@ -570,6 +579,15 @@ static ssize_t device_write(struct file *file, const char __user *buf, goto out_free; } + if (kbuf->cmd == DLM_USER_LOCK && + !(kbuf->i.lock.flags & DLM_LKF_CONVERT)) { + if (kbuf->i.lock.namelen > name_payload || + kbuf->i.lock.namelen > DLM_RESNAME_MAXLEN) { + error = -EINVAL; + goto out_free; + } + } + error = -EINVAL; switch (kbuf->cmd) @@ -910,6 +928,10 @@ static int ctl_device_close(struct inode *inode, struct file *file) static int monitor_device_open(struct inode *inode, struct file *file) { + /* dlm_controld is the only expected opener; last close stops LS. */ + if (!capable(CAP_SYS_ADMIN)) + return -EPERM; + atomic_inc(&dlm_monitor_opened); dlm_monitor_unused = 0; return 0; @@ -958,6 +980,7 @@ static struct miscdevice monitor_device = { .name = "dlm-monitor", .fops = &monitor_device_fops, .minor = MISC_DYNAMIC_MINOR, + .mode = 0600, }; int __init dlm_user_init(void) diff --git a/fs/ecryptfs/inode.c b/fs/ecryptfs/inode.c index 525297c7ebd8..48e520960d66 100644 --- a/fs/ecryptfs/inode.c +++ b/fs/ecryptfs/inode.c @@ -266,7 +266,7 @@ out: * Returns zero on success; non-zero on error condition */ static int -ecryptfs_create(struct mnt_idmap *idmap, +ecryptfs_create(const struct mnt_idmap *idmap, struct inode *directory_inode, struct dentry *ecryptfs_dentry, umode_t mode) { @@ -462,7 +462,7 @@ static int ecryptfs_unlink(struct inode *dir, struct dentry *dentry) return ecryptfs_do_unlink(dir, dentry, d_inode(dentry)); } -static int ecryptfs_symlink(struct mnt_idmap *idmap, +static int ecryptfs_symlink(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *dentry, const char *symname) { @@ -503,7 +503,7 @@ out_lock: return rc; } -static struct dentry *ecryptfs_mkdir(struct mnt_idmap *idmap, struct inode *dir, +static struct dentry *ecryptfs_mkdir(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *dentry, umode_t mode) { int rc; @@ -562,7 +562,7 @@ static int ecryptfs_rmdir(struct inode *dir, struct dentry *dentry) } static int -ecryptfs_mknod(struct mnt_idmap *idmap, struct inode *dir, +ecryptfs_mknod(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *dentry, umode_t mode, dev_t dev) { int rc; @@ -590,7 +590,7 @@ out: } static int -ecryptfs_rename(struct mnt_idmap *idmap, struct inode *old_dir, +ecryptfs_rename(const struct mnt_idmap *idmap, struct inode *old_dir, struct dentry *old_dentry, struct inode *new_dir, struct dentry *new_dentry, unsigned int flags) { @@ -849,7 +849,7 @@ int ecryptfs_truncate(struct dentry *dentry, loff_t new_length) } static int -ecryptfs_permission(struct mnt_idmap *idmap, struct inode *inode, +ecryptfs_permission(const struct mnt_idmap *idmap, struct inode *inode, int mask) { return inode_permission(&nop_mnt_idmap, @@ -869,7 +869,7 @@ ecryptfs_permission(struct mnt_idmap *idmap, struct inode *inode, * All other metadata changes will be passed right to the lower filesystem, * and we will just update our inode to look like the lower. */ -static int ecryptfs_setattr(struct mnt_idmap *idmap, +static int ecryptfs_setattr(const struct mnt_idmap *idmap, struct dentry *dentry, struct iattr *ia) { struct inode *inode = d_inode(dentry); @@ -939,7 +939,7 @@ out: return rc; } -static int ecryptfs_getattr_link(struct mnt_idmap *idmap, +static int ecryptfs_getattr_link(const struct mnt_idmap *idmap, const struct path *path, struct kstat *stat, u32 request_mask, unsigned int flags) { @@ -965,7 +965,7 @@ static int ecryptfs_getattr_link(struct mnt_idmap *idmap, return rc; } -static int ecryptfs_getattr(struct mnt_idmap *idmap, +static int ecryptfs_getattr(const struct mnt_idmap *idmap, const struct path *path, struct kstat *stat, u32 request_mask, unsigned int flags) { @@ -1078,7 +1078,7 @@ static int ecryptfs_fileattr_get(struct dentry *dentry, struct file_kattr *fa) return vfs_fileattr_get(ecryptfs_dentry_to_lower(dentry), fa); } -static int ecryptfs_fileattr_set(struct mnt_idmap *idmap, +static int ecryptfs_fileattr_set(const struct mnt_idmap *idmap, struct dentry *dentry, struct file_kattr *fa) { struct dentry *lower_dentry = ecryptfs_dentry_to_lower(dentry); @@ -1090,14 +1090,14 @@ static int ecryptfs_fileattr_set(struct mnt_idmap *idmap, return rc; } -static struct posix_acl *ecryptfs_get_acl(struct mnt_idmap *idmap, +static struct posix_acl *ecryptfs_get_acl(const struct mnt_idmap *idmap, struct dentry *dentry, int type) { return vfs_get_acl(idmap, ecryptfs_dentry_to_lower(dentry), posix_acl_xattr_name(type)); } -static int ecryptfs_set_acl(struct mnt_idmap *idmap, +static int ecryptfs_set_acl(const struct mnt_idmap *idmap, struct dentry *dentry, struct posix_acl *acl, int type) { @@ -1158,7 +1158,7 @@ static int ecryptfs_xattr_get(const struct xattr_handler *handler, } static int ecryptfs_xattr_set(const struct xattr_handler *handler, - struct mnt_idmap *idmap, + const struct mnt_idmap *idmap, struct dentry *dentry, struct inode *inode, const char *name, const void *value, size_t size, int flags) diff --git a/fs/efivarfs/inode.c b/fs/efivarfs/inode.c index f0d009555fc6..07602cd5d33c 100644 --- a/fs/efivarfs/inode.c +++ b/fs/efivarfs/inode.c @@ -74,7 +74,7 @@ static bool efivarfs_valid_name(const char *str, int len) return uuid_is_valid(s); } -static int efivarfs_create(struct mnt_idmap *idmap, struct inode *dir, +static int efivarfs_create(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *dentry, umode_t mode) { struct inode *inode = NULL; @@ -150,7 +150,7 @@ efivarfs_fileattr_get(struct dentry *dentry, struct file_kattr *fa) } static int -efivarfs_fileattr_set(struct mnt_idmap *idmap, +efivarfs_fileattr_set(const struct mnt_idmap *idmap, struct dentry *dentry, struct file_kattr *fa) { unsigned int i_flags = 0; @@ -170,7 +170,7 @@ efivarfs_fileattr_set(struct mnt_idmap *idmap, } /* copy of simple_setattr except that it doesn't do i_size updates */ -static int efivarfs_setattr(struct mnt_idmap *idmap, struct dentry *dentry, +static int efivarfs_setattr(const struct mnt_idmap *idmap, struct dentry *dentry, struct iattr *iattr) { struct inode *inode = d_inode(dentry); diff --git a/fs/erofs/inode.c b/fs/erofs/inode.c index 45afe5c50de8..26ea3790ff21 100644 --- a/fs/erofs/inode.c +++ b/fs/erofs/inode.c @@ -311,7 +311,7 @@ struct inode *erofs_iget(struct super_block *sb, erofs_nid_t nid) return inode; } -int erofs_getattr(struct mnt_idmap *idmap, const struct path *path, +int erofs_getattr(const struct mnt_idmap *idmap, const struct path *path, struct kstat *stat, u32 request_mask, unsigned int query_flags) { diff --git a/fs/erofs/internal.h b/fs/erofs/internal.h index 12e3a5b80a5a..ab817091bd29 100644 --- a/fs/erofs/internal.h +++ b/fs/erofs/internal.h @@ -417,7 +417,7 @@ void erofs_onlinefolio_init(struct folio *folio); void erofs_onlinefolio_split(struct folio *folio); void erofs_onlinefolio_end(struct folio *folio, int err, bool dirty); struct inode *erofs_iget(struct super_block *sb, erofs_nid_t nid); -int erofs_getattr(struct mnt_idmap *idmap, const struct path *path, +int erofs_getattr(const struct mnt_idmap *idmap, const struct path *path, struct kstat *stat, u32 request_mask, unsigned int query_flags); int erofs_namei(struct inode *dir, const struct qstr *name, diff --git a/fs/eventfd.c b/fs/eventfd.c index 9d33a02757d5..52426795752e 100644 --- a/fs/eventfd.c +++ b/fs/eventfd.c @@ -403,8 +403,8 @@ static int do_eventfd(unsigned int count, int flags) FD_PREPARE(fdf, flags, anon_inode_getfile_fmode("[eventfd]", &eventfd_fops, ctx, flags, FMODE_NOWAIT)); - if (fdf.err) - return fdf.err; + if (fdf->fd < 0) + return fdf->fd; ctx->id = ida_alloc(&eventfd_ida, GFP_KERNEL); retain_and_null_ptr(ctx); diff --git a/fs/eventpoll.c b/fs/eventpoll.c index e0c4bf88a838..f48b829a710f 100644 --- a/fs/eventpoll.c +++ b/fs/eventpoll.c @@ -2514,11 +2514,11 @@ static int do_epoll_create(int flags) FD_PREPARE(fdf, O_RDWR | (flags & O_CLOEXEC), anon_inode_getfile("[eventpoll]", &eventpoll_fops, ep, O_RDWR | (flags & O_CLOEXEC))); - if (fdf.err) { + if (fdf->fd < 0) { ep_clear_and_put(ep); - return fdf.err; + return fdf->fd; } - ep->file = fd_prepare_file(fdf); + ep->file = fdf->file; return fd_publish(fdf); } diff --git a/fs/exec.c b/fs/exec.c index a5269b5e00df..33a1e4689e49 100644 --- a/fs/exec.c +++ b/fs/exec.c @@ -1136,6 +1136,7 @@ static void posixtimer_exec(struct task_struct *me) int begin_new_exec(struct linux_binprm * bprm) { struct task_struct *me = current; + struct files_struct *files = NULL; int retval; /* A pending PT_INTERP substitution this format cannot consume. */ @@ -1160,6 +1161,13 @@ int begin_new_exec(struct linux_binprm * bprm) */ bprm->point_of_no_return = true; + /* + * Cancel any io_uring activity across execve. This runs task work + * that may still create an io-wq worker, so do it while de_thread() + * can still zap it. + */ + io_uring_task_cancel(); + /* Make this the only thread in the thread group */ retval = de_thread(me); if (retval) @@ -1176,15 +1184,13 @@ int begin_new_exec(struct linux_binprm * bprm) /* see the comment in check_unsafe_exec() */ current->fs->in_exec = 0; - /* - * Cancel any io_uring activity across execve - */ - io_uring_task_cancel(); /* Ensure the files table is not shared. */ - retval = unshare_files(); + retval = unshare_fd(CLONE_FILES, &files); if (retval) goto out; + if (files) + switch_files_struct(me, files); /* * We have to apply CLOEXEC before we change whether the process is @@ -1192,13 +1198,13 @@ int begin_new_exec(struct linux_binprm * bprm) * trying to access the should-be-closed file descriptors of a process * undergoing exec(2). * - * This can block on filesystem ->flush() handlers, including waiting - * for FUSE daemons, so do it before exec_mmap takes the - * exec_update_lock. + * This can block on filesystem ->flush() and ->release() handlers, + * including waiting for FUSE daemons, so do it before exec_mmap + * takes the exec_update_lock. * This must happen after the point of no return, and after unsharing * the FD table. */ - do_close_on_exec(me->files); + close_cloexec_files(me->files); /* * Must be called _before_ exec_mmap() as bprm->mm is @@ -1359,7 +1365,7 @@ EXPORT_SYMBOL(begin_new_exec); void would_dump(struct linux_binprm *bprm, struct file *file) { struct inode *inode = file_inode(file); - struct mnt_idmap *idmap = file_mnt_idmap(file); + const struct mnt_idmap *idmap = file_mnt_idmap(file); if (inode_permission(idmap, inode, MAY_READ) < 0) { struct user_namespace *old, *user_ns; bprm->interp_flags |= BINPRM_FLAGS_ENFORCE_NONDUMP; @@ -1643,7 +1649,7 @@ static void check_unsafe_exec(struct linux_binprm *bprm) static void bprm_fill_uid(struct linux_binprm *bprm, struct file *file) { /* Handle suid and sgid on files */ - struct mnt_idmap *idmap; + const struct mnt_idmap *idmap; struct inode *inode = file_inode(file); unsigned int mode; vfsuid_t vfsuid; diff --git a/fs/exfat/dir.c b/fs/exfat/dir.c index fe73b1380c5d..46514b13bebd 100644 --- a/fs/exfat/dir.c +++ b/fs/exfat/dir.c @@ -678,11 +678,54 @@ enum exfat_validate_dentry_mode { ES_MODE_GET_BENIGN_SEC_ENTRY, }; -static bool exfat_validate_entry(unsigned int type, - enum exfat_validate_dentry_mode *mode) +static bool exfat_validate_vendor_alloc(struct super_block *sb, + struct exfat_dentry *ep) +{ + struct exfat_sb_info *sbi = EXFAT_SB(sb); + u8 flags = ep->dentry.vendor_alloc.flags; + u32 start_clu = le32_to_cpu(ep->dentry.vendor_alloc.start_clu); + u64 size = le64_to_cpu(ep->dentry.vendor_alloc.size); + u64 max_size = exfat_cluster_to_bytes(sbi, + (u64)EXFAT_DATA_CLUSTER_COUNT(sbi)); + u64 num_clusters; + + /* AllocationPossible is required for Vendor Allocation entries. */ + if (!(flags & ALLOC_POSSIBLE)) + return false; + + /* The null GUID does not identify a valid vendor allocation. */ + if (!memchr_inv(ep->dentry.vendor_alloc.vendor_guid, 0, + sizeof(ep->dentry.vendor_alloc.vendor_guid))) + return false; + + if (!start_clu) + return !size && !(flags & (ALLOC_NO_FAT_CHAIN ^ ALLOC_FAT_CHAIN)); + + if (!is_valid_cluster(sbi, start_clu) || size > max_size) + return false; + + if ((flags & ALLOC_NO_FAT_CHAIN) == ALLOC_NO_FAT_CHAIN) { + if (!size) + return false; + + num_clusters = DIV_ROUND_UP_ULL(size, sbi->cluster_size); + if (num_clusters > sbi->num_clusters - start_clu) + return false; + } + + return true; +} + +static bool exfat_validate_entry(struct super_block *sb, + struct exfat_dentry *ep, enum exfat_validate_dentry_mode *mode) { + unsigned int type = exfat_get_entry_type(ep); + if (type == TYPE_UNUSED || type == TYPE_DELETED) return false; + if (type == TYPE_VENDOR_ALLOC && + !exfat_validate_vendor_alloc(sb, ep)) + return false; switch (*mode) { case ES_MODE_GET_FILE_ENTRY: @@ -836,7 +879,7 @@ int exfat_get_dentry_set(struct exfat_entry_set_cache *es, /* validate cached dentries */ for (i = ES_IDX_STREAM; i < es->num_entries; i++) { ep = exfat_get_dentry_cached(es, i); - if (!exfat_validate_entry(exfat_get_entry_type(ep), &mode)) + if (!exfat_validate_entry(sb, ep, &mode)) goto put_es; } return 0; @@ -1266,7 +1309,8 @@ static int exfat_get_volume_label_dentry(struct super_block *sb, es->bh = es->__bh; es->bh[0] = bh; es->num_bh = 1; - es->start_off = exfat_dentries_to_bytes(i) % sb->s_blocksize; + es->start_off = exfat_dentries_to_bytes(i) & + ((u32)sb->s_blocksize - 1); return 0; } diff --git a/fs/exfat/exfat_fs.h b/fs/exfat/exfat_fs.h index a9131fe03302..5f258e96fce9 100644 --- a/fs/exfat/exfat_fs.h +++ b/fs/exfat/exfat_fs.h @@ -232,6 +232,7 @@ struct exfat_sb_info { unsigned int num_FAT_sectors; /* num of FAT sectors */ unsigned int root_dir; /* root dir cluster */ unsigned int dentries_per_clu; /* num of dentries per cluster */ + unsigned int dentries_per_clu_bits; unsigned int vol_flags; /* volume flags */ unsigned int vol_flags_persistent; /* volume flags to retain */ struct buffer_head *boot_bh; /* buffer_head of BOOT sector */ @@ -555,9 +556,9 @@ int exfat_trim_fs(struct inode *inode, struct fstrim_range *range); /* file.c */ extern const struct file_operations exfat_file_operations; int __exfat_truncate(struct inode *inode); -int exfat_setattr(struct mnt_idmap *idmap, struct dentry *dentry, +int exfat_setattr(const struct mnt_idmap *idmap, struct dentry *dentry, struct iattr *attr); -int exfat_getattr(struct mnt_idmap *idmap, const struct path *path, +int exfat_getattr(const struct mnt_idmap *idmap, const struct path *path, struct kstat *stat, unsigned int request_mask, unsigned int query_flags); struct file_kattr; diff --git a/fs/exfat/fatent.c b/fs/exfat/fatent.c index a8b11e2ce43f..3c8bdc131f6f 100644 --- a/fs/exfat/fatent.c +++ b/fs/exfat/fatent.c @@ -427,19 +427,22 @@ int exfat_alloc_cluster(struct inode *inode, unsigned int num_alloc, struct super_block *sb = inode->i_sb; struct exfat_sb_info *sbi = EXFAT_SB(sb); + mutex_lock(&sbi->bitmap_lock); + total_cnt = EXFAT_DATA_CLUSTER_COUNT(sbi); if (unlikely(total_cnt < sbi->used_clusters)) { exfat_fs_error_ratelimit(sb, "%s: invalid used clusters(t:%u,u:%u)\n", __func__, total_cnt, sbi->used_clusters); - return -EIO; + ret = -EIO; + goto unlock; } - if (num_alloc > total_cnt - sbi->used_clusters) - return -ENOSPC; - - mutex_lock(&sbi->bitmap_lock); + if (num_alloc > total_cnt - sbi->used_clusters) { + ret = -ENOSPC; + goto unlock; + } hint_clu = p_chain->dir; /* find new cluster */ @@ -509,15 +512,15 @@ int exfat_alloc_cluster(struct inode *inode, unsigned int num_alloc, } } p_chain->size++; + sbi->used_clusters++; last_clu = new_clu; if (p_chain->size == num_alloc) { done: sbi->clu_srch_ptr = hint_clu; - sbi->used_clusters += p_chain->size; - mutex_unlock(&sbi->bitmap_lock); - return 0; + ret = 0; + goto unlock; } hint_clu = new_clu + 1; diff --git a/fs/exfat/file.c b/fs/exfat/file.c index a2a9ee1a2004..3867e78c2312 100644 --- a/fs/exfat/file.c +++ b/fs/exfat/file.c @@ -143,7 +143,7 @@ error: return err; } -static bool exfat_allow_set_time(struct mnt_idmap *idmap, +static bool exfat_allow_set_time(const struct mnt_idmap *idmap, struct exfat_sb_info *sbi, struct inode *inode) { mode_t allow_utime = sbi->options.allow_utime; @@ -319,7 +319,7 @@ write_size: mutex_unlock(&sbi->s_lock); } -int exfat_getattr(struct mnt_idmap *idmap, const struct path *path, +int exfat_getattr(const struct mnt_idmap *idmap, const struct path *path, struct kstat *stat, unsigned int request_mask, unsigned int query_flags) { @@ -347,7 +347,7 @@ int exfat_fileattr_get(struct dentry *dentry, struct file_kattr *fa) return 0; } -int exfat_setattr(struct mnt_idmap *idmap, struct dentry *dentry, +int exfat_setattr(const struct mnt_idmap *idmap, struct dentry *dentry, struct iattr *attr) { struct exfat_sb_info *sbi = EXFAT_SB(dentry->d_sb); diff --git a/fs/exfat/iomap.c b/fs/exfat/iomap.c index 8911aa84a730..533cdb4f0929 100644 --- a/fs/exfat/iomap.c +++ b/fs/exfat/iomap.c @@ -157,6 +157,27 @@ const struct iomap_ops exfat_iomap_ops = { .iomap_next = exfat_iomap_next, }; +#ifdef CONFIG_SWAP +static int exfat_swap_iomap_begin(struct inode *inode, loff_t offset, + loff_t length, unsigned int flags, struct iomap *iomap, + struct iomap *srcmap) +{ + /* + * Swap activation needs the physical mappings of preallocated + * ranges. Do not report the VDL tail as a hole. + */ + return __exfat_iomap_begin(inode, offset, length, + flags & ~IOMAP_REPORT, iomap, false); +} + +static DEFINE_IOMAP_ITER_NEXT(exfat_swap_iomap_next, + exfat_swap_iomap_begin); + +static const struct iomap_ops exfat_swap_iomap_ops = { + .iomap_next = exfat_swap_iomap_next, +}; +#endif + /* * exfat_write_iomap_end - Update the state after write * @@ -275,5 +296,10 @@ const struct iomap_read_ops exfat_iomap_bio_read_ops = { int exfat_iomap_swap_activate(struct swap_info_struct *sis, struct file *file, sector_t *span) { - return iomap_swapfile_activate(sis, file, span, &exfat_iomap_ops); +#ifdef CONFIG_SWAP + return iomap_swapfile_activate(sis, file, span, + &exfat_swap_iomap_ops); +#else + return -EIO; +#endif } diff --git a/fs/exfat/misc.c b/fs/exfat/misc.c index 6f11a96a4ffa..dfd0bbf31c94 100644 --- a/fs/exfat/misc.c +++ b/fs/exfat/misc.c @@ -187,7 +187,7 @@ int exfat_update_bhs(struct buffer_head **bhs, int nr_bhs, int sync) for (i = 0; i < nr_bhs && sync; i++) { wait_on_buffer(bhs[i]); - if (!err && !buffer_uptodate(bhs[i])) + if (!err && buffer_write_io_error(bhs[i])) err = -EIO; } return err; diff --git a/fs/exfat/namei.c b/fs/exfat/namei.c index a4dc83b5949c..d116c89d724e 100644 --- a/fs/exfat/namei.c +++ b/fs/exfat/namei.c @@ -386,7 +386,7 @@ int exfat_find_empty_entry(struct inode *inode, } p_dir->dir = exfat_sector_to_cluster(sbi, es->bh[0]->b_blocknr); - p_dir->size -= dentry / sbi->dentries_per_clu; + p_dir->size -= dentry >> sbi->dentries_per_clu_bits; return dentry & (sbi->dentries_per_clu - 1); } @@ -552,7 +552,7 @@ out: return ret; } -static int exfat_create(struct mnt_idmap *idmap, struct inode *dir, +static int exfat_create(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *dentry, umode_t mode) { struct super_block *sb = dir->i_sb; @@ -638,14 +638,15 @@ static int exfat_find(struct inode *dir, const struct qstr *qname, /* adjust cdir to the optimized value */ cdir.dir = hint_opt.clu; if (cdir.flags & ALLOC_NO_FAT_CHAIN) - cdir.size -= dentry / sbi->dentries_per_clu; + cdir.size -= dentry >> sbi->dentries_per_clu_bits; dentry = hint_opt.eidx; info->dir = cdir; info->entry = dentry; info->num_subdirs = 0; - if (exfat_get_dentry_set(&es, sb, &cdir, dentry, ES_2_ENTRIES)) + /* Validate the complete set, including recognized benign entries. */ + if (exfat_get_dentry_set(&es, sb, &cdir, dentry, ES_ALL_ENTRIES)) return -EIO; ep = exfat_get_dentry_cached(&es, ES_IDX_FILE); ep2 = exfat_get_dentry_cached(&es, ES_IDX_STREAM); @@ -825,7 +826,7 @@ unlock: return err; } -static struct dentry *exfat_mkdir(struct mnt_idmap *idmap, struct inode *dir, +static struct dentry *exfat_mkdir(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *dentry, umode_t mode) { struct super_block *sb = dir->i_sb; @@ -1263,7 +1264,7 @@ out: return ret; } -static int exfat_rename(struct mnt_idmap *idmap, +static int exfat_rename(const struct mnt_idmap *idmap, struct inode *old_dir, struct dentry *old_dentry, struct inode *new_dir, struct dentry *new_dentry, unsigned int flags) diff --git a/fs/exfat/super.c b/fs/exfat/super.c index a9ea36ba2693..4943cef97741 100644 --- a/fs/exfat/super.c +++ b/fs/exfat/super.c @@ -18,6 +18,7 @@ #include <linux/nls.h> #include <linux/buffer_head.h> #include <linux/magic.h> +#include <linux/math64.h> #include "exfat_raw.h" #include "exfat_fs.h" @@ -69,16 +70,28 @@ static int exfat_statfs(struct dentry *dentry, struct kstatfs *buf) return 0; } +static inline __u8 exfat_calc_perc_in_use(const struct exfat_sb_info *sbi) +{ + return (__u8)mul_u64_u32_div(sbi->used_clusters, 100, + EXFAT_DATA_CLUSTER_COUNT(sbi)); +} + static int exfat_set_vol_flags(struct super_block *sb, unsigned short new_flags) { struct exfat_sb_info *sbi = EXFAT_SB(sb); struct boot_sector *p_boot = (struct boot_sector *)sbi->boot_bh->b_data; + __u8 new_piu; /* retain persistent-flags */ new_flags |= sbi->vol_flags_persistent; + if (new_flags & VOLUME_DIRTY) + new_piu = 0xFF; + else + new_piu = exfat_calc_perc_in_use(sbi); + /* flags are not changed */ - if (sbi->vol_flags == new_flags) + if (sbi->vol_flags == new_flags && new_piu == p_boot->percent_in_use) return 0; sbi->vol_flags = new_flags; @@ -90,6 +103,7 @@ static int exfat_set_vol_flags(struct super_block *sb, unsigned short new_flags) return 0; p_boot->vol_flags = cpu_to_le16(new_flags); + p_boot->percent_in_use = new_piu; set_buffer_uptodate(sbi->boot_bh); mark_buffer_dirty(sbi->boot_bh); @@ -505,8 +519,8 @@ static int exfat_read_boot_sector(struct super_block *sb) EXFAT_RESERVED_CLUSTERS; sbi->root_dir = le32_to_cpu(p_boot->root_cluster); - sbi->dentries_per_clu = 1 << - (sbi->cluster_size_bits - DENTRY_SIZE_BITS); + sbi->dentries_per_clu_bits = sbi->cluster_size_bits - DENTRY_SIZE_BITS; + sbi->dentries_per_clu = 1 << sbi->dentries_per_clu_bits; sbi->vol_flags = le16_to_cpu(p_boot->vol_flags); sbi->vol_flags_persistent = sbi->vol_flags & (VOLUME_DIRTY | MEDIA_FAILURE); @@ -775,9 +789,13 @@ static int exfat_reconfigure(struct fs_context *fc) fc->sb_flags |= SB_NODIRATIME; sync_filesystem(sb); - mutex_lock(&sbi->s_lock); - exfat_clear_volume_dirty(sb); - mutex_unlock(&sbi->s_lock); + + if ((fc->sb_flags & (SB_FORCE | SB_RDONLY)) == SB_RDONLY && + !sb_rdonly(sb)) { + mutex_lock(&sbi->s_lock); + exfat_clear_volume_dirty(sb); + mutex_unlock(&sbi->s_lock); + } if (new_opts->allow_utime == (unsigned short)-1) new_opts->allow_utime = ~new_opts->fs_dmask & 0022; diff --git a/fs/ext2/Makefile b/fs/ext2/Makefile index 8860948ef9ca..33db2e9dc908 100644 --- a/fs/ext2/Makefile +++ b/fs/ext2/Makefile @@ -3,6 +3,8 @@ # Makefile for the linux ext2-filesystem routines. # +CONTEXT_ANALYSIS := y + obj-$(CONFIG_EXT2_FS) += ext2.o ext2-y := balloc.o dir.o file.o ialloc.o inode.o \ diff --git a/fs/ext2/acl.c b/fs/ext2/acl.c index 7e54c31589c7..b2746657fc53 100644 --- a/fs/ext2/acl.c +++ b/fs/ext2/acl.c @@ -219,7 +219,7 @@ __ext2_set_acl(struct inode *inode, struct posix_acl *acl, int type) * inode->i_mutex: down */ int -ext2_set_acl(struct mnt_idmap *idmap, struct dentry *dentry, +ext2_set_acl(const struct mnt_idmap *idmap, struct dentry *dentry, struct posix_acl *acl, int type) { int error; diff --git a/fs/ext2/acl.h b/fs/ext2/acl.h index 4a8443a2b8ec..e68bc3545608 100644 --- a/fs/ext2/acl.h +++ b/fs/ext2/acl.h @@ -56,7 +56,7 @@ static inline int ext2_acl_count(size_t size) /* acl.c */ extern struct posix_acl *ext2_get_acl(struct inode *inode, int type, bool rcu); -extern int ext2_set_acl(struct mnt_idmap *idmap, struct dentry *dentry, +extern int ext2_set_acl(const struct mnt_idmap *idmap, struct dentry *dentry, struct posix_acl *acl, int type); extern int ext2_init_acl (struct inode *, struct inode *); diff --git a/fs/ext2/balloc.c b/fs/ext2/balloc.c index adf0f31fbddd..80acc1e19387 100644 --- a/fs/ext2/balloc.c +++ b/fs/ext2/balloc.c @@ -334,6 +334,7 @@ search_reserve_window(struct rb_root *root, ext2_fsblk_t goal) */ void ext2_rsv_window_add(struct super_block *sb, struct ext2_reserve_window_node *rsv) + __must_hold(&EXT2_SB(sb)->s_rsv_window_lock) { struct rb_root *root = &EXT2_SB(sb)->s_rsv_window_root; struct rb_node *node = &rsv->rsv_node; @@ -373,6 +374,7 @@ void ext2_rsv_window_add(struct super_block *sb, */ static void rsv_window_remove(struct super_block *sb, struct ext2_reserve_window_node *rsv) + __must_hold(&EXT2_SB(sb)->s_rsv_window_lock) { rsv->rsv_start = EXT2_RESERVE_WINDOW_NOT_ALLOCATED; rsv->rsv_end = EXT2_RESERVE_WINDOW_NOT_ALLOCATED; @@ -414,6 +416,7 @@ static inline int rsv_is_empty(struct ext2_reserve_window *rsv) * Needs truncate_mutex protection prior to calling this function. */ void ext2_init_block_alloc_info(struct inode *inode) + __must_hold(&EXT2_I(inode)->truncate_mutex) { struct ext2_inode_info *ei = EXT2_I(inode); struct ext2_block_alloc_info *block_i; @@ -758,6 +761,7 @@ static int find_next_reservable_window( struct super_block * sb, ext2_fsblk_t start_block, ext2_fsblk_t last_block) + __must_hold(&EXT2_SB(sb)->s_rsv_window_lock) { struct rb_node *next; struct ext2_reserve_window_node *rsv, *prev; diff --git a/fs/ext2/ext2.h b/fs/ext2/ext2.h index 5642451bf191..7bdada93dd06 100644 --- a/fs/ext2/ext2.h +++ b/fs/ext2/ext2.h @@ -77,8 +77,8 @@ struct ext2_sb_info { unsigned long s_gdb_count; /* Number of group descriptor blocks */ unsigned long s_desc_per_block; /* Number of group descriptors per block */ unsigned long s_groups_count; /* Number of groups in the fs */ - unsigned long s_overhead_last; /* Last calculated overhead */ - unsigned long s_blocks_last; /* Last seen block count */ + unsigned long s_overhead_last __guarded_by(&s_lock); /* Last calculated overhead */ + unsigned long s_blocks_last __guarded_by(&s_lock); /* Last seen block count */ struct buffer_head * s_sbh; /* Buffer containing the super block */ struct ext2_super_block * s_es; /* Pointer to the super block in the buffer */ struct buffer_head ** s_group_desc; @@ -86,14 +86,14 @@ struct ext2_sb_info { unsigned long s_sb_block; kuid_t s_resuid; kgid_t s_resgid; - unsigned short s_mount_state; + unsigned short s_mount_state __guarded_by(&s_lock); unsigned short s_pad; int s_addr_per_block_bits; int s_desc_per_block_bits; int s_inode_size; int s_first_ino; spinlock_t s_next_gen_lock; - u32 s_next_generation; + u32 s_next_generation __guarded_by(&s_next_gen_lock); unsigned long s_dir_count; u8 *s_debts; struct percpu_counter s_freeblocks_counter; @@ -102,7 +102,7 @@ struct ext2_sb_info { struct blockgroup_lock *s_blockgroup_lock; /* root of the per fs reservation window tree */ spinlock_t s_rsv_window_lock; - struct rb_root s_rsv_window_root; + struct rb_root s_rsv_window_root __guarded_by(&s_rsv_window_lock); struct ext2_reserve_window_node s_rsv_window_head; /* * s_lock protects against concurrent modifications of s_mount_state, @@ -710,8 +710,10 @@ extern struct ext2_group_desc * ext2_get_group_desc(struct super_block * sb, struct buffer_head ** bh); extern void ext2_discard_reservation (struct inode *); extern int ext2_should_retry_alloc(struct super_block *sb, int *retries); -extern void ext2_init_block_alloc_info(struct inode *); -extern void ext2_rsv_window_add(struct super_block *sb, struct ext2_reserve_window_node *rsv); +extern void ext2_init_block_alloc_info(struct inode *inode) + __must_hold(&EXT2_I(inode)->truncate_mutex); +extern void ext2_rsv_window_add(struct super_block *sb, struct ext2_reserve_window_node *rsv) + __must_hold(&EXT2_SB(sb)->s_rsv_window_lock); /* dir.c */ int ext2_add_link(struct dentry *, struct inode *); @@ -739,8 +741,8 @@ extern int ext2_sync_inode_metadata(struct inode *, struct writeback_control *); extern void ext2_evict_inode(struct inode *); void ext2_write_failed(struct address_space *mapping, loff_t to); extern int ext2_get_block(struct inode *, sector_t, struct buffer_head *, int); -extern int ext2_setattr (struct mnt_idmap *, struct dentry *, struct iattr *); -extern int ext2_getattr (struct mnt_idmap *, const struct path *, +extern int ext2_setattr (const struct mnt_idmap *, struct dentry *, struct iattr *); +extern int ext2_getattr (const struct mnt_idmap *, const struct path *, struct kstat *, u32, unsigned int); extern void ext2_set_inode_flags(struct inode *inode); extern int ext2_fiemap(struct inode *inode, struct fiemap_extent_info *fieinfo, @@ -748,7 +750,7 @@ extern int ext2_fiemap(struct inode *inode, struct fiemap_extent_info *fieinfo, /* ioctl.c */ extern int ext2_fileattr_get(struct dentry *dentry, struct file_kattr *fa); -extern int ext2_fileattr_set(struct mnt_idmap *idmap, +extern int ext2_fileattr_set(const struct mnt_idmap *idmap, struct dentry *dentry, struct file_kattr *fa); extern long ext2_ioctl(struct file *, unsigned int, unsigned long); extern long ext2_compat_ioctl(struct file *, unsigned int, unsigned long); @@ -761,7 +763,8 @@ extern __printf(3, 4) void ext2_error(struct super_block *, const char *, const char *, ...); extern __printf(3, 4) void ext2_msg(struct super_block *, const char *, const char *, ...); -extern void ext2_update_dynamic_rev (struct super_block *sb); +extern void ext2_update_dynamic_rev(struct super_block *sb) + __must_hold(&EXT2_SB(sb)->s_lock); extern void ext2_sync_super(struct super_block *sb, struct ext2_super_block *es, int wait); diff --git a/fs/ext2/file.c b/fs/ext2/file.c index b9020df7d89e..67fe423c3828 100644 --- a/fs/ext2/file.c +++ b/fs/ext2/file.c @@ -135,6 +135,13 @@ static ssize_t ext2_dio_write_iter(struct kiocb *iocb, struct iov_iter *from) int ret2; iocb->ki_flags &= ~IOCB_DIRECT; + + /* + * Prevent concurrent direct I/O and buffered I/O to the same file + * range. Wait for in-flight DIO to finish before dirtying pages. + */ + inode_dio_wait(inode); + pos = iocb->ki_pos; status = generic_perform_write(iocb, from); if (unlikely(status < 0)) { diff --git a/fs/ext2/inode.c b/fs/ext2/inode.c index a9245f0cda4d..12ac4cfe1500 100644 --- a/fs/ext2/inode.c +++ b/fs/ext2/inode.c @@ -327,6 +327,7 @@ static ext2_fsblk_t ext2_find_near(struct inode *inode, Indirect *ind) static inline ext2_fsblk_t ext2_find_goal(struct inode *inode, long block, Indirect *partial) + __must_hold(&EXT2_I(inode)->truncate_mutex) { struct ext2_block_alloc_info *block_i; @@ -397,6 +398,7 @@ ext2_blks_to_allocate(Indirect * branch, int k, unsigned long blks, static int ext2_alloc_blocks(struct inode *inode, ext2_fsblk_t goal, int indirect_blks, int blks, ext2_fsblk_t new_blocks[4], int *err) + __must_hold(&EXT2_I(inode)->truncate_mutex) { int target, i; unsigned long count = 0; @@ -477,6 +479,7 @@ failed_out: static int ext2_alloc_branch(struct inode *inode, int indirect_blks, int *blks, ext2_fsblk_t goal, int *offsets, Indirect *branch) + __must_hold(&EXT2_I(inode)->truncate_mutex) { int blocksize = inode->i_sb->s_blocksize; int i, n = 0; @@ -558,6 +561,7 @@ failed: */ static void ext2_splice_branch(struct inode *inode, long block, Indirect *where, int num, int blks) + __must_hold(&EXT2_I(inode)->truncate_mutex) { int i; struct ext2_block_alloc_info *block_i; @@ -1002,6 +1006,7 @@ static Indirect *ext2_find_shared(struct inode *inode, int offsets[4], Indirect chain[4], __le32 *top) + __must_hold(&EXT2_I(inode)->truncate_mutex) { Indirect *partial, *p; int k, err; @@ -1057,6 +1062,7 @@ no_top: * appropriately. */ static inline void ext2_free_data(struct inode *inode, __le32 *p, __le32 *q) + __must_hold(&EXT2_I(inode)->truncate_mutex) { ext2_fsblk_t block_to_free = 0, count = 0; ext2_fsblk_t nr; @@ -1097,6 +1103,7 @@ static inline void ext2_free_data(struct inode *inode, __le32 *p, __le32 *q) * appropriately. */ static void ext2_free_branches(struct inode *inode, __le32 *p, __le32 *q, int depth) + __must_hold(&EXT2_I(inode)->truncate_mutex) { struct buffer_head * bh; ext2_fsblk_t nr; @@ -1591,7 +1598,7 @@ out: return err; } -int ext2_getattr(struct mnt_idmap *idmap, const struct path *path, +int ext2_getattr(const struct mnt_idmap *idmap, const struct path *path, struct kstat *stat, u32 request_mask, unsigned int query_flags) { struct inode *inode = d_inode(path->dentry); @@ -1617,7 +1624,7 @@ int ext2_getattr(struct mnt_idmap *idmap, const struct path *path, return 0; } -int ext2_setattr(struct mnt_idmap *idmap, struct dentry *dentry, +int ext2_setattr(const struct mnt_idmap *idmap, struct dentry *dentry, struct iattr *iattr) { struct inode *inode = d_inode(dentry); diff --git a/fs/ext2/ioctl.c b/fs/ext2/ioctl.c index c3fea55b8efa..f2218455fa47 100644 --- a/fs/ext2/ioctl.c +++ b/fs/ext2/ioctl.c @@ -27,7 +27,7 @@ int ext2_fileattr_get(struct dentry *dentry, struct file_kattr *fa) return 0; } -int ext2_fileattr_set(struct mnt_idmap *idmap, +int ext2_fileattr_set(const struct mnt_idmap *idmap, struct dentry *dentry, struct file_kattr *fa) { struct inode *inode = d_inode(dentry); diff --git a/fs/ext2/namei.c b/fs/ext2/namei.c index 8666233ec63b..bfb6a463a95e 100644 --- a/fs/ext2/namei.c +++ b/fs/ext2/namei.c @@ -97,7 +97,7 @@ struct dentry *ext2_get_parent(struct dentry *child) * If the create succeeds, we fill in the inode information * with d_instantiate(). */ -static int ext2_create (struct mnt_idmap * idmap, +static int ext2_create (const struct mnt_idmap * idmap, struct inode * dir, struct dentry * dentry, umode_t mode) { @@ -117,7 +117,7 @@ static int ext2_create (struct mnt_idmap * idmap, return ext2_add_nondir(dentry, inode); } -static int ext2_tmpfile(struct mnt_idmap *idmap, struct inode *dir, +static int ext2_tmpfile(const struct mnt_idmap *idmap, struct inode *dir, struct file *file, umode_t mode) { struct inode *inode = ext2_new_inode(dir, mode, NULL); @@ -131,7 +131,7 @@ static int ext2_tmpfile(struct mnt_idmap *idmap, struct inode *dir, return finish_open_simple(file, 0); } -static int ext2_mknod (struct mnt_idmap * idmap, struct inode * dir, +static int ext2_mknod (const struct mnt_idmap * idmap, struct inode * dir, struct dentry *dentry, umode_t mode, dev_t rdev) { struct inode * inode; @@ -152,7 +152,7 @@ static int ext2_mknod (struct mnt_idmap * idmap, struct inode * dir, return err; } -static int ext2_symlink (struct mnt_idmap * idmap, struct inode * dir, +static int ext2_symlink (const struct mnt_idmap * idmap, struct inode * dir, struct dentry * dentry, const char * symname) { struct super_block * sb = dir->i_sb; @@ -223,7 +223,7 @@ static int ext2_link (struct dentry * old_dentry, struct inode * dir, return err; } -static struct dentry *ext2_mkdir(struct mnt_idmap * idmap, +static struct dentry *ext2_mkdir(const struct mnt_idmap * idmap, struct inode * dir, struct dentry * dentry, umode_t mode) { @@ -316,7 +316,7 @@ static int ext2_rmdir (struct inode * dir, struct dentry *dentry) return err; } -static int ext2_rename (struct mnt_idmap * idmap, +static int ext2_rename (const struct mnt_idmap * idmap, struct inode * old_dir, struct dentry * old_dentry, struct inode * new_dir, struct dentry * new_dentry, unsigned int flags) diff --git a/fs/ext2/super.c b/fs/ext2/super.c index a40f530872a4..fb217ef2d69f 100644 --- a/fs/ext2/super.c +++ b/fs/ext2/super.c @@ -127,6 +127,7 @@ void ext2_msg(struct super_block *sb, const char *prefix, * This must be called with sbi->s_lock held. */ void ext2_update_dynamic_rev(struct super_block *sb) + __must_hold(&EXT2_SB(sb)->s_lock) { struct ext2_super_block *es = EXT2_SB(sb)->s_es; @@ -633,6 +634,7 @@ static int ext2_parse_param(struct fs_context *fc, struct fs_parameter *param) static int ext2_setup_super (struct super_block * sb, struct ext2_super_block * es, int read_only) + __must_hold(&EXT2_SB(sb)->s_lock) { int res = 0; struct ext2_sb_info *sbi = EXT2_SB(sb); @@ -894,7 +896,7 @@ static int ext2_fill_super(struct super_block *sb, struct fs_context *fc) sb->s_fs_info = sbi; sbi->s_sb_block = sb_block; - spin_lock_init(&sbi->s_lock); + guard(spinlock_init)(&EXT2_SB(sb)->s_lock); ret = -EINVAL; /* @@ -1126,23 +1128,26 @@ static int ext2_fill_super(struct super_block *sb, struct fs_context *fc) goto failed_mount2; } sbi->s_gdb_count = db_count; - sbi->s_next_generation = get_random_u32(); spin_lock_init(&sbi->s_next_gen_lock); + scoped_guard(spinlock, &sbi->s_next_gen_lock) + sbi->s_next_generation = get_random_u32(); /* per filesystem reservation list head & lock */ spin_lock_init(&sbi->s_rsv_window_lock); - sbi->s_rsv_window_root = RB_ROOT; - /* - * Add a single, static dummy reservation to the start of the - * reservation window list --- it gives us a placeholder for - * append-at-start-of-list which makes the allocation logic - * _much_ simpler. - */ - sbi->s_rsv_window_head.rsv_start = EXT2_RESERVE_WINDOW_NOT_ALLOCATED; - sbi->s_rsv_window_head.rsv_end = EXT2_RESERVE_WINDOW_NOT_ALLOCATED; - sbi->s_rsv_window_head.rsv_alloc_hit = 0; - sbi->s_rsv_window_head.rsv_goal_size = 0; - ext2_rsv_window_add(sb, &sbi->s_rsv_window_head); + scoped_guard(spinlock, &EXT2_SB(sb)->s_rsv_window_lock) { + sbi->s_rsv_window_root = RB_ROOT; + /* + * Add a single, static dummy reservation to the start of the + * reservation window list --- it gives us a placeholder for + * append-at-start-of-list which makes the allocation logic + * _much_ simpler. + */ + sbi->s_rsv_window_head.rsv_start = EXT2_RESERVE_WINDOW_NOT_ALLOCATED; + sbi->s_rsv_window_head.rsv_end = EXT2_RESERVE_WINDOW_NOT_ALLOCATED; + sbi->s_rsv_window_head.rsv_alloc_hit = 0; + sbi->s_rsv_window_head.rsv_goal_size = 0; + ext2_rsv_window_add(sb, &sbi->s_rsv_window_head); + } err = percpu_counter_init(&sbi->s_freeblocks_counter, ext2_count_free_blocks(sb), GFP_KERNEL); diff --git a/fs/ext2/xattr.c b/fs/ext2/xattr.c index 9b68c490ab26..8f608930a48c 100644 --- a/fs/ext2/xattr.c +++ b/fs/ext2/xattr.c @@ -769,7 +769,7 @@ ext2_xattr_set2(struct inode *inode, struct buffer_head *old_bh, if (IS_SYNC(inode)) { sync_dirty_buffer(new_bh); error = -EIO; - if (buffer_req(new_bh) && !buffer_uptodate(new_bh)) + if (buffer_write_io_error(new_bh)) goto cleanup; } } diff --git a/fs/ext2/xattr_security.c b/fs/ext2/xattr_security.c index db47b8ab153e..ade074354258 100644 --- a/fs/ext2/xattr_security.c +++ b/fs/ext2/xattr_security.c @@ -19,7 +19,7 @@ ext2_xattr_security_get(const struct xattr_handler *handler, static int ext2_xattr_security_set(const struct xattr_handler *handler, - struct mnt_idmap *idmap, + const struct mnt_idmap *idmap, struct dentry *unused, struct inode *inode, const char *name, const void *value, size_t size, int flags) diff --git a/fs/ext2/xattr_trusted.c b/fs/ext2/xattr_trusted.c index 995f931228ce..0f12d634d6d0 100644 --- a/fs/ext2/xattr_trusted.c +++ b/fs/ext2/xattr_trusted.c @@ -26,7 +26,7 @@ ext2_xattr_trusted_get(const struct xattr_handler *handler, static int ext2_xattr_trusted_set(const struct xattr_handler *handler, - struct mnt_idmap *idmap, + const struct mnt_idmap *idmap, struct dentry *unused, struct inode *inode, const char *name, const void *value, size_t size, int flags) diff --git a/fs/ext2/xattr_user.c b/fs/ext2/xattr_user.c index dd1507231081..48002c033e9c 100644 --- a/fs/ext2/xattr_user.c +++ b/fs/ext2/xattr_user.c @@ -30,7 +30,7 @@ ext2_xattr_user_get(const struct xattr_handler *handler, static int ext2_xattr_user_set(const struct xattr_handler *handler, - struct mnt_idmap *idmap, + const struct mnt_idmap *idmap, struct dentry *unused, struct inode *inode, const char *name, const void *value, size_t size, int flags) diff --git a/fs/ext4/acl.c b/fs/ext4/acl.c index 3bffe862f954..59fac55a2426 100644 --- a/fs/ext4/acl.c +++ b/fs/ext4/acl.c @@ -225,7 +225,7 @@ __ext4_set_acl(handle_t *handle, struct inode *inode, int type, } int -ext4_set_acl(struct mnt_idmap *idmap, struct dentry *dentry, +ext4_set_acl(const struct mnt_idmap *idmap, struct dentry *dentry, struct posix_acl *acl, int type) { handle_t *handle; diff --git a/fs/ext4/acl.h b/fs/ext4/acl.h index 0c5a79c3b5d4..a14838c5bc42 100644 --- a/fs/ext4/acl.h +++ b/fs/ext4/acl.h @@ -56,7 +56,7 @@ static inline int ext4_acl_count(size_t size) /* acl.c */ struct posix_acl *ext4_get_acl(struct inode *inode, int type, bool rcu); -int ext4_set_acl(struct mnt_idmap *idmap, struct dentry *dentry, +int ext4_set_acl(const struct mnt_idmap *idmap, struct dentry *dentry, struct posix_acl *acl, int type); extern int ext4_init_acl(handle_t *, struct inode *, struct inode *); diff --git a/fs/ext4/ext4.h b/fs/ext4/ext4.h index 724a27e8be61..cbc59d03ca81 100644 --- a/fs/ext4/ext4.h +++ b/fs/ext4/ext4.h @@ -3044,7 +3044,7 @@ extern int ext4fs_dirhash(const struct inode *dir, const char *name, int len, /* ialloc.c */ extern int ext4_mark_inode_used(struct super_block *sb, int ino); -extern struct inode *__ext4_new_inode(struct mnt_idmap *, handle_t *, +extern struct inode *__ext4_new_inode(const struct mnt_idmap *, handle_t *, struct inode *, umode_t, const struct qstr *qstr, __u32 goal, uid_t *owner, __u32 i_flags, @@ -3179,14 +3179,14 @@ extern struct inode *__ext4_iget(struct super_block *sb, unsigned long ino, extern int ext4_write_inode(struct inode *, struct writeback_control *); extern int ext4_sync_inode_metadata(struct inode *, struct writeback_control *); -extern int ext4_setattr(struct mnt_idmap *, struct dentry *, +extern int ext4_setattr(const struct mnt_idmap *, struct dentry *, struct iattr *); extern u32 ext4_dio_alignment(struct inode *inode); -extern int ext4_getattr(struct mnt_idmap *, const struct path *, +extern int ext4_getattr(const struct mnt_idmap *, const struct path *, struct kstat *, u32, unsigned int); extern void ext4_evict_inode(struct inode *); extern void ext4_clear_inode(struct inode *); -extern int ext4_file_getattr(struct mnt_idmap *, const struct path *, +extern int ext4_file_getattr(const struct mnt_idmap *, const struct path *, struct kstat *, u32, unsigned int); extern void ext4_dirty_inode(struct inode *, int); extern int ext4_change_inode_journal_flag(struct inode *, int); @@ -3246,7 +3246,7 @@ extern int ext4_ind_remove_space(handle_t *handle, struct inode *inode, /* ioctl.c */ extern long ext4_ioctl(struct file *, unsigned int, unsigned long); extern long ext4_compat_ioctl(struct file *, unsigned int, unsigned long); -int ext4_fileattr_set(struct mnt_idmap *idmap, +int ext4_fileattr_set(const struct mnt_idmap *idmap, struct dentry *dentry, struct file_kattr *fa); int ext4_fileattr_get(struct dentry *dentry, struct file_kattr *fa); extern void ext4_reset_inode_seed(struct inode *inode); diff --git a/fs/ext4/ext4_jbd2.c b/fs/ext4/ext4_jbd2.c index 53ddedb52a6f..c241f50b97bc 100644 --- a/fs/ext4/ext4_jbd2.c +++ b/fs/ext4/ext4_jbd2.c @@ -421,7 +421,7 @@ int __ext4_handle_dirty_metadata(const char *where, unsigned int line, } if (inode && inode_needs_sync(inode)) { sync_dirty_buffer(bh); - if (buffer_req(bh) && !buffer_uptodate(bh)) { + if (buffer_write_io_error(bh)) { ext4_error_inode_err(inode, where, line, bh->b_blocknr, EIO, "IO error syncing itable block"); diff --git a/fs/ext4/ialloc.c b/fs/ext4/ialloc.c index a5831fc536db..529623103ae7 100644 --- a/fs/ext4/ialloc.c +++ b/fs/ext4/ialloc.c @@ -930,7 +930,7 @@ static int ext4_xattr_credits_for_new_inode(struct inode *dir, mode_t mode, * For other inodes, search forward from the parent directory's block * group to find a free inode. */ -struct inode *__ext4_new_inode(struct mnt_idmap *idmap, +struct inode *__ext4_new_inode(const struct mnt_idmap *idmap, handle_t *handle, struct inode *dir, umode_t mode, const struct qstr *qstr, __u32 goal, uid_t *owner, __u32 i_flags, diff --git a/fs/ext4/inode.c b/fs/ext4/inode.c index 26f0f9714f03..cb68bf50a3d6 100644 --- a/fs/ext4/inode.c +++ b/fs/ext4/inode.c @@ -6006,7 +6006,7 @@ static void ext4_wait_for_tail_page_commit(struct inode *inode) * * Called with inode->i_rwsem down. */ -int ext4_setattr(struct mnt_idmap *idmap, struct dentry *dentry, +int ext4_setattr(const struct mnt_idmap *idmap, struct dentry *dentry, struct iattr *attr) { struct inode *inode = d_inode(dentry); @@ -6263,7 +6263,7 @@ u32 ext4_dio_alignment(struct inode *inode) return 1; /* use the iomap defaults */ } -int ext4_getattr(struct mnt_idmap *idmap, const struct path *path, +int ext4_getattr(const struct mnt_idmap *idmap, const struct path *path, struct kstat *stat, u32 request_mask, unsigned int query_flags) { struct inode *inode = d_inode(path->dentry); @@ -6332,7 +6332,7 @@ int ext4_getattr(struct mnt_idmap *idmap, const struct path *path, return 0; } -int ext4_file_getattr(struct mnt_idmap *idmap, +int ext4_file_getattr(const struct mnt_idmap *idmap, const struct path *path, struct kstat *stat, u32 request_mask, unsigned int query_flags) { diff --git a/fs/ext4/ioctl.c b/fs/ext4/ioctl.c index c8387e6a2c6e..0a54b00e5be5 100644 --- a/fs/ext4/ioctl.c +++ b/fs/ext4/ioctl.c @@ -373,7 +373,7 @@ void ext4_reset_inode_seed(struct inode *inode) * */ static long swap_inode_boot_loader(struct super_block *sb, - struct mnt_idmap *idmap, + const struct mnt_idmap *idmap, struct inode *inode) { handle_t *handle; @@ -1008,7 +1008,7 @@ int ext4_fileattr_get(struct dentry *dentry, struct file_kattr *fa) return 0; } -int ext4_fileattr_set(struct mnt_idmap *idmap, +int ext4_fileattr_set(const struct mnt_idmap *idmap, struct dentry *dentry, struct file_kattr *fa) { struct inode *inode = d_inode(dentry); @@ -1539,7 +1539,7 @@ static long __ext4_ioctl(struct file *filp, unsigned int cmd, unsigned long arg) { struct inode *inode = file_inode(filp); struct super_block *sb = inode->i_sb; - struct mnt_idmap *idmap = file_mnt_idmap(filp); + const struct mnt_idmap *idmap = file_mnt_idmap(filp); ext4_debug("cmd = %u, arg = %lu\n", cmd, arg); diff --git a/fs/ext4/mmp.c b/fs/ext4/mmp.c index 7ce361484b38..4b18ddef468d 100644 --- a/fs/ext4/mmp.c +++ b/fs/ext4/mmp.c @@ -49,7 +49,7 @@ static int write_mmp_block_thawed(struct super_block *sb, bh_submit(bh, REQ_OP_WRITE | REQ_SYNC | REQ_META | REQ_PRIO, bh_end_write); wait_on_buffer(bh); - if (unlikely(!buffer_uptodate(bh))) + if (unlikely(buffer_write_io_error(bh))) return -EIO; return 0; } diff --git a/fs/ext4/namei.c b/fs/ext4/namei.c index a6386c1d237f..6e0630a49e48 100644 --- a/fs/ext4/namei.c +++ b/fs/ext4/namei.c @@ -2812,7 +2812,7 @@ static int ext4_add_nondir(handle_t *handle, * If the create succeeds, we fill in the inode information * with d_instantiate(). */ -static int ext4_create(struct mnt_idmap *idmap, struct inode *dir, +static int ext4_create(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *dentry, umode_t mode) { handle_t *handle; @@ -2847,7 +2847,7 @@ retry: return err; } -static int ext4_mknod(struct mnt_idmap *idmap, struct inode *dir, +static int ext4_mknod(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *dentry, umode_t mode, dev_t rdev) { handle_t *handle; @@ -2881,7 +2881,7 @@ retry: return err; } -static int ext4_tmpfile(struct mnt_idmap *idmap, struct inode *dir, +static int ext4_tmpfile(const struct mnt_idmap *idmap, struct inode *dir, struct file *file, umode_t mode) { handle_t *handle; @@ -2994,7 +2994,7 @@ out: return err; } -static struct dentry *ext4_mkdir(struct mnt_idmap *idmap, struct inode *dir, +static struct dentry *ext4_mkdir(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *dentry, umode_t mode) { handle_t *handle; @@ -3360,7 +3360,7 @@ out: return err; } -static int ext4_symlink(struct mnt_idmap *idmap, struct inode *dir, +static int ext4_symlink(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *dentry, const char *symname) { handle_t *handle; @@ -3753,7 +3753,7 @@ static void ext4_update_dir_count(handle_t *handle, struct ext4_renament *ent) } } -static struct inode *ext4_whiteout_for_rename(struct mnt_idmap *idmap, +static struct inode *ext4_whiteout_for_rename(const struct mnt_idmap *idmap, struct ext4_renament *ent, int credits, handle_t **h) { @@ -3796,7 +3796,7 @@ retry: * while new_{dentry,inode) refers to the destination dentry/inode * This comes from rename(const char *oldpath, const char *newpath) */ -static int ext4_rename(struct mnt_idmap *idmap, struct inode *old_dir, +static int ext4_rename(const struct mnt_idmap *idmap, struct inode *old_dir, struct dentry *old_dentry, struct inode *new_dir, struct dentry *new_dentry, unsigned int flags) { @@ -4191,7 +4191,7 @@ end_rename: return retval; } -static int ext4_rename2(struct mnt_idmap *idmap, +static int ext4_rename2(const struct mnt_idmap *idmap, struct inode *old_dir, struct dentry *old_dentry, struct inode *new_dir, struct dentry *new_dentry, unsigned int flags) diff --git a/fs/ext4/symlink.c b/fs/ext4/symlink.c index b612262719ed..e680d1e45b47 100644 --- a/fs/ext4/symlink.c +++ b/fs/ext4/symlink.c @@ -55,7 +55,7 @@ static const char *ext4_encrypted_get_link(struct dentry *dentry, return paddr; } -static int ext4_encrypted_symlink_getattr(struct mnt_idmap *idmap, +static int ext4_encrypted_symlink_getattr(const struct mnt_idmap *idmap, const struct path *path, struct kstat *stat, u32 request_mask, unsigned int query_flags) diff --git a/fs/ext4/xattr_hurd.c b/fs/ext4/xattr_hurd.c index 8a5842e4cd95..a3ecbff72b10 100644 --- a/fs/ext4/xattr_hurd.c +++ b/fs/ext4/xattr_hurd.c @@ -32,7 +32,7 @@ ext4_xattr_hurd_get(const struct xattr_handler *handler, static int ext4_xattr_hurd_set(const struct xattr_handler *handler, - struct mnt_idmap *idmap, + const struct mnt_idmap *idmap, struct dentry *unused, struct inode *inode, const char *name, const void *value, size_t size, int flags) diff --git a/fs/ext4/xattr_security.c b/fs/ext4/xattr_security.c index 776cf11d24ca..af5b8a93fed1 100644 --- a/fs/ext4/xattr_security.c +++ b/fs/ext4/xattr_security.c @@ -23,7 +23,7 @@ ext4_xattr_security_get(const struct xattr_handler *handler, static int ext4_xattr_security_set(const struct xattr_handler *handler, - struct mnt_idmap *idmap, + const struct mnt_idmap *idmap, struct dentry *unused, struct inode *inode, const char *name, const void *value, size_t size, int flags) diff --git a/fs/ext4/xattr_trusted.c b/fs/ext4/xattr_trusted.c index 9811eb0ab276..458e1982ef83 100644 --- a/fs/ext4/xattr_trusted.c +++ b/fs/ext4/xattr_trusted.c @@ -30,7 +30,7 @@ ext4_xattr_trusted_get(const struct xattr_handler *handler, static int ext4_xattr_trusted_set(const struct xattr_handler *handler, - struct mnt_idmap *idmap, + const struct mnt_idmap *idmap, struct dentry *unused, struct inode *inode, const char *name, const void *value, size_t size, int flags) diff --git a/fs/ext4/xattr_user.c b/fs/ext4/xattr_user.c index 4b70bf4e7626..ad35215f6610 100644 --- a/fs/ext4/xattr_user.c +++ b/fs/ext4/xattr_user.c @@ -31,7 +31,7 @@ ext4_xattr_user_get(const struct xattr_handler *handler, static int ext4_xattr_user_set(const struct xattr_handler *handler, - struct mnt_idmap *idmap, + const struct mnt_idmap *idmap, struct dentry *unused, struct inode *inode, const char *name, const void *value, size_t size, int flags) diff --git a/fs/f2fs/Makefile b/fs/f2fs/Makefile index 8a7322d229e4..fbf49c30b066 100644 --- a/fs/f2fs/Makefile +++ b/fs/f2fs/Makefile @@ -3,7 +3,7 @@ obj-$(CONFIG_F2FS_FS) += f2fs.o f2fs-y := dir.o file.o inode.o namei.o hash.o super.o inline.o f2fs-y += checkpoint.o gc.o data.o node.o segment.o recovery.o -f2fs-y += shrinker.o extent_cache.o sysfs.o +f2fs-y += shrinker.o extent_cache.o sysfs.o cache.o f2fs-$(CONFIG_F2FS_STAT_FS) += debug.o f2fs-$(CONFIG_F2FS_FS_XATTR) += xattr.o f2fs-$(CONFIG_F2FS_FS_POSIX_ACL) += acl.o diff --git a/fs/f2fs/acl.c b/fs/f2fs/acl.c index d3253549173e..a7485bc38252 100644 --- a/fs/f2fs/acl.c +++ b/fs/f2fs/acl.c @@ -181,7 +181,7 @@ fail: } static struct posix_acl *__f2fs_get_acl(struct inode *inode, int type, - struct folio *dfolio) + struct f2fs_cached_block *entry) { int name_index = F2FS_XATTR_INDEX_POSIX_ACL_DEFAULT; void *value = NULL; @@ -191,13 +191,13 @@ static struct posix_acl *__f2fs_get_acl(struct inode *inode, int type, if (type == ACL_TYPE_ACCESS) name_index = F2FS_XATTR_INDEX_POSIX_ACL_ACCESS; - retval = f2fs_getxattr(inode, name_index, "", NULL, 0, dfolio); + retval = f2fs_getxattr(inode, name_index, "", NULL, 0, entry); if (retval > 0) { value = f2fs_kmalloc(F2FS_I_SB(inode), retval, GFP_F2FS_ZERO); if (!value) return ERR_PTR(-ENOMEM); retval = f2fs_getxattr(inode, name_index, "", value, - retval, dfolio); + retval, entry); } if (retval > 0) @@ -219,7 +219,7 @@ struct posix_acl *f2fs_get_acl(struct inode *inode, int type, bool rcu) return __f2fs_get_acl(inode, type, NULL); } -static int f2fs_acl_update_mode(struct mnt_idmap *idmap, +static int f2fs_acl_update_mode(const struct mnt_idmap *idmap, struct inode *inode, umode_t *mode_p, struct posix_acl **acl) { @@ -240,9 +240,9 @@ static int f2fs_acl_update_mode(struct mnt_idmap *idmap, return 0; } -static int __f2fs_set_acl(struct mnt_idmap *idmap, +static int __f2fs_set_acl(const struct mnt_idmap *idmap, struct inode *inode, int type, - struct posix_acl *acl, struct folio *ifolio) + struct posix_acl *acl, struct f2fs_cached_block *ientry) { int name_index; void *value = NULL; @@ -253,7 +253,7 @@ static int __f2fs_set_acl(struct mnt_idmap *idmap, switch (type) { case ACL_TYPE_ACCESS: name_index = F2FS_XATTR_INDEX_POSIX_ACL_ACCESS; - if (acl && !ifolio) { + if (acl && !ientry) { error = f2fs_acl_update_mode(idmap, inode, &mode, &acl); if (error) return error; @@ -279,7 +279,7 @@ static int __f2fs_set_acl(struct mnt_idmap *idmap, } } - error = f2fs_setxattr(inode, name_index, "", value, size, ifolio, 0); + error = f2fs_setxattr(inode, name_index, "", value, size, ientry, 0); kfree(value); if (!error) @@ -289,7 +289,7 @@ static int __f2fs_set_acl(struct mnt_idmap *idmap, return error; } -int f2fs_set_acl(struct mnt_idmap *idmap, struct dentry *dentry, +int f2fs_set_acl(const struct mnt_idmap *idmap, struct dentry *dentry, struct posix_acl *acl, int type) { struct inode *inode = d_inode(dentry); @@ -374,7 +374,7 @@ static int f2fs_acl_create_masq(struct posix_acl *acl, umode_t *mode_p) static int f2fs_acl_create(struct inode *dir, umode_t *mode, struct posix_acl **default_acl, struct posix_acl **acl, - struct folio *dfolio) + struct f2fs_cached_block *entry) { struct posix_acl *p; struct posix_acl *clone; @@ -386,7 +386,7 @@ static int f2fs_acl_create(struct inode *dir, umode_t *mode, if (S_ISLNK(*mode) || !IS_POSIXACL(dir)) return 0; - p = __f2fs_get_acl(dir, ACL_TYPE_DEFAULT, dfolio); + p = __f2fs_get_acl(dir, ACL_TYPE_DEFAULT, entry); if (!p || p == ERR_PTR(-EOPNOTSUPP)) { *mode &= ~current_umask(); return 0; @@ -423,13 +423,13 @@ release_acl: return ret; } -int f2fs_init_acl(struct inode *inode, struct inode *dir, struct folio *ifolio, - struct folio *dfolio) +int f2fs_init_acl(struct inode *inode, struct inode *dir, struct f2fs_cached_block *ientry, + struct f2fs_cached_block *dentry) { struct posix_acl *default_acl = NULL, *acl = NULL; int error; - error = f2fs_acl_create(dir, &inode->i_mode, &default_acl, &acl, dfolio); + error = f2fs_acl_create(dir, &inode->i_mode, &default_acl, &acl, dentry); if (error) return error; @@ -437,7 +437,7 @@ int f2fs_init_acl(struct inode *inode, struct inode *dir, struct folio *ifolio, if (default_acl) { error = __f2fs_set_acl(NULL, inode, ACL_TYPE_DEFAULT, - default_acl, ifolio); + default_acl, ientry); posix_acl_release(default_acl); } else { inode->i_default_acl = NULL; @@ -445,7 +445,7 @@ int f2fs_init_acl(struct inode *inode, struct inode *dir, struct folio *ifolio, if (acl) { if (!error) error = __f2fs_set_acl(NULL, inode, ACL_TYPE_ACCESS, - acl, ifolio); + acl, ientry); posix_acl_release(acl); } else { inode->i_acl = NULL; diff --git a/fs/f2fs/acl.h b/fs/f2fs/acl.h index 20e87e63c089..b1085efcc05e 100644 --- a/fs/f2fs/acl.h +++ b/fs/f2fs/acl.h @@ -34,16 +34,18 @@ struct f2fs_acl_header { #ifdef CONFIG_F2FS_FS_POSIX_ACL struct posix_acl *f2fs_get_acl(struct inode *, int, bool); -int f2fs_set_acl(struct mnt_idmap *, struct dentry *, +int f2fs_set_acl(const struct mnt_idmap *, struct dentry *, struct posix_acl *, int); -int f2fs_init_acl(struct inode *, struct inode *, struct folio *ifolio, - struct folio *dfolio); +int f2fs_init_acl(struct inode *inode, struct inode *dir, + struct f2fs_cached_block *ientry, + struct f2fs_cached_block *dentry); #else #define f2fs_get_acl NULL #define f2fs_set_acl NULL static inline int f2fs_init_acl(struct inode *inode, struct inode *dir, - struct folio *ifolio, struct folio *dfolio) + struct f2fs_cached_block *ientry, + struct f2fs_cached_block *dentry) { return 0; } diff --git a/fs/f2fs/cache.c b/fs/f2fs/cache.c new file mode 100644 index 000000000000..38fc5eb17f92 --- /dev/null +++ b/fs/f2fs/cache.c @@ -0,0 +1,720 @@ +// SPDX-License-Identifier: GPL-2.0 +/* + * Copyright (c) 2026 Google LLC + * Author: Chao Yu <chaseyu@google.com> + */ +#include <linux/fs.h> +#include <linux/f2fs_fs.h> +#include <linux/radix-tree.h> +#include <linux/slab.h> +#include <linux/list.h> +#include <linux/pagemap.h> +#include <linux/kthread.h> +#include <linux/freezer.h> +#include <linux/delay.h> +#include "f2fs.h" +#include "cache.h" +#include "node.h" +#include <trace/events/f2fs.h> +#include "segment.h" + +void f2fs_cache_wait_writeback_cond(struct f2fs_cached_block *entry, + enum page_type type) +{ + /* in case the entry was truncated or on-going shrink */ + if (!entry->cache) + return; + + if (!f2fs_cache_test_writeback(entry)) + return; + + /* submit cached bio */ + f2fs_submit_merged_write_cache(entry->cache->sbi, entry, 0, type); + + wait_on_bit_io(&entry->state, F2FS_BLOCK_WRITEBACK, + TASK_UNINTERRUPTIBLE); +} + +void f2fs_cache_wait_writeback(struct f2fs_cached_block *entry) +{ + /* in case the entry was truncated or on-going shrink */ + if (!entry->cache) + return; + + f2fs_cache_wait_writeback_cond(entry, + IS_META_CACHE(entry->cache) ? META : NODE); +} + +void f2fs_cache_update_tag(struct f2fs_cached_block *entry, + unsigned int clear_from, unsigned int set_to) +{ + struct f2fs_cached_block_list *cache = entry->cache; + unsigned long flags; + + spin_lock_irqsave(&cache->tree_lock, flags); + if (clear_from != F2FS_CACHE_TAG_NONE) + radix_tree_tag_clear(&cache->root, entry->index, clear_from); + if (set_to != F2FS_CACHE_TAG_NONE) + radix_tree_tag_set(&cache->root, entry->index, set_to); + spin_unlock_irqrestore(&cache->tree_lock, flags); +} + +bool f2fs_mark_cache_dirty(struct f2fs_cached_block *entry) +{ + struct f2fs_cached_block_list *cache = entry->cache; + + f2fs_cache_set_uptodate(entry); + +#ifdef CONFIG_F2FS_CHECK_FS + if (f2fs_is_node_cache(entry) && IS_INODE(cache->sbi, entry)) + f2fs_inode_chksum_set(cache->sbi, entry); +#endif + + if (!f2fs_cache_test_and_set_dirty(entry)) { + enum count_type type = IS_META_CACHE(cache) ? + F2FS_DIRTY_META : F2FS_DIRTY_NODES; + + trace_f2fs_cache_set_dirty(entry, + IS_META_CACHE(cache) ? META : NODE); + f2fs_cache_update_tag(entry, F2FS_CACHE_TAG_NONE, + F2FS_CACHE_TAG_DIRTY); + inc_cache_count(cache->sbi, type); + return true; + } + + return false; +} + +void f2fs_drop_cache_dirty(struct f2fs_cached_block *entry) +{ + + struct f2fs_cached_block_list *cache = entry->cache; + enum count_type type = IS_META_CACHE(cache) ? + F2FS_DIRTY_META : F2FS_DIRTY_NODES; + + f2fs_cache_clear_uptodate(entry); + + if (!f2fs_cache_test_and_clear_dirty(entry)) + return; + + f2fs_cache_update_tag(entry, F2FS_CACHE_TAG_DIRTY, + F2FS_CACHE_TAG_NONE); + dec_cache_count(cache->sbi, type); +} + +void f2fs_start_cache_writeback(struct f2fs_cached_block *entry) +{ + f2fs_cache_set_writeback(entry); + f2fs_cache_update_tag(entry, F2FS_CACHE_TAG_DIRTY, + F2FS_CACHE_TAG_WRITEBACK); +} + +void f2fs_end_cache_writeback(struct f2fs_cached_block *entry) +{ + /* + * should call f2fs_cache_update_tag() before clearing writeback bit, + * in case f2fs_truncate_cache() set entry->cache to NULL. + */ + f2fs_cache_update_tag(entry, F2FS_CACHE_TAG_WRITEBACK, + F2FS_CACHE_TAG_NONE); + clear_and_wake_up_bit(F2FS_BLOCK_WRITEBACK, &entry->state); +} + +static int f2fs_cache_refcount(struct f2fs_cached_block *entry) +{ + return atomic_read(&entry->refcount); +} + +static void f2fs_do_free_cache(struct f2fs_cached_block *entry) +{ + kfree(entry->data); + kfree(entry); +} + +static void f2fs_free_cache(struct f2fs_cached_block *entry) +{ + WARN_ON_ONCE(!list_empty(&entry->list)); + WARN_ON_ONCE(f2fs_cache_refcount(entry)); + f2fs_do_free_cache(entry); +} + +void f2fs_cache_get(struct f2fs_cached_block *entry) +{ + atomic_inc(&entry->refcount); +} + +static bool f2fs_cache_put(struct f2fs_cached_block *entry) +{ + WARN_ON_ONCE(!f2fs_cache_refcount(entry)); + if (atomic_dec_and_test(&entry->refcount)) { + f2fs_free_cache(entry); + return true; + } + return false; +} + +static struct f2fs_cached_block *f2fs_create_cache( + struct f2fs_cached_block_list *cache, + unsigned long index, bool nofail) +{ + struct f2fs_sb_info *sbi = cache->sbi; + struct f2fs_cached_block *entry; + unsigned int flags = GFP_NOFS; + + if (index == ULONG_MAX) + return ERR_PTR(-ERANGE); + + if (nofail) { + flags |= __GFP_NOFAIL; + entry = kzalloc_obj(*entry, flags); + entry->data = kmalloc(sbi->blocksize, flags); + } else { + entry = f2fs_kzalloc(sbi, sizeof(*entry), flags); + if (!entry) + return ERR_PTR(-ENOMEM); + + entry->data = f2fs_kmalloc(sbi, sbi->blocksize, flags); + if (!entry->data) { + kfree(entry); + return ERR_PTR(-ENOMEM); + } + } + entry->index = index; + + atomic_set(&entry->refcount, 0); + if (!IS_COMPRESS_CACHE(cache)) + entry->next_entry = NULL; + else + entry->ino = 0; + INIT_LIST_HEAD(&entry->list); + + entry->cache = cache; + + return entry; +} + +static struct f2fs_cached_block *f2fs_insert_cache( + struct f2fs_cached_block_list *cache, + unsigned long index, + struct f2fs_cached_block *new) +{ + struct f2fs_cached_block *e; + int ret; + unsigned long flags; + + ret = radix_tree_preload(GFP_NOFS | __GFP_NOFAIL); + f2fs_bug_on(cache->sbi, ret); + + spin_lock(&cache->list_lock); + spin_lock_irqsave(&cache->tree_lock, flags); + e = radix_tree_lookup(&cache->root, index); + if (!e) { + e = new; + f2fs_bug_on(cache->sbi, f2fs_cache_refcount(e)); + + ret = radix_tree_insert(&cache->root, index, e); + f2fs_bug_on(cache->sbi, ret); + + /* radix tree referenced cache entry */ + f2fs_cache_get(e); + f2fs_bug_on(cache->sbi, !list_empty(&e->list)); + list_add_tail(&e->list, &cache->lru_list); + cache->num_entries++; + } + f2fs_cache_get(e); + spin_unlock_irqrestore(&cache->tree_lock, flags); + spin_unlock(&cache->list_lock); + radix_tree_preload_end(); + + if (new != e) { + f2fs_bug_on(cache->sbi, f2fs_cache_refcount(new)); + f2fs_do_free_cache(new); + } + + return e; +} + +struct f2fs_cached_block *f2fs_find_cache( + struct f2fs_cached_block_list *cache, + unsigned long index, + enum f2fs_cache_request_flag rflag) +{ + struct f2fs_cached_block *entry; + unsigned long flags; + bool access = rflag & F2FS_CACHE_ACCESS; + + spin_lock_irqsave(&cache->tree_lock, flags); + entry = radix_tree_lookup(&cache->root, index); + if (entry) { + f2fs_bug_on(cache->sbi, !f2fs_cache_refcount(entry)); + f2fs_cache_get(entry); + if (access && !f2fs_cache_test_referenced(entry)) + f2fs_cache_set_referenced(entry); + } else { + entry = ERR_PTR(-ENOENT); + } + spin_unlock_irqrestore(&cache->tree_lock, flags); + + return entry; +} + +struct f2fs_cached_block *f2fs_grab_cache( + struct f2fs_cached_block_list *cache, + unsigned long index, int flags) + +{ + struct f2fs_cached_block *entry, *new; + bool create = flags & F2FS_CACHE_CREATE; + bool nofail = flags & F2FS_CACHE_NOFAIL; + bool lock = flags & F2FS_CACHE_LOCK; + +repeat: + entry = f2fs_find_cache(cache, index, F2FS_CACHE_ACCESS); + if (!IS_ERR(entry)) + goto found; + + if (!create) + return ERR_PTR(-ENOENT); + + new = f2fs_create_cache(cache, index, nofail); + if (IS_ERR(new)) + return new; + + entry = f2fs_insert_cache(cache, index, new); +found: + if (lock) { + f2fs_lock_cache(entry); + /* has been truncated */ + if (entry->cache != cache) { + f2fs_put_cache(entry, true); + goto repeat; + } + } + return entry; +} + +bool f2fs_trylock_cache(struct f2fs_cached_block *entry) +{ + return !test_and_set_bit(F2FS_BLOCK_LOCKED, &entry->state); +} + +void f2fs_lock_cache(struct f2fs_cached_block *entry) +{ + wait_on_bit_lock(&entry->state, F2FS_BLOCK_LOCKED, + TASK_UNINTERRUPTIBLE); +} + +void f2fs_unlock_cache(struct f2fs_cached_block *entry) +{ + clear_and_wake_up_bit(F2FS_BLOCK_LOCKED, &entry->state); +} + +bool f2fs_put_cache(struct f2fs_cached_block *entry, bool unlock) +{ + if (IS_ERR_OR_NULL(entry)) + return false; + if (unlock) + f2fs_unlock_cache(entry); + return f2fs_cache_put(entry); +} + +unsigned int f2fs_cache_gang_lookup(struct f2fs_cached_block_list *cache, + struct f2fs_cached_block **entries, + pgoff_t *index, unsigned long end) +{ + unsigned long flags; + unsigned int max_nr = min((unsigned long)F2FS_ONSTACK_CACHES, end - *index); + int nr, i; + + if (*index >= end || *index == ULONG_MAX) + return 0; + + spin_lock_irqsave(&cache->tree_lock, flags); + nr = radix_tree_gang_lookup(&cache->root, (void **)entries, + *index, max_nr); + if (!nr) + goto out_unlock; + + for (i = 0; i < nr; i++) { + struct f2fs_cached_block *entry = entries[i]; + + if (entry->index >= end) { + nr = i; + break; + } + f2fs_cache_get(entry); + } + if (nr) + *index = entries[nr - 1]->index + 1; +out_unlock: + spin_unlock_irqrestore(&cache->tree_lock, flags); + return nr; +} + +unsigned int f2fs_cache_gang_lookup_tag(struct f2fs_cached_block_list *cache, + struct f2fs_cached_block **entries, + pgoff_t *index, unsigned int max_nr, + int tag) +{ + unsigned long flags; + int nr, i; + + if (*index == ULONG_MAX) + return 0; + + spin_lock_irqsave(&cache->tree_lock, flags); + nr = radix_tree_gang_lookup_tag(&cache->root, (void **)entries, + *index, max_nr, tag); + if (!nr) + goto out; + + for (i = 0; i < nr; i++) + f2fs_cache_get(entries[i]); + *index = entries[nr - 1]->index + 1; +out: + spin_unlock_irqrestore(&cache->tree_lock, flags); + return nr; +} + +void f2fs_cache_gang_release(struct f2fs_cached_block **entries, + unsigned int nr_entries) +{ + int i; + + for (i = 0; i < nr_entries; i++) + f2fs_put_cache(entries[i], false); +} + +void f2fs_cache_wait_on_all_writeback(struct f2fs_cached_block_list *cache) +{ + unsigned long index = 0; + struct f2fs_cached_block *entries[F2FS_ONSTACK_CACHES]; + int nr, i; + +next: + nr = f2fs_cache_gang_lookup_tag(cache, entries, &index, + F2FS_ONSTACK_CACHES, F2FS_CACHE_TAG_WRITEBACK); + if (!nr) + return; + + for (i = 0; i < nr; i++) + f2fs_cache_wait_writeback(entries[i]); + f2fs_cache_gang_release(entries, nr); + + cond_resched(); + goto next; +} + +static void f2fs_do_truncate_cache(struct f2fs_cached_block *entry, + bool drop_dirty) +{ + struct f2fs_cached_block_list *cache = entry->cache; + unsigned long flags; + + if (drop_dirty) + goto drop_it; + + if (f2fs_cache_test_dirty(entry) || + f2fs_cache_test_writeback(entry)) + return; + +drop_it: + f2fs_cache_wait_writeback(entry); + f2fs_drop_cache_dirty(entry); + + spin_lock(&cache->list_lock); + spin_lock_irqsave(&cache->tree_lock, flags); + + f2fs_bug_on(cache->sbi, !entry->cache); + if (!radix_tree_delete(&cache->root, entry->index)) + f2fs_bug_on(cache->sbi, !entry->cache); + + entry->cache = NULL; + cache->num_entries--; + + atomic_dec(&entry->refcount); + f2fs_bug_on(cache->sbi, !f2fs_cache_refcount(entry)); + + f2fs_bug_on(cache->sbi, list_empty(&entry->list)); + list_del_init(&entry->list); + + spin_unlock_irqrestore(&cache->tree_lock, flags); + spin_unlock(&cache->list_lock); +} + +void f2fs_truncate_locked_cache(struct f2fs_cached_block *entry, + bool drop_dirty) +{ + if (!entry->cache) + return; + f2fs_do_truncate_cache(entry, drop_dirty); +} + +void f2fs_truncate_cache(struct f2fs_cached_block *entry, bool drop_dirty) +{ + f2fs_lock_cache(entry); + f2fs_truncate_locked_cache(entry, drop_dirty); + f2fs_unlock_cache(entry); +} + +static void f2fs_drop_cache(struct f2fs_cached_block_list *cache, + block_t blkaddr, bool drop_dirty) +{ + struct f2fs_cached_block *entry; + + entry = f2fs_find_cache(cache, blkaddr, 0); + if (IS_ERR(entry)) + return; + + f2fs_truncate_cache(entry, drop_dirty); + f2fs_put_cache(entry, false); +} + +void f2fs_drop_cache_range(struct f2fs_cached_block_list *cache, + unsigned long start, unsigned long len, bool drop_dirty) +{ + unsigned long index = start; + unsigned long end = (ULONG_MAX - start < len) ? + ULONG_MAX : (start + len); + struct f2fs_cached_block *entries[F2FS_ONSTACK_CACHES]; + int nr, i; + + if (len == 1) + return f2fs_drop_cache(cache, index, drop_dirty); + +next: + nr = f2fs_cache_gang_lookup(cache, entries, &index, end); + if (!nr) + return; + + for (i = 0; i < nr; i++) + f2fs_truncate_cache(entries[i], drop_dirty); + f2fs_cache_gang_release(entries, nr); + + if (index < end) { + cond_resched(); + goto next; + } +} + +int f2fs_init_cache(struct f2fs_sb_info *sbi, + struct f2fs_cached_block_list *cache, + enum f2fs_cache_type type) +{ + cache->sbi = sbi; + cache->type = type; + INIT_RADIX_TREE(&cache->root, GFP_ATOMIC); + spin_lock_init(&cache->tree_lock); + spin_lock_init(&cache->list_lock); + INIT_LIST_HEAD(&cache->lru_list); + cache->num_entries = 0; + + return 0; +} + +void f2fs_destroy_cache(struct f2fs_cached_block_list *cache) +{ + struct list_head *head = &cache->lru_list; + struct f2fs_cached_block *entry; + unsigned long flags; + + f2fs_cache_wait_on_all_writeback(cache); +next: + spin_lock(&cache->list_lock); + if (list_empty(head)) { + spin_unlock(&cache->list_lock); + return; + } + entry = list_first_entry(head, struct f2fs_cached_block, list); + + spin_lock_irqsave(&cache->tree_lock, flags); + radix_tree_delete(&cache->root, entry->index); + cache->num_entries--; + list_del_init(&entry->list); + spin_unlock_irqrestore(&cache->tree_lock, flags); + + spin_unlock(&cache->list_lock); + + /* wait on read cache IO */ + f2fs_lock_cache(entry); + /* wait on write cache IO */ + f2fs_cache_wait_writeback(entry); + f2fs_bug_on(cache->sbi, f2fs_cache_test_dirty(entry)); + f2fs_bug_on(cache->sbi, f2fs_cache_test_writeback(entry)); + f2fs_bug_on(cache->sbi, !list_empty(&entry->list)); + f2fs_bug_on(cache->sbi, f2fs_cache_refcount(entry) != 1); + f2fs_put_cache(entry, true); + goto next; +} + +static unsigned long f2fs_do_shrink_cache(struct f2fs_cached_block_list *cache, + unsigned long nr_to_scan) +{ + struct f2fs_cached_block *entry, *next; + LIST_HEAD(dispose_list); + LIST_HEAD(keep_list); + unsigned long freed = 0; + unsigned long scanned = 0; + + /* Phase 1: Isolate candidate entries from LRU list into dispose_list */ + spin_lock(&cache->list_lock); + list_for_each_entry_safe(entry, next, &cache->lru_list, list) { + if (scanned >= cache->num_entries) + break; + if (scanned++ >= nr_to_scan) + break; + + /* If accessed, give it a second chance to rotate to tail */ + if (f2fs_cache_test_and_clear_referenced(entry)) { + list_move_tail(&entry->list, &cache->lru_list); + continue; + } + + if (f2fs_cache_test_dirty(entry) || + f2fs_cache_test_writeback(entry) || + f2fs_cache_test_locked(entry)) + continue; + + if (f2fs_cache_refcount(entry) != 1) + continue; + + list_move_tail(&entry->list, &dispose_list); + } + spin_unlock(&cache->list_lock); + + /* Phase 2: Process isolated candidates one by one */ + while (1) { + spin_lock(&cache->list_lock); + entry = list_first_entry_or_null(&dispose_list, + struct f2fs_cached_block, list); + if (!entry) { + spin_unlock(&cache->list_lock); + break; + } + f2fs_cache_get(entry); + list_move_tail(&entry->list, &keep_list); + spin_unlock(&cache->list_lock); + + if (!f2fs_trylock_cache(entry)) { + f2fs_put_cache(entry, false); + continue; + } + + /* the entry has been truncated */ + if (!entry->cache) { + f2fs_put_cache(entry, true); + continue; + } + /* + * at least there are shrinker, radix tree and another user + * has referenced the entry. + */ + if (f2fs_cache_refcount(entry) >= 3) { + f2fs_put_cache(entry, true); + continue; + } + + f2fs_do_truncate_cache(entry, false); + + if (f2fs_put_cache(entry, true)) + freed++; + } + + /* Phase 3: Splice un-reclaimed entries back onto cache->lru_list */ + if (!list_empty(&keep_list)) { + spin_lock(&cache->list_lock); + list_splice_tail(&keep_list, &cache->lru_list); + spin_unlock(&cache->list_lock); + } + + return freed; +} + +unsigned long f2fs_shrink_cache(struct f2fs_sb_info *sbi, + unsigned long nr_to_scan) +{ + unsigned long freed; + + freed = f2fs_do_shrink_cache(META_CACHE(sbi), nr_to_scan); + if (freed >= nr_to_scan) + return freed; + + freed += f2fs_do_shrink_cache(NODE_CACHE(sbi), nr_to_scan - freed); + if (freed >= nr_to_scan) + return freed; + + freed += f2fs_do_shrink_cache(COMPRESS_CACHE(sbi), nr_to_scan - freed); + return freed; +} + +static int f2fs_cache_writeback_kthread(void *data) +{ + struct f2fs_sb_info *sbi = data; + struct f2fs_cache_kthread *cache_thread = &sbi->cache_thread; + wait_queue_head_t *wq = &cache_thread->cache_wb_wq; + + set_freezable(); + + while (!kthread_should_stop()) { + unsigned int interval = cache_thread->cache_wb_interval; + + wait_event_freezable_timeout(*wq, + kthread_should_stop(), + msecs_to_jiffies(interval)); + + if (kthread_should_stop()) + break; + + if (f2fs_readonly(sbi->sb)) + continue; + + if (f2fs_cp_error(sbi)) + continue; + + if (unlikely(freezing(current))) + continue; + + if (!sb_start_write_trylock(sbi->sb)) + continue; + + f2fs_write_meta_caches(sbi); + f2fs_write_node_caches(sbi); + + sb_end_write(sbi->sb); + } + return 0; +} + +int f2fs_start_cache_wb_thread(struct f2fs_sb_info *sbi) +{ + struct f2fs_cache_kthread *cache_thread = &sbi->cache_thread; + struct task_struct *task; + dev_t dev = sbi->sb->s_dev; + char name[36]; + + if (cache_thread->cache_wb_task) + return 0; + + init_waitqueue_head(&cache_thread->cache_wb_wq); + cache_thread->cache_wb_interval = DEF_DIRTY_CACHE_TIMEOUT; + snprintf(name, sizeof(name), "f2fs_writeback-%u:%u", + MAJOR(dev), MINOR(dev)); + + task = kthread_run(f2fs_cache_writeback_kthread, sbi, "%s", name); + if (IS_ERR(task)) + return PTR_ERR(task); + + cache_thread->cache_wb_task = task; + return 0; +} + +void f2fs_stop_cache_wb_thread(struct f2fs_sb_info *sbi) +{ + struct f2fs_cache_kthread *cache_thread = &sbi->cache_thread; + + if (!cache_thread->cache_wb_task) + return; + + kthread_stop(cache_thread->cache_wb_task); + cache_thread->cache_wb_task = NULL; +} diff --git a/fs/f2fs/cache.h b/fs/f2fs/cache.h new file mode 100644 index 000000000000..6c4db910d767 --- /dev/null +++ b/fs/f2fs/cache.h @@ -0,0 +1,240 @@ +/* SPDX-License-Identifier: GPL-2.0 */ +/* + * Copyright (c) 2026 Google LLC + * Author: Chao Yu <chaseyu@google.com> + */ +#ifndef _LINUX_F2FS_CACHE_H +#define _LINUX_F2FS_CACHE_H + +#include <linux/pagemap.h> +#include <linux/mm.h> +#include <linux/list.h> +#include <linux/radix-tree.h> +#include <linux/spinlock.h> +#include <linux/wait.h> +#include <linux/types.h> + +struct f2fs_rwsem; +struct f2fs_io_info; +enum page_type; + +/* Represents a single cached block (meta, node or compress) */ +struct f2fs_cached_block { + struct list_head list; /* LRU list head */ + struct f2fs_cached_block_list *cache; /* parent cache list */ + union { + struct f2fs_cached_block *next_entry;/* chain for merged BIO */ + nid_t ino; /* inode number for compress cache */ + }; + unsigned long index; /* key in radix tree, (meta/compress: pba, node: nid) */ + unsigned long state; /* cache entry state (e.g., Dirty, UpToDate) */ + void *data; /* blocksize-aligned memory (4KB or 16KB) */ + atomic_t refcount; /* reference count */ +}; + +struct f2fs_sb_info; + +enum f2fs_cache_type { + F2FS_META_CACHE, + F2FS_NODE_CACHE, + F2FS_COMPRESS_CACHE, + NR_CACHE_TYPES, +}; + +/* Main cache control structure (per sb_info) */ +struct f2fs_cached_block_list { + struct f2fs_sb_info *sbi; /* Pointer to f2fs_sb_info */ + struct radix_tree_root root; /* Radix tree for cache lookup */ + spinlock_t tree_lock; /* Lock for radix tree */ + struct list_head lru_list; /* Single global LRU list */ + spinlock_t list_lock; /* Lock for LRU list */ + enum f2fs_cache_type type; /* Cache type (Node, Meta, Compress) */ + unsigned long num_entries; /* Current number of entries */ +}; + +#define IS_META_CACHE(cache) ((cache)->type == F2FS_META_CACHE) +#define IS_NODE_CACHE(cache) ((cache)->type == F2FS_NODE_CACHE) +#define IS_COMPRESS_CACHE(cache) ((cache)->type == F2FS_COMPRESS_CACHE) + +/* Flags for f2fs_cached_block state */ +enum f2fs_cached_state { + F2FS_BLOCK_LOCKED, /* cache entry is locked */ + F2FS_BLOCK_UPTODATE, /* cache data is valid */ + F2FS_BLOCK_DIRTY, /* cache data is dirty, need to writeback the data */ + F2FS_BLOCK_WRITEBACK, /* cache data is writeback state */ + F2FS_BLOCK_REFERENCED, /* cache was accessed recently, shrinker will skip it for once */ + F2FS_BLOCK_INLINE_DATA, /* indicate inline data */ +}; + +enum { + __F2FS_CACHE_CREATE, /* create the cache if there is no cache entry */ + __F2FS_CACHE_LOCK, /* get and lock the cache entry */ + __F2FS_CACHE_NOFAIL, /* do not allow failure */ + __F2FS_CACHE_ACCESS, /* give a chance to add referenced tag */ +}; + +enum f2fs_cache_request_flag { + F2FS_CACHE_CREATE = 1 << __F2FS_CACHE_CREATE, + F2FS_CACHE_LOCK = 1 << __F2FS_CACHE_LOCK, + F2FS_CACHE_NOFAIL = 1 << __F2FS_CACHE_NOFAIL, + F2FS_CACHE_ACCESS = 1 << __F2FS_CACHE_ACCESS, +}; + +#define F2FS_CACHE_LOCK_CREATE (F2FS_CACHE_LOCK | F2FS_CACHE_CREATE) + +#define F2FS_ONSTACK_CACHES (32) + +#define F2FS_CACHE_FLAG_TEST_FUNC(name, flagname) \ +static inline bool f2fs_cache_test_##name( \ + const struct f2fs_cached_block *entry) \ +{ \ + return test_bit(F2FS_BLOCK_##flagname, &entry->state); \ +} \ + +#define F2FS_CACHE_FLAG_SET_FUNC(name, flagname) \ +static inline void f2fs_cache_set_##name( \ + struct f2fs_cached_block *entry) \ +{ \ + set_bit(F2FS_BLOCK_##flagname, &entry->state); \ +} \ + +#define F2FS_CACHE_FLAG_CLEAR_FUNC(name, flagname) \ +static inline void f2fs_cache_clear_##name( \ + struct f2fs_cached_block *entry) \ +{ \ + clear_bit(F2FS_BLOCK_##flagname, &entry->state); \ +} \ + +#define F2FS_CACHE_FLAG_TEST_AND_SET_FUNC(name, flagname) \ +static inline bool f2fs_cache_test_and_set_##name( \ + struct f2fs_cached_block *entry) \ +{ \ + return test_and_set_bit(F2FS_BLOCK_##flagname, &entry->state); \ +} \ + +#define F2FS_CACHE_FLAG_TEST_AND_CLEAR_FUNC(name, flagname) \ +static inline bool f2fs_cache_test_and_clear_##name( \ + struct f2fs_cached_block *entry) \ +{ \ + return test_and_clear_bit(F2FS_BLOCK_##flagname, &entry->state);\ +} \ + +F2FS_CACHE_FLAG_TEST_FUNC(locked, LOCKED); +F2FS_CACHE_FLAG_SET_FUNC(locked, LOCKED); +F2FS_CACHE_FLAG_CLEAR_FUNC(locked, LOCKED); + +F2FS_CACHE_FLAG_TEST_FUNC(uptodate, UPTODATE); +F2FS_CACHE_FLAG_SET_FUNC(uptodate, UPTODATE); +F2FS_CACHE_FLAG_CLEAR_FUNC(uptodate, UPTODATE); + +F2FS_CACHE_FLAG_TEST_FUNC(dirty, DIRTY); +F2FS_CACHE_FLAG_SET_FUNC(dirty, DIRTY); +F2FS_CACHE_FLAG_CLEAR_FUNC(dirty, DIRTY); +F2FS_CACHE_FLAG_TEST_AND_SET_FUNC(dirty, DIRTY); +F2FS_CACHE_FLAG_TEST_AND_CLEAR_FUNC(dirty, DIRTY); + +F2FS_CACHE_FLAG_TEST_FUNC(writeback, WRITEBACK); +F2FS_CACHE_FLAG_SET_FUNC(writeback, WRITEBACK); +F2FS_CACHE_FLAG_CLEAR_FUNC(writeback, WRITEBACK); + +F2FS_CACHE_FLAG_TEST_FUNC(inline, INLINE_DATA); +F2FS_CACHE_FLAG_SET_FUNC(inline, INLINE_DATA); +F2FS_CACHE_FLAG_CLEAR_FUNC(inline, INLINE_DATA); + +F2FS_CACHE_FLAG_TEST_FUNC(referenced, REFERENCED); +F2FS_CACHE_FLAG_SET_FUNC(referenced, REFERENCED); +F2FS_CACHE_FLAG_TEST_AND_CLEAR_FUNC(referenced, REFERENCED); + +static inline void *cache_address(const struct f2fs_cached_block *entry) +{ + return entry->data; +} + +#define CACHED_NODE(entry) ((struct f2fs_node *)(cache_address(entry))) + +static inline struct folio *cache_folio(const struct f2fs_cached_block *entry) +{ + return virt_to_folio(entry->data); +} + +int f2fs_init_cache(struct f2fs_sb_info *sbi, + struct f2fs_cached_block_list *cache, + enum f2fs_cache_type type); +void f2fs_destroy_cache(struct f2fs_cached_block_list *cache); +void f2fs_cache_get(struct f2fs_cached_block *entry); +struct f2fs_cached_block *f2fs_find_cache( + struct f2fs_cached_block_list *cache, + unsigned long index, + enum f2fs_cache_request_flag rflag); +#define F2FS_CACHE_TAG_NONE 0 +#define F2FS_CACHE_TAG_DIRTY 1 +#define F2FS_CACHE_TAG_WRITEBACK 2 + +bool f2fs_trylock_cache(struct f2fs_cached_block *entry); +void f2fs_lock_cache(struct f2fs_cached_block *entry); +void f2fs_unlock_cache(struct f2fs_cached_block *entry); +bool f2fs_put_cache(struct f2fs_cached_block *entry, bool unlock); +bool f2fs_mark_cache_dirty(struct f2fs_cached_block *entry); +void f2fs_drop_cache_dirty(struct f2fs_cached_block *entry); +void f2fs_start_cache_writeback(struct f2fs_cached_block *entry); +void f2fs_end_cache_writeback(struct f2fs_cached_block *entry); +unsigned int f2fs_cache_gang_lookup(struct f2fs_cached_block_list *cache, + struct f2fs_cached_block **entries, + pgoff_t *index, unsigned long end); +unsigned int f2fs_cache_gang_lookup_tag(struct f2fs_cached_block_list *cache, + struct f2fs_cached_block **results, pgoff_t *first_index, + unsigned int max_items, int tag); +void f2fs_cache_gang_release(struct f2fs_cached_block **entries, + unsigned int nr_entries); +void f2fs_cache_wait_on_all_writeback(struct f2fs_cached_block_list *cache); +void f2fs_cache_wait_writeback_cond(struct f2fs_cached_block *entry, + enum page_type type); +void f2fs_cache_wait_writeback(struct f2fs_cached_block *entry); +void f2fs_cache_update_tag(struct f2fs_cached_block *entry, + unsigned int clear_from, unsigned int set_to); +struct f2fs_cached_block *f2fs_grab_cache(struct f2fs_cached_block_list *cache, + unsigned long index, int flags); +void f2fs_truncate_locked_cache(struct f2fs_cached_block *entry, + bool drop_dirty); +void f2fs_truncate_cache(struct f2fs_cached_block *entry, bool drop_dirty); +void f2fs_drop_cache_range(struct f2fs_cached_block_list *cache, + unsigned long start, unsigned long len, bool drop_dirty); + +#define META_CACHE(sbi) (&(sbi)->meta_blocks) +#define NODE_CACHE(sbi) (&(sbi)->node_blocks) +#define COMPRESS_CACHE(sbi) (&(sbi)->compress_blocks) + +#define f2fs_find_meta_cache(sbi, blkaddr) \ + f2fs_find_cache(META_CACHE(sbi), blkaddr, 0) +#define f2fs_invalidate_meta_caches(sbi, start, len) \ + f2fs_drop_cache_range(META_CACHE(sbi), start, len, false) +#define f2fs_truncate_meta_caches(sbi, start, len) \ + f2fs_drop_cache_range(META_CACHE(sbi), start, len, true) + +#define f2fs_grab_node_cache(sbi, blkaddr) \ + f2fs_grab_cache(NODE_CACHE(sbi), blkaddr, \ + F2FS_CACHE_LOCK_CREATE) +#define f2fs_find_node_cache(sbi, blkaddr) \ + f2fs_find_cache(NODE_CACHE(sbi), blkaddr, F2FS_CACHE_ACCESS) +#define f2fs_invalidate_node_cache(sbi, blkaddr) \ + f2fs_drop_cache_range(NODE_CACHE(sbi), blkaddr, 1, false) +#define f2fs_truncate_node_caches(sbi, start, len) \ + f2fs_drop_cache_range(NODE_CACHE(sbi), start, len, true) + +unsigned long f2fs_shrink_cache(struct f2fs_sb_info *sbi, + unsigned long nr_to_scan); + +#define DEF_DIRTY_CACHE_TIMEOUT 5000 +#define MIN_DIRTY_CACHE_TIMEOUT 100 +#define MAX_DIRTY_CACHE_TIMEOUT 30000 + +struct f2fs_cache_kthread { + struct task_struct *cache_wb_task; + wait_queue_head_t cache_wb_wq; + unsigned int cache_wb_interval; +}; + +int f2fs_start_cache_wb_thread(struct f2fs_sb_info *sbi); +void f2fs_stop_cache_wb_thread(struct f2fs_sb_info *sbi); + +#endif /* _LINUX_F2FS_CACHE_H */ diff --git a/fs/f2fs/checkpoint.c b/fs/f2fs/checkpoint.c index 4b59f30ef45d..7b89bb838946 100644 --- a/fs/f2fs/checkpoint.c +++ b/fs/f2fs/checkpoint.c @@ -7,7 +7,6 @@ */ #include <linux/fs.h> #include <linux/bio.h> -#include <linux/mpage.h> #include <linux/writeback.h> #include <linux/blkdev.h> #include <linux/f2fs_fs.h> @@ -17,6 +16,7 @@ #include <linux/delayacct.h> #include <linux/ioprio.h> #include <linux/math64.h> +#include <linux/freezer.h> #include "f2fs.h" #include "node.h" @@ -234,29 +234,29 @@ static struct kmem_cache *ino_entry_slab; struct kmem_cache *f2fs_inode_entry_slab; /* - * We guarantee no failure on the returned page. + * We guarantee no failure on the returned cache. */ -struct folio *f2fs_grab_meta_folio(struct f2fs_sb_info *sbi, pgoff_t index) +struct f2fs_cached_block *f2fs_grab_meta_cache(struct f2fs_sb_info *sbi, + pgoff_t index) { - struct address_space *mapping = META_MAPPING(sbi); - struct folio *folio; + struct f2fs_cached_block *entry; repeat: - folio = f2fs_grab_cache_folio(mapping, index, false); - if (IS_ERR(folio)) { + entry = f2fs_grab_cache(META_CACHE(sbi), index, + F2FS_CACHE_LOCK_CREATE); + if (IS_ERR(entry)) { cond_resched(); goto repeat; } - f2fs_folio_wait_writeback(folio, META, true, true); - if (!folio_test_uptodate(folio)) - folio_mark_uptodate(folio); - return folio; + f2fs_cache_wait_writeback(entry); + if (!f2fs_cache_test_uptodate(entry)) + f2fs_cache_set_uptodate(entry); + return entry; } -static struct folio *__get_meta_folio(struct f2fs_sb_info *sbi, pgoff_t index, - bool is_meta) +static struct f2fs_cached_block *__get_meta_cache(struct f2fs_sb_info *sbi, + pgoff_t index, bool is_meta) { - struct address_space *mapping = META_MAPPING(sbi); - struct folio *folio; + struct f2fs_cached_block *entry; struct f2fs_io_info fio = { .sbi = sbi, .type = META, @@ -266,70 +266,75 @@ static struct folio *__get_meta_folio(struct f2fs_sb_info *sbi, pgoff_t index, .new_blkaddr = index, .encrypted_page = NULL, .is_por = !is_meta ? 1 : 0, + .is_cache = 1, }; int err; if (unlikely(!is_meta)) fio.op_flags &= ~REQ_META; repeat: - folio = f2fs_grab_cache_folio(mapping, index, false); - if (IS_ERR(folio)) { + entry = f2fs_grab_cache(META_CACHE(sbi), index, + F2FS_CACHE_LOCK_CREATE); + if (IS_ERR(entry)) { cond_resched(); goto repeat; } - if (folio_test_uptodate(folio)) + if (f2fs_cache_test_uptodate(entry)) goto out; - fio.folio = folio; + fio.cache_entry = entry; - err = f2fs_submit_page_bio(&fio); + err = f2fs_submit_cache_read(&fio); if (err) { - f2fs_folio_put(folio, true); + f2fs_put_cache(entry, true); return ERR_PTR(err); } - f2fs_update_iostat(sbi, NULL, FS_META_READ_IO, F2FS_BLKSIZE); + f2fs_update_iostat(sbi, NULL, FS_META_READ_IO, F2FS_BLKSIZE(sbi)); - folio_lock(folio); - if (unlikely(!is_meta_folio(folio))) { - f2fs_folio_put(folio, true); + f2fs_lock_cache(entry); + if (unlikely(!f2fs_is_meta_cache(entry))) { + f2fs_put_cache(entry, true); goto repeat; } - if (unlikely(!folio_test_uptodate(folio))) { - f2fs_handle_page_eio(sbi, folio, META); - f2fs_folio_put(folio, true); + if (unlikely(!f2fs_cache_test_uptodate(entry))) { + f2fs_handle_page_eio(sbi, entry->index, META); + f2fs_put_cache(entry, true); return ERR_PTR(-EIO); } out: - return folio; + return entry; } -struct folio *f2fs_get_meta_folio(struct f2fs_sb_info *sbi, pgoff_t index) +struct f2fs_cached_block *f2fs_get_meta_cache(struct f2fs_sb_info *sbi, + pgoff_t index) { - return __get_meta_folio(sbi, index, true); + return __get_meta_cache(sbi, index, true); } -struct folio *f2fs_get_meta_folio_retry(struct f2fs_sb_info *sbi, pgoff_t index) +struct f2fs_cached_block *f2fs_get_meta_cache_retry(struct f2fs_sb_info *sbi, + pgoff_t index) { - struct folio *folio; + struct f2fs_cached_block *entry; int count = 0; retry: - folio = __get_meta_folio(sbi, index, true); - if (IS_ERR(folio)) { - if (PTR_ERR(folio) == -EIO && + entry = __get_meta_cache(sbi, index, true); + if (IS_ERR(entry)) { + if (PTR_ERR(entry) == -EIO && ++count <= DEFAULT_RETRY_IO_COUNT) goto retry; f2fs_stop_checkpoint(sbi, false, STOP_CP_REASON_META_PAGE); } - return folio; + return entry; } /* for POR only */ -struct folio *f2fs_get_tmp_folio(struct f2fs_sb_info *sbi, pgoff_t index) +struct f2fs_cached_block *f2fs_get_tmp_cache(struct f2fs_sb_info *sbi, + pgoff_t index) { - return __get_meta_folio(sbi, index, false); + return __get_meta_cache(sbi, index, false); } static bool __is_bitmap_valid(struct f2fs_sb_info *sbi, block_t blkaddr, @@ -445,11 +450,12 @@ bool f2fs_is_valid_blkaddr_raw(struct f2fs_sb_info *sbi, } /* - * Readahead CP/NAT/SIT/SSA/POR pages + * Readahead CP/NAT/SIT/SSA/POR blocks */ -int f2fs_ra_meta_pages(struct f2fs_sb_info *sbi, block_t start, int nrpages, - int type, bool sync) +int f2fs_ra_meta_caches(struct f2fs_sb_info *sbi, block_t start, + int nrblocks, int type, bool sync) { + struct f2fs_cached_block_list *cache = META_CACHE(sbi); block_t blkno = start; struct f2fs_io_info fio = { .sbi = sbi, @@ -459,6 +465,7 @@ int f2fs_ra_meta_pages(struct f2fs_sb_info *sbi, block_t start, int nrpages, .encrypted_page = NULL, .in_list = 0, .is_por = (type == META_POR) ? 1 : 0, + .is_cache = 1, }; struct blk_plug plug; int err; @@ -467,8 +474,8 @@ int f2fs_ra_meta_pages(struct f2fs_sb_info *sbi, block_t start, int nrpages, fio.op_flags &= ~REQ_META; blk_start_plug(&plug); - for (; nrpages-- > 0; blkno++) { - struct folio *folio; + for (; nrblocks-- > 0; blkno++) { + struct f2fs_cached_block *entry; if (!f2fs_is_valid_blkaddr(sbi, blkno, type)) goto out; @@ -476,18 +483,18 @@ int f2fs_ra_meta_pages(struct f2fs_sb_info *sbi, block_t start, int nrpages, switch (type) { case META_NAT: if (unlikely(blkno >= - NAT_BLOCK_OFFSET(NM_I(sbi)->max_nid))) + NAT_BLOCK_OFFSET(sbi, NM_I(sbi)->max_nid))) blkno = 0; /* get nat block addr */ fio.new_blkaddr = current_nat_addr(sbi, - blkno * NAT_ENTRY_PER_BLOCK); + blkno * NAT_ENTRY_PER_BLOCK(sbi)); break; case META_SIT: if (unlikely(blkno >= TOTAL_SEGS(sbi))) goto out; /* get sit block addr */ fio.new_blkaddr = current_sit_addr(sbi, - blkno * SIT_ENTRY_PER_BLOCK); + blkno * SIT_ENTRY_PER_BLOCK(sbi)); break; case META_SSA: case META_CP: @@ -495,62 +502,63 @@ int f2fs_ra_meta_pages(struct f2fs_sb_info *sbi, block_t start, int nrpages, fio.new_blkaddr = blkno; break; default: - BUG(); + f2fs_bug_on(sbi, 1); } - folio = f2fs_grab_cache_folio(META_MAPPING(sbi), - fio.new_blkaddr, false); - if (IS_ERR(folio)) + entry = f2fs_grab_cache(cache, fio.new_blkaddr, + F2FS_CACHE_LOCK_CREATE); + if (IS_ERR(entry)) continue; - if (folio_test_uptodate(folio)) { - f2fs_folio_put(folio, true); + if (f2fs_cache_test_uptodate(entry)) { + f2fs_put_cache(entry, true); continue; } - fio.folio = folio; - err = f2fs_submit_page_bio(&fio); - f2fs_folio_put(folio, err ? true : false); + fio.cache_entry = entry; + err = f2fs_submit_cache_read(&fio); + f2fs_put_cache(entry, err ? true : false); if (!err) f2fs_update_iostat(sbi, NULL, FS_META_READ_IO, - F2FS_BLKSIZE); + F2FS_BLKSIZE(sbi)); } out: blk_finish_plug(&plug); return blkno - start; } -void f2fs_ra_meta_pages_cond(struct f2fs_sb_info *sbi, pgoff_t index, - unsigned int ra_blocks) +void f2fs_ra_meta_caches_cond(struct f2fs_sb_info *sbi, pgoff_t index, + unsigned int ra_blocks) { - struct folio *folio; + struct f2fs_cached_block *entry; bool readahead = false; if (ra_blocks == RECOVERY_MIN_RA_BLOCKS) return; - folio = filemap_get_folio(META_MAPPING(sbi), index); - if (IS_ERR(folio) || !folio_test_uptodate(folio)) + entry = f2fs_find_cache(META_CACHE(sbi), index, 0); + if (IS_ERR(entry) || !f2fs_cache_test_uptodate(entry)) readahead = true; - f2fs_folio_put(folio, false); + f2fs_put_cache(entry, false); if (readahead) - f2fs_ra_meta_pages(sbi, index, ra_blocks, META_POR, true); + f2fs_ra_meta_caches(sbi, index, ra_blocks, META_POR, true); } -static bool __f2fs_write_meta_folio(struct folio *folio, - struct writeback_control *wbc, +static bool __f2fs_write_meta_cache(struct f2fs_cached_block *entry, enum iostat_type io_type) { - struct f2fs_sb_info *sbi = F2FS_F_SB(folio); + struct f2fs_sb_info *sbi = entry->cache->sbi; - trace_f2fs_writepage(folio, META); + trace_f2fs_write_cache(entry, META); if (unlikely(f2fs_cp_error(sbi))) { if (is_sbi_flag_set(sbi, SBI_IS_CLOSE)) { - folio_clear_uptodate(folio); - dec_page_count(sbi, F2FS_DIRTY_META); - folio_unlock(folio); + f2fs_cache_clear_uptodate(entry); + dec_cache_count(sbi, F2FS_DIRTY_META); + f2fs_cache_update_tag(entry, F2FS_CACHE_TAG_DIRTY, + F2FS_CACHE_TAG_NONE); + f2fs_unlock_cache(entry); return true; } goto redirty_out; @@ -558,10 +566,10 @@ static bool __f2fs_write_meta_folio(struct folio *folio, if (unlikely(is_sbi_flag_set(sbi, SBI_POR_DOING))) goto redirty_out; - f2fs_do_write_meta_page(sbi, folio, io_type); - dec_page_count(sbi, F2FS_DIRTY_META); + f2fs_do_write_meta_cache(sbi, entry, io_type); + dec_cache_count(sbi, F2FS_DIRTY_META); - folio_unlock(folio); + f2fs_unlock_cache(entry); if (unlikely(f2fs_cp_error(sbi))) f2fs_submit_merged_write(sbi, META); @@ -569,101 +577,91 @@ static bool __f2fs_write_meta_folio(struct folio *folio, return true; redirty_out: - folio_redirty_for_writepage(wbc, folio); + f2fs_cache_set_dirty(entry); return false; } -static int f2fs_write_meta_pages(struct address_space *mapping, - struct writeback_control *wbc) +void f2fs_write_meta_caches(struct f2fs_sb_info *sbi) { - struct f2fs_sb_info *sbi = F2FS_M_SB(mapping); struct f2fs_lock_context lc; - long diff, written; + long nr_to_write = LONG_MAX; if (unlikely(is_sbi_flag_set(sbi, SBI_POR_DOING))) - goto skip_write; + return; - /* collect a number of dirty meta pages and write together */ - if (wbc->sync_mode != WB_SYNC_ALL && - get_pages(sbi, F2FS_DIRTY_META) < - nr_pages_to_skip(sbi, META)) - goto skip_write; + /* collect a number of dirty meta caches and write together */ + if (get_nr_caches(sbi, F2FS_DIRTY_META) < + nr_caches_to_skip(sbi, META)) + return; - /* if locked failed, cp will flush dirty pages instead */ + /* if locked failed, cp will flush dirty caches instead */ if (!f2fs_down_write_trylock_trace(&sbi->cp_global_sem, &lc)) - goto skip_write; + return; - trace_f2fs_writepages(mapping->host, wbc, META); - diff = nr_pages_to_write(sbi, META, wbc); - written = f2fs_sync_meta_pages(sbi, wbc->nr_to_write, FS_META_IO); + nr_to_write = adjust_flush_cache_number(sbi, META); + f2fs_sync_meta_caches(sbi, nr_to_write, false, FS_META_IO); f2fs_up_write_trace(&sbi->cp_global_sem, &lc); - wbc->nr_to_write = max((long)0, wbc->nr_to_write - written - diff); - return 0; - -skip_write: - wbc->pages_skipped += get_pages(sbi, F2FS_DIRTY_META); - trace_f2fs_writepages(mapping->host, wbc, META); - return 0; } -long f2fs_sync_meta_pages(struct f2fs_sb_info *sbi, long nr_to_write, - enum iostat_type io_type) +long f2fs_sync_meta_caches(struct f2fs_sb_info *sbi, long nr_to_write, + bool sync, enum iostat_type io_type) { - struct address_space *mapping = META_MAPPING(sbi); pgoff_t index = 0, prev = ULONG_MAX; - struct folio_batch fbatch; + struct f2fs_cached_block *entries[F2FS_ONSTACK_CACHES]; long nwritten = 0; - int nr_folios; - struct writeback_control wbc = {}; + int nr; struct blk_plug plug; + bool background = nr_to_write != LONG_MAX; - folio_batch_init(&fbatch); + trace_f2fs_write_caches(sbi, nr_to_write, 0, META); blk_start_plug(&plug); - while ((nr_folios = filemap_get_folios_tag(mapping, &index, - (pgoff_t)-1, - PAGECACHE_TAG_DIRTY, &fbatch))) { + while ((nr = f2fs_cache_gang_lookup_tag(META_CACHE(sbi), entries, + &index, F2FS_ONSTACK_CACHES, F2FS_CACHE_TAG_DIRTY))) { int i; - for (i = 0; i < nr_folios; i++) { - struct folio *folio = fbatch.folios[i]; + for (i = 0; i < nr; i++) { + struct f2fs_cached_block *entry = entries[i]; - if (nr_to_write != LONG_MAX && i != 0 && - folio->index != prev + - folio_nr_pages(fbatch.folios[i-1])) { - folio_batch_release(&fbatch); + if (background && unlikely(freezing(current))) { + f2fs_cache_gang_release(entries, nr); goto stop; } - folio_lock(folio); + if (background && i != 0 && + entry->index != prev + 1) { + f2fs_cache_gang_release(entries, nr); + goto stop; + } - if (unlikely(!is_meta_folio(folio))) { + f2fs_lock_cache(entry); + + if (unlikely(!f2fs_is_meta_cache(entry))) { continue_unlock: - folio_unlock(folio); + f2fs_unlock_cache(entry); continue; } - if (!folio_test_dirty(folio)) { + if (!f2fs_cache_test_dirty(entry)) { /* someone wrote it for us */ goto continue_unlock; } - f2fs_folio_wait_writeback(folio, META, true, true); + f2fs_cache_wait_writeback(entry); - if (!folio_clear_dirty_for_io(folio)) + if (!f2fs_cache_test_and_clear_dirty(entry)) goto continue_unlock; - if (!__f2fs_write_meta_folio(folio, &wbc, - io_type)) { - folio_unlock(folio); + if (!__f2fs_write_meta_cache(entry, io_type)) { + f2fs_unlock_cache(entry); break; } - nwritten += folio_nr_pages(folio); - prev = folio->index; + nwritten++; + prev = entry->index; if (unlikely(nwritten >= nr_to_write)) break; } - folio_batch_release(&fbatch); + f2fs_cache_gang_release(entries, nr); cond_resched(); } stop: @@ -672,32 +670,11 @@ stop: blk_finish_plug(&plug); - return nwritten; -} - -static bool f2fs_dirty_meta_folio(struct address_space *mapping, - struct folio *folio) -{ - trace_f2fs_set_page_dirty(folio, META); + trace_f2fs_write_caches(sbi, nr_to_write, nwritten, META); - if (!folio_test_uptodate(folio)) - folio_mark_uptodate(folio); - if (filemap_dirty_folio(mapping, folio)) { - inc_page_count(F2FS_M_SB(mapping), F2FS_DIRTY_META); - folio_set_f2fs_reference(folio); - return true; - } - return false; + return nwritten; } -const struct address_space_operations f2fs_meta_aops = { - .writepages = f2fs_write_meta_pages, - .dirty_folio = f2fs_dirty_meta_folio, - .invalidate_folio = f2fs_invalidate_folio, - .release_folio = f2fs_release_folio, - .migrate_folio = filemap_migrate_folio, -}; - static void __add_ino_entry(struct f2fs_sb_info *sbi, nid_t ino, unsigned int devidx, int type) { @@ -825,15 +802,6 @@ static void __clear_ino_bitmap(struct f2fs_sb_info *sbi, nid_t ino, int type) spin_unlock(&im->ino_lock); } -static void f2fs_wait_for_inode_record(struct f2fs_sb_info *sbi, int mode) -{ - if (mode != APPEND_INO && mode != UPDATE_INO) - return; - - /* Let's wait for some pending updates for APPEND_INO and UPDATE_INO. */ - flush_workqueue(sbi->evict_wq); -} - static void __f2fs_add_ino_entry(struct f2fs_sb_info *sbi, nid_t ino, unsigned int devidx, int type) { @@ -887,8 +855,6 @@ void f2fs_release_ino_entry(struct f2fs_sb_info *sbi, bool all) for (i = all ? ORPHAN_INO : FLUSH_INO; i <= FLUSH_INO; i++) { struct inode_management *im = &sbi->im[i]; - f2fs_wait_for_inode_record(sbi, i); - spin_lock(&im->ino_lock); list_for_each_entry_safe(e, tmp, &im->ino_list, list) { list_del(&e->list); @@ -899,6 +865,9 @@ void f2fs_release_ino_entry(struct f2fs_sb_info *sbi, bool all) spin_unlock(&im->ino_lock); } + /* Wait for pending APPEND/UPDATE inode state updates. */ + flush_workqueue(sbi->evict_wq); + for (i = APPEND_INO; i < MAX_INO_ENTRY; i++) { struct inode_management *im = &sbi->im[i]; @@ -964,7 +933,7 @@ void f2fs_add_orphan_inode(struct inode *inode) { /* add new orphan ino entry into list */ f2fs_add_ino_entry(F2FS_I_SB(inode), inode->i_ino, ORPHAN_INO); - f2fs_update_inode_page(inode); + f2fs_update_inode_cache(inode); } void f2fs_remove_orphan_inode(struct f2fs_sb_info *sbi, nid_t ino) @@ -1037,28 +1006,30 @@ int f2fs_recover_orphan_inodes(struct f2fs_sb_info *sbi) start_blk = __start_cp_addr(sbi) + 1 + __cp_payload(sbi); orphan_blocks = __start_sum_addr(sbi) - 1 - __cp_payload(sbi); - f2fs_ra_meta_pages(sbi, start_blk, orphan_blocks, META_CP, true); + f2fs_ra_meta_caches(sbi, start_blk, orphan_blocks, META_CP, true); for (i = 0; i < orphan_blocks; i++) { - struct folio *folio; + struct f2fs_cached_block *entry; struct f2fs_orphan_block *orphan_blk; + struct f2fs_orphan_footer *footer; unsigned int entry_count; - folio = f2fs_get_meta_folio(sbi, start_blk + i); - if (IS_ERR(folio)) { - err = PTR_ERR(folio); + entry = f2fs_get_meta_cache(sbi, start_blk + i); + if (IS_ERR(entry)) { + err = PTR_ERR(entry); goto out; } - orphan_blk = folio_address(folio); - entry_count = le32_to_cpu(orphan_blk->entry_count); - if (entry_count > F2FS_ORPHANS_PER_BLOCK) { + orphan_blk = cache_address(entry); + footer = f2fs_orphan_footer(orphan_blk, sbi); + entry_count = le32_to_cpu(footer->entry_count); + if (entry_count > F2FS_ORPHANS_PER_BLOCK(sbi)) { f2fs_err(sbi, "invalid orphan inode entry count %u", entry_count); set_sbi_flag(sbi, SBI_NEED_FSCK); f2fs_handle_error(sbi, ERROR_INCONSISTENT_ORPHAN); err = -EFSCORRUPTED; - f2fs_folio_put(folio, true); + f2fs_put_cache(entry, true); goto out; } @@ -1067,11 +1038,11 @@ int f2fs_recover_orphan_inodes(struct f2fs_sb_info *sbi) err = recover_orphan_inode(sbi, ino); if (err) { - f2fs_folio_put(folio, true); + f2fs_put_cache(entry, true); goto out; } } - f2fs_folio_put(folio, true); + f2fs_put_cache(entry, true); } /* clear Orphan Flag */ clear_ckpt_flags(sbi, CP_ORPHAN_PRESENT_FLAG); @@ -1085,96 +1056,99 @@ static void write_orphan_inodes(struct f2fs_sb_info *sbi, block_t start_blk) { struct list_head *head; struct f2fs_orphan_block *orphan_blk = NULL; + struct f2fs_orphan_footer *footer = NULL; unsigned int nentries = 0; unsigned short index = 1; unsigned short orphan_blocks; - struct folio *folio = NULL; struct ino_entry *orphan = NULL; struct inode_management *im = &sbi->im[ORPHAN_INO]; + struct f2fs_cached_block *entry = NULL; - orphan_blocks = GET_ORPHAN_BLOCKS(im->ino_num); + orphan_blocks = GET_ORPHAN_BLOCKS(sbi, im->ino_num); /* * we don't need to do spin_lock(&im->ino_lock) here, since all the * orphan inode operations are covered under f2fs_lock_op(). - * And, spin_lock should be avoided due to page operations below. + * And, spin_lock should be avoided due to cache operations below. */ head = &im->ino_list; /* loop for each orphan inode entry and write them in journal block */ list_for_each_entry(orphan, head, list) { - if (!folio) { - folio = f2fs_grab_meta_folio(sbi, start_blk++); - orphan_blk = folio_address(folio); - memset(orphan_blk, 0, sizeof(*orphan_blk)); + if (!entry) { + entry = f2fs_grab_meta_cache(sbi, start_blk++); + orphan_blk = cache_address(entry); + footer = f2fs_orphan_footer(orphan_blk, sbi); + memset(orphan_blk, 0, sbi->blocksize); } orphan_blk->ino[nentries++] = cpu_to_le32(orphan->ino); - if (nentries == F2FS_ORPHANS_PER_BLOCK) { + if (nentries == F2FS_ORPHANS_PER_BLOCK(sbi)) { /* - * an orphan block is full of 1020 entries, + * an orphan block is full, * then we need to flush current orphan blocks * and bring another one in memory */ - orphan_blk->blk_addr = cpu_to_le16(index); - orphan_blk->blk_count = cpu_to_le16(orphan_blocks); - orphan_blk->entry_count = cpu_to_le32(nentries); - folio_mark_dirty(folio); - f2fs_folio_put(folio, true); + footer->blk_addr = cpu_to_le16(index); + footer->blk_count = cpu_to_le16(orphan_blocks); + footer->entry_count = cpu_to_le32(nentries); + f2fs_mark_cache_dirty(entry); + f2fs_put_cache(entry, true); index++; nentries = 0; - folio = NULL; + entry = NULL; } } - if (folio) { - orphan_blk->blk_addr = cpu_to_le16(index); - orphan_blk->blk_count = cpu_to_le16(orphan_blocks); - orphan_blk->entry_count = cpu_to_le32(nentries); - folio_mark_dirty(folio); - f2fs_folio_put(folio, true); + if (entry) { + footer->blk_addr = cpu_to_le16(index); + footer->blk_count = cpu_to_le16(orphan_blocks); + footer->entry_count = cpu_to_le32(nentries); + f2fs_mark_cache_dirty(entry); + f2fs_put_cache(entry, true); } } -static __u32 f2fs_checkpoint_chksum(struct f2fs_checkpoint *ckpt) +static __u32 f2fs_checkpoint_chksum(struct f2fs_sb_info *sbi, + struct f2fs_checkpoint *ckpt) { unsigned int chksum_ofs = le32_to_cpu(ckpt->checksum_offset); __u32 chksum; chksum = f2fs_crc32(ckpt, chksum_ofs); - if (chksum_ofs < CP_CHKSUM_OFFSET) { + if (chksum_ofs < CP_CHKSUM_OFFSET(sbi)) { chksum_ofs += sizeof(chksum); chksum = f2fs_chksum(chksum, (__u8 *)ckpt + chksum_ofs, - F2FS_BLKSIZE - chksum_ofs); + F2FS_BLKSIZE(sbi) - chksum_ofs); } return chksum; } static int get_checkpoint_version(struct f2fs_sb_info *sbi, block_t cp_addr, - struct f2fs_checkpoint **cp_block, struct folio **cp_folio, + struct f2fs_checkpoint **cp_block, struct f2fs_cached_block **cp_entry, unsigned long long *version) { size_t crc_offset = 0; __u32 crc; - *cp_folio = f2fs_get_meta_folio(sbi, cp_addr); - if (IS_ERR(*cp_folio)) - return PTR_ERR(*cp_folio); + *cp_entry = f2fs_get_meta_cache(sbi, cp_addr); + if (IS_ERR(*cp_entry)) + return PTR_ERR(*cp_entry); - *cp_block = folio_address(*cp_folio); + *cp_block = cache_address(*cp_entry); crc_offset = le32_to_cpu((*cp_block)->checksum_offset); if (crc_offset < CP_MIN_CHKSUM_OFFSET || - crc_offset > CP_CHKSUM_OFFSET) { - f2fs_folio_put(*cp_folio, true); + crc_offset > CP_CHKSUM_OFFSET(sbi)) { + f2fs_put_cache(*cp_entry, true); f2fs_warn(sbi, "invalid crc_offset: %zu", crc_offset); return -EINVAL; } - crc = f2fs_checkpoint_chksum(*cp_block); + crc = f2fs_checkpoint_chksum(sbi, *cp_block); if (crc != cur_cp_crc(*cp_block)) { - f2fs_folio_put(*cp_folio, true); + f2fs_put_cache(*cp_entry, true); f2fs_warn(sbi, "invalid crc value"); return -EINVAL; } @@ -1183,17 +1157,17 @@ static int get_checkpoint_version(struct f2fs_sb_info *sbi, block_t cp_addr, return 0; } -static struct folio *validate_checkpoint(struct f2fs_sb_info *sbi, +static struct f2fs_cached_block *validate_checkpoint(struct f2fs_sb_info *sbi, block_t cp_addr, unsigned long long *version) { - struct folio *cp_folio_1 = NULL, *cp_folio_2 = NULL; + struct f2fs_cached_block *cp_entry_1 = NULL, *cp_entry_2 = NULL; struct f2fs_checkpoint *cp_block = NULL; unsigned long long cur_version = 0, pre_version = 0; unsigned int cp_blocks; int err; err = get_checkpoint_version(sbi, cp_addr, &cp_block, - &cp_folio_1, version); + &cp_entry_1, version); if (err) return NULL; @@ -1208,19 +1182,19 @@ static struct folio *validate_checkpoint(struct f2fs_sb_info *sbi, cp_addr += cp_blocks - 1; err = get_checkpoint_version(sbi, cp_addr, &cp_block, - &cp_folio_2, version); + &cp_entry_2, version); if (err) goto invalid_cp; cur_version = *version; if (cur_version == pre_version) { *version = cur_version; - f2fs_folio_put(cp_folio_2, true); - return cp_folio_1; + f2fs_put_cache(cp_entry_2, true); + return cp_entry_1; } - f2fs_folio_put(cp_folio_2, true); + f2fs_put_cache(cp_entry_2, true); invalid_cp: - f2fs_folio_put(cp_folio_1, true); + f2fs_put_cache(cp_entry_1, true); return NULL; } @@ -1228,7 +1202,7 @@ int f2fs_get_valid_checkpoint(struct f2fs_sb_info *sbi) { struct f2fs_checkpoint *cp_block; struct f2fs_super_block *fsb = sbi->raw_super; - struct folio *cp1, *cp2, *cur_folio; + struct f2fs_cached_block *cp1 = NULL, *cp2 = NULL, *cur_entry; unsigned long blk_size = sbi->blocksize; unsigned long long cp1_version = 0, cp2_version = 0; unsigned long long cp_start_blk_no; @@ -1255,22 +1229,22 @@ int f2fs_get_valid_checkpoint(struct f2fs_sb_info *sbi) if (cp1 && cp2) { if (ver_after(cp2_version, cp1_version)) - cur_folio = cp2; + cur_entry = cp2; else - cur_folio = cp1; + cur_entry = cp1; } else if (cp1) { - cur_folio = cp1; + cur_entry = cp1; } else if (cp2) { - cur_folio = cp2; + cur_entry = cp2; } else { err = -EFSCORRUPTED; goto fail_no_cp; } - cp_block = folio_address(cur_folio); + cp_block = cache_address(cur_entry); memcpy(sbi->ckpt, cp_block, blk_size); - if (cur_folio == cp1) + if (cur_entry == cp1) sbi->cur_cp_pack = 1; else sbi->cur_cp_pack = 2; @@ -1285,30 +1259,31 @@ int f2fs_get_valid_checkpoint(struct f2fs_sb_info *sbi) goto done; cp_blk_no = le32_to_cpu(fsb->cp_blkaddr); - if (cur_folio == cp2) + if (cur_entry == cp2) cp_blk_no += BIT(le32_to_cpu(fsb->log_blocks_per_seg)); for (i = 1; i < cp_blks; i++) { + struct f2fs_cached_block *cur_entry_payload; void *sit_bitmap_ptr; unsigned char *ckpt = (unsigned char *)sbi->ckpt; - cur_folio = f2fs_get_meta_folio(sbi, cp_blk_no + i); - if (IS_ERR(cur_folio)) { - err = PTR_ERR(cur_folio); + cur_entry_payload = f2fs_get_meta_cache(sbi, cp_blk_no + i); + if (IS_ERR(cur_entry_payload)) { + err = PTR_ERR(cur_entry_payload); goto free_fail_no_cp; } - sit_bitmap_ptr = folio_address(cur_folio); + sit_bitmap_ptr = cache_address(cur_entry_payload); memcpy(ckpt + i * blk_size, sit_bitmap_ptr, blk_size); - f2fs_folio_put(cur_folio, true); + f2fs_put_cache(cur_entry_payload, true); } done: - f2fs_folio_put(cp1, true); - f2fs_folio_put(cp2, true); + f2fs_put_cache(cp1, true); + f2fs_put_cache(cp2, true); return 0; free_fail_no_cp: - f2fs_folio_put(cp1, true); - f2fs_folio_put(cp2, true); + f2fs_put_cache(cp1, true); + f2fs_put_cache(cp2, true); fail_no_cp: kvfree(sbi->ckpt); return err; @@ -1384,12 +1359,12 @@ int f2fs_sync_dirty_inodes(struct f2fs_sb_info *sbi, enum inode_type type, unsigned long ino = 0; trace_f2fs_sync_dirty_inodes_enter(sbi->sb, is_dir, - get_pages(sbi, is_dir ? + get_nr_caches(sbi, is_dir ? F2FS_DIRTY_DENTS : F2FS_DIRTY_DATA)); retry: if (unlikely(f2fs_cp_error(sbi))) { trace_f2fs_sync_dirty_inodes_exit(sbi->sb, is_dir, - get_pages(sbi, is_dir ? + get_nr_caches(sbi, is_dir ? F2FS_DIRTY_DENTS : F2FS_DIRTY_DATA)); return -EIO; } @@ -1400,11 +1375,12 @@ retry: if (list_empty(head)) { spin_unlock(&sbi->inode_lock[type]); trace_f2fs_sync_dirty_inodes_exit(sbi->sb, is_dir, - get_pages(sbi, is_dir ? + get_nr_caches(sbi, is_dir ? F2FS_DIRTY_DENTS : F2FS_DIRTY_DATA)); return 0; } fi = list_first_entry(head, struct f2fs_inode_info, dirty_list); + list_move_tail(&fi->dirty_list, head); inode = igrab(&fi->vfs_inode); spin_unlock(&sbi->inode_lock[type]); if (inode) { @@ -1427,11 +1403,6 @@ retry: else ino = cur_ino; } else { - /* - * We should submit bio, since it exists several - * writebacking dentry pages in the freeing inode. - */ - f2fs_submit_merged_write(sbi, DATA); cond_resched(); } goto retry; @@ -1442,7 +1413,7 @@ static int f2fs_sync_inode_meta(struct f2fs_sb_info *sbi) struct list_head *head = &sbi->inode_list[DIRTY_META]; struct inode *inode; struct f2fs_inode_info *fi; - s64 total = get_pages(sbi, F2FS_DIRTY_IMETA); + s64 total = get_nr_caches(sbi, F2FS_DIRTY_IMETA); while (total--) { if (unlikely(f2fs_cp_error(sbi))) @@ -1455,6 +1426,7 @@ static int f2fs_sync_inode_meta(struct f2fs_sb_info *sbi) } fi = list_first_entry(head, struct f2fs_inode_info, gdirty_list); + list_move_tail(&fi->gdirty_list, head); inode = igrab(&fi->vfs_inode); spin_unlock(&sbi->inode_lock[DIRTY_META]); if (inode) { @@ -1462,8 +1434,10 @@ static int f2fs_sync_inode_meta(struct f2fs_sb_info *sbi) /* it's on eviction */ if (is_inode_flag_set(inode, FI_DIRTY_INODE)) - f2fs_update_inode_page(inode); + f2fs_update_inode_cache(inode); iput(inode); + } else { + cond_resched(); } } return 0; @@ -1503,7 +1477,7 @@ static bool __need_flush_quota(struct f2fs_sb_info *sbi) } else if (is_sbi_flag_set(sbi, SBI_QUOTA_NEED_FLUSH)) { clear_sbi_flag(sbi, SBI_QUOTA_NEED_FLUSH); ret = true; - } else if (get_pages(sbi, F2FS_DIRTY_QDATA)) { + } else if (get_nr_caches(sbi, F2FS_DIRTY_QDATA)) { ret = true; } f2fs_up_write(&sbi->quota_sem); @@ -1515,14 +1489,10 @@ static bool __need_flush_quota(struct f2fs_sb_info *sbi) */ static int block_operations(struct f2fs_sb_info *sbi) { - struct writeback_control wbc = { - .sync_mode = WB_SYNC_ALL, - .nr_to_write = LONG_MAX, - }; int err = 0, cnt = 0; /* - * Let's flush inline_data in dirty node pages. + * Let's flush inline_data in dirty node caches. */ f2fs_flush_inline_data(sbi); @@ -1551,7 +1521,7 @@ retry_flush_quotas: retry_flush_dents: /* write all the dirty dentry pages */ - if (get_pages(sbi, F2FS_DIRTY_DENTS)) { + if (get_nr_caches(sbi, F2FS_DIRTY_DENTS)) { f2fs_unlock_all(sbi); err = f2fs_sync_dirty_inodes(sbi, DIR_INODE, true); if (err) @@ -1561,12 +1531,12 @@ retry_flush_dents: } /* - * POR: we should ensure that there are no dirty node pages + * POR: we should ensure that there are no dirty node caches * until finishing nat/sit flush. inode->i_blocks can be updated. */ f2fs_down_write(&sbi->node_change); - if (get_pages(sbi, F2FS_DIRTY_IMETA)) { + if (get_nr_caches(sbi, F2FS_DIRTY_IMETA)) { f2fs_up_write(&sbi->node_change); f2fs_unlock_all(sbi); err = f2fs_sync_inode_meta(sbi); @@ -1579,10 +1549,11 @@ retry_flush_dents: retry_flush_nodes: f2fs_down_write(&sbi->node_write); - if (get_pages(sbi, F2FS_DIRTY_NODES)) { + if (get_nr_caches(sbi, F2FS_DIRTY_NODES)) { f2fs_up_write(&sbi->node_write); atomic_inc(&sbi->wb_sync_req[NODE]); - err = f2fs_sync_node_pages(sbi, &wbc, false, FS_CP_NODE_IO); + err = f2fs_writeback_node_caches(sbi, LONG_MAX, + true, false, FS_CP_NODE_IO); atomic_dec(&sbi->wb_sync_req[NODE]); if (err) { f2fs_up_write(&sbi->node_change); @@ -1608,12 +1579,12 @@ static void unblock_operations(struct f2fs_sb_info *sbi) f2fs_unlock_all(sbi); } -void f2fs_wait_on_all_pages(struct f2fs_sb_info *sbi, int type) +void f2fs_sync_dirty_data(struct f2fs_sb_info *sbi, int type) { DEFINE_WAIT(wait); for (;;) { - if (!get_pages(sbi, type)) + if (!get_nr_caches(sbi, type)) break; if (unlikely(f2fs_cp_error(sbi) && @@ -1621,7 +1592,7 @@ void f2fs_wait_on_all_pages(struct f2fs_sb_info *sbi, int type) break; if (type == F2FS_DIRTY_META) - f2fs_sync_meta_pages(sbi, LONG_MAX, FS_CP_META_IO); + f2fs_sync_meta_caches(sbi, LONG_MAX, true, FS_META_IO); else if (type == F2FS_WB_CP_DATA) f2fs_submit_merged_write(sbi, DATA); @@ -1700,31 +1671,24 @@ static void update_ckpt_flags(struct f2fs_sb_info *sbi, struct cp_control *cpc) static void commit_checkpoint(struct f2fs_sb_info *sbi, void *src, block_t blk_addr) { - struct writeback_control wbc = {}; + struct f2fs_cached_block *entry = f2fs_grab_meta_cache(sbi, blk_addr); - /* - * filemap_get_folios_tag and folio_lock again will take - * some extra time. Therefore, f2fs_update_meta_pages and - * f2fs_sync_meta_pages are combined in this function. - */ - struct folio *folio = f2fs_grab_meta_folio(sbi, blk_addr); + memcpy(cache_address(entry), src, F2FS_BLKSIZE(sbi)); - memcpy(folio_address(folio), src, PAGE_SIZE); - - folio_mark_dirty(folio); - if (unlikely(!folio_clear_dirty_for_io(folio))) + f2fs_mark_cache_dirty(entry); + if (unlikely(!f2fs_cache_test_and_clear_dirty(entry))) f2fs_bug_on(sbi, 1); - /* writeout cp pack 2 page */ - if (unlikely(!__f2fs_write_meta_folio(folio, &wbc, FS_CP_META_IO))) { + /* writeout cp pack 2 cache */ + if (unlikely(!__f2fs_write_meta_cache(entry, FS_CP_META_IO))) { if (f2fs_cp_error(sbi)) { - f2fs_folio_put(folio, true); + f2fs_put_cache(entry, true); return; } f2fs_bug_on(sbi, true); } - f2fs_folio_put(folio, false); + f2fs_put_cache(entry, false); /* submit checkpoint (with barrier if NOBARRIER is not set) */ f2fs_submit_merged_write(sbi, META_FLUSH); @@ -1792,8 +1756,8 @@ static int do_checkpoint(struct f2fs_sb_info *sbi, struct cp_control *cpc) u64 kbytes_written; int err; - /* Flush all the NAT/SIT pages */ - f2fs_sync_meta_pages(sbi, LONG_MAX, FS_CP_META_IO); + /* Flush all the NAT/SIT caches */ + f2fs_sync_meta_caches(sbi, LONG_MAX, true, FS_CP_META_IO); stat_cp_time(cpc, CP_TIME_SYNC_META); @@ -1816,7 +1780,7 @@ static int do_checkpoint(struct f2fs_sb_info *sbi, struct cp_control *cpc) } /* 2 cp + n data seg summary + orphan inode blocks */ - data_sum_blocks = f2fs_npages_for_summary_flush(sbi, false); + data_sum_blocks = f2fs_nblocks_for_summary_flush(sbi, false); spin_lock_irqsave(&sbi->cp_lock, flags); if (data_sum_blocks < NR_CURSEG_DATA_TYPE) __set_ckpt_flags(ckpt, CP_COMPACT_SUM_FLAG); @@ -1824,7 +1788,7 @@ static int do_checkpoint(struct f2fs_sb_info *sbi, struct cp_control *cpc) __clear_ckpt_flags(ckpt, CP_COMPACT_SUM_FLAG); spin_unlock_irqrestore(&sbi->cp_lock, flags); - orphan_blocks = GET_ORPHAN_BLOCKS(orphan_num); + orphan_blocks = GET_ORPHAN_BLOCKS(sbi, orphan_num); ckpt->cp_pack_start_sum = cpu_to_le32(1 + cp_payload_blks + orphan_blocks); @@ -1844,7 +1808,7 @@ static int do_checkpoint(struct f2fs_sb_info *sbi, struct cp_control *cpc) get_sit_bitmap(sbi, __bitmap_ptr(sbi, SIT_BITMAP)); get_nat_bitmap(sbi, __bitmap_ptr(sbi, NAT_BITMAP)); - crc32 = f2fs_checkpoint_chksum(ckpt); + crc32 = f2fs_checkpoint_chksum(sbi, ckpt); *((__le32 *)((unsigned char *)ckpt + le32_to_cpu(ckpt->checksum_offset))) = cpu_to_le32(crc32); @@ -1861,16 +1825,16 @@ static int do_checkpoint(struct f2fs_sb_info *sbi, struct cp_control *cpc) blk = start_blk + BLKS_PER_SEG(sbi) - nm_i->nat_bits_blocks; for (i = 0; i < nm_i->nat_bits_blocks; i++) - f2fs_update_meta_page(sbi, nm_i->nat_bits + - F2FS_BLK_TO_BYTES(i), blk + i); + f2fs_update_meta_block(sbi, nm_i->nat_bits + + F2FS_BLK_TO_BYTES(sbi, i), blk + i); } /* write out checkpoint buffer at block 0 */ - f2fs_update_meta_page(sbi, ckpt, start_blk++); + f2fs_update_meta_block(sbi, ckpt, start_blk++); for (i = 1; i < 1 + cp_payload_blks; i++) - f2fs_update_meta_page(sbi, (char *)ckpt + i * F2FS_BLKSIZE, - start_blk++); + f2fs_update_meta_block(sbi, (char *)ckpt + + i * F2FS_BLKSIZE(sbi), start_blk++); if (orphan_num) { write_orphan_inodes(sbi, start_blk); @@ -1891,16 +1855,16 @@ static int do_checkpoint(struct f2fs_sb_info *sbi, struct cp_control *cpc) start_blk += NR_CURSEG_NODE_TYPE; } - /* Here, we have one bio having CP pack except cp pack 2 page */ - f2fs_sync_meta_pages(sbi, LONG_MAX, FS_CP_META_IO); + /* Here, we have one bio having CP pack except cp pack 2 blocks */ + f2fs_sync_meta_caches(sbi, LONG_MAX, true, FS_CP_META_IO); stat_cp_time(cpc, CP_TIME_SYNC_CP_META); - /* Wait for all dirty meta pages to be submitted for IO */ - f2fs_wait_on_all_pages(sbi, F2FS_DIRTY_META); + /* Wait for all dirty meta blocks to be submitted for IO */ + f2fs_sync_dirty_data(sbi, F2FS_DIRTY_META); stat_cp_time(cpc, CP_TIME_WAIT_DIRTY_META); - /* wait for previous submitted meta pages writeback */ - f2fs_wait_on_all_pages(sbi, F2FS_WB_CP_DATA); + /* wait for previous submitted meta blocks writeback */ + f2fs_sync_dirty_data(sbi, F2FS_WB_CP_DATA); stat_cp_time(cpc, CP_TIME_WAIT_CP_DATA); /* flush all device cache */ @@ -1909,20 +1873,20 @@ static int do_checkpoint(struct f2fs_sb_info *sbi, struct cp_control *cpc) return err; stat_cp_time(cpc, CP_TIME_FLUSH_DEVICE); - /* barrier and flush checkpoint cp pack 2 page if it can */ + /* barrier and flush checkpoint cp pack 2 blocks if it can */ commit_checkpoint(sbi, ckpt, start_blk); - f2fs_wait_on_all_pages(sbi, F2FS_WB_CP_DATA); + f2fs_sync_dirty_data(sbi, F2FS_WB_CP_DATA); stat_cp_time(cpc, CP_TIME_WAIT_LAST_CP); /* - * invalidate intermediate page cache borrowed from meta inode which are - * used for migration of encrypted, verity or compressed inode's blocks. + * invalidate intermediate meta cache used for migration of + * encrypted, verity or compressed inode's blocks. */ if (f2fs_sb_has_encrypt(sbi) || f2fs_sb_has_verity(sbi) || - f2fs_sb_has_compression(sbi)) - f2fs_bug_on(sbi, - invalidate_inode_pages2_range(META_MAPPING(sbi), - MAIN_BLKADDR(sbi), MAX_BLKADDR(sbi) - 1)); + f2fs_sb_has_compression(sbi)) { + f2fs_truncate_meta_caches(sbi, MAIN_BLKADDR(sbi), + MAX_BLKADDR(sbi) - MAIN_BLKADDR(sbi)); + } f2fs_release_ino_entry(sbi, false); @@ -1939,14 +1903,14 @@ static int do_checkpoint(struct f2fs_sb_info *sbi, struct cp_control *cpc) __set_cp_next_pack(sbi); /* - * redirty superblock if metadata like node page or inode cache is + * redirty superblock if metadata like node caches or inode cache is * updated during writing checkpoint. */ - if (get_pages(sbi, F2FS_DIRTY_NODES) || - get_pages(sbi, F2FS_DIRTY_IMETA)) + if (get_nr_caches(sbi, F2FS_DIRTY_NODES) || + get_nr_caches(sbi, F2FS_DIRTY_IMETA)) set_sbi_flag(sbi, SBI_IS_DIRTY); - f2fs_bug_on(sbi, get_pages(sbi, F2FS_DIRTY_DENTS)); + f2fs_bug_on(sbi, get_nr_caches(sbi, F2FS_DIRTY_DENTS)); return unlikely(f2fs_cp_error(sbi)) ? -EIO : 0; } @@ -2080,7 +2044,7 @@ void f2fs_init_ino_entry_info(struct f2fs_sb_info *sbi) sbi->max_orphans = (BLKS_PER_SEG(sbi) - F2FS_CP_PACKS - NR_CURSEG_PERSIST_TYPE - __cp_payload(sbi)) * - F2FS_ORPHANS_PER_BLOCK; + F2FS_ORPHANS_PER_BLOCK(sbi); } int __init f2fs_create_checkpoint_caches(void) diff --git a/fs/f2fs/compress.c b/fs/f2fs/compress.c index 09d9b8d0fdcc..f10c75380027 100644 --- a/fs/f2fs/compress.c +++ b/fs/f2fs/compress.c @@ -807,7 +807,7 @@ void f2fs_end_read_compressed_page(struct folio *folio, bool failed, struct decompress_io_ctx *dic = folio->private; struct f2fs_sb_info *sbi = dic->sbi; - dec_page_count(sbi, F2FS_RD_DATA); + dec_cache_count(sbi, F2FS_RD_DATA); if (failed) WRITE_ONCE(dic->failed, true); @@ -909,7 +909,7 @@ bool f2fs_sanity_check_cluster(struct dnode_of_data *dn) } for (i = 1, count = 1; i < cluster_size; i++, count++) { - block_t blkaddr = data_blkaddr(dn->inode, dn->node_folio, + block_t blkaddr = data_blkaddr(dn->inode, dn->node_entry, dn->ofs_in_node + i); /* [COMPR_ADDR, ..., COMPR_ADDR] */ @@ -950,7 +950,7 @@ static int __f2fs_get_cluster_blocks(struct inode *inode, int count, i; for (i = 0, count = 0; i < cluster_size; i++) { - block_t blkaddr = data_blkaddr(dn->inode, dn->node_folio, + block_t blkaddr = data_blkaddr(dn->inode, dn->node_entry, dn->ofs_in_node + i); if (__is_valid_data_blkaddr(blkaddr)) @@ -1152,11 +1152,11 @@ retry: goto release_and_retry; } - f2fs_folio_wait_writeback(folio, DATA, true, true); + f2fs_folio_wait_writeback(folio, true, true); f2fs_compress_ctx_add_page(cc, folio); if (!folio_test_uptodate(folio)) { - f2fs_handle_page_eio(sbi, folio, DATA); + f2fs_handle_page_eio(sbi, folio->index, DATA); release_and_retry: f2fs_put_rpages(cc); f2fs_unlock_rpages(cc, i + 1); @@ -1228,6 +1228,7 @@ int f2fs_truncate_partial_cluster(struct inode *inode, u64 from, bool lock) int i; int err; +repeat: err = f2fs_is_compressed_cluster(inode, start_idx); if (err < 0) return err; @@ -1239,12 +1240,11 @@ int f2fs_truncate_partial_cluster(struct inode *inode, u64 from, bool lock) /* truncate compressed cluster */ err = f2fs_prepare_compress_overwrite(inode, &pagep, start_idx, &fsdata); - - /* should not be a normal cluster */ - f2fs_bug_on(F2FS_I_SB(inode), err == 0); - - if (err <= 0) + if (err < 0) return err; + else if (err == 0) + /* the cluster became non-compressed one due to race case */ + goto repeat; rpages = fsdata; @@ -1328,7 +1328,7 @@ static int f2fs_write_compressed_pages(struct compress_ctx *cc, goto out_unlock_op; for (i = 0; i < cc->cluster_size; i++) { - if (data_blkaddr(dn.inode, dn.node_folio, + if (data_blkaddr(dn.inode, dn.node_entry, dn.ofs_in_node + i) == NULL_ADDR) goto out_put_dnode; } @@ -1360,10 +1360,10 @@ static int f2fs_write_compressed_pages(struct compress_ctx *cc, page_folio(cc->rpages[i + 1])->index, cic); fio.compressed_page = cc->cpages[i]; - fio.old_blkaddr = data_blkaddr(dn.inode, dn.node_folio, + fio.old_blkaddr = data_blkaddr(dn.inode, dn.node_entry, dn.ofs_in_node + i + 1); - /* wait for GCed page writeback via META_MAPPING */ + /* wait for GCed page writeback via generic cache */ f2fs_wait_on_block_writeback(inode, fio.old_blkaddr); } @@ -1477,7 +1477,7 @@ void f2fs_compress_write_end_io(struct bio *bio, struct folio *folio) f2fs_compress_free_page(page); if (atomic_dec_return(&cic->pending_pages)) { - dec_page_count(sbi, type); + dec_cache_count(sbi, type); return; } @@ -1494,12 +1494,12 @@ void f2fs_compress_write_end_io(struct bio *bio, struct folio *folio) kmem_cache_free(cic_entry_slab, cic); /* - * Make sure dec_page_count() is the last access to sbi. + * Make sure dec_cache_count() is the last access to sbi. * Once it drops the F2FS_WB_CP_DATA counter to zero, the * unmount thread can proceed to destroy sbi and * sbi->page_array_slab. */ - dec_page_count(sbi, type); + dec_cache_count(sbi, type); } static int f2fs_write_raw_pages(struct compress_ctx *cc, @@ -1551,7 +1551,7 @@ continue_unlock: if (folio_test_writeback(folio)) { if (wbc->sync_mode == WB_SYNC_NONE) goto continue_unlock; - f2fs_folio_wait_writeback(folio, DATA, true, true); + f2fs_folio_wait_writeback(folio, true, true); } if (!folio_clear_dirty_for_io(folio)) @@ -1883,14 +1883,14 @@ void f2fs_put_folio_dic(struct folio *folio, bool in_task) unsigned int f2fs_cluster_blocks_are_contiguous(struct dnode_of_data *dn, unsigned int ofs_in_node) { - bool compressed = data_blkaddr(dn->inode, dn->node_folio, + bool compressed = data_blkaddr(dn->inode, dn->node_entry, ofs_in_node) == COMPRESS_ADDR; int i = compressed ? 1 : 0; - block_t first_blkaddr = data_blkaddr(dn->inode, dn->node_folio, + block_t first_blkaddr = data_blkaddr(dn->inode, dn->node_entry, ofs_in_node + i); for (i += 1; i < F2FS_I(dn->inode)->i_cluster_size; i++) { - block_t blkaddr = data_blkaddr(dn->inode, dn->node_folio, + block_t blkaddr = data_blkaddr(dn->inode, dn->node_entry, ofs_in_node + i); if (!__is_valid_data_blkaddr(blkaddr)) @@ -1902,30 +1902,19 @@ unsigned int f2fs_cluster_blocks_are_contiguous(struct dnode_of_data *dn, return compressed ? i - 1 : i; } -const struct address_space_operations f2fs_compress_aops = { - .release_folio = f2fs_release_folio, - .invalidate_folio = f2fs_invalidate_folio, - .migrate_folio = filemap_migrate_folio, -}; - -struct address_space *COMPRESS_MAPPING(struct f2fs_sb_info *sbi) -{ - return sbi->compress_inode->i_mapping; -} - void f2fs_invalidate_compress_pages_range(struct f2fs_sb_info *sbi, block_t blkaddr, unsigned int len) { - if (!sbi->compress_inode) + if (!test_opt(sbi, COMPRESS_CACHE)) return; - invalidate_mapping_pages(COMPRESS_MAPPING(sbi), blkaddr, blkaddr + len - 1); + + f2fs_drop_cache_range(COMPRESS_CACHE(sbi), blkaddr, len, false); } static void f2fs_cache_compressed_page(struct f2fs_sb_info *sbi, struct folio *folio, nid_t ino, block_t blkaddr) { - struct folio *cfolio; - int ret; + struct f2fs_cached_block *entry; if (!test_opt(sbi, COMPRESS_CACHE)) return; @@ -1933,52 +1922,48 @@ static void f2fs_cache_compressed_page(struct f2fs_sb_info *sbi, if (!f2fs_is_valid_blkaddr(sbi, blkaddr, DATA_GENERIC_ENHANCE_READ)) return; - if (!f2fs_available_free_memory(sbi, COMPRESS_PAGE)) + if (!f2fs_available_free_memory(sbi, COMPRESS_BLOCK)) return; - cfolio = filemap_get_folio(COMPRESS_MAPPING(sbi), blkaddr); - if (!IS_ERR(cfolio)) { - f2fs_folio_put(cfolio, false); + entry = f2fs_find_cache(COMPRESS_CACHE(sbi), blkaddr, + F2FS_CACHE_ACCESS); + if (!IS_ERR(entry)) { + f2fs_put_cache(entry, false); return; } - cfolio = filemap_alloc_folio(__GFP_NOWARN | __GFP_IO, 0, NULL); - if (!cfolio) + entry = f2fs_grab_cache(COMPRESS_CACHE(sbi), blkaddr, + F2FS_CACHE_LOCK_CREATE); + if (IS_ERR(entry)) return; - ret = filemap_add_folio(COMPRESS_MAPPING(sbi), cfolio, - blkaddr, GFP_NOFS); - if (ret) { - f2fs_folio_put(cfolio, false); - return; - } - - folio_set_f2fs_data(cfolio, ino); - - memcpy(folio_address(cfolio), folio_address(folio), PAGE_SIZE); - folio_mark_uptodate(cfolio); - f2fs_folio_put(cfolio, true); + entry->ino = ino; + memcpy(cache_address(entry), folio_address(folio), sbi->blocksize); + f2fs_cache_set_uptodate(entry); + f2fs_put_cache(entry, true); } bool f2fs_load_compressed_folio(struct f2fs_sb_info *sbi, struct folio *folio, block_t blkaddr) { - struct folio *cfolio; + struct f2fs_cached_block *entry; bool hitted = false; if (!test_opt(sbi, COMPRESS_CACHE)) return false; - cfolio = f2fs_filemap_get_folio(COMPRESS_MAPPING(sbi), - blkaddr, FGP_LOCK | FGP_NOWAIT, GFP_NOFS); - if (!IS_ERR(cfolio)) { - if (folio_test_uptodate(cfolio)) { + entry = f2fs_find_cache(COMPRESS_CACHE(sbi), blkaddr, + F2FS_CACHE_ACCESS); + if (!IS_ERR(entry)) { + f2fs_lock_cache(entry); + if (f2fs_is_compress_cache(entry) && + f2fs_cache_test_uptodate(entry)) { atomic_inc(&sbi->compress_page_hit); memcpy(folio_address(folio), - folio_address(cfolio), folio_size(folio)); + cache_address(entry), folio_size(folio)); hitted = true; } - f2fs_folio_put(cfolio, true); + f2fs_put_cache(entry, true); } return hitted; @@ -1986,71 +1971,49 @@ bool f2fs_load_compressed_folio(struct f2fs_sb_info *sbi, struct folio *folio, void f2fs_invalidate_compress_pages(struct f2fs_sb_info *sbi, nid_t ino) { - struct address_space *mapping = COMPRESS_MAPPING(sbi); - struct folio_batch fbatch; - pgoff_t index = 0; - pgoff_t end = MAX_BLKADDR(sbi); + struct f2fs_cached_block_list *cache = COMPRESS_CACHE(sbi); + struct f2fs_cached_block *entries[F2FS_ONSTACK_CACHES]; + pgoff_t index = 0, end = ULONG_MAX; + int nr; + int i; - if (!mapping->nrpages) + if (!cache->num_entries) return; +next: + nr = f2fs_cache_gang_lookup(cache, entries, &index, end); + if (!nr) + return; + for (i = 0; i < nr; i++) { + struct f2fs_cached_block *entry = entries[i]; - folio_batch_init(&fbatch); - - do { - unsigned int nr, i; - - nr = filemap_get_folios(mapping, &index, end - 1, &fbatch); - if (!nr) - break; - - for (i = 0; i < nr; i++) { - struct folio *folio = fbatch.folios[i]; + index = entry->index + 1; - folio_lock(folio); - if (folio->mapping != mapping) { - folio_unlock(folio); - continue; - } + f2fs_lock_cache(entry); + if (unlikely(!f2fs_is_compress_cache(entry))) + goto unlock; + if (entry->ino != ino) + goto unlock; - if (ino != folio_get_f2fs_data(folio)) { - folio_unlock(folio); - continue; - } + f2fs_truncate_locked_cache(entry, false); +unlock: + f2fs_unlock_cache(entry); + } + f2fs_cache_gang_release(entries, nr); - generic_error_remove_folio(mapping, folio); - folio_unlock(folio); - } - folio_batch_release(&fbatch); + if (index < end) { cond_resched(); - } while (index < end); + goto next; + } } -int f2fs_init_compress_inode(struct f2fs_sb_info *sbi) +void f2fs_init_compress_cache_context(struct f2fs_sb_info *sbi) { - struct inode *inode; - if (!test_opt(sbi, COMPRESS_CACHE)) - return 0; - - inode = f2fs_iget(sbi->sb, F2FS_COMPRESS_INO(sbi)); - if (IS_ERR(inode)) - return PTR_ERR(inode); - sbi->compress_inode = inode; + return; sbi->compress_percent = COMPRESS_PERCENT; sbi->compress_watermark = COMPRESS_WATERMARK; - atomic_set(&sbi->compress_page_hit, 0); - - return 0; -} - -void f2fs_destroy_compress_inode(struct f2fs_sb_info *sbi) -{ - if (!sbi->compress_inode) - return; - iput(sbi->compress_inode); - sbi->compress_inode = NULL; } int f2fs_init_page_array_cache(struct f2fs_sb_info *sbi) diff --git a/fs/f2fs/data.c b/fs/f2fs/data.c index ca8232a9095f..29c4a81eb947 100644 --- a/fs/f2fs/data.c +++ b/fs/f2fs/data.c @@ -41,11 +41,6 @@ struct f2fs_folio_state { unsigned int read_pages_pending; }; -struct f2fs_bio { - struct work_struct work; - struct bio bio; -}; - #define F2FS_BIO_POOL_SIZE NR_CURSEG_TYPE int __init f2fs_init_bioset(void) @@ -63,14 +58,10 @@ bool f2fs_is_cp_guaranteed(const struct folio *folio) { struct address_space *mapping = folio->mapping; struct inode *inode; - struct f2fs_sb_info *sbi; inode = mapping->host; - sbi = F2FS_I_SB(inode); - if (inode->i_ino == F2FS_META_INO(sbi) || - inode->i_ino == F2FS_NODE_INO(sbi) || - S_ISDIR(inode->i_mode)) + if (S_ISDIR(inode->i_mode)) return true; if ((S_ISREG(inode->i_mode) && IS_NOQUOTA(inode)) || @@ -79,23 +70,6 @@ bool f2fs_is_cp_guaranteed(const struct folio *folio) return false; } -static enum count_type __read_io_type(struct folio *folio) -{ - struct address_space *mapping = folio->mapping; - - if (mapping) { - struct inode *inode = mapping->host; - struct f2fs_sb_info *sbi = F2FS_I_SB(inode); - - if (inode->i_ino == F2FS_META_INO(sbi)) - return F2FS_RD_META; - - if (inode->i_ino == F2FS_NODE_INO(sbi)) - return F2FS_RD_NODE; - } - return F2FS_RD_DATA; -} - /* postprocessing steps for read bios */ enum bio_post_read_step { #ifdef CONFIG_F2FS_FS_COMPRESSION @@ -169,13 +143,7 @@ static void f2fs_finish_read_bio(struct bio *bio, bool in_task) } while (nr_pages--) - dec_page_count(F2FS_F_SB(folio), __read_io_type(folio)); - - if (bio->bi_status == BLK_STS_OK && - F2FS_F_SB(folio)->node_inode && is_node_folio(folio) && - f2fs_sanity_check_node_footer(F2FS_F_SB(folio), - folio, folio->index, NODE_TYPE_REGULAR, true)) - bio->bi_status = BLK_STS_IOERR; + dec_cache_count(F2FS_F_SB(folio), F2FS_RD_DATA); if (finished) folio_end_read(folio, bio->bi_status == BLK_STS_OK); @@ -359,21 +327,13 @@ static void f2fs_write_end_bio(struct bio *bio) } } - if (is_node_folio(folio)) { - f2fs_sanity_check_node_footer(sbi, folio, - folio->index, NODE_TYPE_REGULAR, true); - f2fs_bug_on(sbi, folio->index != nid_of_node(folio)); - } - if (f2fs_in_warm_node_list(folio)) - f2fs_del_fsync_node_entry(sbi, folio); - - dec_page_count(sbi, type); + dec_cache_count(sbi, type); /* * we should access sbi before folio_end_writeback() to * avoid racing w/ kill_f2fs_super() */ - if (type == F2FS_WB_CP_DATA && !get_pages(sbi, type) && + if (type == F2FS_WB_CP_DATA && !get_nr_caches(sbi, type) && wq_has_sleeper(&sbi->cp_wait)) wake_up(&sbi->cp_wait); @@ -410,6 +370,79 @@ static void f2fs_write_end_io(struct bio *bio) } } +static void f2fs_cache_read_end_io(struct bio *bio) +{ + struct f2fs_cached_block *entry = F2FS_BIO(bio)->entry; + struct f2fs_sb_info *sbi = entry->cache->sbi; + enum count_type io_type = IS_META_CACHE(entry->cache) ? + F2FS_RD_META : F2FS_RD_NODE; + struct f2fs_cached_block *next; + + iostat_update_and_unbind_ctx(bio); + + if (time_to_inject(sbi, FAULT_READ_IO)) + bio->bi_status = BLK_STS_IOERR; + + while (entry) { + next = entry->next_entry; + entry->next_entry = NULL; + + if (bio->bi_status == BLK_STS_OK && + f2fs_is_node_cache(entry) && + f2fs_sanity_check_node_footer(sbi, entry, + entry->index, NODE_TYPE_REGULAR, true)) + bio->bi_status = BLK_STS_IOERR; + + if (bio->bi_status == BLK_STS_OK) + f2fs_cache_set_uptodate(entry); + + dec_cache_count(sbi, io_type); + + f2fs_unlock_cache(entry); + entry = next; + } + bio_put(bio); +} + +static void f2fs_cache_write_end_io(struct bio *bio) +{ + struct f2fs_cached_block *entry = F2FS_BIO(bio)->entry; + struct f2fs_sb_info *sbi = entry->cache->sbi; + struct f2fs_cached_block *next; + + iostat_update_and_unbind_ctx(bio); + + if (time_to_inject(sbi, FAULT_WRITE_IO)) + bio->bi_status = BLK_STS_IOERR; + + if (bio->bi_status != BLK_STS_OK) + f2fs_stop_checkpoint(sbi, true, + STOP_CP_REASON_WRITE_FAIL); + + while (entry) { + next = entry->next_entry; + entry->next_entry = NULL; + + if (f2fs_is_node_cache(entry)) { + f2fs_sanity_check_node_footer(sbi, entry, + entry->index, NODE_TYPE_REGULAR, true); + f2fs_bug_on(sbi, entry->index != nid_of_node(sbi, entry)); + } + if (f2fs_in_warm_node_list(sbi, entry)) + f2fs_del_fsync_node_entry(sbi, entry); + + dec_cache_count(sbi, F2FS_WB_CP_DATA); + + if (!get_nr_caches(sbi, F2FS_WB_CP_DATA) && + wq_has_sleeper(&sbi->cp_wait)) + wake_up(&sbi->cp_wait); + + f2fs_end_cache_writeback(entry); + entry = next; + } + bio_put(bio); +} + #ifdef CONFIG_BLK_DEV_ZONED static void f2fs_zone_write_end_io(struct bio *bio) { @@ -417,7 +450,10 @@ static void f2fs_zone_write_end_io(struct bio *bio) bio->bi_private = io->bi_private; complete(&io->zone_wait); - f2fs_write_end_io(bio); + if (f2fs_is_cache_bio(bio)) + f2fs_cache_write_end_io(bio); + else + f2fs_write_end_io(bio); } #endif @@ -439,7 +475,7 @@ struct block_device *f2fs_target_device(struct f2fs_sb_info *sbi, } if (sector) - *sector = SECTOR_FROM_BLOCK(blk_addr); + *sector = SECTOR_FROM_BLOCK(sbi, blk_addr); return bdev; } @@ -504,12 +540,21 @@ static struct bio *__bio_alloc(struct f2fs_io_info *fio, int npages) fio->op | fio->op_flags | f2fs_io_flags(fio), GFP_NOIO, &f2fs_bioset); bio->bi_iter.bi_sector = sector; + F2FS_BIO(bio)->entry = NULL; + bio->bi_private = NULL; if (is_read_io(fio->op)) { - bio->bi_end_io = f2fs_read_end_io; - bio->bi_private = NULL; + if (fio->is_cache) + bio->bi_end_io = f2fs_cache_read_end_io; + else + bio->bi_end_io = f2fs_read_end_io; } else { - bio->bi_end_io = f2fs_write_end_io; - bio->bi_private = sbi; + if (fio->is_cache) { + bio->bi_end_io = f2fs_cache_write_end_io; + } else { + bio->bi_end_io = f2fs_write_end_io; + bio->bi_private = sbi; + } + bio->bi_write_hint = f2fs_io_type_to_rw_hint(sbi, fio->type, fio->temp); bio->bi_write_stream = f2fs_io_type_to_write_stream(bdev, fio->type, @@ -532,7 +577,7 @@ static void f2fs_set_bio_crypt_ctx(struct bio *bio, const struct inode *inode, * The f2fs garbage collector sets ->encrypted_page when it wants to * read/write raw data without encryption. */ - if (!fio || !fio->encrypted_page) + if (!fio || (!fio->encrypted_page && !fio->is_cache)) fscrypt_set_bio_crypt_ctx(bio, inode, (loff_t)first_idx << inode->i_blkbits, gfp_mask); @@ -593,16 +638,19 @@ static void __submit_merged_bio(struct f2fs_bio_info *io) } static bool __has_merged_page(struct bio *bio, struct inode *inode, - struct folio *folio, nid_t ino) + struct folio *folio) { struct folio_iter fi; if (!bio) return false; - if (!inode && !folio && !ino) + if (!inode && !folio) return true; + if (f2fs_is_cache_bio(bio)) + return false; + bio_for_each_folio_all(fi, bio) { struct folio *target = fi.folio; @@ -616,8 +664,6 @@ static bool __has_merged_page(struct bio *bio, struct inode *inode, return true; if (folio && folio == target) return true; - if (ino && ino == ino_of_node(target)) - return true; } return false; @@ -684,26 +730,25 @@ unlock_out: f2fs_up_write_trace(&io->io_rwsem, &lc); } -static void __submit_merged_write_cond(struct f2fs_sb_info *sbi, +static void __submit_merged_data_write_cond(struct f2fs_sb_info *sbi, struct inode *inode, struct folio *folio, - nid_t ino, enum page_type type, bool writeback) + bool writeback) { enum temp_type temp; bool ret = true; - bool force = !inode && !folio && !ino; + bool force = !inode && !folio; for (temp = HOT; temp < NR_TEMP_TYPE; temp++) { - if (!force) { - enum page_type btype = PAGE_TYPE_OF_BIO(type); - struct f2fs_bio_info *io = sbi->write_io[btype] + temp; + if (!force) { + struct f2fs_bio_info *io = sbi->write_io[DATA] + temp; struct f2fs_lock_context lc; f2fs_down_read_trace(&io->io_rwsem, &lc); - ret = __has_merged_page(io->bio, inode, folio, ino); + ret = __has_merged_page(io->bio, inode, folio); f2fs_up_read_trace(&io->io_rwsem, &lc); } if (ret) { - __f2fs_submit_merged_write(sbi, type, temp); + __f2fs_submit_merged_write(sbi, DATA, temp); /* * For waitting writebck case, if the bio owned by the * folio is already submitted, we do not need to submit @@ -712,29 +757,86 @@ static void __submit_merged_write_cond(struct f2fs_sb_info *sbi, if (writeback) break; } - - /* TODO: use HOT temp only for meta pages now. */ - if (type >= META) - break; } } -void f2fs_submit_merged_write(struct f2fs_sb_info *sbi, enum page_type type) +void f2fs_submit_merged_write_cond(struct f2fs_sb_info *sbi, + struct inode *inode, struct folio *folio) { - __submit_merged_write_cond(sbi, NULL, NULL, 0, type, false); + __submit_merged_data_write_cond(sbi, inode, folio, false); } -void f2fs_submit_merged_write_cond(struct f2fs_sb_info *sbi, - struct inode *inode, struct folio *folio, +void f2fs_submit_merged_write_folio(struct f2fs_sb_info *sbi, + struct folio *folio) +{ + __submit_merged_data_write_cond(sbi, NULL, folio, true); +} + +static bool __has_merged_cache(struct f2fs_sb_info *sbi, struct bio *bio, + struct f2fs_cached_block *target, nid_t ino) +{ + struct f2fs_cached_block *entry; + + if (!bio) + return false; + + entry = F2FS_BIO(bio)->entry; + + while (entry) { + if (target && entry == target) + return true; + + if (ino) { + if (likely(f2fs_is_node_cache(entry))) { + if (ino_of_node(sbi, entry) == ino) + return true; + } else { + WARN_ON_ONCE(1); + } + } + entry = entry->next_entry; + } + return false; +} + +bool f2fs_submit_merged_write_cache(struct f2fs_sb_info *sbi, + struct f2fs_cached_block *entry, nid_t ino, enum page_type type) { - __submit_merged_write_cond(sbi, inode, folio, ino, type, false); + enum temp_type temp; + bool ret = false; + bool force = !entry && !ino; + + for (temp = HOT; temp < NR_TEMP_TYPE; temp++) { + enum page_type btype = PAGE_TYPE_OF_BIO(type); + struct f2fs_bio_info *io = sbi->write_io[btype] + temp; + struct f2fs_lock_context lc; + bool merged = true; + + if (!force) { + f2fs_down_read_trace(&io->io_rwsem, &lc); + merged = __has_merged_cache(sbi, io->bio, entry, ino); + f2fs_up_read_trace(&io->io_rwsem, &lc); + } + + if (merged) { + __f2fs_submit_merged_write(sbi, type, temp); + ret = true; + } + + /* TODO: use HOT temp only for meta pages now. */ + if (type >= META) + break; + } + return ret; } -void f2fs_submit_merged_write_folio(struct f2fs_sb_info *sbi, - struct folio *folio, enum page_type type) +void f2fs_submit_merged_write(struct f2fs_sb_info *sbi, enum page_type type) { - __submit_merged_write_cond(sbi, NULL, folio, 0, type, true); + if (type == DATA) + __submit_merged_data_write_cond(sbi, NULL, NULL, false); + else + f2fs_submit_merged_write_cache(sbi, NULL, 0, type); } void f2fs_flush_merged_writes(struct f2fs_sb_info *sbi) @@ -772,8 +874,8 @@ int f2fs_submit_page_bio(struct f2fs_io_info *fio) if (fio->io_wbc && !is_read_io(fio->op)) wbc_account_cgroup_owner(fio->io_wbc, fio_folio, PAGE_SIZE); - inc_page_count(fio->sbi, is_read_io(fio->op) ? - __read_io_type(data_folio) : WB_DATA_TYPE(fio->folio, false)); + inc_cache_count(fio->sbi, is_read_io(fio->op) ? + F2FS_RD_DATA : WB_DATA_TYPE(fio->folio, false)); if (is_read_io(bio_op(bio))) f2fs_submit_read_bio(fio->sbi, bio, fio->type); @@ -800,6 +902,8 @@ static bool io_type_is_mergeable(struct f2fs_bio_info *io, if (io->fio.op != fio->op) return false; + if (io->fio.is_cache != fio->is_cache) + return false; return (io->fio.op_flags & mask) == (fio->op_flags & mask); } @@ -907,8 +1011,8 @@ void f2fs_submit_merged_ipu_write(struct f2fs_sb_info *sbi, if (target) found = (target == be->bio); else - found = __has_merged_page(be->bio, NULL, - folio, 0); + found = __has_merged_page(be->bio, + NULL, folio); if (found) break; } @@ -924,8 +1028,8 @@ void f2fs_submit_merged_ipu_write(struct f2fs_sb_info *sbi, if (target) found = (target == be->bio); else - found = __has_merged_page(be->bio, NULL, - folio, 0); + found = __has_merged_page(be->bio, + NULL, folio); if (found) { target = be->bio; del_bio_entry(be); @@ -1003,7 +1107,7 @@ alloc_new: if (fio->io_wbc) wbc_account_cgroup_owner(fio->io_wbc, folio, folio_size(folio)); - inc_page_count(fio->sbi, WB_DATA_TYPE(folio, false)); + inc_cache_count(fio->sbi, WB_DATA_TYPE(folio, false)); *fio->last_block = fio->new_blkaddr; *fio->bio = bio; @@ -1031,6 +1135,33 @@ static bool is_end_zone_blkaddr(struct f2fs_sb_info *sbi, block_t blkaddr) f2fs_blkz_is_seq(sbi, devi, blkaddr) && (blkaddr % sbi->blocks_per_blkz == sbi->blocks_per_blkz - 1); } + +static void f2fs_wait_zone_io_completion(struct f2fs_sb_info *sbi, + struct f2fs_bio_info *io, enum page_type btype) +{ + if (f2fs_sb_has_blkzoned(sbi) && btype < META && io->zone_pending_bio) { + wait_for_completion_io(&io->zone_wait); + bio_put(io->zone_pending_bio); + io->zone_pending_bio = NULL; + io->bi_private = NULL; + } +} + +static void f2fs_submit_zone_io(struct f2fs_sb_info *sbi, + struct f2fs_io_info *fio, struct f2fs_bio_info *io, + enum page_type btype) +{ + if (f2fs_sb_has_blkzoned(sbi) && btype < META && + is_end_zone_blkaddr(sbi, fio->new_blkaddr)) { + bio_get(io->bio); + reinit_completion(&io->zone_wait); + io->bi_private = io->bio->bi_private; + io->bio->bi_private = io; + io->bio->bi_end_io = f2fs_zone_write_end_io; + io->zone_pending_bio = io->bio; + __submit_merged_bio(io); + } +} #endif void f2fs_submit_page_write(struct f2fs_io_info *fio) @@ -1047,14 +1178,8 @@ void f2fs_submit_page_write(struct f2fs_io_info *fio) f2fs_down_write_trace(&io->io_rwsem, &lc); next: #ifdef CONFIG_BLK_DEV_ZONED - if (f2fs_sb_has_blkzoned(sbi) && btype < META && io->zone_pending_bio) { - wait_for_completion_io(&io->zone_wait); - bio_put(io->zone_pending_bio); - io->zone_pending_bio = NULL; - io->bi_private = NULL; - } + f2fs_wait_zone_io_completion(sbi, io, btype); #endif - if (fio->in_list) { spin_lock(&io->io_lock); if (list_empty(&io->io_list)) { @@ -1080,7 +1205,7 @@ next: fio->submitted = 1; type = WB_DATA_TYPE(bio_folio, fio->compressed_page); - inc_page_count(sbi, type); + inc_cache_count(sbi, type); if (io->bio && (!io_is_mergeable(sbi, io->bio, io, fio, io->last_block_in_bio, @@ -1109,16 +1234,109 @@ alloc_new: trace_f2fs_submit_folio_write(fio->folio, fio); #ifdef CONFIG_BLK_DEV_ZONED - if (f2fs_sb_has_blkzoned(sbi) && btype < META && - is_end_zone_blkaddr(sbi, fio->new_blkaddr)) { - bio_get(io->bio); - reinit_completion(&io->zone_wait); - io->bi_private = io->bio->bi_private; - io->bio->bi_private = io; - io->bio->bi_end_io = f2fs_zone_write_end_io; - io->zone_pending_bio = io->bio; + f2fs_submit_zone_io(sbi, fio, io, btype); +#endif + + if (fio->in_list) + goto next; +out: + if (is_sbi_flag_set(sbi, SBI_IS_SHUTDOWN) || + !f2fs_is_checkpoint_ready(sbi)) __submit_merged_bio(io); + f2fs_up_write_trace(&io->io_rwsem, &lc); +} + +static void f2fs_bio_add_cache(struct f2fs_io_info *fio, struct bio *bio) +{ + struct f2fs_bio *fbio = F2FS_BIO(bio); + struct f2fs_cached_block *head = fbio->entry; + struct f2fs_cached_block *new = fio->cache_entry; + + new->next_entry = head; + fbio->entry = new; +} + +int f2fs_submit_cache_read(struct f2fs_io_info *fio) +{ + struct f2fs_sb_info *sbi = fio->sbi; + struct f2fs_cached_block *entry = fio->cache_entry; + struct bio *bio; + enum count_type io_type = IS_META_CACHE(entry->cache) ? + F2FS_RD_META : F2FS_RD_NODE; + + if (!f2fs_is_valid_blkaddr(fio->sbi, fio->new_blkaddr, + fio->is_por ? META_POR : (__is_meta_io(fio) ? + META_GENERIC : DATA_GENERIC_ENHANCE))) + return -EFSCORRUPTED; + + bio = __bio_alloc(fio, 1); + + bio_add_virt_nofail(bio, cache_address(entry), sbi->blocksize); + f2fs_bio_add_cache(fio, bio); + inc_cache_count(sbi, io_type); + + f2fs_submit_read_bio(sbi, bio, fio->type); + return 0; +} + +void f2fs_submit_cache_write(struct f2fs_io_info *fio) +{ + struct f2fs_sb_info *sbi = fio->sbi; + enum page_type btype = PAGE_TYPE_OF_BIO(fio->type); + struct f2fs_bio_info *io = sbi->write_io[btype] + fio->temp; + struct f2fs_lock_context lc; + struct folio *folio; + + f2fs_bug_on(sbi, is_read_io(fio->op)); + + f2fs_down_write_trace(&io->io_rwsem, &lc); +next: +#ifdef CONFIG_BLK_DEV_ZONED + f2fs_wait_zone_io_completion(sbi, io, btype); +#endif + if (fio->in_list) { + spin_lock(&io->io_lock); + if (list_empty(&io->io_list)) { + spin_unlock(&io->io_lock); + goto out; + } + fio = list_first_entry(&io->io_list, + struct f2fs_io_info, list); + list_del(&fio->list); + spin_unlock(&io->io_lock); + } + + verify_fio_blkaddr(fio); + + fio->submitted = 1; + inc_cache_count(sbi, F2FS_WB_CP_DATA); + + if (io->bio && + (!io_is_mergeable(sbi, io->bio, io, fio, io->last_block_in_bio, + fio->new_blkaddr))) + __submit_merged_bio(io); +alloc_new: + if (io->bio == NULL) { + io->bio = __bio_alloc(fio, BIO_MAX_VECS); + io->fio = *fio; + } + + folio = cache_folio(fio->cache_entry); + + if (!bio_add_folio(io->bio, folio, sbi->blocksize, + offset_in_folio(folio, cache_address(fio->cache_entry)))) { + f2fs_bug_on(sbi, !F2FS_BIO(io->bio)->entry); + + __submit_merged_bio(io); + goto alloc_new; } + + f2fs_bio_add_cache(fio, io->bio); + + io->last_block_in_bio = fio->new_blkaddr; + +#ifdef CONFIG_BLK_DEV_ZONED + f2fs_submit_zone_io(sbi, fio, io, btype); #endif if (fio->in_list) goto next; @@ -1185,20 +1403,20 @@ static void f2fs_submit_page_read(struct inode *inode, struct fsverity_info *vi, bio = f2fs_grab_read_bio(inode, vi, blkaddr, 1, op_flags, folio->index, for_write); - /* wait for GCed page writeback via META_MAPPING */ + /* wait for GCed page writeback via generic cache */ f2fs_wait_on_block_writeback(inode, blkaddr); if (!bio_add_folio(bio, folio, PAGE_SIZE, 0)) f2fs_bug_on(sbi, 1); - inc_page_count(sbi, F2FS_RD_DATA); - f2fs_update_iostat(sbi, NULL, FS_DATA_READ_IO, F2FS_BLKSIZE); + inc_cache_count(sbi, F2FS_RD_DATA); + f2fs_update_iostat(sbi, NULL, FS_DATA_READ_IO, F2FS_BLKSIZE(sbi)); f2fs_submit_read_bio(sbi, bio, DATA); } static void __set_data_blkaddr(struct dnode_of_data *dn, block_t blkaddr) { - __le32 *addr = get_dnode_addr(dn->inode, dn->node_folio); + __le32 *addr = get_dnode_addr(dn->inode, dn->node_entry); dn->data_blkaddr = blkaddr; addr[dn->ofs_in_node] = cpu_to_le32(dn->data_blkaddr); @@ -1212,9 +1430,9 @@ static void __set_data_blkaddr(struct dnode_of_data *dn, block_t blkaddr) */ void f2fs_set_data_blkaddr(struct dnode_of_data *dn, block_t blkaddr) { - f2fs_folio_wait_writeback(dn->node_folio, NODE, true, true); + f2fs_cache_wait_writeback(dn->node_entry); __set_data_blkaddr(dn, blkaddr); - if (folio_mark_dirty(dn->node_folio)) + if (f2fs_mark_cache_dirty(dn->node_entry)) dn->node_changed = true; } @@ -1242,7 +1460,7 @@ int f2fs_reserve_new_blocks(struct dnode_of_data *dn, blkcnt_t count) trace_f2fs_reserve_new_blocks(dn->inode, dn->nid, dn->ofs_in_node, count); - f2fs_folio_wait_writeback(dn->node_folio, NODE, true, true); + f2fs_cache_wait_writeback(dn->node_entry); for (; count > 0; dn->ofs_in_node++) { block_t blkaddr = f2fs_data_blkaddr(dn); @@ -1253,7 +1471,7 @@ int f2fs_reserve_new_blocks(struct dnode_of_data *dn, blkcnt_t count) } } - if (folio_mark_dirty(dn->node_folio)) + if (f2fs_mark_cache_dirty(dn->node_entry)) dn->node_changed = true; return 0; } @@ -1271,7 +1489,7 @@ int f2fs_reserve_new_block(struct dnode_of_data *dn) int f2fs_reserve_block(struct dnode_of_data *dn, pgoff_t index) { - bool need_put = dn->inode_folio ? false : true; + bool need_put = dn->inode_entry ? false : true; int err; err = f2fs_get_dnode_of_data(dn, index, ALLOC_NODE); @@ -1437,11 +1655,11 @@ struct folio *f2fs_get_lock_data_folio(struct inode *inode, pgoff_t index, * * Also, caller should grab and release a rwsem by calling f2fs_lock_op() and * f2fs_unlock_op(). - * Note that, ifolio is set only by make_empty_dir, and if any error occur, - * ifolio should be released by this function. + * Note that, ientry is set only by make_empty_dir, and if any error occur, + * ientry should be released by this function. */ struct folio *f2fs_get_new_data_folio(struct inode *inode, - struct folio *ifolio, pgoff_t index, bool new_i_size) + struct f2fs_cached_block *ientry, pgoff_t index, bool new_i_size) { struct address_space *mapping = inode->i_mapping; struct folio *folio; @@ -1451,20 +1669,20 @@ struct folio *f2fs_get_new_data_folio(struct inode *inode, folio = f2fs_grab_cache_folio(mapping, index, true); if (IS_ERR(folio)) { /* - * before exiting, we should make sure ifolio will be released + * before exiting, we should make sure ientry will be released * if any error occur. */ - f2fs_folio_put(ifolio, true); + f2fs_put_cache(ientry, true); return ERR_PTR(-ENOMEM); } - set_new_dnode(&dn, inode, ifolio, NULL, 0); + set_new_dnode(&dn, inode, ientry, NULL, 0); err = f2fs_reserve_block(&dn, index); if (err) { f2fs_folio_put(folio, true); return ERR_PTR(err); } - if (!ifolio) + if (!ientry) f2fs_put_dnode(&dn); if (folio_test_uptodate(folio)) @@ -1477,8 +1695,8 @@ struct folio *f2fs_get_new_data_folio(struct inode *inode, } else { f2fs_folio_put(folio, true); - /* if ifolio exists, blkaddr should be NEW_ADDR */ - f2fs_bug_on(F2FS_I_SB(inode), ifolio); + /* if ientry exists, blkaddr should be NEW_ADDR */ + f2fs_bug_on(F2FS_I_SB(inode), ientry); folio = f2fs_get_lock_data_folio(inode, index, true); if (IS_ERR(folio)) return folio; @@ -1515,7 +1733,7 @@ static int __allocate_data_block(struct dnode_of_data *dn, int seg_type) set_summary(&sum, dn->nid, dn->ofs_in_node, ni.version); old_blkaddr = dn->data_blkaddr; - err = f2fs_allocate_data_block(sbi, NULL, old_blkaddr, + err = f2fs_allocate_data_block(sbi, old_blkaddr, &dn->data_blkaddr, &sum, seg_type, NULL); if (err) { if (old_blkaddr == NULL_ADDR) @@ -1727,7 +1945,7 @@ next_dnode: start_pgofs = pgofs; prealloc = 0; last_ofs_in_node = ofs_in_node = dn.ofs_in_node; - end_offset = ADDRS_PER_PAGE(dn.node_folio, inode); + end_offset = ADDRS_PER_PAGE(dn.node_entry, inode); next_block: blkaddr = f2fs_data_blkaddr(&dn); @@ -1942,12 +2160,12 @@ static bool __f2fs_overwrite_io(struct inode *inode, loff_t pos, size_t len, if (pos + len > i_size_read(inode)) return false; - map.m_lblk = F2FS_BYTES_TO_BLK(pos); + map.m_lblk = F2FS_BYTES_TO_BLK(F2FS_I_SB(inode), pos); map.m_next_pgofs = NULL; map.m_next_extent = NULL; map.m_seg_type = NO_CHECK_TYPE; map.m_may_create = false; - last_lblk = F2FS_BLK_ALIGN(pos + len); + last_lblk = F2FS_BLK_ALIGN(F2FS_I_SB(inode), pos + len); while (map.m_lblk < last_lblk) { map.m_len = last_lblk - map.m_lblk; @@ -1978,27 +2196,27 @@ static int f2fs_xattr_fiemap(struct inode *inode, if (f2fs_has_inline_xattr(inode)) { int offset; - struct folio *folio = f2fs_grab_cache_folio(NODE_MAPPING(sbi), - inode->i_ino, false); + struct f2fs_cached_block *entry = + f2fs_grab_node_cache(sbi, inode->i_ino); - if (IS_ERR(folio)) - return PTR_ERR(folio); + if (IS_ERR(entry)) + return PTR_ERR(entry); err = f2fs_get_node_info(sbi, inode->i_ino, &ni, false); if (err) { - f2fs_folio_put(folio, true); + f2fs_put_cache(entry, true); return err; } - phys = F2FS_BLK_TO_BYTES(ni.blk_addr); + phys = F2FS_BLK_TO_BYTES(sbi, ni.blk_addr); offset = offsetof(struct f2fs_inode, i_addr) + - sizeof(__le32) * (DEF_ADDRS_PER_INODE - + sizeof(__le32) * (DEF_ADDRS_PER_INODE(sbi) - get_inline_xattr_addrs(inode)); phys += offset; len = inline_xattr_size(inode); - f2fs_folio_put(folio, true); + f2fs_put_cache(entry, true); flags = FIEMAP_EXTENT_DATA_INLINE | FIEMAP_EXTENT_NOT_ALIGNED; @@ -2012,22 +2230,22 @@ static int f2fs_xattr_fiemap(struct inode *inode, } if (xnid) { - struct folio *folio = f2fs_grab_cache_folio(NODE_MAPPING(sbi), - xnid, false); + struct f2fs_cached_block *entry = + f2fs_grab_node_cache(sbi, xnid); - if (IS_ERR(folio)) - return PTR_ERR(folio); + if (IS_ERR(entry)) + return PTR_ERR(entry); err = f2fs_get_node_info(sbi, xnid, &ni, false); if (err) { - f2fs_folio_put(folio, true); + f2fs_put_cache(entry, true); return err; } - phys = F2FS_BLK_TO_BYTES(ni.blk_addr); + phys = F2FS_BLK_TO_BYTES(sbi, ni.blk_addr); len = inode->i_sb->s_blocksize; - f2fs_folio_put(folio, true); + f2fs_put_cache(entry, true); flags = FIEMAP_EXTENT_LAST; } @@ -2043,6 +2261,7 @@ static int f2fs_xattr_fiemap(struct inode *inode, int f2fs_fiemap(struct inode *inode, struct fiemap_extent_info *fieinfo, u64 start, u64 len) { + struct f2fs_sb_info *sbi = F2FS_I_SB(inode); struct f2fs_map_blocks map; sector_t start_blk, last_blk, blk_len, max_len; pgoff_t next_pgofs; @@ -2066,7 +2285,7 @@ int f2fs_fiemap(struct inode *inode, struct fiemap_extent_info *fieinfo, inode_lock_shared(inode); - maxbytes = F2FS_BLK_TO_BYTES(max_file_blocks(inode)); + maxbytes = F2FS_BLK_TO_BYTES(sbi, max_file_blocks(sbi, inode)); if (start > maxbytes) { ret = -EFBIG; goto out; @@ -2086,10 +2305,10 @@ int f2fs_fiemap(struct inode *inode, struct fiemap_extent_info *fieinfo, goto out; } - start_blk = F2FS_BYTES_TO_BLK(start); - last_blk = F2FS_BYTES_TO_BLK(start + len - 1); + start_blk = F2FS_BYTES_TO_BLK(sbi, start); + last_blk = F2FS_BYTES_TO_BLK(sbi, start + len - 1); blk_len = last_blk - start_blk + 1; - max_len = F2FS_BYTES_TO_BLK(maxbytes) - start_blk; + max_len = F2FS_BYTES_TO_BLK(sbi, maxbytes) - start_blk; next: memset(&map, 0, sizeof(map)); @@ -2111,7 +2330,7 @@ next: if (!compr_cluster && !(map.m_flags & F2FS_MAP_FLAGS)) { start_blk = next_pgofs; - if (F2FS_BLK_TO_BYTES(start_blk) < maxbytes) + if (F2FS_BLK_TO_BYTES(sbi, start_blk) < maxbytes) goto prep_next; flags |= FIEMAP_EXTENT_LAST; @@ -2159,14 +2378,14 @@ skip_fill: } else if (compr_appended) { unsigned int appended_blks = cluster_size - count_in_cluster + 1; - size += F2FS_BLK_TO_BYTES(appended_blks); + size += F2FS_BLK_TO_BYTES(sbi, appended_blks); start_blk += appended_blks; compr_cluster = false; } else { - logical = F2FS_BLK_TO_BYTES(start_blk); + logical = F2FS_BLK_TO_BYTES(sbi, start_blk); phys = __is_valid_data_blkaddr(map.m_pblk) ? - F2FS_BLK_TO_BYTES(map.m_pblk) : 0; - size = F2FS_BLK_TO_BYTES(map.m_len); + F2FS_BLK_TO_BYTES(sbi, map.m_pblk) : 0; + size = F2FS_BLK_TO_BYTES(sbi, map.m_len); flags = 0; if (compr_cluster) { @@ -2174,13 +2393,13 @@ skip_fill: count_in_cluster += map.m_len; if (count_in_cluster == cluster_size) { compr_cluster = false; - size += F2FS_BLKSIZE; + size += F2FS_BLKSIZE(sbi); } } else if (map.m_flags & F2FS_MAP_DELALLOC) { flags = FIEMAP_EXTENT_UNWRITTEN; } - start_blk += F2FS_BYTES_TO_BLK(size); + start_blk += F2FS_BYTES_TO_BLK(sbi, size); } prep_next: @@ -2200,7 +2419,8 @@ out: static inline loff_t f2fs_readpage_limit(struct inode *inode) { if (IS_ENABLED(CONFIG_FS_VERITY) && IS_VERITY(inode)) - return F2FS_BLK_TO_BYTES(max_file_blocks(inode)); + return F2FS_BLK_TO_BYTES(F2FS_I_SB(inode), + max_file_blocks(F2FS_I_SB(inode), inode)); return i_size_read(inode); } @@ -2218,7 +2438,7 @@ static int f2fs_read_single_page(struct inode *inode, struct fsverity_info *vi, struct readahead_control *rac) { struct bio *bio = *bio_ret; - const unsigned int blocksize = F2FS_BLKSIZE; + const unsigned int blocksize = F2FS_BLKSIZE(F2FS_I_SB(inode)); sector_t block_in_file; sector_t last_block; sector_t last_block_in_file; @@ -2228,7 +2448,8 @@ static int f2fs_read_single_page(struct inode *inode, struct fsverity_info *vi, block_in_file = (sector_t)index; last_block = block_in_file + nr_pages; - last_block_in_file = F2FS_BYTES_TO_BLK(f2fs_readpage_limit(inode) + + last_block_in_file = F2FS_BYTES_TO_BLK(F2FS_I_SB(inode), + f2fs_readpage_limit(inode) + blocksize - 1); if (last_block > last_block_in_file) last_block = last_block_in_file; @@ -2304,9 +2525,9 @@ submit_and_realloc: if (!bio_add_folio(bio, folio, blocksize, 0)) goto submit_and_realloc; - inc_page_count(F2FS_I_SB(inode), F2FS_RD_DATA); + inc_cache_count(F2FS_I_SB(inode), F2FS_RD_DATA); f2fs_update_iostat(F2FS_I_SB(inode), NULL, FS_DATA_READ_IO, - F2FS_BLKSIZE); + F2FS_BLKSIZE(F2FS_I_SB(inode))); *last_block_in_bio = block_nr; out: *bio_ret = bio; @@ -2324,7 +2545,7 @@ int f2fs_read_multi_pages(struct compress_ctx *cc, struct bio **bio_ret, struct bio *bio = *bio_ret; unsigned int start_idx = cc->cluster_idx << cc->log_cluster_size; sector_t last_block_in_file; - const unsigned int blocksize = F2FS_BLKSIZE; + const unsigned int blocksize = F2FS_BLKSIZE(sbi); struct decompress_io_ctx *dic = NULL; struct extent_info ei = {}; bool from_dnode = true; @@ -2339,7 +2560,8 @@ int f2fs_read_multi_pages(struct compress_ctx *cc, struct bio **bio_ret, f2fs_bug_on(sbi, f2fs_cluster_is_empty(cc)); - last_block_in_file = F2FS_BYTES_TO_BLK(f2fs_readpage_limit(inode) + + last_block_in_file = F2FS_BYTES_TO_BLK(sbi, + f2fs_readpage_limit(inode) + blocksize - 1); /* get rid of pages beyond EOF */ @@ -2386,7 +2608,7 @@ skip_reading_dnode: for (i = 1; i < cc->cluster_size; i++) { block_t blkaddr; - blkaddr = from_dnode ? data_blkaddr(dn.inode, dn.node_folio, + blkaddr = from_dnode ? data_blkaddr(dn.inode, dn.node_entry, dn.ofs_in_node + i) : ei.blk + i - 1; @@ -2420,7 +2642,7 @@ skip_reading_dnode: block_t blkaddr; struct bio_post_read_ctx *ctx; - blkaddr = from_dnode ? data_blkaddr(dn.inode, dn.node_folio, + blkaddr = from_dnode ? data_blkaddr(dn.inode, dn.node_entry, dn.ofs_in_node + i + 1) : ei.blk + i; @@ -2455,8 +2677,9 @@ submit_and_realloc: ctx->enabled_steps |= STEP_DECOMPRESS; refcount_inc(&dic->refcnt); - inc_page_count(sbi, F2FS_RD_DATA); - f2fs_update_iostat(sbi, inode, FS_DATA_READ_IO, F2FS_BLKSIZE); + inc_cache_count(sbi, F2FS_RD_DATA); + f2fs_update_iostat(sbi, inode, FS_DATA_READ_IO, + F2FS_BLKSIZE(sbi)); *last_block_in_bio = blkaddr; } @@ -2630,14 +2853,14 @@ submit_and_realloc: */ f2fs_wait_on_block_writeback(inode, block_nr); - if (!bio_add_folio(bio, folio, F2FS_BLKSIZE, - offset << PAGE_SHIFT)) + if (!bio_add_folio(bio, folio, F2FS_BLKSIZE(F2FS_I_SB(inode)), + offset << PAGE_SHIFT)) goto submit_and_realloc; folio_in_bio = true; - inc_page_count(F2FS_I_SB(inode), F2FS_RD_DATA); + inc_cache_count(F2FS_I_SB(inode), F2FS_RD_DATA); f2fs_update_iostat(F2FS_I_SB(inode), NULL, FS_DATA_READ_IO, - F2FS_BLKSIZE); + F2FS_BLKSIZE(F2FS_I_SB(inode))); last_block_in_bio = block_nr; } trace_f2fs_read_folio(folio, DATA); @@ -2909,10 +3132,12 @@ bool f2fs_should_update_outplace(struct inode *inode, struct f2fs_io_info *fio) return true; if (f2fs_used_in_atomic_write(inode)) return true; - /* rewrite low ratio compress data w/ OPU mode to avoid fragmentation */ - if (f2fs_compressed_file(inode) && - F2FS_OPTION(sbi).compress_mode == COMPR_MODE_USER && - is_inode_flag_set(inode, FI_ENABLE_COMPRESS)) + /* + * rewrite low ratio compress data w/ OPU mode to avoid fragmentation. + * If IO comes from compressed write path and fallback to raw write, + * force out‑place to prevent metadata‑data inconsistency. + */ + if (f2fs_compressed_file(inode)) return true; /* swap file is migrating in aligned write mode */ @@ -3001,7 +3226,7 @@ got_it: goto out_writepage; } - /* wait for GCed page writeback via META_MAPPING */ + /* wait for GCed page writeback via generic cache */ if (fio->meta_gc) f2fs_wait_on_block_writeback(inode, fio->old_blkaddr); @@ -3410,7 +3635,7 @@ continue_unlock: if (folio_test_writeback(folio)) { if (wbc->sync_mode == WB_SYNC_NONE) goto continue_unlock; - f2fs_folio_wait_writeback(folio, DATA, true, true); + f2fs_folio_wait_writeback(folio, true, true); } if (!folio_clear_dirty_for_io(folio)) @@ -3492,8 +3717,8 @@ next: mapping->writeback_index = done_index; if (nwritten) - f2fs_submit_merged_write_cond(F2FS_M_SB(mapping), mapping->host, - NULL, 0, DATA); + f2fs_submit_merged_write_cond(F2FS_M_SB(mapping), + mapping->host, NULL); /* submit cached bio of IPU write */ if (bio) f2fs_submit_merged_ipu_write(sbi, &bio, NULL); @@ -3549,7 +3774,7 @@ static inline void update_skipped_write(struct f2fs_sb_info *sbi, if (is_sbi_flag_set(sbi, SBI_ENABLE_CHECKPOINT) && skipped && wbc->sync_mode == WB_SYNC_ALL) - atomic_add(skipped, &sbi->nr_pages[F2FS_SKIPPED_WRITE]); + atomic_add(skipped, &sbi->nr_caches[F2FS_SKIPPED_WRITE]); } static int __f2fs_write_data_pages(struct address_space *mapping, @@ -3572,7 +3797,7 @@ static int __f2fs_write_data_pages(struct address_space *mapping, if ((S_ISDIR(inode->i_mode) || IS_NOQUOTA(inode)) && wbc->sync_mode == WB_SYNC_NONE && - get_dirty_pages(inode) < nr_pages_to_skip(sbi, DATA) && + get_dirty_pages(inode) < nr_caches_to_skip(sbi, DATA) && f2fs_available_free_memory(sbi, DIRTY_DENTS)) goto skip_write; @@ -3671,7 +3896,7 @@ static int prepare_write_begin(struct f2fs_sb_info *sbi, pgoff_t index = folio->index; struct dnode_of_data dn; struct f2fs_lock_context lc; - struct folio *ifolio; + struct f2fs_cached_block *ientry; bool locked = false; int flag = F2FS_GET_BLOCK_PRE_AIO; int err = 0; @@ -3690,10 +3915,11 @@ static int prepare_write_begin(struct f2fs_sb_info *sbi, /* f2fs_lock_op avoids race between write CP and convert_inline_page */ if (f2fs_has_inline_data(inode)) { - if (pos + len > MAX_INLINE_DATA(inode)) + if (pos + len > MAX_INLINE_DATA(inode)) { flag = F2FS_GET_BLOCK_DEFAULT; - f2fs_map_lock(sbi, &lc, flag); - locked = true; + f2fs_map_lock(sbi, &lc, flag); + locked = true; + } } else if ((pos & PAGE_MASK) >= i_size_read(inode)) { f2fs_map_lock(sbi, &lc, flag); locked = true; @@ -3701,20 +3927,20 @@ static int prepare_write_begin(struct f2fs_sb_info *sbi, restart: /* check inline_data */ - ifolio = f2fs_get_inode_folio(sbi, inode->i_ino); - if (IS_ERR(ifolio)) { - err = PTR_ERR(ifolio); + ientry = f2fs_get_inode_cache(sbi, inode->i_ino); + if (IS_ERR(ientry)) { + err = PTR_ERR(ientry); goto unlock_out; } - set_new_dnode(&dn, inode, ifolio, ifolio, 0); + set_new_dnode(&dn, inode, ientry, ientry, 0); if (f2fs_has_inline_data(inode)) { if (pos + len <= MAX_INLINE_DATA(inode)) { - f2fs_do_read_inline_data(folio, ifolio); + f2fs_do_read_inline_data(folio, ientry); set_inode_flag(inode, FI_DATA_EXIST); if (inode->i_nlink) - folio_set_f2fs_inline(ifolio); + f2fs_cache_set_inline(ientry); goto out; } err = f2fs_convert_inline_folio(&dn, folio); @@ -3761,14 +3987,14 @@ static int __find_data_block(struct inode *inode, pgoff_t index, block_t *blk_addr) { struct dnode_of_data dn; - struct folio *ifolio; + struct f2fs_cached_block *ientry; int err = 0; - ifolio = f2fs_get_inode_folio(F2FS_I_SB(inode), inode->i_ino); - if (IS_ERR(ifolio)) - return PTR_ERR(ifolio); + ientry = f2fs_get_inode_cache(F2FS_I_SB(inode), inode->i_ino); + if (IS_ERR(ientry)) + return PTR_ERR(ientry); - set_new_dnode(&dn, inode, ifolio, ifolio, 0); + set_new_dnode(&dn, inode, ientry, ientry, 0); if (!f2fs_lookup_read_extent_cache_block(inode, index, &dn.data_blkaddr)) { @@ -3790,17 +4016,17 @@ static int __reserve_data_block(struct inode *inode, pgoff_t index, struct f2fs_sb_info *sbi = F2FS_I_SB(inode); struct dnode_of_data dn; struct f2fs_lock_context lc; - struct folio *ifolio; + struct f2fs_cached_block *ientry; int err = 0; f2fs_map_lock(sbi, &lc, F2FS_GET_BLOCK_PRE_AIO); - ifolio = f2fs_get_inode_folio(sbi, inode->i_ino); - if (IS_ERR(ifolio)) { - err = PTR_ERR(ifolio); + ientry = f2fs_get_inode_cache(sbi, inode->i_ino); + if (IS_ERR(ientry)) { + err = PTR_ERR(ientry); goto unlock_out; } - set_new_dnode(&dn, inode, ifolio, ifolio, 0); + set_new_dnode(&dn, inode, ientry, ientry, 0); if (!f2fs_lookup_read_extent_cache_block(dn.inode, index, &dn.data_blkaddr)) @@ -3956,7 +4182,7 @@ repeat: } } - f2fs_folio_wait_writeback(folio, DATA, false, true); + f2fs_folio_wait_writeback(folio, false, true); if (len == folio_size(folio) || folio_test_uptodate(folio)) return 0; @@ -4073,14 +4299,8 @@ void f2fs_invalidate_folio(struct folio *folio, size_t offset, size_t length) return; if (folio_test_dirty(folio)) { - if (inode->i_ino == F2FS_META_INO(sbi)) { - dec_page_count(sbi, F2FS_DIRTY_META); - } else if (inode->i_ino == F2FS_NODE_INO(sbi)) { - dec_page_count(sbi, F2FS_DIRTY_NODES); - } else { - inode_dec_dirty_pages(inode); - f2fs_remove_dirty_inode(inode); - } + inode_dec_dirty_pages(inode); + f2fs_remove_dirty_inode(inode); } if (offset || length != folio_size(folio)) @@ -4151,6 +4371,7 @@ static sector_t f2fs_bmap_compress(struct inode *inode, sector_t block) static sector_t f2fs_bmap(struct address_space *mapping, sector_t block) { struct inode *inode = mapping->host; + struct f2fs_sb_info *sbi = F2FS_I_SB(inode); sector_t blknr = 0; if (f2fs_has_inline_data(inode)) @@ -4161,7 +4382,7 @@ static sector_t f2fs_bmap(struct address_space *mapping, sector_t block) filemap_write_and_wait(mapping); /* Block number less than F2FS MAX BLOCKS */ - if (unlikely(block >= max_file_blocks(inode))) + if (unlikely(block >= max_file_blocks(sbi, inode))) goto out; if (f2fs_compressed_file(inode)) { @@ -4277,7 +4498,7 @@ static int check_swap_activate(struct swap_info_struct *sis, * to be very smart. */ cur_lblock = 0; - last_lblock = F2FS_BYTES_TO_BLK(i_size_read(inode)); + last_lblock = F2FS_BYTES_TO_BLK(sbi, i_size_read(inode)); while (cur_lblock < last_lblock && cur_lblock < sis->max) { struct f2fs_map_blocks map; @@ -4363,7 +4584,8 @@ retry: out: if (not_aligned) f2fs_warn(sbi, "Swapfile (%u) is not align to section: 1) creat(), 2) ioctl(F2FS_IOC_SET_PIN_FILE), 3) fallocate(%lu * N)", - not_aligned, blks_per_sec * F2FS_BLKSIZE); + not_aligned, + (unsigned long)blks_per_sec * F2FS_BLKSIZE(sbi)); return ret; } @@ -4532,12 +4754,14 @@ static int f2fs_iomap_begin(struct inode *inode, loff_t offset, loff_t length, unsigned int flags, struct iomap *iomap, struct iomap *srcmap) { + struct f2fs_sb_info *sbi = F2FS_I_SB(inode); struct f2fs_map_blocks map = { NULL, }; pgoff_t next_pgofs = 0; int err; - map.m_lblk = F2FS_BYTES_TO_BLK(offset); - map.m_len = F2FS_BYTES_TO_BLK(offset + length - 1) - map.m_lblk + 1; + map.m_lblk = F2FS_BYTES_TO_BLK(sbi, offset); + map.m_len = F2FS_BYTES_TO_BLK(sbi, offset + length - 1) - + map.m_lblk + 1; map.m_next_pgofs = &next_pgofs; map.m_seg_type = f2fs_rw_hint_to_seg_type(F2FS_I_SB(inode), inode->i_write_hint); @@ -4558,7 +4782,7 @@ static int f2fs_iomap_begin(struct inode *inode, loff_t offset, loff_t length, if (err) return err; - iomap->offset = F2FS_BLK_TO_BYTES(map.m_lblk); + iomap->offset = F2FS_BLK_TO_BYTES(sbi, map.m_lblk); /* * Sometimes I/O to an encrypted file has to be broken up to guarantee @@ -4578,11 +4802,11 @@ static int f2fs_iomap_begin(struct inode *inode, loff_t offset, loff_t length, if (WARN_ON_ONCE(map.m_pblk == NEW_ADDR)) return -EINVAL; - iomap->length = F2FS_BLK_TO_BYTES(map.m_len); + iomap->length = F2FS_BLK_TO_BYTES(sbi, map.m_len); iomap->type = IOMAP_MAPPED; iomap->flags |= IOMAP_F_MERGED; iomap->bdev = map.m_bdev; - iomap->addr = F2FS_BLK_TO_BYTES(map.m_pblk); + iomap->addr = F2FS_BLK_TO_BYTES(sbi, map.m_pblk); if (flags & IOMAP_WRITE && map.m_last_pblk) iomap->private = (void *)map.m_last_pblk; @@ -4591,11 +4815,11 @@ static int f2fs_iomap_begin(struct inode *inode, loff_t offset, loff_t length, return -ENOTBLK; if (map.m_pblk == NULL_ADDR) { - iomap->length = F2FS_BLK_TO_BYTES(next_pgofs) - + iomap->length = F2FS_BLK_TO_BYTES(sbi, next_pgofs) - iomap->offset; iomap->type = IOMAP_HOLE; } else if (map.m_pblk == NEW_ADDR) { - iomap->length = F2FS_BLK_TO_BYTES(map.m_len); + iomap->length = F2FS_BLK_TO_BYTES(sbi, map.m_len); iomap->type = IOMAP_UNWRITTEN; } else { f2fs_bug_on(F2FS_I_SB(inode), 1); diff --git a/fs/f2fs/debug.c b/fs/f2fs/debug.c index ff379aff4472..6fe606e2c70d 100644 --- a/fs/f2fs/debug.c +++ b/fs/f2fs/debug.c @@ -156,12 +156,12 @@ static void update_general_status(struct f2fs_sb_info *sbi) si->allocated_data_blocks = atomic64_read(&sbi->allocated_data_blocks); /* validation check of the segment numbers */ - si->ndirty_node = get_pages(sbi, F2FS_DIRTY_NODES); - si->ndirty_dent = get_pages(sbi, F2FS_DIRTY_DENTS); - si->ndirty_meta = get_pages(sbi, F2FS_DIRTY_META); - si->ndirty_data = get_pages(sbi, F2FS_DIRTY_DATA); - si->ndirty_qdata = get_pages(sbi, F2FS_DIRTY_QDATA); - si->ndirty_imeta = get_pages(sbi, F2FS_DIRTY_IMETA); + si->ndirty_node = get_nr_caches(sbi, F2FS_DIRTY_NODES); + si->ndirty_dent = get_nr_caches(sbi, F2FS_DIRTY_DENTS); + si->ndirty_meta = get_nr_caches(sbi, F2FS_DIRTY_META); + si->ndirty_data = get_nr_caches(sbi, F2FS_DIRTY_DATA); + si->ndirty_qdata = get_nr_caches(sbi, F2FS_DIRTY_QDATA); + si->ndirty_imeta = get_nr_caches(sbi, F2FS_DIRTY_IMETA); si->ndirty_dirs = sbi->ndirty_inode[DIR_INODE]; si->ndirty_files = sbi->ndirty_inode[FILE_INODE]; si->ndonate_files = sbi->donate_files; @@ -169,13 +169,13 @@ static void update_general_status(struct f2fs_sb_info *sbi) si->ndirty_all = sbi->ndirty_inode[DIRTY_META]; si->aw_cnt = atomic_read(&sbi->atomic_files); si->max_aw_cnt = atomic_read(&sbi->max_aw_cnt); - si->nr_dio_read = get_pages(sbi, F2FS_DIO_READ); - si->nr_dio_write = get_pages(sbi, F2FS_DIO_WRITE); - si->nr_wb_cp_data = get_pages(sbi, F2FS_WB_CP_DATA); - si->nr_wb_data = get_pages(sbi, F2FS_WB_DATA); - si->nr_rd_data = get_pages(sbi, F2FS_RD_DATA); - si->nr_rd_node = get_pages(sbi, F2FS_RD_NODE); - si->nr_rd_meta = get_pages(sbi, F2FS_RD_META); + si->nr_dio_read = get_nr_caches(sbi, F2FS_DIO_READ); + si->nr_dio_write = get_nr_caches(sbi, F2FS_DIO_WRITE); + si->nr_wb_cp_data = get_nr_caches(sbi, F2FS_WB_CP_DATA); + si->nr_wb_data = get_nr_caches(sbi, F2FS_WB_DATA); + si->nr_rd_data = get_nr_caches(sbi, F2FS_RD_DATA); + si->nr_rd_node = get_nr_caches(sbi, F2FS_RD_NODE); + si->nr_rd_meta = get_nr_caches(sbi, F2FS_RD_META); if (SM_I(sbi)->fcc_info) { si->nr_flushed = atomic_read(&SM_I(sbi)->fcc_info->issued_flush); @@ -222,13 +222,11 @@ static void update_general_status(struct f2fs_sb_info *sbi) si->free_secs = free_sections(sbi); si->prefree_count = prefree_segments(sbi); si->dirty_count = dirty_segments(sbi); - if (sbi->node_inode) - si->node_pages = NODE_MAPPING(sbi)->nrpages; - if (sbi->meta_inode) - si->meta_pages = META_MAPPING(sbi)->nrpages; + si->node_caches = NODE_CACHE(sbi)->num_entries; + si->meta_caches = META_CACHE(sbi)->num_entries; #ifdef CONFIG_F2FS_FS_COMPRESSION - if (sbi->compress_inode) { - si->compress_pages = COMPRESS_MAPPING(sbi)->nrpages; + if (test_opt(sbi, COMPRESS_CACHE)) { + si->compress_pages = COMPRESS_CACHE(sbi)->num_entries; si->compress_page_hit = atomic_read(&sbi->compress_page_hit); } #endif @@ -343,9 +341,9 @@ static void update_mem_info(struct f2fs_sb_info *sbi) /* build nm */ si->base_mem += sizeof(struct f2fs_nm_info); si->base_mem += __bitmap_size(sbi, NAT_BITMAP); - si->base_mem += F2FS_BLK_TO_BYTES(NM_I(sbi)->nat_bits_blocks); + si->base_mem += F2FS_BLK_TO_BYTES(sbi, NM_I(sbi)->nat_bits_blocks); si->base_mem += NM_I(sbi)->nat_blocks * - f2fs_bitmap_size(NAT_ENTRY_PER_BLOCK); + f2fs_bitmap_size(NAT_ENTRY_PER_BLOCK(sbi)); si->base_mem += NM_I(sbi)->nat_blocks / 8; si->base_mem += NM_I(sbi)->nat_blocks * sizeof(unsigned short); @@ -382,22 +380,36 @@ get_cache: si->cache_mem += si->ext_mem[i]; } - si->page_mem = 0; - if (sbi->node_inode) { - unsigned long npages = NODE_MAPPING(sbi)->nrpages; - - si->page_mem += (unsigned long long)npages << PAGE_SHIFT; - } - if (sbi->meta_inode) { - unsigned long npages = META_MAPPING(sbi)->nrpages; - - si->page_mem += (unsigned long long)npages << PAGE_SHIFT; - } + si->cache_entry_mem[F2FS_META_CACHE] = + (unsigned long long)META_CACHE(sbi)->num_entries * + sizeof(struct f2fs_cached_block); + si->cache_data_mem[F2FS_META_CACHE] = + (unsigned long long)META_CACHE(sbi)->num_entries * sbi->blocksize; + + si->cache_entry_mem[F2FS_NODE_CACHE] = + (unsigned long long)NODE_CACHE(sbi)->num_entries * + sizeof(struct f2fs_cached_block); + si->cache_data_mem[F2FS_NODE_CACHE] = + (unsigned long long)NODE_CACHE(sbi)->num_entries * sbi->blocksize; + + si->cache_mem += si->cache_entry_mem[F2FS_META_CACHE] + + si->cache_entry_mem[F2FS_NODE_CACHE]; + si->page_mem = si->cache_data_mem[F2FS_META_CACHE] + + si->cache_data_mem[F2FS_NODE_CACHE]; #ifdef CONFIG_F2FS_FS_COMPRESSION - if (sbi->compress_inode) { - unsigned long npages = COMPRESS_MAPPING(sbi)->nrpages; - - si->page_mem += (unsigned long long)npages << PAGE_SHIFT; + if (test_opt(sbi, COMPRESS_CACHE)) { + si->cache_entry_mem[F2FS_COMPRESS_CACHE] = + (unsigned long long)COMPRESS_CACHE(sbi)->num_entries * + sizeof(struct f2fs_cached_block); + si->cache_data_mem[F2FS_COMPRESS_CACHE] = + (unsigned long long)COMPRESS_CACHE(sbi)->num_entries * + sbi->blocksize; + + si->cache_mem += si->cache_entry_mem[F2FS_COMPRESS_CACHE]; + si->page_mem += si->cache_data_mem[F2FS_COMPRESS_CACHE]; + } else { + si->cache_entry_mem[F2FS_COMPRESS_CACHE] = 0; + si->cache_data_mem[F2FS_COMPRESS_CACHE] = 0; } #endif } @@ -700,7 +712,7 @@ static int stat_show(struct seq_file *s, void *v) si->aw_cnt, si->max_aw_cnt); seq_printf(s, " - compress: %4d, hit:%8d\n", si->compress_pages, si->compress_page_hit); seq_printf(s, " - nodes: %4d in %4d\n", - si->ndirty_node, si->node_pages); + si->ndirty_node, si->node_caches); seq_printf(s, " - dents: %4d in dirs:%4d (%4d)\n", si->ndirty_dent, si->ndirty_dirs, si->ndirty_all); seq_printf(s, " - data: %4d in files:%4d\n", @@ -708,7 +720,7 @@ static int stat_show(struct seq_file *s, void *v) seq_printf(s, " - quota data: %4d in quota files:%4d\n", si->ndirty_qdata, si->nquota_files); seq_printf(s, " - meta: %4d in %4d\n", - si->ndirty_meta, si->meta_pages); + si->ndirty_meta, si->meta_caches); seq_printf(s, " - imeta: %4d\n", si->ndirty_imeta); seq_printf(s, " - fsync mark: %4lld\n", @@ -756,6 +768,18 @@ static int stat_show(struct seq_file *s, void *v) si->ext_mem[EX_READ] >> 10); seq_printf(s, " - block age extent cache: %llu KB\n", si->ext_mem[EX_BLOCK_AGE] >> 10); + seq_printf(s, " - meta entry: %llu KB, meta cache: %llu KB\n", + si->cache_entry_mem[F2FS_META_CACHE] >> 10, + si->cache_data_mem[F2FS_META_CACHE] >> 10); + seq_printf(s, " - node entry: %llu KB, node cache: %llu KB\n", + si->cache_entry_mem[F2FS_NODE_CACHE] >> 10, + si->cache_data_mem[F2FS_NODE_CACHE] >> 10); +#ifdef CONFIG_F2FS_FS_COMPRESSION + if (test_opt(sbi, COMPRESS_CACHE)) + seq_printf(s, " - compress entry: %llu KB, compress cache: %llu KB\n", + si->cache_entry_mem[F2FS_COMPRESS_CACHE] >> 10, + si->cache_data_mem[F2FS_COMPRESS_CACHE] >> 10); +#endif seq_printf(s, " - paged : %llu KB\n", si->page_mem >> 10); } diff --git a/fs/f2fs/dir.c b/fs/f2fs/dir.c index fd0e2cd31a81..97f7fef6b068 100644 --- a/fs/f2fs/dir.c +++ b/fs/f2fs/dir.c @@ -195,7 +195,7 @@ static struct f2fs_dir_entry *find_in_block(struct inode *dir, int *max_slots, bool use_hash) { - struct f2fs_dentry_block *dentry_blk; + void *dentry_blk; struct f2fs_dentry_ptr d; dentry_blk = folio_address(dentry_folio); @@ -282,7 +282,7 @@ found: static struct f2fs_dir_entry *find_in_level(struct inode *dir, unsigned int level, const struct f2fs_filename *fname, - struct folio **res_folio, + void **dentry_block, bool use_hash) { int s = GET_DENTRY_SLOTS(fname->disk_name.len); @@ -313,7 +313,7 @@ start_find_bucket: bidx = next_pgofs; continue; } else { - *res_folio = dentry_folio; + *dentry_block = dentry_folio; break; } } @@ -321,11 +321,11 @@ start_find_bucket: de = find_in_block(dir, dentry_folio, fname, &max_slots, use_hash); if (IS_ERR(de)) { f2fs_folio_put(dentry_folio, false); - *res_folio = ERR_CAST(de); + *dentry_block = ERR_CAST(de); de = NULL; break; } else if (de) { - *res_folio = dentry_folio; + *dentry_block = dentry_folio; break; } @@ -352,7 +352,7 @@ start_find_bucket: struct f2fs_dir_entry *__f2fs_find_entry(struct inode *dir, const struct f2fs_filename *fname, - struct folio **res_folio) + void **dentry_block) { unsigned long npages = dir_blocks(dir); struct f2fs_dir_entry *de = NULL; @@ -360,13 +360,13 @@ struct f2fs_dir_entry *__f2fs_find_entry(struct inode *dir, unsigned int level; bool use_hash = true; - *res_folio = NULL; + *dentry_block = NULL; #if IS_ENABLED(CONFIG_UNICODE) start_find_entry: #endif if (f2fs_has_inline_dentry(dir)) { - de = f2fs_find_in_inline_dir(dir, fname, res_folio, use_hash); + de = f2fs_find_in_inline_dir(dir, fname, dentry_block, use_hash); goto out; } @@ -382,8 +382,8 @@ start_find_entry: } for (level = 0; level < max_depth; level++) { - de = find_in_level(dir, level, fname, res_folio, use_hash); - if (de || IS_ERR(*res_folio)) + de = find_in_level(dir, level, fname, dentry_block, use_hash); + if (de || IS_ERR(*dentry_block)) break; } @@ -408,7 +408,7 @@ out: * Entry is guaranteed to be valid. */ struct f2fs_dir_entry *f2fs_find_entry(struct inode *dir, - const struct qstr *child, struct folio **res_folio) + const struct qstr *child, void **dentry_block) { struct f2fs_dir_entry *de = NULL; struct f2fs_filename fname; @@ -417,67 +417,78 @@ struct f2fs_dir_entry *f2fs_find_entry(struct inode *dir, err = f2fs_setup_filename(dir, child, 1, &fname); if (err) { if (err == -ENOENT) - *res_folio = NULL; + *dentry_block = NULL; else - *res_folio = ERR_PTR(err); + *dentry_block = ERR_PTR(err); return NULL; } - de = __f2fs_find_entry(dir, &fname, res_folio); + de = __f2fs_find_entry(dir, &fname, dentry_block); f2fs_free_filename(&fname); return de; } -struct f2fs_dir_entry *f2fs_parent_dir(struct inode *dir, struct folio **f) +struct f2fs_dir_entry *f2fs_parent_dir(struct inode *dir, void **dentry_block) { - return f2fs_find_entry(dir, &dotdot_name, f); + return f2fs_find_entry(dir, &dotdot_name, dentry_block); } ino_t f2fs_inode_by_name(struct inode *dir, const struct qstr *qstr, - struct folio **folio) + void **dentry_block) { ino_t res = 0; struct f2fs_dir_entry *de; - de = f2fs_find_entry(dir, qstr, folio); + de = f2fs_find_entry(dir, qstr, dentry_block); if (de) { res = le32_to_cpu(de->ino); - f2fs_folio_put(*folio, false); + f2fs_put_dentry_block(*dentry_block, false); } return res; } void f2fs_set_link(struct inode *dir, struct f2fs_dir_entry *de, - struct folio *folio, struct inode *inode) + void *dentry_block, struct inode *inode) { - enum page_type type = f2fs_has_inline_dentry(dir) ? NODE : DATA; + if (f2fs_dentry_is_cache(dentry_block)) { + struct f2fs_cached_block *entry = + f2fs_dentry_cache(dentry_block); + + f2fs_lock_cache(entry); + f2fs_cache_wait_writeback(entry); + de->ino = cpu_to_le32(inode->i_ino); + de->file_type = fs_umode_to_ftype(inode->i_mode); + f2fs_mark_cache_dirty(entry); + } else { + struct folio *folio = f2fs_dentry_folio(dentry_block); - folio_lock(folio); - f2fs_folio_wait_writeback(folio, type, true, true); - de->ino = cpu_to_le32(inode->i_ino); - de->file_type = fs_umode_to_ftype(inode->i_mode); - folio_mark_dirty(folio); + folio_lock(folio); + f2fs_folio_wait_writeback(folio, true, true); + de->ino = cpu_to_le32(inode->i_ino); + de->file_type = fs_umode_to_ftype(inode->i_mode); + folio_mark_dirty(folio); + } inode_set_mtime_to_ts(dir, inode_set_ctime_current(dir)); f2fs_mark_inode_dirty_sync(dir, true); - f2fs_folio_put(folio, true); + f2fs_put_dentry_block(dentry_block, true); } static void init_dent_inode(struct inode *dir, struct inode *inode, const struct f2fs_filename *fname, - struct folio *ifolio) + struct f2fs_cached_block *ientry) { struct f2fs_inode *ri; if (!fname) /* tmpfile case? */ return; - f2fs_folio_wait_writeback(ifolio, NODE, true, true); + f2fs_cache_wait_writeback(ientry); /* copy name info. to this inode folio */ - ri = F2FS_INODE(ifolio); + ri = F2FS_INODE(ientry); ri->i_namelen = cpu_to_le32(fname->disk_name.len); memcpy(ri->i_name, fname->disk_name.name, fname->disk_name.len); if (IS_ENCRYPTED(dir)) { @@ -498,7 +509,7 @@ static void init_dent_inode(struct inode *dir, struct inode *inode, file_lost_pino(inode); } } - folio_mark_dirty(ifolio); + f2fs_mark_cache_dirty(ientry); } void f2fs_do_make_empty_dir(struct inode *inode, struct inode *parent, @@ -515,22 +526,22 @@ void f2fs_do_make_empty_dir(struct inode *inode, struct inode *parent, } static int make_empty_dir(struct inode *inode, - struct inode *parent, struct folio *folio) + struct inode *parent, struct f2fs_cached_block *ientry) { struct folio *dentry_folio; - struct f2fs_dentry_block *dentry_blk; + void *dentry_blk; struct f2fs_dentry_ptr d; if (f2fs_has_inline_dentry(inode)) - return f2fs_make_empty_inline_dir(inode, parent, folio); + return f2fs_make_empty_inline_dir(inode, parent, ientry); - dentry_folio = f2fs_get_new_data_folio(inode, folio, 0, true); + dentry_folio = f2fs_get_new_data_folio(inode, ientry, 0, true); if (IS_ERR(dentry_folio)) return PTR_ERR(dentry_folio); dentry_blk = folio_address(dentry_folio); - make_dentry_ptr_block(NULL, &d, dentry_blk); + make_dentry_ptr_block(inode, &d, dentry_blk); f2fs_do_make_empty_dir(inode, parent, &d); folio_mark_dirty(dentry_folio); @@ -538,50 +549,50 @@ static int make_empty_dir(struct inode *inode, return 0; } -struct folio *f2fs_init_inode_metadata(struct inode *inode, struct inode *dir, - const struct f2fs_filename *fname, struct folio *dfolio) +struct f2fs_cached_block *f2fs_init_inode_metadata(struct inode *inode, struct inode *dir, + const struct f2fs_filename *fname, struct f2fs_cached_block *dentry) { - struct folio *folio; + struct f2fs_cached_block *ientry; int err; if (is_inode_flag_set(inode, FI_NEW_INODE)) { - folio = f2fs_new_inode_folio(inode); - if (IS_ERR(folio)) - return folio; + ientry = f2fs_new_inode_cache(inode); + if (IS_ERR(ientry)) + return ientry; if (S_ISDIR(inode->i_mode)) { /* in order to handle error case */ - folio_get(folio); - err = make_empty_dir(inode, dir, folio); + f2fs_cache_get(ientry); + err = make_empty_dir(inode, dir, ientry); if (err) { - folio_lock(folio); + f2fs_lock_cache(ientry); goto put_error; } - folio_put(folio); + f2fs_put_cache(ientry, false); } - err = f2fs_init_acl(inode, dir, folio, dfolio); + err = f2fs_init_acl(inode, dir, ientry, dentry); if (err) goto put_error; err = f2fs_init_security(inode, dir, fname ? fname->usr_fname : NULL, - folio); + ientry); if (err) goto put_error; if (IS_ENCRYPTED(inode)) { - err = fscrypt_set_context(inode, folio); + err = fscrypt_set_context(inode, ientry); if (err) goto put_error; } } else { - folio = f2fs_get_inode_folio(F2FS_I_SB(dir), inode->i_ino); - if (IS_ERR(folio)) - return folio; + ientry = f2fs_get_inode_cache(F2FS_I_SB(dir), inode->i_ino); + if (IS_ERR(ientry)) + return ientry; } - init_dent_inode(dir, inode, fname, folio); + init_dent_inode(dir, inode, fname, ientry); /* * This file should be checkpointed during fsync. @@ -598,12 +609,12 @@ struct folio *f2fs_init_inode_metadata(struct inode *inode, struct inode *dir, f2fs_remove_orphan_inode(F2FS_I_SB(dir), inode->i_ino); f2fs_i_links_write(inode, true); } - return folio; + return ientry; put_error: clear_nlink(inode); - f2fs_update_inode(inode, folio); - f2fs_folio_put(folio, true); + f2fs_update_inode(inode, ientry); + f2fs_put_cache(ientry, true); return ERR_PTR(err); } @@ -645,14 +656,14 @@ next: goto next; } -bool f2fs_has_enough_room(struct inode *dir, struct folio *ifolio, +bool f2fs_has_enough_room(struct inode *dir, struct f2fs_cached_block *ientry, const struct f2fs_filename *fname) { struct f2fs_dentry_ptr d; unsigned int bit_pos; int slots = GET_DENTRY_SLOTS(fname->disk_name.len); - make_dentry_ptr_inline(dir, &d, inline_data_addr(dir, ifolio)); + make_dentry_ptr_inline(dir, &d, inline_data_addr(dir, ientry)); bit_pos = f2fs_room_for_filename(d.bitmap, slots, d.max); @@ -690,9 +701,9 @@ int f2fs_add_regular_entry(struct inode *dir, const struct f2fs_filename *fname, unsigned long bidx, block; unsigned int nbucket, nblock; struct folio *dentry_folio = NULL; - struct f2fs_dentry_block *dentry_blk = NULL; + void *dentry_blk = NULL; struct f2fs_dentry_ptr d; - struct folio *folio = NULL; + struct f2fs_cached_block *entry = NULL; int slots, err = 0; level = 0; @@ -727,9 +738,9 @@ start: return PTR_ERR(dentry_folio); dentry_blk = folio_address(dentry_folio); - bit_pos = f2fs_room_for_filename(&dentry_blk->dentry_bitmap, - slots, NR_DENTRY_IN_BLOCK); - if (bit_pos < NR_DENTRY_IN_BLOCK) + make_dentry_ptr_block(dir, &d, dentry_blk); + bit_pos = f2fs_room_for_filename(d.bitmap, slots, d.max); + if (bit_pos < d.max) goto add_dentry; f2fs_folio_put(dentry_folio, true); @@ -739,18 +750,17 @@ start: ++level; goto start; add_dentry: - f2fs_folio_wait_writeback(dentry_folio, DATA, true, true); + f2fs_folio_wait_writeback(dentry_folio, true, true); if (inode) { f2fs_down_write(&F2FS_I(inode)->i_sem); - folio = f2fs_init_inode_metadata(inode, dir, fname, NULL); - if (IS_ERR(folio)) { - err = PTR_ERR(folio); + entry = f2fs_init_inode_metadata(inode, dir, fname, NULL); + if (IS_ERR(entry)) { + err = PTR_ERR(entry); goto fail; } } - make_dentry_ptr_block(NULL, &d, dentry_blk); f2fs_update_dentry(ino, mode, &d, &fname->disk_name, fname->hash, bit_pos); @@ -761,9 +771,9 @@ add_dentry: /* synchronize inode page's data from inode cache */ if (is_inode_flag_set(inode, FI_NEW_INODE)) - f2fs_update_inode(inode, folio); + f2fs_update_inode(inode, entry); - f2fs_folio_put(folio, true); + f2fs_put_cache(entry, true); } f2fs_update_parent_metadata(dir, inode, current_depth); @@ -805,7 +815,7 @@ int f2fs_do_add_link(struct inode *dir, const struct qstr *name, struct inode *inode, nid_t ino, umode_t mode) { struct f2fs_filename fname; - struct folio *folio = NULL; + void *dentry_blk = NULL; struct f2fs_dir_entry *de = NULL; int err; @@ -821,14 +831,14 @@ int f2fs_do_add_link(struct inode *dir, const struct qstr *name, * consistency more. */ if (current != F2FS_I(dir)->task) { - de = __f2fs_find_entry(dir, &fname, &folio); + de = __f2fs_find_entry(dir, &fname, &dentry_blk); F2FS_I(dir)->task = NULL; } if (de) { - f2fs_folio_put(folio, false); + f2fs_put_dentry_block(dentry_blk, false); err = -EEXIST; - } else if (IS_ERR(folio)) { - err = PTR_ERR(folio); + } else if (IS_ERR(dentry_blk)) { + err = PTR_ERR(dentry_blk); } else { err = f2fs_add_dentry(dir, &fname, inode, ino, mode); } @@ -839,16 +849,16 @@ int f2fs_do_add_link(struct inode *dir, const struct qstr *name, int f2fs_do_tmpfile(struct inode *inode, struct inode *dir, struct f2fs_filename *fname) { - struct folio *folio; + struct f2fs_cached_block *ientry; int err = 0; f2fs_down_write(&F2FS_I(inode)->i_sem); - folio = f2fs_init_inode_metadata(inode, dir, fname, NULL); - if (IS_ERR(folio)) { - err = PTR_ERR(folio); + ientry = f2fs_init_inode_metadata(inode, dir, fname, NULL); + if (IS_ERR(ientry)) { + err = PTR_ERR(ientry); goto fail; } - f2fs_folio_put(folio, true); + f2fs_put_cache(ientry, true); clear_inode_flag(inode, FI_NEW_INODE); f2fs_update_time(F2FS_I_SB(inode), REQ_TIME); @@ -884,13 +894,15 @@ void f2fs_drop_nlink(struct inode *dir, struct inode *inode) * It only removes the dentry from the dentry page, corresponding name * entry in name page does not need to be touched during deletion. */ -void f2fs_delete_entry(struct f2fs_dir_entry *dentry, struct folio *folio, +void f2fs_delete_entry(struct f2fs_dir_entry *dentry, void *dentry_block, struct inode *dir, struct inode *inode) { - struct f2fs_dentry_block *dentry_blk; + void *dentry_blk; + struct f2fs_dentry_ptr d; + struct folio *folio; unsigned int bit_pos; int slots = GET_DENTRY_SLOTS(le16_to_cpu(dentry->name_len)); - pgoff_t index = folio->index; + pgoff_t index; int i; f2fs_update_time(F2FS_I_SB(dir), REQ_TIME); @@ -899,24 +911,27 @@ void f2fs_delete_entry(struct f2fs_dir_entry *dentry, struct folio *folio, f2fs_add_ino_entry(F2FS_I_SB(dir), dir->i_ino, TRANS_DIR_INO); if (f2fs_has_inline_dentry(dir)) - return f2fs_delete_inline_entry(dentry, folio, dir, inode); + return f2fs_delete_inline_entry(dentry, + f2fs_dentry_cache(dentry_block), dir, inode); + + folio = f2fs_dentry_folio(dentry_block); + index = folio->index; folio_lock(folio); - f2fs_folio_wait_writeback(folio, DATA, true, true); + f2fs_folio_wait_writeback(folio, true, true); dentry_blk = folio_address(folio); - bit_pos = dentry - dentry_blk->dentry; + make_dentry_ptr_block(dir, &d, dentry_blk); + bit_pos = dentry - d.dentry; for (i = 0; i < slots; i++) - __clear_bit_le(bit_pos + i, &dentry_blk->dentry_bitmap); + __clear_bit_le(bit_pos + i, d.bitmap); /* Let's check and deallocate this dentry page */ - bit_pos = find_next_bit_le(&dentry_blk->dentry_bitmap, - NR_DENTRY_IN_BLOCK, - 0); + bit_pos = find_next_bit_le(d.bitmap, d.max, 0); folio_mark_dirty(folio); - if (bit_pos == NR_DENTRY_IN_BLOCK && - !f2fs_truncate_hole(dir, index, index + 1)) { + if (bit_pos == d.max && + !f2fs_truncate_hole(dir, index, index + 1)) { f2fs_clear_page_cache_dirty_tag(folio); folio_clear_dirty_for_io(folio); folio_clear_uptodate(folio); @@ -938,7 +953,8 @@ bool f2fs_empty_dir(struct inode *dir) { unsigned long bidx = 0; unsigned int bit_pos; - struct f2fs_dentry_block *dentry_blk; + void *dentry_blk; + struct f2fs_dentry_ptr d; unsigned long nblock = dir_blocks(dir); if (f2fs_has_inline_dentry(dir)) @@ -959,17 +975,16 @@ bool f2fs_empty_dir(struct inode *dir) } dentry_blk = folio_address(dentry_folio); + make_dentry_ptr_block(dir, &d, dentry_blk); if (bidx == 0) bit_pos = 2; else bit_pos = 0; - bit_pos = find_next_bit_le(&dentry_blk->dentry_bitmap, - NR_DENTRY_IN_BLOCK, - bit_pos); + bit_pos = find_next_bit_le(d.bitmap, d.max, bit_pos); f2fs_folio_put(dentry_folio, false); - if (bit_pos < NR_DENTRY_IN_BLOCK) + if (bit_pos < d.max) return false; bidx++; @@ -1051,7 +1066,7 @@ int f2fs_fill_dentries(struct dir_context *ctx, struct f2fs_dentry_ptr *d, } if (readdir_ra) - f2fs_ra_node_page(sbi, le32_to_cpu(de->ino)); + f2fs_ra_node_cache(sbi, le32_to_cpu(de->ino)); ctx->pos = start_pos + bit_pos; found_valid_dirent = true; @@ -1066,10 +1081,12 @@ static int f2fs_readdir(struct file *file, struct dir_context *ctx) { struct inode *inode = file_inode(file); unsigned long npages = dir_blocks(inode); - struct f2fs_dentry_block *dentry_blk = NULL; + void *dentry_blk = NULL; struct file_ra_state *ra = &file->f_ra; loff_t start_pos = ctx->pos; - unsigned int n = ((unsigned long)ctx->pos / NR_DENTRY_IN_BLOCK); + struct f2fs_sb_info *sbi = F2FS_I_SB(inode); + unsigned int entries = sbi->dentries_per_block; + unsigned int n = (unsigned long)ctx->pos / entries; struct f2fs_dentry_ptr d; struct fscrypt_str fstr = FSTR_INIT(NULL, 0); int err = 0; @@ -1089,7 +1106,7 @@ static int f2fs_readdir(struct file *file, struct dir_context *ctx) goto out_free; } - for (; n < npages; ctx->pos = n * NR_DENTRY_IN_BLOCK) { + for (; n < npages; ctx->pos = n * entries) { struct folio *dentry_folio; pgoff_t next_pgofs; @@ -1122,7 +1139,7 @@ static int f2fs_readdir(struct file *file, struct dir_context *ctx) make_dentry_ptr_block(inode, &d, dentry_blk); err = f2fs_fill_dentries(ctx, &d, - n * NR_DENTRY_IN_BLOCK, &fstr); + n * entries, &fstr); f2fs_folio_put(dentry_folio, false); if (err) break; diff --git a/fs/f2fs/extent_cache.c b/fs/f2fs/extent_cache.c index 37cf9fa8d537..5a944428aca0 100644 --- a/fs/f2fs/extent_cache.c +++ b/fs/f2fs/extent_cache.c @@ -20,10 +20,10 @@ #include "segment.h" #include <trace/events/f2fs.h> -bool sanity_check_extent_cache(struct inode *inode, struct folio *ifolio) +bool sanity_check_extent_cache(struct inode *inode, struct f2fs_cached_block *ientry) { struct f2fs_sb_info *sbi = F2FS_I_SB(inode); - struct f2fs_extent *i_ext = &F2FS_INODE(ifolio)->i_ext; + struct f2fs_extent *i_ext = &F2FS_INODE(ientry)->i_ext; struct extent_info ei; int devi; @@ -416,11 +416,11 @@ static void __drop_largest_extent(struct extent_tree *et, } } -void f2fs_init_read_extent_tree(struct inode *inode, struct folio *ifolio) +void f2fs_init_read_extent_tree(struct inode *inode, struct f2fs_cached_block *ientry) { struct f2fs_sb_info *sbi = F2FS_I_SB(inode); struct extent_tree_info *eti = &sbi->extent_tree[EX_READ]; - struct f2fs_extent *i_ext = &F2FS_INODE(ifolio)->i_ext; + struct f2fs_extent *i_ext = &F2FS_INODE(ientry)->i_ext; struct extent_tree *et; struct extent_node *en; struct extent_info ei = {0}; @@ -428,9 +428,9 @@ void f2fs_init_read_extent_tree(struct inode *inode, struct folio *ifolio) if (!__may_extent_tree(inode, EX_READ)) { /* drop largest read extent */ if (i_ext->len) { - f2fs_folio_wait_writeback(ifolio, NODE, true, true); + f2fs_cache_wait_writeback(ientry); i_ext->len = 0; - folio_mark_dirty(ifolio); + f2fs_mark_cache_dirty(ientry); } set_inode_flag(inode, FI_NO_EXTENT); return; @@ -904,11 +904,12 @@ static int __get_new_block_age(struct inode *inode, struct extent_info *ei, struct extent_info tei = *ei; /* only fofs and len are valid */ /* - * When I/O is not aligned to a PAGE_SIZE, update will happen to the last - * file block even in seq write. So don't record age for newly last file - * block here. + * When I/O is not aligned to the filesystem block size, update will + * happen to the last file block even in seq write. Do not record the age + * of the new last file block here. */ - if ((f_size >> PAGE_SHIFT) == ei->fofs && f_size & (PAGE_SIZE - 1) && + if ((f_size >> sbi->log_blocksize) == ei->fofs && + f_size & F2FS_BLKSIZE_MASK(sbi) && blkaddr == NEW_ADDR) return -EINVAL; @@ -956,8 +957,8 @@ static void __update_extent_cache(struct dnode_of_data *dn, enum extent_type typ if (!__may_extent_tree(dn->inode, type)) return; - ei.fofs = f2fs_start_bidx_of_node(ofs_of_node(dn->node_folio), dn->inode) + - dn->ofs_in_node; + ei.fofs = f2fs_start_bidx_of_node(ofs_of_node(F2FS_I_SB(dn->inode), + dn->node_entry), dn->inode) + dn->ofs_in_node; ei.len = 1; if (type == EX_READ) { diff --git a/fs/f2fs/f2fs.h b/fs/f2fs/f2fs.h index 85937de3d701..dd9b4e4a88fc 100644 --- a/fs/f2fs/f2fs.h +++ b/fs/f2fs/f2fs.h @@ -222,6 +222,8 @@ struct f2fs_rwsem { #endif }; +#include "cache.h" + struct f2fs_mount_info { unsigned long long opt; block_t root_reserved_blocks; /* root reserved blocks */ @@ -416,7 +418,7 @@ struct inode_entry { struct fsync_node_entry { struct list_head list; /* list head */ - struct folio *folio; /* warm node folio pointer */ + struct f2fs_cached_block *entry; /* warm node cache entry pointer */ unsigned int seq_id; /* sequence id */ }; @@ -466,6 +468,7 @@ struct discard_entry { #define MAX_PLIST_NUM 512 #define plist_idx(blk_num) ((blk_num) >= MAX_PLIST_NUM ? \ (MAX_PLIST_NUM - 1) : ((blk_num) - 1)) +#define MAX_DISCARD_DROP_COUNT 512 enum { D_PREP, /* initial */ @@ -602,8 +605,9 @@ static inline int update_sits_in_cursum(struct f2fs_journal *journal, int i) #define DEF_INLINE_RESERVED_SIZE 1 static inline int get_extra_isize(struct inode *inode); static inline int get_inline_xattr_addrs(struct inode *inode); +static inline unsigned int cur_addrs_per_inode(struct inode *inode); #define MAX_INLINE_DATA(inode) (sizeof(__le32) * \ - (CUR_ADDRS_PER_INODE(inode) - \ + (cur_addrs_per_inode(inode) - \ get_inline_xattr_addrs(inode) - \ DEF_INLINE_RESERVED_SIZE)) @@ -669,17 +673,6 @@ struct f2fs_dentry_ptr { int nr_bitmap; }; -static inline void make_dentry_ptr_block(struct inode *inode, - struct f2fs_dentry_ptr *d, struct f2fs_dentry_block *t) -{ - d->inode = inode; - d->max = NR_DENTRY_IN_BLOCK; - d->nr_bitmap = SIZE_OF_DENTRY_BITMAP; - d->bitmap = t->dentry_bitmap; - d->dentry = t->dentry; - d->filename = t->filename; -} - static inline void make_dentry_ptr_inline(struct inode *inode, struct f2fs_dentry_ptr *d, void *t) { @@ -1081,7 +1074,7 @@ struct f2fs_nm_info { nid_t next_scan_nid; /* the next nid to be scanned */ nid_t max_rf_node_blocks; /* max # of nodes for recovery */ unsigned int ram_thresh; /* control the memory footprint */ - unsigned int ra_nid_pages; /* # of nid pages to be readaheaded */ + unsigned int ra_nid_blocks; /* # of nid blocks to be readaheaded */ unsigned int dirty_nats_ratio; /* control dirty nats ratio threshold */ /* NAT cache management */ @@ -1123,11 +1116,11 @@ struct f2fs_nm_info { */ struct dnode_of_data { struct inode *inode; /* vfs inode pointer */ - struct folio *inode_folio; /* its inode folio, NULL is possible */ - struct folio *node_folio; /* cached direct node folio */ + struct f2fs_cached_block *inode_entry; /* generic cache inode entry */ + struct f2fs_cached_block *node_entry; /* generic cache node entry */ nid_t nid; /* node id of the direct node block */ unsigned int ofs_in_node; /* data offset in the node page */ - bool inode_folio_locked; /* inode folio is locked or not */ + bool inode_entry_locked; /* inode entry is locked or not */ bool node_changed; /* is node block changed */ char cur_level; /* level of hole node page */ char max_level; /* level of current page located */ @@ -1135,12 +1128,12 @@ struct dnode_of_data { }; static inline void set_new_dnode(struct dnode_of_data *dn, struct inode *inode, - struct folio *ifolio, struct folio *nfolio, nid_t nid) + struct f2fs_cached_block *ientry, struct f2fs_cached_block *nentry, nid_t nid) { memset(dn, 0, sizeof(*dn)); dn->inode = inode; - dn->inode_folio = ifolio; - dn->node_folio = nfolio; + dn->inode_entry = ientry; + dn->node_entry = nentry; dn->nid = nid; } @@ -1368,8 +1361,10 @@ struct f2fs_io_info { unsigned int in_list:1; /* indicate fio is in io_list */ unsigned int is_por:1; /* indicate IO is from recovery or not */ unsigned int meta_gc:1; /* require meta inode GC */ + unsigned int is_cache:1; /* indicate IO is from internal cache */ enum iostat_type io_type; /* io type */ struct writeback_control *io_wbc; /* writeback control */ + struct f2fs_cached_block *cache_entry; struct bio **bio; /* bio for ipu */ sector_t *last_block; /* last block number in bio */ }; @@ -1616,10 +1611,9 @@ static inline void f2fs_clear_bit(unsigned int nr, char *addr); * | bit0 = 1 | bit1 | bit2 | ... | bit MAX | private data .... | * bit 0 F2FS_FOLIO_PRIVATE_NOT_POINTER * bit 1 F2FS_FOLIO_PRIVATE_ONGOING_MIGRATION - * bit 2 F2FS_FOLIO_PRIVATE_INLINE_INODE - * bit 3 F2FS_FOLIO_PRIVATE_REF_RESOURCE - * bit 4 F2FS_FOLIO_PRIVATE_ATOMIC_WRITE - * bit 5- f2fs private data + * bit 2 F2FS_FOLIO_PRIVATE_REF_RESOURCE + * bit 3 F2FS_FOLIO_PRIVATE_ATOMIC_WRITE + * bit 4- f2fs private data * * Layout B: lowest bit should be 0 * folio->private is a wrapped pointer. @@ -1627,7 +1621,6 @@ static inline void f2fs_clear_bit(unsigned int nr, char *addr); enum { F2FS_FOLIO_PRIVATE_NOT_POINTER, /* private contains non-pointer data */ F2FS_FOLIO_PRIVATE_ONGOING_MIGRATION, /* data page which is on-going migrating */ - F2FS_FOLIO_PRIVATE_INLINE_INODE, /* inode page contains inline data */ F2FS_FOLIO_PRIVATE_REF_RESOURCE, /* dirty page has referenced resources */ F2FS_FOLIO_PRIVATE_ATOMIC_WRITE, /* data page from atomic write path */ F2FS_FOLIO_PRIVATE_MAX @@ -1779,6 +1772,12 @@ struct f2fs_gc_kthread { unsigned int boost_gc_greedy; }; +struct f2fs_bio { + struct work_struct work; + struct f2fs_cached_block *entry; + struct bio bio; +}; + struct f2fs_sb_info { struct super_block *sb; /* pointer to VFS super block */ struct proc_dir_entry *s_proc; /* proc entry */ @@ -1798,7 +1797,6 @@ struct f2fs_sb_info { /* for node-related operations */ struct f2fs_nm_info *nm_info; /* node manager */ - struct inode *node_inode; /* cache node blocks */ /* for segment-related operations */ struct f2fs_sm_info *sm_info; /* segment manager */ @@ -1816,7 +1814,6 @@ struct f2fs_sb_info { struct f2fs_checkpoint *ckpt; /* raw checkpoint pointer */ int cur_cp_pack; /* remain current cp pack */ spinlock_t cp_lock; /* for flag in ckpt */ - struct inode *meta_inode; /* cache meta blocks */ struct f2fs_rwsem cp_global_sem; /* checkpoint procedure lock */ struct f2fs_rwsem cp_rwsem; /* blocking FS operations */ struct f2fs_rwsem node_write; /* locking node writes */ @@ -1859,9 +1856,17 @@ struct f2fs_sb_info { unsigned int log_sectors_per_block; /* log2 sectors per block */ unsigned int log_blocksize; /* log2 block size */ unsigned int blocksize; /* block size */ + unsigned int nat_entries_per_block; /* NAT entries in a block */ + unsigned int addrs_per_inode; /* addresses in an inode block */ + unsigned int addrs_per_block; /* addresses in a direct node block */ + unsigned int nids_per_block; /* node IDs in an indirect node block */ + unsigned int sit_entries_per_block; /* SIT entries in a block */ + unsigned int orphans_per_block; /* orphan inodes in a block */ + unsigned int dentries_per_block; /* dentries in a block */ + unsigned int dentry_bitmap_size; /* dentry bitmap size in bytes */ + unsigned int dentry_reserved_size; /* dentry reserved bytes */ unsigned int root_ino_num; /* root inode number*/ unsigned int node_ino_num; /* node inode number*/ - unsigned int meta_ino_num; /* meta inode number*/ unsigned int log_blocks_per_seg; /* log2 blocks per segment */ unsigned int blocks_per_seg; /* blocks per segment */ unsigned int segs_per_sec; /* segments per section */ @@ -1897,8 +1902,8 @@ struct f2fs_sb_info { struct f2fs_rwsem quota_sem; /* blocking cp for flags */ struct task_struct *umount_lock_holder; /* s_umount lock holder */ - /* # of pages, see count_type */ - atomic_t nr_pages[NR_COUNT_TYPE]; + /* # of cache entries, see count_type */ + atomic_t nr_caches[NR_COUNT_TYPE]; /* # of allocated blocks */ struct percpu_counter alloc_valid_block_count; /* # of node block writes as roll forward recovery */ @@ -2069,7 +2074,6 @@ struct f2fs_sb_info { u32 compr_new_inode; /* For compressed block cache */ - struct inode *compress_inode; /* cache compressed blocks */ unsigned int compress_percent; /* cache page percentage */ unsigned int compress_watermark; /* cache page watermark */ atomic_t compress_page_hit; /* cache hit count */ @@ -2094,6 +2098,14 @@ struct f2fs_sb_info { #ifdef CONFIG_DEBUG_LOCK_ALLOC struct lock_class_key cp_global_sem_key; #endif + + /* f2fs internal cache */ + struct f2fs_cached_block_list meta_blocks; + struct f2fs_cached_block_list node_blocks; + struct f2fs_cached_block_list compress_blocks; + + /* internal cache flush thread */ + struct f2fs_cache_kthread cache_thread; }; /* Definitions to access f2fs_sb_info */ @@ -2247,6 +2259,47 @@ static inline struct f2fs_sb_info *F2FS_F_SB(const struct folio *folio) return F2FS_M_SB(folio->mapping); } +#define SIT_ENTRY_PER_BLOCK(sbi) ((sbi)->sit_entries_per_block) +#define NAT_ENTRY_PER_BLOCK(sbi) ((sbi)->nat_entries_per_block) +#define DEF_ADDRS_PER_INODE(sbi) ((sbi)->addrs_per_inode) +#define DEF_ADDRS_PER_BLOCK(sbi) ((sbi)->addrs_per_block) +#define NIDS_PER_BLOCK(sbi) ((sbi)->nids_per_block) +#define F2FS_ORPHANS_PER_BLOCK(sbi) ((sbi)->orphans_per_block) +#define GET_ORPHAN_BLOCKS(sbi, n) DIV_ROUND_UP((n), \ + F2FS_ORPHANS_PER_BLOCK(sbi)) +#define CP_CHKSUM_OFFSET(sbi) (F2FS_BLKSIZE(sbi) - sizeof(__le32)) + +#define NODE_DIR1_BLOCK(sbi) (DEF_ADDRS_PER_INODE(sbi) + 1) +#define NODE_DIR2_BLOCK(sbi) (DEF_ADDRS_PER_INODE(sbi) + 2) +#define NODE_IND1_BLOCK(sbi) (DEF_ADDRS_PER_INODE(sbi) + 3) +#define NODE_IND2_BLOCK(sbi) (DEF_ADDRS_PER_INODE(sbi) + 4) +#define NODE_DIND_BLOCK(sbi) (DEF_ADDRS_PER_INODE(sbi) + 5) + +static inline struct f2fs_orphan_footer * +f2fs_orphan_footer(void *orphan_block, struct f2fs_sb_info *sbi) +{ + return (struct f2fs_orphan_footer *) + ((char *)orphan_block + sbi->blocksize - + sizeof(struct f2fs_orphan_footer)); +} + +static inline void make_dentry_ptr_block(struct inode *inode, + struct f2fs_dentry_ptr *d, void *t) +{ + struct f2fs_sb_info *sbi = F2FS_I_SB(inode); + unsigned int entries = sbi->dentries_per_block; + unsigned int bitmap_size = sbi->dentry_bitmap_size; + unsigned int reserved_size = sbi->dentry_reserved_size; + + d->inode = inode; + d->max = entries; + d->nr_bitmap = bitmap_size; + d->bitmap = t; + d->dentry = t + bitmap_size + reserved_size; + d->filename = t + bitmap_size + reserved_size + + SIZE_OF_DIR_ENTRY * entries; +} + static inline struct f2fs_super_block *F2FS_RAW_SUPER(struct f2fs_sb_info *sbi) { return (struct f2fs_super_block *)(sbi->raw_super); @@ -2267,14 +2320,30 @@ static inline struct f2fs_checkpoint *F2FS_CKPT(struct f2fs_sb_info *sbi) return (struct f2fs_checkpoint *)(sbi->ckpt); } -static inline struct f2fs_node *F2FS_NODE(const struct folio *folio) +static inline struct node_footer *F2FS_NODE_FOOTER(struct f2fs_sb_info *sbi, + const struct f2fs_cached_block *entry) +{ + return (struct node_footer *)(cache_address(entry) + + F2FS_BLKSIZE(sbi) - sizeof(struct node_footer)); +} + +static inline struct f2fs_node *F2FS_NODE(const struct f2fs_cached_block *entry) { - return (struct f2fs_node *)folio_address(folio); + return (struct f2fs_node *)CACHED_NODE(entry); } -static inline struct f2fs_inode *F2FS_INODE(const struct folio *folio) +static inline struct f2fs_inode *F2FS_INODE(const struct f2fs_cached_block *entry) { - return &((struct f2fs_node *)folio_address(folio))->i; + return &CACHED_NODE(entry)->i; +} + +static inline __le32 *F2FS_INODE_NIDS(struct f2fs_sb_info *sbi, + const struct f2fs_cached_block *entry) +{ + return (__le32 *)(cache_address(entry) + + F2FS_BLKSIZE(sbi) - + sizeof(struct node_footer) - + SIZE_OF_I_NID); } static inline struct f2fs_nm_info *NM_I(struct f2fs_sb_info *sbi) @@ -2302,24 +2371,29 @@ static inline struct dirty_seglist_info *DIRTY_I(struct f2fs_sb_info *sbi) return (struct dirty_seglist_info *)(SM_I(sbi)->dirty_info); } -static inline struct address_space *META_MAPPING(struct f2fs_sb_info *sbi) +static inline bool f2fs_is_meta_cache(struct f2fs_cached_block *entry) { - return sbi->meta_inode->i_mapping; + return entry->cache && entry->cache == META_CACHE(entry->cache->sbi); } -static inline struct address_space *NODE_MAPPING(struct f2fs_sb_info *sbi) +static inline bool f2fs_is_node_cache(struct f2fs_cached_block *entry) { - return sbi->node_inode->i_mapping; + return entry->cache && entry->cache == NODE_CACHE(entry->cache->sbi); } -static inline bool is_meta_folio(struct folio *folio) +static inline bool f2fs_is_compress_cache(struct f2fs_cached_block *entry) { - return folio->mapping == META_MAPPING(F2FS_F_SB(folio)); + return entry->cache && entry->cache == COMPRESS_CACHE(entry->cache->sbi); } -static inline bool is_node_folio(struct folio *folio) +static inline struct f2fs_bio *F2FS_BIO(struct bio *bio) { - return folio->mapping == NODE_MAPPING(F2FS_F_SB(folio)); + return container_of(bio, struct f2fs_bio, bio); +} + +static inline bool f2fs_is_cache_bio(struct bio *bio) +{ + return F2FS_BIO(bio)->entry != NULL; } static inline bool is_sbi_flag_set(struct f2fs_sb_info *sbi, unsigned int type) @@ -2563,7 +2637,8 @@ static inline int F2FS_HAS_BLOCKS(struct inode *inode) { block_t xattr_block = F2FS_I(inode)->i_xattr_nid ? 1 : 0; - return (inode->i_blocks >> F2FS_LOG_SECTORS_PER_BLOCK) > xattr_block; + return (inode->i_blocks >> + F2FS_LOG_SECTORS_PER_BLOCK(F2FS_I_SB(inode))) > xattr_block; } static inline bool f2fs_has_xattr_block(unsigned int ofs) @@ -2713,17 +2788,14 @@ static inline void folio_clear_f2fs_##name(struct folio *folio) \ } F2FS_FOLIO_PRIVATE_GET_FUNC(nonpointer, NOT_POINTER); -F2FS_FOLIO_PRIVATE_GET_FUNC(inline, INLINE_INODE); F2FS_FOLIO_PRIVATE_GET_FUNC(gcing, ONGOING_MIGRATION); F2FS_FOLIO_PRIVATE_GET_FUNC(atomic, ATOMIC_WRITE); F2FS_FOLIO_PRIVATE_SET_FUNC(reference, REF_RESOURCE); -F2FS_FOLIO_PRIVATE_SET_FUNC(inline, INLINE_INODE); F2FS_FOLIO_PRIVATE_SET_FUNC(gcing, ONGOING_MIGRATION); F2FS_FOLIO_PRIVATE_SET_FUNC(atomic, ATOMIC_WRITE); F2FS_FOLIO_PRIVATE_CLEAR_FUNC(reference, REF_RESOURCE); -F2FS_FOLIO_PRIVATE_CLEAR_FUNC(inline, INLINE_INODE); F2FS_FOLIO_PRIVATE_CLEAR_FUNC(gcing, ONGOING_MIGRATION); F2FS_FOLIO_PRIVATE_CLEAR_FUNC(atomic, ATOMIC_WRITE); @@ -2751,7 +2823,7 @@ static inline void dec_valid_block_count(struct f2fs_sb_info *sbi, struct inode *inode, block_t count) { - blkcnt_t sectors = count << F2FS_LOG_SECTORS_PER_BLOCK; + blkcnt_t sectors = count << F2FS_LOG_SECTORS_PER_BLOCK(sbi); spin_lock(&sbi->stat_lock); if (unlikely(sbi->total_valid_block_count < count)) { @@ -2778,9 +2850,9 @@ static inline void dec_valid_block_count(struct f2fs_sb_info *sbi, f2fs_i_blocks_write(inode, count, false, true); } -static inline void inc_page_count(struct f2fs_sb_info *sbi, int count_type) +static inline void inc_cache_count(struct f2fs_sb_info *sbi, int count_type) { - atomic_inc(&sbi->nr_pages[count_type]); + atomic_inc(&sbi->nr_caches[count_type]); if (count_type == F2FS_DIRTY_DENTS || count_type == F2FS_DIRTY_NODES || @@ -2793,15 +2865,15 @@ static inline void inc_page_count(struct f2fs_sb_info *sbi, int count_type) static inline void inode_inc_dirty_pages(struct inode *inode) { atomic_inc(&F2FS_I(inode)->dirty_pages); - inc_page_count(F2FS_I_SB(inode), S_ISDIR(inode->i_mode) ? + inc_cache_count(F2FS_I_SB(inode), S_ISDIR(inode->i_mode) ? F2FS_DIRTY_DENTS : F2FS_DIRTY_DATA); if (IS_NOQUOTA(inode)) - inc_page_count(F2FS_I_SB(inode), F2FS_DIRTY_QDATA); + inc_cache_count(F2FS_I_SB(inode), F2FS_DIRTY_QDATA); } -static inline void dec_page_count(struct f2fs_sb_info *sbi, int count_type) +static inline void dec_cache_count(struct f2fs_sb_info *sbi, int count_type) { - atomic_dec(&sbi->nr_pages[count_type]); + atomic_dec(&sbi->nr_caches[count_type]); } static inline void inode_dec_dirty_pages(struct inode *inode) @@ -2811,10 +2883,10 @@ static inline void inode_dec_dirty_pages(struct inode *inode) return; atomic_dec(&F2FS_I(inode)->dirty_pages); - dec_page_count(F2FS_I_SB(inode), S_ISDIR(inode->i_mode) ? + dec_cache_count(F2FS_I_SB(inode), S_ISDIR(inode->i_mode) ? F2FS_DIRTY_DENTS : F2FS_DIRTY_DATA); if (IS_NOQUOTA(inode)) - dec_page_count(F2FS_I_SB(inode), F2FS_DIRTY_QDATA); + dec_cache_count(F2FS_I_SB(inode), F2FS_DIRTY_QDATA); } static inline void inc_atomic_write_cnt(struct inode *inode) @@ -2839,9 +2911,9 @@ static inline void release_atomic_write_cnt(struct inode *inode) fi->atomic_write_cnt = 0; } -static inline s64 get_pages(struct f2fs_sb_info *sbi, int count_type) +static inline s64 get_nr_caches(struct f2fs_sb_info *sbi, int count_type) { - return atomic_read(&sbi->nr_pages[count_type]); + return atomic_read(&sbi->nr_caches[count_type]); } static inline int get_dirty_pages(struct inode *inode) @@ -2851,7 +2923,7 @@ static inline int get_dirty_pages(struct inode *inode) static inline int get_blocktype_secs(struct f2fs_sb_info *sbi, int block_type) { - return div_u64(get_pages(sbi, block_type) + BLKS_PER_SEC(sbi) - 1, + return div_u64(get_nr_caches(sbi, block_type) + BLKS_PER_SEC(sbi) - 1, BLKS_PER_SEC(sbi)); } @@ -2903,7 +2975,7 @@ static inline void *__bitmap_ptr(struct f2fs_sb_info *sbi, int flag) if (flag == NAT_BITMAP) return tmp_ptr; else - return (unsigned char *)ckpt + F2FS_BLKSIZE; + return (unsigned char *)ckpt + F2FS_BLKSIZE(sbi); } else { offset = (flag == NAT_BITMAP) ? le32_to_cpu(ckpt->sit_ver_bitmap_bytesize) : 0; @@ -3133,14 +3205,51 @@ static inline void f2fs_put_page(struct page *page, bool unlock) f2fs_folio_put(page_folio(page), unlock); } +#define F2FS_DENTRY_TAG_CACHE 1UL +#define F2FS_DENTRY_TAG_MASK 1UL + +static inline void *f2fs_cache_make_dentry_block(struct f2fs_cached_block *entry) +{ + if (IS_ERR_OR_NULL(entry)) + return entry; + return (void *)((unsigned long)entry | F2FS_DENTRY_TAG_CACHE); +} + +static inline bool f2fs_dentry_is_cache(void *dentry_block) +{ + return ((unsigned long)dentry_block & F2FS_DENTRY_TAG_MASK) == + F2FS_DENTRY_TAG_CACHE; +} + +static inline struct f2fs_cached_block *f2fs_dentry_cache(void *dentry_block) +{ + return (struct f2fs_cached_block *) + ((unsigned long)dentry_block & ~F2FS_DENTRY_TAG_MASK); +} + +static inline struct folio *f2fs_dentry_folio(void *dentry_block) +{ + return (struct folio *)dentry_block; +} + +static inline void f2fs_put_dentry_block(void *dentry_block, bool unlock) +{ + if (IS_ERR_OR_NULL(dentry_block)) + return; + if (f2fs_dentry_is_cache(dentry_block)) + f2fs_put_cache(f2fs_dentry_cache(dentry_block), unlock); + else + f2fs_folio_put(f2fs_dentry_folio(dentry_block), unlock); +} + static inline void f2fs_put_dnode(struct dnode_of_data *dn) { - if (dn->node_folio) - f2fs_folio_put(dn->node_folio, true); - if (dn->inode_folio && dn->node_folio != dn->inode_folio) - f2fs_folio_put(dn->inode_folio, false); - dn->node_folio = NULL; - dn->inode_folio = NULL; + if (dn->node_entry) + f2fs_put_cache(dn->node_entry, true); + if (dn->inode_entry && dn->node_entry != dn->inode_entry) + f2fs_put_cache(dn->inode_entry, false); + dn->node_entry = NULL; + dn->inode_entry = NULL; } static inline struct kmem_cache *f2fs_kmem_cache_create(const char *name, @@ -3174,11 +3283,13 @@ static inline void *f2fs_kmem_cache_alloc(struct kmem_cache *cachep, static inline bool is_inflight_io(struct f2fs_sb_info *sbi, int type) { - if (get_pages(sbi, F2FS_RD_DATA) || get_pages(sbi, F2FS_RD_NODE) || - get_pages(sbi, F2FS_RD_META) || get_pages(sbi, F2FS_WB_DATA) || - get_pages(sbi, F2FS_WB_CP_DATA) || - get_pages(sbi, F2FS_DIO_READ) || - get_pages(sbi, F2FS_DIO_WRITE)) + if (get_nr_caches(sbi, F2FS_RD_DATA) || + get_nr_caches(sbi, F2FS_RD_NODE) || + get_nr_caches(sbi, F2FS_RD_META) || + get_nr_caches(sbi, F2FS_WB_DATA) || + get_nr_caches(sbi, F2FS_WB_CP_DATA) || + get_nr_caches(sbi, F2FS_DIO_READ) || + get_nr_caches(sbi, F2FS_DIO_WRITE)) return true; if (type != DISCARD_TIME && SM_I(sbi) && SM_I(sbi)->dcc_info && @@ -3193,7 +3304,8 @@ static inline bool is_inflight_io(struct f2fs_sb_info *sbi, int type) static inline bool is_inflight_read_io(struct f2fs_sb_info *sbi) { - return get_pages(sbi, F2FS_RD_DATA) || get_pages(sbi, F2FS_DIO_READ); + return get_nr_caches(sbi, F2FS_RD_DATA) || + get_nr_caches(sbi, F2FS_DIO_READ); } static inline bool is_idle(struct f2fs_sb_info *sbi, int type) @@ -3229,13 +3341,11 @@ static inline void f2fs_radix_tree_insert(struct radix_tree_root *root, cond_resched(); } -#define RAW_IS_INODE(p) ((p)->footer.nid == (p)->footer.ino) - -static inline bool IS_INODE(const struct folio *folio) +static inline bool IS_INODE(struct f2fs_sb_info *sbi, const struct f2fs_cached_block *entry) { - struct f2fs_node *p = F2FS_NODE(folio); + struct node_footer *footer = F2FS_NODE_FOOTER(sbi, entry); - return RAW_IS_INODE(p); + return footer->nid == footer->ino; } static inline int offset_in_addr(struct f2fs_inode *i) @@ -3244,38 +3354,45 @@ static inline int offset_in_addr(struct f2fs_inode *i) (le16_to_cpu(i->i_extra_isize) / sizeof(__le32)) : 0; } -static inline __le32 *blkaddr_in_node(struct f2fs_node *node) +static inline __le32 *blkaddr_in_node(struct f2fs_sb_info *sbi, + const struct f2fs_cached_block *entry) { - return RAW_IS_INODE(node) ? node->i.i_addr : node->dn.addr; + struct f2fs_node *node = F2FS_NODE(entry); + + return IS_INODE(sbi, entry) ? node->i.i_addr : node->dn.addr; } static inline int f2fs_has_extra_attr(struct inode *inode); + static inline unsigned int get_dnode_base(struct inode *inode, - struct folio *node_folio) + const struct f2fs_cached_block *entry) { - if (!IS_INODE(node_folio)) + if (!IS_INODE(inode ? F2FS_I_SB(inode) : entry->cache->sbi, entry)) return 0; return inode ? get_extra_isize(inode) : - offset_in_addr(&F2FS_NODE(node_folio)->i); + offset_in_addr(&CACHED_NODE(entry)->i); } static inline __le32 *get_dnode_addr(struct inode *inode, - struct folio *node_folio) + const struct f2fs_cached_block *entry) { - return blkaddr_in_node(F2FS_NODE(node_folio)) + - get_dnode_base(inode, node_folio); + struct f2fs_sb_info *sbi = inode ? F2FS_I_SB(inode) : entry->cache->sbi; + + return blkaddr_in_node(sbi, entry) + + get_dnode_base(inode, entry); } static inline block_t data_blkaddr(struct inode *inode, - struct folio *node_folio, unsigned int offset) + const struct f2fs_cached_block *entry, + unsigned int offset) { - return le32_to_cpu(*(get_dnode_addr(inode, node_folio) + offset)); + return le32_to_cpu(*(get_dnode_addr(inode, entry) + offset)); } static inline block_t f2fs_data_blkaddr(struct dnode_of_data *dn) { - return data_blkaddr(dn->inode, dn->node_folio, dn->ofs_in_node); + return data_blkaddr(dn->inode, dn->node_entry, dn->ofs_in_node); } static inline int f2fs_test_bit(unsigned int nr, char *addr) @@ -3578,20 +3695,26 @@ static inline bool f2fs_need_compress_data(struct inode *inode) static inline unsigned int addrs_per_page(struct inode *inode, bool is_inode) { - unsigned int addrs = is_inode ? (CUR_ADDRS_PER_INODE(inode) - - get_inline_xattr_addrs(inode)) : DEF_ADDRS_PER_BLOCK; + unsigned int addrs = is_inode ? (cur_addrs_per_inode(inode) - + get_inline_xattr_addrs(inode)) : + DEF_ADDRS_PER_BLOCK(F2FS_I_SB(inode)); if (f2fs_compressed_file(inode)) return ALIGN_DOWN(addrs, F2FS_I(inode)->i_cluster_size); return addrs; } -static inline -void *inline_xattr_addr(struct inode *inode, const struct folio *folio) +static inline unsigned int cur_addrs_per_inode(struct inode *inode) { - struct f2fs_inode *ri = F2FS_INODE(folio); + return DEF_ADDRS_PER_INODE(F2FS_I_SB(inode)) - get_extra_isize(inode); +} + +static inline void *inline_xattr_addr(struct inode *inode, + const struct f2fs_cached_block *entry) +{ + struct f2fs_inode *ri = F2FS_INODE(entry); - return (void *)&(ri->i_addr[DEF_ADDRS_PER_INODE - + return (void *)&(ri->i_addr[DEF_ADDRS_PER_INODE(F2FS_I_SB(inode)) - get_inline_xattr_addrs(inode)]); } @@ -3636,9 +3759,10 @@ static inline bool f2fs_is_cow_file(struct inode *inode) return is_inode_flag_set(inode, FI_COW_FILE); } -static inline void *inline_data_addr(struct inode *inode, struct folio *folio) +static inline void *inline_data_addr(struct inode *inode, + const struct f2fs_cached_block *entry) { - __le32 *addr = get_dnode_addr(inode, folio); + __le32 *addr = get_dnode_addr(inode, entry); return (void *)(addr + DEF_INLINE_RESERVED_SIZE); } @@ -3827,9 +3951,9 @@ int f2fs_sync_file(struct file *file, loff_t start, loff_t end, int datasync); int f2fs_do_truncate_blocks(struct inode *inode, u64 from, bool lock); int f2fs_truncate_blocks(struct inode *inode, u64 from, bool lock); int f2fs_truncate(struct inode *inode); -int f2fs_getattr(struct mnt_idmap *idmap, const struct path *path, +int f2fs_getattr(const struct mnt_idmap *idmap, const struct path *path, struct kstat *stat, u32 request_mask, unsigned int flags); -int f2fs_setattr(struct mnt_idmap *idmap, struct dentry *dentry, +int f2fs_setattr(const struct mnt_idmap *idmap, struct dentry *dentry, struct iattr *attr); int f2fs_truncate_hole(struct inode *inode, pgoff_t pg_start, pgoff_t pg_end); void f2fs_truncate_data_blocks_range(struct dnode_of_data *dn, int count); @@ -3837,7 +3961,7 @@ int f2fs_do_shutdown(struct f2fs_sb_info *sbi, unsigned int flag, bool readonly, bool need_lock); int f2fs_precache_extents(struct inode *inode); int f2fs_fileattr_get(struct dentry *dentry, struct file_kattr *fa); -int f2fs_fileattr_set(struct mnt_idmap *idmap, +int f2fs_fileattr_set(const struct mnt_idmap *idmap, struct dentry *dentry, struct file_kattr *fa); long f2fs_ioctl(struct file *filp, unsigned int cmd, unsigned long arg); long f2fs_compat_ioctl(struct file *file, unsigned int cmd, unsigned long arg); @@ -3848,13 +3972,13 @@ int f2fs_pin_file_control(struct inode *inode, bool inc); * inode.c */ void f2fs_set_inode_flags(struct inode *inode); -bool f2fs_inode_chksum_verify(struct f2fs_sb_info *sbi, struct folio *folio); -void f2fs_inode_chksum_set(struct f2fs_sb_info *sbi, struct folio *folio); +bool f2fs_inode_chksum_verify(struct f2fs_sb_info *sbi, struct f2fs_cached_block *entry); +void f2fs_inode_chksum_set(struct f2fs_sb_info *sbi, struct f2fs_cached_block *entry); struct inode *f2fs_iget(struct super_block *sb, unsigned long ino); struct inode *f2fs_iget_retry(struct super_block *sb, unsigned long ino); int f2fs_try_to_free_nats(struct f2fs_sb_info *sbi, int nr_shrink); -void f2fs_update_inode(struct inode *inode, struct folio *node_folio); -void f2fs_update_inode_page(struct inode *inode); +void f2fs_update_inode(struct inode *inode, struct f2fs_cached_block *entry); +void f2fs_update_inode_cache(struct inode *inode); int f2fs_write_inode(struct inode *inode, struct writeback_control *wbc); void f2fs_remove_donate_inode(struct inode *inode); void f2fs_evict_inode(struct inode *inode); @@ -3869,7 +3993,7 @@ void f2fs_destroy_evict_inode_work(void); int f2fs_update_extension_list(struct f2fs_sb_info *sbi, const char *name, bool hot, bool set); struct dentry *f2fs_get_parent(struct dentry *child); -int f2fs_get_tmpfile(struct mnt_idmap *idmap, struct inode *dir, +int f2fs_get_tmpfile(const struct mnt_idmap *idmap, struct inode *dir, struct inode **new_inode); /* @@ -3903,22 +4027,22 @@ int f2fs_fill_dentries(struct dir_context *ctx, struct f2fs_dentry_ptr *d, unsigned int start_pos, struct fscrypt_str *fstr); void f2fs_do_make_empty_dir(struct inode *inode, struct inode *parent, struct f2fs_dentry_ptr *d); -struct folio *f2fs_init_inode_metadata(struct inode *inode, struct inode *dir, - const struct f2fs_filename *fname, struct folio *dfolio); +struct f2fs_cached_block *f2fs_init_inode_metadata(struct inode *inode, struct inode *dir, + const struct f2fs_filename *fname, struct f2fs_cached_block *dentry); void f2fs_update_parent_metadata(struct inode *dir, struct inode *inode, unsigned int current_depth); int f2fs_room_for_filename(const void *bitmap, int slots, int max_slots); void f2fs_drop_nlink(struct inode *dir, struct inode *inode); struct f2fs_dir_entry *__f2fs_find_entry(struct inode *dir, - const struct f2fs_filename *fname, struct folio **res_folio); + const struct f2fs_filename *fname, void **dentry_block); struct f2fs_dir_entry *f2fs_find_entry(struct inode *dir, - const struct qstr *child, struct folio **res_folio); -struct f2fs_dir_entry *f2fs_parent_dir(struct inode *dir, struct folio **f); + const struct qstr *child, void **dentry_block); +struct f2fs_dir_entry *f2fs_parent_dir(struct inode *dir, void **dentry_block); ino_t f2fs_inode_by_name(struct inode *dir, const struct qstr *qstr, - struct folio **folio); + void **dentry_block); void f2fs_set_link(struct inode *dir, struct f2fs_dir_entry *de, - struct folio *folio, struct inode *inode); -bool f2fs_has_enough_room(struct inode *dir, struct folio *ifolio, + void *dentry_blk, struct inode *inode); +bool f2fs_has_enough_room(struct inode *dir, struct f2fs_cached_block *ientry, const struct f2fs_filename *fname); void f2fs_update_dentry(nid_t ino, umode_t mode, struct f2fs_dentry_ptr *d, const struct fscrypt_str *name, f2fs_hash_t name_hash, @@ -3929,7 +4053,7 @@ int f2fs_add_dentry(struct inode *dir, const struct f2fs_filename *fname, struct inode *inode, nid_t ino, umode_t mode); int f2fs_do_add_link(struct inode *dir, const struct qstr *name, struct inode *inode, nid_t ino, umode_t mode); -void f2fs_delete_entry(struct f2fs_dir_entry *dentry, struct folio *folio, +void f2fs_delete_entry(struct f2fs_dir_entry *dentry, void *dentry_blk, struct inode *dir, struct inode *inode); int f2fs_do_tmpfile(struct inode *inode, struct inode *dir, struct f2fs_filename *fname); @@ -3951,7 +4075,7 @@ void f2fs_inode_synced(struct inode *inode); int f2fs_dquot_initialize(struct inode *inode); int f2fs_enable_quota_files(struct f2fs_sb_info *sbi, bool rdonly); int f2fs_do_quota_sync(struct super_block *sb, int type); -loff_t max_file_blocks(struct inode *inode); +loff_t max_file_blocks(struct f2fs_sb_info *sbi, struct inode *inode); void f2fs_quota_off_umount(struct super_block *sb); void f2fs_save_errors(struct f2fs_sb_info *sbi, unsigned char flag); void f2fs_handle_error(struct f2fs_sb_info *sbi, unsigned char error); @@ -3972,9 +4096,9 @@ enum node_type; int f2fs_check_nid_range(struct f2fs_sb_info *sbi, nid_t nid); bool f2fs_available_free_memory(struct f2fs_sb_info *sbi, int type); -bool f2fs_in_warm_node_list(struct folio *folio); +bool f2fs_in_warm_node_list(struct f2fs_sb_info *sbi, struct f2fs_cached_block *entry); void f2fs_init_fsync_node_info(struct f2fs_sb_info *sbi); -void f2fs_del_fsync_node_entry(struct f2fs_sb_info *sbi, struct folio *folio); +void f2fs_del_fsync_node_entry(struct f2fs_sb_info *sbi, struct f2fs_cached_block *entry); void f2fs_reset_fsync_node_info(struct f2fs_sb_info *sbi); bool f2fs_need_dentry_mark(struct f2fs_sb_info *sbi, nid_t nid); bool f2fs_is_checkpointed_node(struct f2fs_sb_info *sbi, nid_t nid); @@ -3985,37 +4109,37 @@ pgoff_t f2fs_get_next_page_offset(struct dnode_of_data *dn, pgoff_t pgofs); int f2fs_get_dnode_of_data(struct dnode_of_data *dn, pgoff_t index, int mode); int f2fs_truncate_inode_blocks(struct inode *inode, pgoff_t from); int f2fs_truncate_xattr_node(struct inode *inode); -int f2fs_wait_on_node_pages_writeback(struct f2fs_sb_info *sbi, +int f2fs_wait_on_node_caches_writeback(struct f2fs_sb_info *sbi, unsigned int seq_id); -int f2fs_remove_inode_page(struct inode *inode); -struct folio *f2fs_new_inode_folio(struct inode *inode); -struct folio *f2fs_new_node_folio(struct dnode_of_data *dn, unsigned int ofs); -void f2fs_ra_node_page(struct f2fs_sb_info *sbi, nid_t nid); -struct folio *f2fs_get_node_folio(struct f2fs_sb_info *sbi, pgoff_t nid, +int f2fs_write_node_caches(struct f2fs_sb_info *sbi); +int f2fs_remove_inode_cache(struct inode *inode); +struct f2fs_cached_block *f2fs_new_inode_cache(struct inode *inode); +struct f2fs_cached_block *f2fs_new_node_cache(struct dnode_of_data *dn, unsigned int ofs); +void f2fs_ra_node_cache(struct f2fs_sb_info *sbi, nid_t nid); +struct f2fs_cached_block *f2fs_get_node_cache(struct f2fs_sb_info *sbi, pgoff_t nid, enum node_type node_type); int f2fs_sanity_check_node_footer(struct f2fs_sb_info *sbi, - struct folio *folio, pgoff_t nid, + struct f2fs_cached_block *entry, pgoff_t nid, enum node_type ntype, bool in_irq); -struct folio *f2fs_get_inode_folio(struct f2fs_sb_info *sbi, pgoff_t ino); -struct folio *f2fs_get_xnode_folio(struct f2fs_sb_info *sbi, pgoff_t xnid); -int f2fs_write_single_node_folio(struct folio *node_folio, int sync_mode, +struct f2fs_cached_block *f2fs_get_inode_cache(struct f2fs_sb_info *sbi, pgoff_t ino); +struct f2fs_cached_block *f2fs_get_xnode_cache(struct f2fs_sb_info *sbi, pgoff_t xnid); +int f2fs_write_node_cache(struct f2fs_cached_block *entry, int sync_mode, bool mark_dirty, enum iostat_type io_type); -int f2fs_move_node_folio(struct folio *node_folio, int gc_type); +int f2fs_move_node_cache(struct f2fs_cached_block *entry, int gc_type); void f2fs_flush_inline_data(struct f2fs_sb_info *sbi); -int f2fs_fsync_node_pages(struct f2fs_sb_info *sbi, struct inode *inode, - struct writeback_control *wbc, bool atomic, - unsigned int *seq_id); -int f2fs_sync_node_pages(struct f2fs_sb_info *sbi, - struct writeback_control *wbc, - bool do_balance, enum iostat_type io_type); +int f2fs_fsync_node_caches(struct f2fs_sb_info *sbi, struct inode *inode, + bool atomic, unsigned int *seq_id); +int f2fs_writeback_node_caches(struct f2fs_sb_info *sbi, long nr_to_write, + bool sync, bool do_balance, enum iostat_type io_type); int f2fs_build_free_nids(struct f2fs_sb_info *sbi, bool sync, bool mount); bool f2fs_alloc_nid(struct f2fs_sb_info *sbi, nid_t *nid); void f2fs_alloc_nid_done(struct f2fs_sb_info *sbi, nid_t nid); void f2fs_alloc_nid_failed(struct f2fs_sb_info *sbi, nid_t nid); int f2fs_try_to_free_nids(struct f2fs_sb_info *sbi, int nr_shrink); -int f2fs_recover_inline_xattr(struct inode *inode, struct folio *folio); -int f2fs_recover_xattr_data(struct inode *inode, struct folio *folio); -int f2fs_recover_inode_page(struct f2fs_sb_info *sbi, struct folio *folio); +int f2fs_recover_inline_xattr(struct inode *inode, struct f2fs_cached_block *entry); +int f2fs_recover_xattr_data(struct inode *inode, struct f2fs_cached_block *entry); +int f2fs_recover_inode_cache(struct f2fs_sb_info *sbi, + struct f2fs_cached_block *entry); int f2fs_restore_node_summary(struct f2fs_sb_info *sbi, unsigned int segno, struct f2fs_summary_block *sum); int f2fs_flush_nat_entries(struct f2fs_sb_info *sbi, struct cp_control *cpc); @@ -4043,6 +4167,8 @@ void f2fs_reserve_device_alias(struct f2fs_sb_info *sbi, block_t addr, bool f2fs_is_checkpointed_data(struct f2fs_sb_info *sbi, block_t blkaddr); int f2fs_start_discard_thread(struct f2fs_sb_info *sbi); void f2fs_drop_discard_cmd(struct f2fs_sb_info *sbi); +void f2fs_drop_discard_cmd_range(struct f2fs_sb_info *sbi, + block_t start, block_t len); void f2fs_stop_discard_thread(struct f2fs_sb_info *sbi); bool f2fs_issue_discard_timeout(struct f2fs_sb_info *sbi, bool need_check); void f2fs_clear_prefree_segments(struct f2fs_sb_info *sbi, @@ -4051,7 +4177,7 @@ void f2fs_dirty_to_prefree(struct f2fs_sb_info *sbi); block_t f2fs_get_unusable_blocks(struct f2fs_sb_info *sbi); int f2fs_disable_cp_again(struct f2fs_sb_info *sbi, block_t unusable); void f2fs_release_discard_addrs(struct f2fs_sb_info *sbi); -int f2fs_npages_for_summary_flush(struct f2fs_sb_info *sbi, bool for_ra); +int f2fs_nblocks_for_summary_flush(struct f2fs_sb_info *sbi, bool for_ra); bool f2fs_segment_has_free_slot(struct f2fs_sb_info *sbi, int segno); int f2fs_init_inmem_curseg(struct f2fs_sb_info *sbi); int f2fs_reinit_atgc_curseg(struct f2fs_sb_info *sbi); @@ -4065,12 +4191,14 @@ int f2fs_allocate_new_segments(struct f2fs_sb_info *sbi); int f2fs_trim_fs(struct f2fs_sb_info *sbi, struct fstrim_range *range); bool f2fs_exist_trim_candidates(struct f2fs_sb_info *sbi, struct cp_control *cpc); -struct folio *f2fs_get_sum_folio(struct f2fs_sb_info *sbi, unsigned int segno); -void f2fs_update_meta_page(struct f2fs_sb_info *sbi, void *src, +struct f2fs_cached_block *f2fs_get_sum_cache(struct f2fs_sb_info *sbi, + unsigned int segno); +void f2fs_update_meta_block(struct f2fs_sb_info *sbi, void *src, block_t blk_addr); -void f2fs_do_write_meta_page(struct f2fs_sb_info *sbi, struct folio *folio, - enum iostat_type io_type); -void f2fs_do_write_node_page(unsigned int nid, struct f2fs_io_info *fio); +void f2fs_do_write_meta_cache(struct f2fs_sb_info *sbi, + struct f2fs_cached_block *entry, + enum iostat_type io_type); +void f2fs_do_write_node_cache(unsigned int nid, struct f2fs_io_info *fio); void f2fs_outplace_write_data(struct dnode_of_data *dn, struct f2fs_io_info *fio); int f2fs_inplace_write_data(struct f2fs_io_info *fio); @@ -4084,16 +4212,13 @@ void f2fs_replace_block(struct f2fs_sb_info *sbi, struct dnode_of_data *dn, bool recover_newaddr); enum temp_type f2fs_get_segment_temp(struct f2fs_sb_info *sbi, enum log_type seg_type); -int f2fs_allocate_data_block(struct f2fs_sb_info *sbi, struct folio *folio, +int f2fs_allocate_data_block(struct f2fs_sb_info *sbi, block_t old_blkaddr, block_t *new_blkaddr, struct f2fs_summary *sum, int type, struct f2fs_io_info *fio); void f2fs_update_device_state(struct f2fs_sb_info *sbi, nid_t ino, block_t blkaddr, unsigned int blkcnt); -void f2fs_folio_wait_writeback(struct folio *folio, enum page_type type, - bool ordered, bool locked); -#define f2fs_wait_on_page_writeback(page, type, ordered, locked) \ - f2fs_folio_wait_writeback(page_folio(page), type, ordered, locked) +void f2fs_folio_wait_writeback(struct folio *folio, bool ordered, bool locked); void f2fs_wait_on_block_writeback(struct inode *inode, block_t blkaddr); void f2fs_wait_on_block_writeback_range(struct inode *inode, block_t blkaddr, block_t len); @@ -4159,20 +4284,21 @@ void f2fs_unlock_op(struct f2fs_sb_info *sbi, struct f2fs_lock_context *lc); void f2fs_stop_checkpoint(struct f2fs_sb_info *sbi, bool end_io, unsigned char reason); void f2fs_flush_ckpt_thread(struct f2fs_sb_info *sbi); -struct folio *f2fs_grab_meta_folio(struct f2fs_sb_info *sbi, pgoff_t index); -struct folio *f2fs_get_meta_folio(struct f2fs_sb_info *sbi, pgoff_t index); -struct folio *f2fs_get_meta_folio_retry(struct f2fs_sb_info *sbi, pgoff_t index); -struct folio *f2fs_get_tmp_folio(struct f2fs_sb_info *sbi, pgoff_t index); +struct f2fs_cached_block *f2fs_grab_meta_cache(struct f2fs_sb_info *sbi, pgoff_t index); +struct f2fs_cached_block *f2fs_get_meta_cache(struct f2fs_sb_info *sbi, pgoff_t index); +struct f2fs_cached_block *f2fs_get_meta_cache_retry(struct f2fs_sb_info *sbi, pgoff_t index); +struct f2fs_cached_block *f2fs_get_tmp_cache(struct f2fs_sb_info *sbi, pgoff_t index); bool f2fs_is_valid_blkaddr(struct f2fs_sb_info *sbi, block_t blkaddr, int type); bool f2fs_is_valid_blkaddr_raw(struct f2fs_sb_info *sbi, block_t blkaddr, int type); -int f2fs_ra_meta_pages(struct f2fs_sb_info *sbi, block_t start, int nrpages, +int f2fs_ra_meta_caches(struct f2fs_sb_info *sbi, block_t start, int nrpages, int type, bool sync); -void f2fs_ra_meta_pages_cond(struct f2fs_sb_info *sbi, pgoff_t index, +void f2fs_ra_meta_caches_cond(struct f2fs_sb_info *sbi, pgoff_t index, unsigned int ra_blocks); -long f2fs_sync_meta_pages(struct f2fs_sb_info *sbi, long nr_to_write, - enum iostat_type io_type); +void f2fs_write_meta_caches(struct f2fs_sb_info *sbi); +long f2fs_sync_meta_caches(struct f2fs_sb_info *sbi, long nr_to_write, + bool sync, enum iostat_type io_type); void f2fs_add_ino_entry(struct f2fs_sb_info *sbi, nid_t ino, int type); void f2fs_remove_ino_entry(struct f2fs_sb_info *sbi, nid_t ino, int type); void f2fs_release_ino_entry(struct f2fs_sb_info *sbi, bool all); @@ -4191,7 +4317,7 @@ void f2fs_update_dirty_folio(struct inode *inode, struct folio *folio); void f2fs_remove_dirty_inode(struct inode *inode); int f2fs_sync_dirty_inodes(struct f2fs_sb_info *sbi, enum inode_type type, bool from_cp); -void f2fs_wait_on_all_pages(struct f2fs_sb_info *sbi, int type); +void f2fs_sync_dirty_data(struct f2fs_sb_info *sbi, int type); u64 f2fs_get_sectors_written(struct f2fs_sb_info *sbi); int f2fs_write_checkpoint(struct f2fs_sb_info *sbi, struct cp_control *cpc); void f2fs_init_ino_entry_info(struct f2fs_sb_info *sbi); @@ -4213,12 +4339,14 @@ void f2fs_destroy_bio_entry_cache(void); void f2fs_submit_read_bio(struct f2fs_sb_info *sbi, struct bio *bio, enum page_type type); int f2fs_init_write_merge_io(struct f2fs_sb_info *sbi); -void f2fs_submit_merged_write(struct f2fs_sb_info *sbi, enum page_type type); void f2fs_submit_merged_write_cond(struct f2fs_sb_info *sbi, - struct inode *inode, struct folio *folio, - nid_t ino, enum page_type type); + struct inode *inode, struct folio *folio); void f2fs_submit_merged_write_folio(struct f2fs_sb_info *sbi, - struct folio *folio, enum page_type type); + struct folio *folio); +bool f2fs_submit_merged_write_cache(struct f2fs_sb_info *sbi, + struct f2fs_cached_block *entry, + nid_t ino, enum page_type type); +void f2fs_submit_merged_write(struct f2fs_sb_info *sbi, enum page_type type); void f2fs_submit_merged_ipu_write(struct f2fs_sb_info *sbi, struct bio **bio, struct folio *folio); void f2fs_submit_all_merged_ipu_writes(struct f2fs_sb_info *sbi); @@ -4226,6 +4354,8 @@ void f2fs_flush_merged_writes(struct f2fs_sb_info *sbi); int f2fs_submit_page_bio(struct f2fs_io_info *fio); int f2fs_merge_page_bio(struct f2fs_io_info *fio); void f2fs_submit_page_write(struct f2fs_io_info *fio); +int f2fs_submit_cache_read(struct f2fs_io_info *fio); +void f2fs_submit_cache_write(struct f2fs_io_info *fio); struct block_device *f2fs_target_device(struct f2fs_sb_info *sbi, block_t blk_addr, sector_t *sector); int f2fs_target_device_index(struct f2fs_sb_info *sbi, block_t blkaddr); @@ -4242,7 +4372,7 @@ struct folio *f2fs_find_data_folio(struct inode *inode, pgoff_t index, struct folio *f2fs_get_lock_data_folio(struct inode *inode, pgoff_t index, bool for_write); struct folio *f2fs_get_new_data_folio(struct inode *inode, - struct folio *ifolio, pgoff_t index, bool new_i_size); + struct f2fs_cached_block *ientry, pgoff_t index, bool new_i_size); int f2fs_do_write_data_page(struct f2fs_io_info *fio); int f2fs_map_blocks(struct inode *inode, struct f2fs_map_blocks *map, int flag); int f2fs_fiemap(struct inode *inode, struct fiemap_extent_info *fieinfo, @@ -4354,7 +4484,7 @@ struct f2fs_stat_info { unsigned int bimodal, avg_vblocks; int util_free, util_valid, util_invalid; int rsvd_segs, overp_segs; - int dirty_count, node_pages, meta_pages, compress_pages; + int dirty_count, node_caches, meta_caches, compress_pages; int compress_page_hit; int prefree_count, free_segs, free_secs; int cp_call_count[MAX_CALL_TYPE], cp_count; @@ -4377,6 +4507,8 @@ struct f2fs_stat_info { unsigned int block_count[2]; unsigned int inplace_count; unsigned long long base_mem, cache_mem, page_mem; + unsigned long long cache_entry_mem[NR_CACHE_TYPES]; + unsigned long long cache_data_mem[NR_CACHE_TYPES]; struct f2fs_dev_stats *dev_stats; }; @@ -4555,8 +4687,6 @@ extern const struct file_operations f2fs_dir_operations; extern const struct file_operations f2fs_file_operations; extern const struct inode_operations f2fs_file_inode_operations; extern const struct address_space_operations f2fs_dblock_aops; -extern const struct address_space_operations f2fs_node_aops; -extern const struct address_space_operations f2fs_meta_aops; extern const struct inode_operations f2fs_dir_inode_operations; extern const struct inode_operations f2fs_symlink_inode_operations; extern const struct inode_operations f2fs_encrypted_symlink_inode_operations; @@ -4567,26 +4697,27 @@ extern struct kmem_cache *f2fs_inode_entry_slab; * inline.c */ bool f2fs_may_inline_data(struct inode *inode); -bool f2fs_sanity_check_inline_data(struct inode *inode, struct folio *ifolio); +bool f2fs_sanity_check_inline_data(struct inode *inode, struct f2fs_cached_block *ientry); bool f2fs_may_inline_dentry(struct inode *inode); -void f2fs_do_read_inline_data(struct folio *folio, struct folio *ifolio); -void f2fs_truncate_inline_inode(struct inode *inode, struct folio *ifolio, +void f2fs_do_read_inline_data(struct folio *folio, struct f2fs_cached_block *ientry); +void f2fs_truncate_inline_inode(struct inode *inode, struct f2fs_cached_block *ientry, u64 from); int f2fs_read_inline_data(struct inode *inode, struct folio *folio); int f2fs_convert_inline_folio(struct dnode_of_data *dn, struct folio *folio); int f2fs_convert_inline_inode(struct inode *inode); int f2fs_try_convert_inline_dir(struct inode *dir, struct dentry *dentry); int f2fs_write_inline_data(struct inode *inode, struct folio *folio); -int f2fs_recover_inline_data(struct inode *inode, struct folio *nfolio); +int f2fs_recover_inline_data(struct inode *inode, struct f2fs_cached_block *entry); struct f2fs_dir_entry *f2fs_find_in_inline_dir(struct inode *dir, - const struct f2fs_filename *fname, struct folio **res_folio, + const struct f2fs_filename *fname, void **dentry_block, bool use_hash); int f2fs_make_empty_inline_dir(struct inode *inode, struct inode *parent, - struct folio *ifolio); + struct f2fs_cached_block *ientry); int f2fs_add_inline_entry(struct inode *dir, const struct f2fs_filename *fname, struct inode *inode, nid_t ino, umode_t mode); void f2fs_delete_inline_entry(struct f2fs_dir_entry *dentry, - struct folio *folio, struct inode *dir, struct inode *inode); + struct f2fs_cached_block *ientry, struct inode *dir, + struct inode *inode); bool f2fs_empty_inline_dir(struct inode *dir); int f2fs_read_inline_dir(struct file *file, struct dir_context *ctx, struct fscrypt_str *fstr); @@ -4609,7 +4740,7 @@ void f2fs_leave_shrinker(struct f2fs_sb_info *sbi); /* * extent_cache.c */ -bool sanity_check_extent_cache(struct inode *inode, struct folio *ifolio); +bool sanity_check_extent_cache(struct inode *inode, struct f2fs_cached_block *ientry); void f2fs_init_extent_tree(struct inode *inode); void f2fs_drop_extent_tree(struct inode *inode); void f2fs_destroy_extent_node(struct inode *inode); @@ -4619,7 +4750,7 @@ int __init f2fs_create_extent_cache(void); void f2fs_destroy_extent_cache(void); /* read extent cache ops */ -void f2fs_init_read_extent_tree(struct inode *inode, struct folio *ifolio); +void f2fs_init_read_extent_tree(struct inode *inode, struct f2fs_cached_block *ientry); bool f2fs_lookup_read_extent_cache(struct inode *inode, pgoff_t pgofs, struct extent_info *ei); bool f2fs_lookup_read_extent_cache_block(struct inode *inode, pgoff_t index, @@ -4741,13 +4872,11 @@ unsigned int f2fs_cluster_blocks_are_contiguous(struct dnode_of_data *dn, int f2fs_init_compress_ctx(struct compress_ctx *cc); void f2fs_destroy_compress_ctx(struct compress_ctx *cc, bool reuse); void f2fs_init_compress_info(struct f2fs_sb_info *sbi); -int f2fs_init_compress_inode(struct f2fs_sb_info *sbi); -void f2fs_destroy_compress_inode(struct f2fs_sb_info *sbi); +void f2fs_init_compress_cache_context(struct f2fs_sb_info *sbi); int f2fs_init_page_array_cache(struct f2fs_sb_info *sbi); void f2fs_destroy_page_array_cache(struct f2fs_sb_info *sbi); int __init f2fs_init_compress_cache(void); void f2fs_destroy_compress_cache(void); -struct address_space *COMPRESS_MAPPING(struct f2fs_sb_info *sbi); void f2fs_invalidate_compress_pages_range(struct f2fs_sb_info *sbi, block_t blkaddr, unsigned int len); bool f2fs_load_compressed_folio(struct f2fs_sb_info *sbi, struct folio *folio, @@ -4796,8 +4925,7 @@ static inline void f2fs_put_folio_dic(struct folio *folio, bool in_task) static inline unsigned int f2fs_cluster_blocks_are_contiguous( struct dnode_of_data *dn, unsigned int ofs_in_node) { return 0; } static inline bool f2fs_sanity_check_cluster(struct dnode_of_data *dn) { return false; } -static inline int f2fs_init_compress_inode(struct f2fs_sb_info *sbi) { return 0; } -static inline void f2fs_destroy_compress_inode(struct f2fs_sb_info *sbi) { } +static inline void f2fs_init_compress_cache_context(struct f2fs_sb_info *sbi) { } static inline int f2fs_init_page_array_cache(struct f2fs_sb_info *sbi) { return 0; } static inline void f2fs_destroy_page_array_cache(struct f2fs_sb_info *sbi) { } static inline int __init f2fs_init_compress_cache(void) { return 0; } @@ -5137,10 +5265,8 @@ static inline void f2fs_schedule_timeout_killable(long timeout, bool io) } static inline void f2fs_handle_page_eio(struct f2fs_sb_info *sbi, - struct folio *folio, enum page_type type) + pgoff_t ofs, enum page_type type) { - pgoff_t ofs = folio->index; - if (unlikely(f2fs_cp_error(sbi))) return; @@ -5175,36 +5301,10 @@ static inline bool f2fs_is_readonly(struct f2fs_sb_info *sbi) return f2fs_sb_has_readonly(sbi) || f2fs_readonly(sbi->sb); } -static inline void f2fs_truncate_meta_inode_pages(struct f2fs_sb_info *sbi, - block_t blkaddr, unsigned int cnt) -{ - bool need_submit = false; - int i = 0; - - do { - struct folio *folio; - - folio = filemap_get_folio(META_MAPPING(sbi), blkaddr + i); - if (!IS_ERR(folio)) { - if (folio_test_writeback(folio)) - need_submit = true; - f2fs_folio_put(folio, false); - } - } while (++i < cnt && !need_submit); - - if (need_submit) - f2fs_submit_merged_write_cond(sbi, sbi->meta_inode, - NULL, 0, DATA); - - truncate_inode_pages_range(META_MAPPING(sbi), - F2FS_BLK_TO_BYTES((loff_t)blkaddr), - F2FS_BLK_END_BYTES((loff_t)(blkaddr + cnt - 1))); -} - static inline void f2fs_invalidate_internal_cache(struct f2fs_sb_info *sbi, block_t blkaddr, unsigned int len) { - f2fs_truncate_meta_inode_pages(sbi, blkaddr, len); + f2fs_truncate_meta_caches(sbi, blkaddr, len); f2fs_invalidate_compress_pages_range(sbi, blkaddr, len); } diff --git a/fs/f2fs/file.c b/fs/f2fs/file.c index edc352569e87..626c6f97b6f8 100644 --- a/fs/f2fs/file.c +++ b/fs/f2fs/file.c @@ -110,7 +110,8 @@ static vm_fault_t f2fs_filemap_fault(struct vm_fault *vmf) ret = filemap_fault(vmf); if (ret & VM_FAULT_LOCKED) f2fs_update_iostat(F2FS_I_SB(inode), inode, - APP_MAPPED_READ_IO, F2FS_BLKSIZE); + APP_MAPPED_READ_IO, + F2FS_BLKSIZE(F2FS_I_SB(inode))); trace_f2fs_filemap_fault(inode, vmf->pgoff, flags, ret); @@ -212,9 +213,9 @@ static vm_fault_t f2fs_vm_page_mkwrite(struct vm_fault *vmf) goto out_sem; } - f2fs_folio_wait_writeback(folio, DATA, false, true); + f2fs_folio_wait_writeback(folio, false, true); - /* wait for GCed page writeback via META_MAPPING */ + /* wait for GCed page writeback via generic cache */ f2fs_wait_on_block_writeback(inode, dn.data_blkaddr); /* @@ -233,7 +234,7 @@ static vm_fault_t f2fs_vm_page_mkwrite(struct vm_fault *vmf) } folio_mark_dirty(folio); - f2fs_update_iostat(sbi, inode, APP_MAPPED_IO, F2FS_BLKSIZE); + f2fs_update_iostat(sbi, inode, APP_MAPPED_IO, F2FS_BLKSIZE(sbi)); f2fs_update_time(sbi, REQ_TIME); out_sem: @@ -307,13 +308,13 @@ static inline enum cp_reason_type need_do_checkpoint(struct inode *inode) static bool need_inode_page_update(struct f2fs_sb_info *sbi, nid_t ino) { - struct folio *i = filemap_get_folio(NODE_MAPPING(sbi), ino); + struct f2fs_cached_block *entry = f2fs_find_node_cache(sbi, ino); bool ret = false; /* But we need to avoid that there are some inode updates */ - if ((!IS_ERR(i) && folio_test_dirty(i)) || + if ((!IS_ERR(entry) && f2fs_cache_test_dirty(entry)) || f2fs_need_inode_block_update(sbi, ino)) ret = true; - f2fs_folio_put(i, false); + f2fs_put_cache(entry, false); return ret; } @@ -339,10 +340,6 @@ static int f2fs_do_sync_file(struct file *file, loff_t start, loff_t end, nid_t ino = inode->i_ino; int ret = 0; enum cp_reason_type cp_reason = 0; - struct writeback_control wbc = { - .sync_mode = WB_SYNC_ALL, - .nr_to_write = LONG_MAX, - }; unsigned int seq_id = 0; if (unlikely(f2fs_readonly(inode->i_sb))) @@ -421,7 +418,7 @@ go_write: } sync_nodes: atomic_inc(&sbi->wb_sync_req[NODE]); - ret = f2fs_fsync_node_pages(sbi, inode, &wbc, atomic, &seq_id); + ret = f2fs_fsync_node_caches(sbi, inode, atomic, &seq_id); atomic_dec(&sbi->wb_sync_req[NODE]); if (ret) goto out; @@ -447,7 +444,7 @@ sync_nodes: * given fsync mark. */ if (!atomic) { - ret = f2fs_wait_on_node_pages_writeback(sbi, seq_id); + ret = f2fs_wait_on_node_caches_writeback(sbi, seq_id); if (ret) goto out; } @@ -484,7 +481,7 @@ static bool __found_offset(struct address_space *mapping, bool compressed_cluster = false; if (f2fs_compressed_file(inode)) { - block_t first_blkaddr = data_blkaddr(dn->inode, dn->node_folio, + block_t first_blkaddr = data_blkaddr(dn->inode, dn->node_entry, ALIGN_DOWN(dn->ofs_in_node, F2FS_I(inode)->i_cluster_size)); compressed_cluster = first_blkaddr == COMPRESS_ADDR; @@ -513,7 +510,8 @@ static bool __found_offset(struct address_space *mapping, static loff_t f2fs_seek_block(struct file *file, loff_t offset, int whence) { struct inode *inode = file->f_mapping->host; - loff_t maxbytes = F2FS_BLK_TO_BYTES(max_file_blocks(inode)); + loff_t maxbytes = F2FS_BLK_TO_BYTES(F2FS_I_SB(inode), + max_file_blocks(F2FS_I_SB(inode), inode)); struct dnode_of_data dn; pgoff_t pgofs, end_offset; loff_t data_ofs = offset; @@ -537,9 +535,10 @@ static loff_t f2fs_seek_block(struct file *file, loff_t offset, int whence) } } - pgofs = (pgoff_t)(offset >> PAGE_SHIFT); + pgofs = F2FS_BYTES_TO_BLK(F2FS_I_SB(inode), offset); - for (; data_ofs < isize; data_ofs = (loff_t)pgofs << PAGE_SHIFT) { + for (; data_ofs < isize; + data_ofs = F2FS_BLK_TO_BYTES(F2FS_I_SB(inode), pgofs)) { set_new_dnode(&dn, inode, NULL, NULL, 0); err = f2fs_get_dnode_of_data(&dn, pgofs, LOOKUP_NODE); if (err && err != -ENOENT) { @@ -554,12 +553,12 @@ static loff_t f2fs_seek_block(struct file *file, loff_t offset, int whence) } } - end_offset = ADDRS_PER_PAGE(dn.node_folio, inode); + end_offset = ADDRS_PER_PAGE(dn.node_entry, inode); /* find data/hole in dnode block */ for (; dn.ofs_in_node < end_offset; dn.ofs_in_node++, pgofs++, - data_ofs = (loff_t)pgofs << PAGE_SHIFT) { + data_ofs = F2FS_BLK_TO_BYTES(F2FS_I_SB(inode), pgofs)) { block_t blkaddr; blkaddr = f2fs_data_blkaddr(&dn); @@ -595,7 +594,8 @@ fail: static loff_t f2fs_llseek(struct file *file, loff_t offset, int whence) { struct inode *inode = file->f_mapping->host; - loff_t maxbytes = F2FS_BLK_TO_BYTES(max_file_blocks(inode)); + loff_t maxbytes = F2FS_BLK_TO_BYTES(F2FS_I_SB(inode), + max_file_blocks(F2FS_I_SB(inode), inode)); switch (whence) { case SEEK_SET: @@ -716,7 +716,7 @@ void f2fs_truncate_data_blocks_range(struct dnode_of_data *dn, int count) block_t blkstart; int blklen = 0; - addr = get_dnode_addr(dn->inode, dn->node_folio) + ofs; + addr = get_dnode_addr(dn->inode, dn->node_entry) + ofs; blkstart = le32_to_cpu(*addr); /* Assumption: truncation starts with cluster */ @@ -780,7 +780,7 @@ next: * once we invalidate valid blkaddr in range [ofs, ofs + count], * we will invalidate all blkaddr in the whole range. */ - fofs = f2fs_start_bidx_of_node(ofs_of_node(dn->node_folio), + fofs = f2fs_start_bidx_of_node(ofs_of_node(sbi, dn->node_entry), dn->inode) + ofs; f2fs_update_read_extent_cache_range(dn, fofs, 0, len); f2fs_update_age_extent_cache_range(dn, fofs, len); @@ -818,7 +818,7 @@ static int truncate_partial_data_page(struct inode *inode, u64 from, if (IS_ERR(folio)) return PTR_ERR(folio) == -ENOENT ? 0 : PTR_ERR(folio); truncate_out: - f2fs_folio_wait_writeback(folio, DATA, true, true); + f2fs_folio_wait_writeback(folio, true, true); folio_zero_segment(folio, offset, folio_size(folio)); /* An encrypted inode should have a key and truncate the last page. */ @@ -836,7 +836,7 @@ int f2fs_do_truncate_blocks(struct inode *inode, u64 from, bool lock) struct f2fs_lock_context lc; pgoff_t free_from; int count = 0, err = 0; - struct folio *ifolio; + struct f2fs_cached_block *ientry; bool truncate_page = false; trace_f2fs_truncate_blocks_enter(inode, from); @@ -846,17 +846,17 @@ int f2fs_do_truncate_blocks(struct inode *inode, u64 from, bool lock) goto out_err; } - free_from = (pgoff_t)F2FS_BLK_ALIGN(from); + free_from = (pgoff_t)F2FS_BLK_ALIGN(sbi, from); - if (free_from >= max_file_blocks(inode)) + if (free_from >= max_file_blocks(sbi, inode)) goto free_partial; if (lock) f2fs_lock_op(sbi, &lc); - ifolio = f2fs_get_inode_folio(sbi, inode->i_ino); - if (IS_ERR(ifolio)) { - err = PTR_ERR(ifolio); + ientry = f2fs_get_inode_cache(sbi, inode->i_ino); + if (IS_ERR(ientry)) { + err = PTR_ERR(ientry); goto out; } @@ -875,18 +875,18 @@ int f2fs_do_truncate_blocks(struct inode *inode, u64 from, bool lock) f2fs_drop_extent_tree(inode); - f2fs_folio_put(ifolio, true); + f2fs_put_cache(ientry, true); goto out; } if (f2fs_has_inline_data(inode)) { - f2fs_truncate_inline_inode(inode, ifolio, from); - f2fs_folio_put(ifolio, true); + f2fs_truncate_inline_inode(inode, ientry, from); + f2fs_put_cache(ientry, true); truncate_page = true; goto out; } - set_new_dnode(&dn, inode, ifolio, NULL, 0); + set_new_dnode(&dn, inode, ientry, NULL, 0); err = f2fs_get_dnode_of_data(&dn, free_from, LOOKUP_NODE_RA); if (err) { if (err == -ENOENT) @@ -894,12 +894,12 @@ int f2fs_do_truncate_blocks(struct inode *inode, u64 from, bool lock) goto out; } - count = ADDRS_PER_PAGE(dn.node_folio, inode); + count = ADDRS_PER_PAGE(dn.node_entry, inode); count -= dn.ofs_in_node; f2fs_bug_on(sbi, count < 0); - if (dn.ofs_in_node || IS_INODE(dn.node_folio)) { + if (dn.ofs_in_node || IS_INODE(sbi, dn.node_entry)) { f2fs_truncate_data_blocks_range(&dn, count); free_from += count; } @@ -959,9 +959,10 @@ int f2fs_truncate_blocks(struct inode *inode, u64 from, bool lock) int f2fs_truncate(struct inode *inode) { + struct f2fs_sb_info *sbi = F2FS_I_SB(inode); int err; - if (unlikely(f2fs_cp_error(F2FS_I_SB(inode)))) + if (unlikely(f2fs_cp_error(sbi))) return -EIO; if (!(S_ISREG(inode->i_mode) || S_ISDIR(inode->i_mode) || @@ -986,8 +987,8 @@ int f2fs_truncate(struct inode *inode) * leak in evict() path. */ truncate_inode_pages_range(inode->i_mapping, - F2FS_BLK_TO_BYTES(0), - F2FS_BLK_END_BYTES(0)); + F2FS_BLK_TO_BYTES(sbi, 0), + F2FS_BLK_END_BYTES(sbi, 0)); return err; } } @@ -1032,7 +1033,7 @@ static bool f2fs_force_buffered_io(struct inode *inode, int rw) return false; } -int f2fs_getattr(struct mnt_idmap *idmap, const struct path *path, +int f2fs_getattr(const struct mnt_idmap *idmap, const struct path *path, struct kstat *stat, u32 request_mask, unsigned int query_flags) { struct inode *inode = d_inode(path->dentry); @@ -1096,7 +1097,7 @@ int f2fs_getattr(struct mnt_idmap *idmap, const struct path *path, } #ifdef CONFIG_F2FS_FS_POSIX_ACL -static void __setattr_copy(struct mnt_idmap *idmap, +static void __setattr_copy(const struct mnt_idmap *idmap, struct inode *inode, const struct iattr *attr) { unsigned int ia_valid = attr->ia_valid; @@ -1121,7 +1122,7 @@ static void __setattr_copy(struct mnt_idmap *idmap, #define __setattr_copy setattr_copy #endif -int f2fs_setattr(struct mnt_idmap *idmap, struct dentry *dentry, +int f2fs_setattr(const struct mnt_idmap *idmap, struct dentry *dentry, struct iattr *attr) { struct inode *inode = d_inode(dentry); @@ -1157,7 +1158,7 @@ int f2fs_setattr(struct mnt_idmap *idmap, struct dentry *dentry, return -EOPNOTSUPP; if (is_inode_flag_set(inode, FI_COMPRESS_RELEASED) && !IS_ALIGNED(attr->ia_size, - F2FS_BLK_TO_BYTES(fi->i_cluster_size))) + F2FS_BLK_TO_BYTES(sbi, fi->i_cluster_size))) return -EINVAL; if (f2fs_is_pinned_file(inode)) { @@ -1173,7 +1174,7 @@ int f2fs_setattr(struct mnt_idmap *idmap, struct dentry *dentry, * pinned file. */ else if (!IS_ALIGNED(attr->ia_size, - F2FS_BLK_TO_BYTES(CAP_BLKS_PER_SEC(sbi)))) + F2FS_BLK_TO_BYTES(sbi, CAP_BLKS_PER_SEC(sbi)))) return -EINVAL; } } @@ -1304,7 +1305,7 @@ static int fill_zero(struct inode *inode, pgoff_t index, if (IS_ERR(folio)) return PTR_ERR(folio); - f2fs_folio_wait_writeback(folio, DATA, true, true); + f2fs_folio_wait_writeback(folio, true, true); folio_zero_range(folio, start, len); folio_mark_dirty(folio); f2fs_folio_put(folio, true); @@ -1330,7 +1331,7 @@ int f2fs_truncate_hole(struct inode *inode, pgoff_t pg_start, pgoff_t pg_end) return err; } - end_offset = ADDRS_PER_PAGE(dn.node_folio, inode); + end_offset = ADDRS_PER_PAGE(dn.node_entry, inode); count = min(end_offset - dn.ofs_in_node, pg_end - pg_start); f2fs_bug_on(F2FS_I_SB(inode), count == 0 || count > end_offset); @@ -1430,7 +1431,7 @@ next_dnode: goto next; } - done = min((pgoff_t)ADDRS_PER_PAGE(dn.node_folio, inode) - + done = min((pgoff_t)ADDRS_PER_PAGE(dn.node_entry, inode) - dn.ofs_in_node, len); for (i = 0; i < done; i++, blkaddr++, do_replace++, dn.ofs_in_node++) { *blkaddr = f2fs_data_blkaddr(&dn); @@ -1519,7 +1520,7 @@ static int __clone_blkaddrs(struct inode *src_inode, struct inode *dst_inode, } ilen = min((pgoff_t) - ADDRS_PER_PAGE(dn.node_folio, dst_inode) - + ADDRS_PER_PAGE(dn.node_entry, dst_inode) - dn.ofs_in_node, len - i); do { dn.data_blkaddr = f2fs_data_blkaddr(&dn); @@ -1557,7 +1558,7 @@ static int __clone_blkaddrs(struct inode *src_inode, struct inode *dst_inode, return PTR_ERR(fdst); } - f2fs_folio_wait_writeback(fdst, DATA, true, true); + f2fs_folio_wait_writeback(fdst, true, true); memcpy_folio(fdst, 0, fsrc, 0, PAGE_SIZE); folio_mark_dirty(fdst); @@ -1667,7 +1668,8 @@ static int f2fs_collapse_range(struct inode *inode, loff_t offset, loff_t len) return -EINVAL; /* collapse range should be aligned to block size of f2fs. */ - if (offset & (F2FS_BLKSIZE - 1) || len & (F2FS_BLKSIZE - 1)) + if (offset & F2FS_BLKSIZE_MASK(F2FS_I_SB(inode)) || + len & F2FS_BLKSIZE_MASK(F2FS_I_SB(inode))) return -EINVAL; ret = f2fs_convert_inline_inode(inode); @@ -1826,7 +1828,7 @@ static int f2fs_zero_range(struct inode *inode, loff_t offset, loff_t len, goto out; } - end_offset = ADDRS_PER_PAGE(dn.node_folio, inode); + end_offset = ADDRS_PER_PAGE(dn.node_entry, inode); end = min(pg_end, end_offset - dn.ofs_in_node + index); ret = f2fs_do_zero_range(&dn, index, end); @@ -1882,7 +1884,8 @@ static int f2fs_insert_range(struct inode *inode, loff_t offset, loff_t len) return -EINVAL; /* insert range should be aligned to block size of f2fs. */ - if (offset & (F2FS_BLKSIZE - 1) || len & (F2FS_BLKSIZE - 1)) + if (offset & F2FS_BLKSIZE_MASK(F2FS_I_SB(inode)) || + len & F2FS_BLKSIZE_MASK(F2FS_I_SB(inode))) return -EINVAL; ret = f2fs_convert_inline_inode(inode); @@ -2361,7 +2364,7 @@ static int f2fs_ioc_getversion(struct file *filp, unsigned long arg) static int f2fs_ioc_start_atomic_write(struct file *filp, bool truncate) { struct inode *inode = file_inode(filp); - struct mnt_idmap *idmap = file_mnt_idmap(filp); + const struct mnt_idmap *idmap = file_mnt_idmap(filp); struct f2fs_inode_info *fi = F2FS_I(inode); struct f2fs_sb_info *sbi = F2FS_I_SB(inode); loff_t isize; @@ -2473,7 +2476,7 @@ out: static int f2fs_ioc_commit_atomic_write(struct file *filp) { struct inode *inode = file_inode(filp); - struct mnt_idmap *idmap = file_mnt_idmap(filp); + const struct mnt_idmap *idmap = file_mnt_idmap(filp); int ret; if (!(filp->f_mode & FMODE_WRITE)) @@ -2508,7 +2511,7 @@ static int f2fs_ioc_commit_atomic_write(struct file *filp) static int f2fs_ioc_abort_atomic_write(struct file *filp) { struct inode *inode = file_inode(filp); - struct mnt_idmap *idmap = file_mnt_idmap(filp); + const struct mnt_idmap *idmap = file_mnt_idmap(filp); int ret; if (!(filp->f_mode & FMODE_WRITE)) @@ -2560,7 +2563,7 @@ int f2fs_do_shutdown(struct f2fs_sb_info *sbi, unsigned int flag, f2fs_stop_checkpoint(sbi, false, STOP_CP_REASON_SHUTDOWN); break; case F2FS_GOING_DOWN_METAFLUSH: - f2fs_sync_meta_pages(sbi, LONG_MAX, FS_META_IO); + f2fs_sync_meta_caches(sbi, LONG_MAX, true, FS_META_IO); f2fs_stop_checkpoint(sbi, false, STOP_CP_REASON_SHUTDOWN); break; case F2FS_GOING_DOWN_NEED_FSCK: @@ -2644,7 +2647,8 @@ static int f2fs_keep_noreuse_range(struct inode *inode, loff_t offset, loff_t len) { struct f2fs_sb_info *sbi = F2FS_I_SB(inode); - u64 max_bytes = F2FS_BLK_TO_BYTES(max_file_blocks(inode)); + u64 max_bytes = F2FS_BLK_TO_BYTES(sbi, + max_file_blocks(sbi, inode)); u64 start, end; int ret = 0; @@ -3125,7 +3129,7 @@ do_map: goto clear_out; } - f2fs_folio_wait_writeback(folio, DATA, true, true); + f2fs_folio_wait_writeback(folio, true, true); folio_mark_dirty(folio); folio_set_f2fs_gcing(folio); @@ -3184,11 +3188,12 @@ static int f2fs_ioc_defragment(struct file *filp, unsigned long arg) return -EFAULT; /* verify alignment of offset & size */ - if (range.start & (F2FS_BLKSIZE - 1) || range.len & (F2FS_BLKSIZE - 1)) + if (range.start & F2FS_BLKSIZE_MASK(sbi) || + range.len & F2FS_BLKSIZE_MASK(sbi)) return -EINVAL; - if (unlikely((range.start + range.len) >> PAGE_SHIFT > - max_file_blocks(inode))) + if (unlikely(F2FS_BYTES_TO_BLK(sbi, range.start + range.len) > + max_file_blocks(sbi, inode))) return -EINVAL; err = mnt_want_write_file(filp); @@ -3270,7 +3275,7 @@ static int f2fs_move_file_range(struct file *file_in, loff_t pos_in, if (src == dst && pos_out > pos_in && pos_out < pos_in + len) goto out_unlock; if (pos_in + len == src->i_size) - len = ALIGN(src->i_size, F2FS_BLKSIZE) - pos_in; + len = ALIGN(src->i_size, F2FS_BLKSIZE(sbi)) - pos_in; if (len == 0) { ret = 0; goto out_unlock; @@ -3288,9 +3293,9 @@ static int f2fs_move_file_range(struct file *file_in, loff_t pos_in, } /* verify the end result is block aligned */ - if (!IS_ALIGNED(pos_in, F2FS_BLKSIZE) || - !IS_ALIGNED(pos_in + len, F2FS_BLKSIZE) || - !IS_ALIGNED(pos_out, F2FS_BLKSIZE)) + if (!IS_ALIGNED(pos_in, F2FS_BLKSIZE(sbi)) || + !IS_ALIGNED(pos_in + len, F2FS_BLKSIZE(sbi)) || + !IS_ALIGNED(pos_out, F2FS_BLKSIZE(sbi))) goto out_unlock; ret = f2fs_convert_inline_inode(src); @@ -3322,9 +3327,9 @@ static int f2fs_move_file_range(struct file *file_in, loff_t pos_in, } f2fs_lock_op(sbi, &lc); - ret = __exchange_data_block(src, dst, F2FS_BYTES_TO_BLK(pos_in), - F2FS_BYTES_TO_BLK(pos_out), - F2FS_BYTES_TO_BLK(len), false); + ret = __exchange_data_block(src, dst, F2FS_BYTES_TO_BLK(sbi, pos_in), + F2FS_BYTES_TO_BLK(sbi, pos_out), + F2FS_BYTES_TO_BLK(sbi, len), false); if (!ret) { if (dst_max_i_size) @@ -3580,7 +3585,7 @@ int f2fs_fileattr_get(struct dentry *dentry, struct file_kattr *fa) return 0; } -int f2fs_fileattr_set(struct mnt_idmap *idmap, +int f2fs_fileattr_set(const struct mnt_idmap *idmap, struct dentry *dentry, struct file_kattr *fa) { struct inode *inode = d_inode(dentry); @@ -3843,10 +3848,11 @@ static int f2fs_ioc_reserve_dev_alias(struct file *filp) write_unlock(&et->lock); clear_inode_flag(inode, FI_NO_EXTENT); + f2fs_drop_discard_cmd_range(sbi, ei.blk, ei.len); f2fs_reserve_device_alias(sbi, ei.blk, ei.len); i_size_write(inode, (loff_t)ei.len << sbi->log_blocksize); - f2fs_update_inode_page(inode); + f2fs_update_inode_cache(inode); spin_lock(&FREE_I(sbi)->segmap_lock); FDEV(f2fs_target_device_index(sbi, ei.blk)).is_reserving = false; @@ -3940,7 +3946,7 @@ static int f2fs_ioc_release_dev_alias(struct file *filp) goto out_inode_unlock; } - f2fs_update_inode_page(inode); + f2fs_update_inode_cache(inode); f2fs_unlock_op(sbi, &lc); f2fs_up_write_trace(&sbi->gc_lock, &glc); @@ -3990,7 +3996,7 @@ int f2fs_precache_extents(struct inode *inode) map.m_next_extent = &m_next_extent; map.m_seg_type = NO_CHECK_TYPE; map.m_may_create = false; - end = F2FS_BLK_ALIGN(i_size_read(inode)); + end = F2FS_BLK_ALIGN(F2FS_I_SB(inode), i_size_read(inode)); while (map.m_lblk < end) { map.m_len = end - map.m_lblk; @@ -4159,7 +4165,7 @@ static int release_compress_blocks(struct dnode_of_data *dn, pgoff_t count) int i; for (i = 0; i < count; i++) { - blkaddr = data_blkaddr(dn->inode, dn->node_folio, + blkaddr = data_blkaddr(dn->inode, dn->node_entry, dn->ofs_in_node + i); if (!__is_valid_data_blkaddr(blkaddr)) @@ -4278,7 +4284,7 @@ static int f2fs_release_compress_blocks(struct file *filp, unsigned long arg) break; } - end_offset = ADDRS_PER_PAGE(dn.node_folio, inode); + end_offset = ADDRS_PER_PAGE(dn.node_entry, inode); count = min(end_offset - dn.ofs_in_node, last_idx - page_idx); count = round_up(count, fi->i_cluster_size); @@ -4329,7 +4335,7 @@ static int reserve_compress_blocks(struct dnode_of_data *dn, pgoff_t count, int i; for (i = 0; i < count; i++) { - blkaddr = data_blkaddr(dn->inode, dn->node_folio, + blkaddr = data_blkaddr(dn->inode, dn->node_entry, dn->ofs_in_node + i); if (!__is_valid_data_blkaddr(blkaddr)) @@ -4346,7 +4352,7 @@ static int reserve_compress_blocks(struct dnode_of_data *dn, pgoff_t count, int ret; for (i = 0; i < cluster_size; i++) { - blkaddr = data_blkaddr(dn->inode, dn->node_folio, + blkaddr = data_blkaddr(dn->inode, dn->node_entry, dn->ofs_in_node + i); if (i == 0) { @@ -4457,7 +4463,7 @@ static int f2fs_reserve_compress_blocks(struct file *filp, unsigned long arg) break; } - end_offset = ADDRS_PER_PAGE(dn.node_folio, inode); + end_offset = ADDRS_PER_PAGE(dn.node_entry, inode); count = min(end_offset - dn.ofs_in_node, last_idx - page_idx); count = round_up(count, fi->i_cluster_size); @@ -4506,8 +4512,8 @@ unlock_inode: static int f2fs_secure_erase(struct block_device *bdev, struct inode *inode, pgoff_t off, block_t block, block_t len, u32 flags) { - sector_t sector = SECTOR_FROM_BLOCK(block); - sector_t nr_sects = SECTOR_FROM_BLOCK(len); + sector_t sector = SECTOR_FROM_BLOCK(F2FS_I_SB(inode), block); + sector_t nr_sects = SECTOR_FROM_BLOCK(F2FS_I_SB(inode), len); int ret = 0; if (flags & F2FS_TRIM_FILE_DISCARD) { @@ -4584,14 +4590,14 @@ static int f2fs_sec_trim_file(struct file *filp, unsigned long arg) to_end = true; } - if (!IS_ALIGNED(range.start, F2FS_BLKSIZE) || - (!to_end && !IS_ALIGNED(end_addr, F2FS_BLKSIZE))) { + if (!IS_ALIGNED(range.start, F2FS_BLKSIZE(sbi)) || + (!to_end && !IS_ALIGNED(end_addr, F2FS_BLKSIZE(sbi)))) { ret = -EINVAL; goto err; } - index = F2FS_BYTES_TO_BLK(range.start); - pg_end = DIV_ROUND_UP(end_addr, F2FS_BLKSIZE); + index = F2FS_BYTES_TO_BLK(sbi, range.start); + pg_end = F2FS_BLK_ALIGN(sbi, end_addr); ret = f2fs_convert_inline_inode(inode); if (ret) @@ -4623,7 +4629,7 @@ static int f2fs_sec_trim_file(struct file *filp, unsigned long arg) goto out; } - end_offset = ADDRS_PER_PAGE(dn.node_folio, inode); + end_offset = ADDRS_PER_PAGE(dn.node_entry, inode); count = min(end_offset - dn.ofs_in_node, pg_end - index); for (i = 0; i < count; i++, index++, dn.ofs_in_node++) { struct block_device *cur_bdev; @@ -4819,7 +4825,7 @@ static int redirty_blocks(struct inode *inode, pgoff_t page_idx, int len) /* It will never fail, when folio has pinned above */ f2fs_bug_on(F2FS_I_SB(inode), IS_ERR(folio)); - f2fs_folio_wait_writeback(folio, DATA, true, true); + f2fs_folio_wait_writeback(folio, true, true); folio_mark_dirty(folio); folio_set_f2fs_gcing(folio); @@ -5171,7 +5177,7 @@ static int f2fs_dio_read_end_io(struct kiocb *iocb, ssize_t size, int error, { struct f2fs_sb_info *sbi = F2FS_I_SB(file_inode(iocb->ki_filp)); - dec_page_count(sbi, F2FS_DIO_READ); + dec_cache_count(sbi, F2FS_DIO_READ); if (error) return error; f2fs_update_iostat(sbi, NULL, APP_DIRECT_READ_IO, size); @@ -5229,13 +5235,13 @@ static ssize_t f2fs_dio_read_iter(struct kiocb *iocb, struct iov_iter *to) * the higher-level function iomap_dio_rw() in order to ensure that the * F2FS_DIO_READ counter will be decremented correctly in all cases. */ - inc_page_count(sbi, F2FS_DIO_READ); + inc_cache_count(sbi, F2FS_DIO_READ); dio = __iomap_dio_rw(iocb, to, &f2fs_iomap_ops, &f2fs_iomap_dio_read_ops, 0, NULL, 0); if (IS_ERR_OR_NULL(dio)) { ret = PTR_ERR_OR_ZERO(dio); if (ret != -EIOCBQUEUED) - dec_page_count(sbi, F2FS_DIO_READ); + dec_cache_count(sbi, F2FS_DIO_READ); } else { ret = iomap_dio_complete(dio); } @@ -5288,7 +5294,7 @@ static ssize_t f2fs_file_read_iter(struct kiocb *iocb, struct iov_iter *to) /* In LFS mode, if there is inflight dio, wait for its completion */ if (f2fs_lfs_mode(F2FS_I_SB(inode)) && - get_pages(F2FS_I_SB(inode), F2FS_DIO_WRITE) && + get_nr_caches(F2FS_I_SB(inode), F2FS_DIO_WRITE) && (!f2fs_is_pinned_file(inode) || !dio)) inode_dio_wait(inode); @@ -5381,7 +5387,8 @@ static int f2fs_preallocate_blocks(struct kiocb *iocb, struct iov_iter *iter, * buffered IO, if DIO meets any holes. */ if (dio && i_size_read(inode) && - (F2FS_BYTES_TO_BLK(pos) < F2FS_BLK_ALIGN(i_size_read(inode)))) + (F2FS_BYTES_TO_BLK(sbi, pos) < + F2FS_BLK_ALIGN(sbi, i_size_read(inode)))) return 0; /* No-wait I/O can't allocate blocks. */ @@ -5402,8 +5409,8 @@ static int f2fs_preallocate_blocks(struct kiocb *iocb, struct iov_iter *iter, } /* Do not preallocate blocks that will be written partially in 4KB. */ - map.m_lblk = F2FS_BLK_ALIGN(pos); - map.m_len = F2FS_BYTES_TO_BLK(pos + count); + map.m_lblk = F2FS_BLK_ALIGN(sbi, pos); + map.m_len = F2FS_BYTES_TO_BLK(sbi, pos + count); if (map.m_len > map.m_lblk) map.m_len -= map.m_lblk; else @@ -5453,7 +5460,7 @@ static int f2fs_dio_write_end_io(struct kiocb *iocb, ssize_t size, int error, { struct f2fs_sb_info *sbi = F2FS_I_SB(file_inode(iocb->ki_filp)); - dec_page_count(sbi, F2FS_DIO_WRITE); + dec_cache_count(sbi, F2FS_DIO_WRITE); if (error) return error; f2fs_update_time(sbi, REQ_TIME); @@ -5564,7 +5571,7 @@ static ssize_t f2fs_dio_write_iter(struct kiocb *iocb, struct iov_iter *from, * the higher-level function iomap_dio_rw() in order to ensure that the * F2FS_DIO_WRITE counter will be decremented correctly in all cases. */ - inc_page_count(sbi, F2FS_DIO_WRITE); + inc_cache_count(sbi, F2FS_DIO_WRITE); dio_flags = 0; if (pos + count > inode->i_size) dio_flags |= IOMAP_DIO_FORCE_WAIT; @@ -5575,7 +5582,7 @@ static ssize_t f2fs_dio_write_iter(struct kiocb *iocb, struct iov_iter *from, if (ret == -ENOTBLK) ret = 0; if (ret != -EIOCBQUEUED) - dec_page_count(sbi, F2FS_DIO_WRITE); + dec_cache_count(sbi, F2FS_DIO_WRITE); } else { ret = iomap_dio_complete(dio); } diff --git a/fs/f2fs/gc.c b/fs/f2fs/gc.c index 0a00180c21dc..926e6976bc22 100644 --- a/fs/f2fs/gc.c +++ b/fs/f2fs/gc.c @@ -1054,7 +1054,7 @@ next_step: for (off = 0; off < usable_blks_in_seg; off++, entry++) { nid_t nid = le32_to_cpu(entry->nid); - struct folio *node_folio; + struct f2fs_cached_block *node_entry; struct node_info ni; int err; @@ -1066,38 +1066,38 @@ next_step: continue; if (phase == 0) { - f2fs_ra_meta_pages(sbi, NAT_BLOCK_OFFSET(nid), 1, + f2fs_ra_meta_caches(sbi, NAT_BLOCK_OFFSET(sbi, nid), 1, META_NAT, true); continue; } if (phase == 1) { - f2fs_ra_node_page(sbi, nid); + f2fs_ra_node_cache(sbi, nid); continue; } /* phase == 2 */ - node_folio = f2fs_get_node_folio(sbi, nid, NODE_TYPE_REGULAR); - if (IS_ERR(node_folio)) + node_entry = f2fs_get_node_cache(sbi, nid, NODE_TYPE_REGULAR); + if (IS_ERR(node_entry)) continue; /* block may become invalid during f2fs_get_node_folio */ if (check_valid_map(sbi, segno, off) == 0) { - f2fs_folio_put(node_folio, true); + f2fs_put_cache(node_entry, true); continue; } if (f2fs_get_node_info(sbi, nid, &ni, false)) { - f2fs_folio_put(node_folio, true); + f2fs_put_cache(node_entry, true); continue; } if (ni.blk_addr != start_addr + off) { - f2fs_folio_put(node_folio, true); + f2fs_put_cache(node_entry, true); continue; } - err = f2fs_move_node_folio(node_folio, gc_type); + err = f2fs_move_node_cache(node_entry, gc_type); if (!err && gc_type == FG_GC) submitted++; stat_inc_node_blk_count(sbi, 1, gc_type); @@ -1123,7 +1123,8 @@ next_step: */ block_t f2fs_start_bidx_of_node(unsigned int node_ofs, struct inode *inode) { - unsigned int indirect_blks = 2 * NIDS_PER_BLOCK + 4; + struct f2fs_sb_info *sbi = F2FS_I_SB(inode); + unsigned int indirect_blks = 2 * NIDS_PER_BLOCK(sbi) + 4; unsigned int bidx; if (node_ofs == 0) @@ -1132,11 +1133,12 @@ block_t f2fs_start_bidx_of_node(unsigned int node_ofs, struct inode *inode) if (node_ofs <= 2) { bidx = node_ofs - 1; } else if (node_ofs <= indirect_blks) { - int dec = (node_ofs - 4) / (NIDS_PER_BLOCK + 1); + int dec = (node_ofs - 4) / (NIDS_PER_BLOCK(sbi) + 1); bidx = node_ofs - 2 - dec; } else { - int dec = (node_ofs - indirect_blks - 3) / (NIDS_PER_BLOCK + 1); + int dec = (node_ofs - indirect_blks - 3) / + (NIDS_PER_BLOCK(sbi) + 1); bidx = node_ofs - 5 - dec; } @@ -1146,7 +1148,7 @@ block_t f2fs_start_bidx_of_node(unsigned int node_ofs, struct inode *inode) static bool is_alive(struct f2fs_sb_info *sbi, struct f2fs_summary *sum, struct node_info *dni, block_t blkaddr, unsigned int *nofs) { - struct folio *node_folio; + struct f2fs_cached_block *node_entry; nid_t nid; unsigned int ofs_in_node, max_addrs, base; block_t source_blkaddr; @@ -1154,12 +1156,12 @@ static bool is_alive(struct f2fs_sb_info *sbi, struct f2fs_summary *sum, nid = le32_to_cpu(sum->nid); ofs_in_node = le16_to_cpu(sum->ofs_in_node); - node_folio = f2fs_get_node_folio(sbi, nid, NODE_TYPE_REGULAR); - if (IS_ERR(node_folio)) + node_entry = f2fs_get_node_cache(sbi, nid, NODE_TYPE_REGULAR); + if (IS_ERR(node_entry)) return false; if (f2fs_get_node_info(sbi, nid, dni, false)) { - f2fs_folio_put(node_folio, true); + f2fs_put_cache(node_entry, true); return false; } @@ -1170,28 +1172,28 @@ static bool is_alive(struct f2fs_sb_info *sbi, struct f2fs_summary *sum, } if (f2fs_check_nid_range(sbi, dni->ino)) { - f2fs_folio_put(node_folio, true); + f2fs_put_cache(node_entry, true); return false; } - if (IS_INODE(node_folio)) { - base = offset_in_addr(F2FS_INODE(node_folio)); - max_addrs = DEF_ADDRS_PER_INODE; + if (IS_INODE(sbi, node_entry)) { + base = offset_in_addr(F2FS_INODE(node_entry)); + max_addrs = DEF_ADDRS_PER_INODE(sbi); } else { base = 0; - max_addrs = DEF_ADDRS_PER_BLOCK; + max_addrs = DEF_ADDRS_PER_BLOCK(sbi); } if (base + ofs_in_node >= max_addrs) { f2fs_err(sbi, "Inconsistent blkaddr offset: base:%u, ofs_in_node:%u, max:%u, ino:%u, nid:%u", base, ofs_in_node, max_addrs, dni->ino, dni->nid); - f2fs_folio_put(node_folio, true); + f2fs_put_cache(node_entry, true); return false; } - *nofs = ofs_of_node(node_folio); - source_blkaddr = data_blkaddr(NULL, node_folio, ofs_in_node); - f2fs_folio_put(node_folio, true); + *nofs = ofs_of_node(sbi, node_entry); + source_blkaddr = data_blkaddr(NULL, node_entry, ofs_in_node); + f2fs_put_cache(node_entry, true); if (source_blkaddr != blkaddr) { #ifdef CONFIG_F2FS_CHECK_FS @@ -1217,7 +1219,8 @@ static int ra_data_block(struct inode *inode, pgoff_t index) struct address_space *mapping = inode->i_mapping; struct inode *atomic_inode = NULL; struct dnode_of_data dn; - struct folio *folio, *efolio; + struct folio *folio; + struct f2fs_cached_block *entry; struct f2fs_io_info fio = { .sbi = sbi, .ino = inode->i_ino, @@ -1227,6 +1230,7 @@ static int ra_data_block(struct inode *inode, pgoff_t index) .op_flags = 0, .encrypted_page = NULL, .in_list = 0, + .is_cache = 1, }; int err = 0; @@ -1281,36 +1285,35 @@ got_it: * don't cache encrypted data into meta inode until previous dirty * data were writebacked to avoid racing between GC and flush. */ - f2fs_folio_wait_writeback(folio, DATA, true, true); + f2fs_folio_wait_writeback(folio, true, true); f2fs_wait_on_block_writeback(inode, dn.data_blkaddr); - efolio = f2fs_filemap_get_folio(META_MAPPING(sbi), dn.data_blkaddr, - FGP_LOCK | FGP_CREAT, GFP_NOFS); - if (IS_ERR(efolio)) { - err = PTR_ERR(efolio); + entry = f2fs_grab_cache(META_CACHE(sbi), dn.data_blkaddr, + F2FS_CACHE_LOCK_CREATE); + if (IS_ERR(entry)) { + err = PTR_ERR(entry); goto put_folio; } - fio.encrypted_page = &efolio->page; + if (f2fs_cache_test_uptodate(entry)) + goto put_cache; - if (folio_test_uptodate(efolio)) - goto put_encrypted_page; - - err = f2fs_submit_page_bio(&fio); + fio.cache_entry = entry; + err = f2fs_submit_cache_read(&fio); if (err) - goto put_encrypted_page; - f2fs_put_page(fio.encrypted_page, false); + goto put_cache; + f2fs_put_cache(entry, false); f2fs_folio_put(folio, true); - f2fs_update_iostat(sbi, inode, FS_DATA_READ_IO, F2FS_BLKSIZE); - f2fs_update_iostat(sbi, NULL, FS_GDATA_READ_IO, F2FS_BLKSIZE); + f2fs_update_iostat(sbi, inode, FS_DATA_READ_IO, F2FS_BLKSIZE(sbi)); + f2fs_update_iostat(sbi, NULL, FS_GDATA_READ_IO, F2FS_BLKSIZE(sbi)); if (atomic_inode) iput(atomic_inode); return 0; -put_encrypted_page: - f2fs_put_page(fio.encrypted_page, true); +put_cache: + f2fs_put_cache(entry, true); put_folio: f2fs_folio_put(folio, true); out_iput: @@ -1320,7 +1323,7 @@ out_iput: } /* - * Move data block via META_MAPPING while keeping locked data page. + * Move data block via meta cache while keeping locked data page. * This can be used to move blocks, aka LBAs, directly on disk. */ static int move_data_block(struct inode *inode, block_t bidx, @@ -1337,11 +1340,13 @@ static int move_data_block(struct inode *inode, block_t bidx, .op_flags = 0, .encrypted_page = NULL, .in_list = 0, + .is_cache = 1, }; struct dnode_of_data dn; struct f2fs_summary sum; struct node_info ni; - struct folio *folio, *mfolio, *efolio; + struct folio *folio; + struct f2fs_cached_block *sentry, *tentry; block_t newaddr; int err = 0; bool lfs_mode = f2fs_lfs_mode(fio.sbi); @@ -1391,7 +1396,7 @@ static int move_data_block(struct inode *inode, block_t bidx, * don't cache encrypted data into meta inode until previous dirty * data were writebacked to avoid racing between GC and flush. */ - f2fs_folio_wait_writeback(folio, DATA, true, true); + f2fs_folio_wait_writeback(folio, true, true); f2fs_wait_on_block_writeback(inode, dn.data_blkaddr); @@ -1406,33 +1411,33 @@ static int move_data_block(struct inode *inode, block_t bidx, if (lfs_mode) f2fs_down_write(&fio.sbi->io_order_lock); - mfolio = f2fs_grab_cache_folio(META_MAPPING(fio.sbi), - fio.old_blkaddr, false); - if (IS_ERR(mfolio)) { - err = PTR_ERR(mfolio); + sentry = f2fs_grab_cache(META_CACHE(fio.sbi), fio.old_blkaddr, + F2FS_CACHE_LOCK_CREATE); + if (IS_ERR(sentry)) { + err = PTR_ERR(sentry); goto up_out; } - fio.encrypted_page = folio_file_page(mfolio, fio.old_blkaddr); + fio.cache_entry = sentry; - /* read source block in mfolio */ - if (!folio_test_uptodate(mfolio)) { - err = f2fs_submit_page_bio(&fio); + /* read source block in cache */ + if (!f2fs_cache_test_uptodate(sentry)) { + err = f2fs_submit_cache_read(&fio); if (err) { - f2fs_folio_put(mfolio, true); + f2fs_put_cache(sentry, true); goto up_out; } f2fs_update_iostat(fio.sbi, inode, FS_DATA_READ_IO, - F2FS_BLKSIZE); + F2FS_BLKSIZE(fio.sbi)); f2fs_update_iostat(fio.sbi, NULL, FS_GDATA_READ_IO, - F2FS_BLKSIZE); + F2FS_BLKSIZE(fio.sbi)); - folio_lock(mfolio); - if (unlikely(!is_meta_folio(mfolio) || - !folio_test_uptodate(mfolio))) { + f2fs_lock_cache(sentry); + if (unlikely(!f2fs_is_meta_cache(sentry) || + !f2fs_cache_test_uptodate(sentry))) { err = -EIO; - f2fs_folio_put(mfolio, true); + f2fs_put_cache(sentry, true); goto up_out; } } @@ -1440,49 +1445,50 @@ static int move_data_block(struct inode *inode, block_t bidx, set_summary(&sum, dn.nid, dn.ofs_in_node, ni.version); /* allocate block address */ - err = f2fs_allocate_data_block(fio.sbi, NULL, fio.old_blkaddr, &newaddr, + err = f2fs_allocate_data_block(fio.sbi, fio.old_blkaddr, &newaddr, &sum, type, NULL); if (err) { - f2fs_folio_put(mfolio, true); + f2fs_put_cache(sentry, true); /* filesystem should shutdown, no need to recovery block */ goto up_out; } - efolio = f2fs_filemap_get_folio(META_MAPPING(fio.sbi), newaddr, - FGP_LOCK | FGP_CREAT, GFP_NOFS); - if (IS_ERR(efolio)) { - err = PTR_ERR(efolio); - f2fs_folio_put(mfolio, true); + tentry = f2fs_grab_cache(META_CACHE(fio.sbi), newaddr, + F2FS_CACHE_LOCK_CREATE); + if (IS_ERR(tentry)) { + err = PTR_ERR(tentry); + f2fs_put_cache(sentry, true); goto recover_block; } - fio.encrypted_page = &efolio->page; + fio.cache_entry = tentry; /* write target block */ - f2fs_wait_on_page_writeback(fio.encrypted_page, DATA, true, true); - memcpy(page_address(fio.encrypted_page), - folio_address(mfolio), PAGE_SIZE); - f2fs_folio_put(mfolio, true); + f2fs_cache_wait_writeback_cond(tentry, DATA); + memcpy(cache_address(tentry), cache_address(sentry), + fio.sbi->blocksize); + f2fs_put_cache(sentry, true); f2fs_invalidate_internal_cache(fio.sbi, fio.old_blkaddr, 1); - set_page_dirty(fio.encrypted_page); - if (clear_page_dirty_for_io(fio.encrypted_page)) - dec_page_count(fio.sbi, F2FS_DIRTY_META); + f2fs_mark_cache_dirty(tentry); + if (f2fs_cache_test_and_clear_dirty(tentry)) + dec_cache_count(fio.sbi, F2FS_DIRTY_META); - set_page_writeback(fio.encrypted_page); + f2fs_start_cache_writeback(tentry); fio.op = REQ_OP_WRITE; fio.op_flags = REQ_SYNC; fio.new_blkaddr = newaddr; - f2fs_submit_page_write(&fio); + f2fs_submit_cache_write(&fio); - f2fs_update_iostat(fio.sbi, NULL, FS_GC_DATA_IO, F2FS_BLKSIZE); + f2fs_update_iostat(fio.sbi, NULL, FS_GC_DATA_IO, + F2FS_BLKSIZE(fio.sbi)); f2fs_update_data_blkaddr(&dn, newaddr); set_inode_flag(inode, FI_APPEND_WRITE); - f2fs_put_page(fio.encrypted_page, true); + f2fs_put_cache(tentry, true); recover_block: if (err) f2fs_do_replace_block(fio.sbi, &sum, newaddr, fio.old_blkaddr, @@ -1543,7 +1549,7 @@ static int move_data_page(struct inode *inode, block_t bidx, int gc_type, bool is_dirty = folio_test_dirty(folio); retry: - f2fs_folio_wait_writeback(folio, DATA, true, true); + f2fs_folio_wait_writeback(folio, true, true); folio_mark_dirty(folio); if (folio_clear_dirty_for_io(folio)) { @@ -1614,13 +1620,13 @@ next_step: continue; if (phase == 0) { - f2fs_ra_meta_pages(sbi, NAT_BLOCK_OFFSET(nid), 1, + f2fs_ra_meta_caches(sbi, NAT_BLOCK_OFFSET(sbi, nid), 1, META_NAT, true); continue; } if (phase == 1) { - f2fs_ra_node_page(sbi, nid); + f2fs_ra_node_cache(sbi, nid); continue; } @@ -1629,7 +1635,7 @@ next_step: continue; if (phase == 2) { - f2fs_ra_node_page(sbi, dni.ino); + f2fs_ra_node_cache(sbi, dni.ino); continue; } @@ -1818,28 +1824,31 @@ static int do_garbage_collect(struct f2fs_sb_info *sbi, sum_blk_cnt = DIV_ROUND_UP(end_segno - segno, sbi->sums_per_block); /* readahead multi ssa blocks those have contiguous address */ if (__is_large_section(sbi)) - f2fs_ra_meta_pages(sbi, GET_SUM_BLOCK(sbi, segno), + f2fs_ra_meta_caches(sbi, GET_SUM_BLOCK(sbi, segno), sum_blk_cnt, META_SSA, true); /* reference all summary page */ while (segno < end_segno) { - struct folio *sum_folio = f2fs_get_sum_folio(sbi, segno); + struct f2fs_cached_block *sum_entry = + f2fs_get_sum_cache(sbi, segno); segno += sbi->sums_per_block; - if (IS_ERR(sum_folio)) { - int err = PTR_ERR(sum_folio); + if (IS_ERR(sum_entry)) { + int err = PTR_ERR(sum_entry); end_segno = segno - sbi->sums_per_block; segno = rounddown(start_segno, sbi->sums_per_block); while (segno < end_segno) { - sum_folio = filemap_get_folio(META_MAPPING(sbi), + sum_entry = f2fs_find_meta_cache(sbi, GET_SUM_BLOCK(sbi, segno)); - folio_put_refs(sum_folio, 2); + f2fs_put_cache(sum_entry, false); + f2fs_put_cache(sum_entry, false); segno += sbi->sums_per_block; } return err; } - folio_unlock(sum_folio); + f2fs_unlock_cache(sum_entry); + } blk_start_plug(&plug); @@ -1847,11 +1856,16 @@ static int do_garbage_collect(struct f2fs_sb_info *sbi, segno = start_segno; while (segno < end_segno) { unsigned int cur_segno; + unsigned int block_end_segno; /* find segment summary of victim */ - struct folio *sum_folio = filemap_get_folio(META_MAPPING(sbi), + struct f2fs_cached_block *sum_entry = + f2fs_find_meta_cache(sbi, GET_SUM_BLOCK(sbi, segno)); - unsigned int block_end_segno = rounddown(segno, sbi->sums_per_block) + + f2fs_bug_on(sbi, IS_ERR(sum_entry)); + + block_end_segno = rounddown(segno, sbi->sums_per_block) + sbi->sums_per_block; if (block_end_segno > end_segno) @@ -1864,8 +1878,8 @@ static int do_garbage_collect(struct f2fs_sb_info *sbi, goto next_block; } - if (!folio_test_uptodate(sum_folio) || - unlikely(f2fs_cp_error(sbi))) + if (!f2fs_cache_test_uptodate(sum_entry) || + unlikely(f2fs_cp_error(sbi))) goto next_block; for (cur_segno = segno; cur_segno < block_end_segno; @@ -1884,7 +1898,7 @@ static int do_garbage_collect(struct f2fs_sb_info *sbi, data_type = (type == SUM_TYPE_DATA) ? DATA : NODE; } - sum = SUM_BLK_PAGE_ADDR(sbi, sum_folio, cur_segno); + sum = SUM_BLK_ENTRY_ADDR(sbi, sum_entry, cur_segno); if (type != GET_SUM_TYPE(sum_footer(sbi, sum))) { f2fs_err(sbi, "Inconsistent segment (%u) type " "[%d, %d] in SIT and SSA", @@ -1926,12 +1940,14 @@ freed: cur_segno + 1 : NULL_SEGNO; if (unlikely(freezing(current))) { - folio_put_refs(sum_folio, 2); + f2fs_put_cache(sum_entry, false); + f2fs_put_cache(sum_entry, false); goto stop; } } next_block: - folio_put_refs(sum_folio, 2); + f2fs_put_cache(sum_entry, false); + f2fs_put_cache(sum_entry, false); segno = block_end_segno; } @@ -1963,9 +1979,9 @@ int f2fs_gc(struct f2fs_sb_info *sbi, struct f2fs_gc_control *gc_control) trace_f2fs_gc_begin(sbi->sb, gc_type, gc_control->no_bg_gc, gc_control->nr_free_secs, - get_pages(sbi, F2FS_DIRTY_NODES), - get_pages(sbi, F2FS_DIRTY_DENTS), - get_pages(sbi, F2FS_DIRTY_IMETA), + get_nr_caches(sbi, F2FS_DIRTY_NODES), + get_nr_caches(sbi, F2FS_DIRTY_DENTS), + get_nr_caches(sbi, F2FS_DIRTY_IMETA), free_sections(sbi), free_segments(sbi), reserved_segments(sbi), @@ -2089,9 +2105,9 @@ stop: f2fs_unpin_all_sections(sbi, true); trace_f2fs_gc_end(sbi->sb, ret, total_freed, total_sec_freed, - get_pages(sbi, F2FS_DIRTY_NODES), - get_pages(sbi, F2FS_DIRTY_DENTS), - get_pages(sbi, F2FS_DIRTY_IMETA), + get_nr_caches(sbi, F2FS_DIRTY_NODES), + get_nr_caches(sbi, F2FS_DIRTY_DENTS), + get_nr_caches(sbi, F2FS_DIRTY_IMETA), free_sections(sbi), free_segments(sbi), reserved_segments(sbi), diff --git a/fs/f2fs/inline.c b/fs/f2fs/inline.c index aec06fb4fd76..273b6428c9f9 100644 --- a/fs/f2fs/inline.c +++ b/fs/f2fs/inline.c @@ -34,27 +34,26 @@ bool f2fs_may_inline_data(struct inode *inode) return !f2fs_post_read_required(inode); } -static bool inode_has_blocks(struct inode *inode, struct folio *ifolio) +static bool inode_has_blocks(struct inode *inode, struct f2fs_cached_block *ientry) { - struct f2fs_inode *ri = F2FS_INODE(ifolio); int i; if (F2FS_HAS_BLOCKS(inode)) return true; for (i = 0; i < DEF_NIDS_PER_INODE; i++) { - if (ri->i_nid[i]) + if (F2FS_INODE_NIDS(F2FS_I_SB(inode), ientry)[i]) return true; } return false; } -bool f2fs_sanity_check_inline_data(struct inode *inode, struct folio *ifolio) +bool f2fs_sanity_check_inline_data(struct inode *inode, struct f2fs_cached_block *ientry) { if (!f2fs_has_inline_data(inode)) return false; - if (inode_has_blocks(inode, ifolio)) + if (inode_has_blocks(inode, ientry)) return false; if (!support_inline_data(inode)) @@ -80,7 +79,7 @@ bool f2fs_may_inline_dentry(struct inode *inode) return true; } -void f2fs_do_read_inline_data(struct folio *folio, struct folio *ifolio) +void f2fs_do_read_inline_data(struct folio *folio, struct f2fs_cached_block *ientry) { struct inode *inode = folio->mapping->host; @@ -92,13 +91,13 @@ void f2fs_do_read_inline_data(struct folio *folio, struct folio *ifolio) folio_zero_segment(folio, MAX_INLINE_DATA(inode), folio_size(folio)); /* Copy the whole inline data block */ - memcpy_to_folio(folio, 0, inline_data_addr(inode, ifolio), + memcpy_to_folio(folio, 0, inline_data_addr(inode, ientry), MAX_INLINE_DATA(inode)); if (!folio_test_uptodate(folio)) folio_mark_uptodate(folio); } -void f2fs_truncate_inline_inode(struct inode *inode, struct folio *ifolio, +void f2fs_truncate_inline_inode(struct inode *inode, struct f2fs_cached_block *ientry, u64 from) { void *addr; @@ -106,11 +105,11 @@ void f2fs_truncate_inline_inode(struct inode *inode, struct folio *ifolio, if (from >= MAX_INLINE_DATA(inode)) return; - addr = inline_data_addr(inode, ifolio); + addr = inline_data_addr(inode, ientry); - f2fs_folio_wait_writeback(ifolio, NODE, true, true); + f2fs_cache_wait_writeback(ientry); memset(addr + from, 0, MAX_INLINE_DATA(inode) - from); - folio_mark_dirty(ifolio); + f2fs_mark_cache_dirty(ientry); if (from == 0) clear_inode_flag(inode, FI_DATA_EXIST); @@ -118,27 +117,27 @@ void f2fs_truncate_inline_inode(struct inode *inode, struct folio *ifolio, int f2fs_read_inline_data(struct inode *inode, struct folio *folio) { - struct folio *ifolio; + struct f2fs_cached_block *ientry; - ifolio = f2fs_get_inode_folio(F2FS_I_SB(inode), inode->i_ino); - if (IS_ERR(ifolio)) { + ientry = f2fs_get_inode_cache(F2FS_I_SB(inode), inode->i_ino); + if (IS_ERR(ientry)) { folio_unlock(folio); - return PTR_ERR(ifolio); + return PTR_ERR(ientry); } if (!f2fs_has_inline_data(inode)) { - f2fs_folio_put(ifolio, true); + f2fs_put_cache(ientry, true); return -EAGAIN; } if (folio->index) folio_zero_segment(folio, 0, folio_size(folio)); else - f2fs_do_read_inline_data(folio, ifolio); + f2fs_do_read_inline_data(folio, ientry); if (!folio_test_uptodate(folio)) folio_mark_uptodate(folio); - f2fs_folio_put(ifolio, true); + f2fs_put_cache(ientry, true); folio_unlock(folio); return 0; } @@ -186,7 +185,7 @@ int f2fs_convert_inline_folio(struct dnode_of_data *dn, struct folio *folio) f2fs_bug_on(F2FS_F_SB(folio), folio_test_writeback(folio)); - f2fs_do_read_inline_data(folio, dn->inode_folio); + f2fs_do_read_inline_data(folio, dn->inode_entry); folio_mark_dirty(folio); /* clear dirty state */ @@ -197,7 +196,7 @@ int f2fs_convert_inline_folio(struct dnode_of_data *dn, struct folio *folio) fio.old_blkaddr = dn->data_blkaddr; set_inode_flag(dn->inode, FI_HOT_DATA); f2fs_outplace_write_data(dn, &fio); - f2fs_folio_wait_writeback(folio, DATA, true, true); + f2fs_folio_wait_writeback(folio, true, true); if (dirty) { inode_dec_dirty_pages(dn->inode); f2fs_remove_dirty_inode(dn->inode); @@ -207,8 +206,8 @@ int f2fs_convert_inline_folio(struct dnode_of_data *dn, struct folio *folio) set_inode_flag(dn->inode, FI_APPEND_WRITE); /* clear inline data and flag after data writeback */ - f2fs_truncate_inline_inode(dn->inode, dn->inode_folio, 0); - folio_clear_f2fs_inline(dn->inode_folio); + f2fs_truncate_inline_inode(dn->inode, dn->inode_entry, 0); + f2fs_cache_clear_inline(dn->inode_entry); clear_out: stat_dec_inline_inode(dn->inode); clear_inode_flag(dn->inode, FI_INLINE_DATA); @@ -221,7 +220,8 @@ int f2fs_convert_inline_inode(struct inode *inode) struct f2fs_sb_info *sbi = F2FS_I_SB(inode); struct dnode_of_data dn; struct f2fs_lock_context lc; - struct folio *ifolio, *folio; + struct f2fs_cached_block *ientry; + struct folio *folio; int err = 0; if (f2fs_hw_is_readonly(sbi) || f2fs_readonly(sbi->sb)) @@ -240,13 +240,13 @@ int f2fs_convert_inline_inode(struct inode *inode) f2fs_lock_op(sbi, &lc); - ifolio = f2fs_get_inode_folio(sbi, inode->i_ino); - if (IS_ERR(ifolio)) { - err = PTR_ERR(ifolio); + ientry = f2fs_get_inode_cache(sbi, inode->i_ino); + if (IS_ERR(ientry)) { + err = PTR_ERR(ientry); goto out; } - set_new_dnode(&dn, inode, ifolio, ifolio, 0); + set_new_dnode(&dn, inode, ientry, ientry, 0); if (f2fs_has_inline_data(inode)) err = f2fs_convert_inline_folio(&dn, folio); @@ -266,35 +266,35 @@ out: int f2fs_write_inline_data(struct inode *inode, struct folio *folio) { struct f2fs_sb_info *sbi = F2FS_I_SB(inode); - struct folio *ifolio; + struct f2fs_cached_block *ientry; - ifolio = f2fs_get_inode_folio(sbi, inode->i_ino); - if (IS_ERR(ifolio)) - return PTR_ERR(ifolio); + ientry = f2fs_get_inode_cache(sbi, inode->i_ino); + if (IS_ERR(ientry)) + return PTR_ERR(ientry); if (!f2fs_has_inline_data(inode)) { - f2fs_folio_put(ifolio, true); + f2fs_put_cache(ientry, true); return -EAGAIN; } f2fs_bug_on(F2FS_I_SB(inode), folio->index); - f2fs_folio_wait_writeback(ifolio, NODE, true, true); - memcpy_from_folio(inline_data_addr(inode, ifolio), + f2fs_cache_wait_writeback(ientry); + memcpy_from_folio(inline_data_addr(inode, ientry), folio, 0, MAX_INLINE_DATA(inode)); - folio_mark_dirty(ifolio); + f2fs_mark_cache_dirty(ientry); f2fs_clear_page_cache_dirty_tag(folio); set_inode_flag(inode, FI_APPEND_WRITE); set_inode_flag(inode, FI_DATA_EXIST); - folio_clear_f2fs_inline(ifolio); - f2fs_folio_put(ifolio, true); + f2fs_cache_clear_inline(ientry); + f2fs_put_cache(ientry, true); return 0; } -int f2fs_recover_inline_data(struct inode *inode, struct folio *nfolio) +int f2fs_recover_inline_data(struct inode *inode, struct f2fs_cached_block *entry) { struct f2fs_sb_info *sbi = F2FS_I_SB(inode); struct f2fs_inode *ri = NULL; @@ -308,39 +308,41 @@ int f2fs_recover_inline_data(struct inode *inode, struct folio *nfolio) * x o -> remove data blocks, and then recover inline_data * x x -> recover data blocks */ - if (IS_INODE(nfolio)) - ri = F2FS_INODE(nfolio); + if (IS_INODE(F2FS_I_SB(inode), entry)) + ri = &CACHED_NODE(entry)->i; if (f2fs_has_inline_data(inode) && ri && (ri->i_inline & F2FS_INLINE_DATA)) { - struct folio *ifolio; + struct f2fs_cached_block *ientry; + process_inline: - ifolio = f2fs_get_inode_folio(sbi, inode->i_ino); - if (IS_ERR(ifolio)) - return PTR_ERR(ifolio); + ientry = f2fs_get_inode_cache(sbi, inode->i_ino); + if (IS_ERR(ientry)) + return PTR_ERR(ientry); - f2fs_folio_wait_writeback(ifolio, NODE, true, true); + f2fs_cache_wait_writeback(ientry); - src_addr = inline_data_addr(inode, nfolio); - dst_addr = inline_data_addr(inode, ifolio); + src_addr = inline_data_addr(inode, entry); + dst_addr = inline_data_addr(inode, ientry); memcpy(dst_addr, src_addr, MAX_INLINE_DATA(inode)); set_inode_flag(inode, FI_INLINE_DATA); set_inode_flag(inode, FI_DATA_EXIST); - folio_mark_dirty(ifolio); - f2fs_folio_put(ifolio, true); + f2fs_mark_cache_dirty(ientry); + f2fs_put_cache(ientry, true); return 1; } if (f2fs_has_inline_data(inode)) { - struct folio *ifolio = f2fs_get_inode_folio(sbi, inode->i_ino); - if (IS_ERR(ifolio)) - return PTR_ERR(ifolio); - f2fs_truncate_inline_inode(inode, ifolio, 0); + struct f2fs_cached_block *ientry = f2fs_get_inode_cache(sbi, inode->i_ino); + + if (IS_ERR(ientry)) + return PTR_ERR(ientry); + f2fs_truncate_inline_inode(inode, ientry, 0); stat_dec_inline_inode(inode); clear_inode_flag(inode, FI_INLINE_DATA); - f2fs_folio_put(ifolio, true); + f2fs_put_cache(ientry, true); } else if (ri && (ri->i_inline & F2FS_INLINE_DATA)) { int ret; @@ -355,50 +357,50 @@ process_inline: struct f2fs_dir_entry *f2fs_find_in_inline_dir(struct inode *dir, const struct f2fs_filename *fname, - struct folio **res_folio, + void **dentry_block, bool use_hash) { struct f2fs_sb_info *sbi = F2FS_SB(dir->i_sb); struct f2fs_dir_entry *de; struct f2fs_dentry_ptr d; - struct folio *ifolio; + struct f2fs_cached_block *ientry; void *inline_dentry; - ifolio = f2fs_get_inode_folio(sbi, dir->i_ino); - if (IS_ERR(ifolio)) { - *res_folio = ifolio; + ientry = f2fs_get_inode_cache(sbi, dir->i_ino); + if (IS_ERR(ientry)) { + *dentry_block = ientry; return NULL; } - inline_dentry = inline_data_addr(dir, ifolio); + inline_dentry = inline_data_addr(dir, ientry); make_dentry_ptr_inline(dir, &d, inline_dentry); de = f2fs_find_target_dentry(&d, fname, NULL, use_hash); - folio_unlock(ifolio); + f2fs_unlock_cache(ientry); if (IS_ERR(de)) { - *res_folio = ERR_CAST(de); + *dentry_block = ERR_CAST(de); de = NULL; } if (de) - *res_folio = ifolio; + *dentry_block = f2fs_cache_make_dentry_block(ientry); else - f2fs_folio_put(ifolio, false); + f2fs_put_cache(ientry, false); return de; } int f2fs_make_empty_inline_dir(struct inode *inode, struct inode *parent, - struct folio *ifolio) + struct f2fs_cached_block *ientry) { struct f2fs_dentry_ptr d; void *inline_dentry; - inline_dentry = inline_data_addr(inode, ifolio); + inline_dentry = inline_data_addr(inode, ientry); make_dentry_ptr_inline(inode, &d, inline_dentry); f2fs_do_make_empty_dir(inode, parent, &d); - folio_mark_dirty(ifolio); + f2fs_mark_cache_dirty(ientry); /* update i_size to MAX_INLINE_DATA */ if (i_size_read(inode) < MAX_INLINE_DATA(inode)) @@ -410,22 +412,23 @@ int f2fs_make_empty_inline_dir(struct inode *inode, struct inode *parent, * NOTE: ipage is grabbed by caller, but if any error occurs, we should * release ipage in this function. */ -static int f2fs_move_inline_dirents(struct inode *dir, struct folio *ifolio, - void *inline_dentry) +static int f2fs_move_inline_dirents(struct inode *dir, + struct f2fs_cached_block *ientry, + void *inline_dentry) { struct folio *folio; struct dnode_of_data dn; - struct f2fs_dentry_block *dentry_blk; + void *dentry_blk; struct f2fs_dentry_ptr src, dst; int err; folio = f2fs_grab_cache_folio(dir->i_mapping, 0, true); if (IS_ERR(folio)) { - f2fs_folio_put(ifolio, true); + f2fs_put_cache(ientry, true); return PTR_ERR(folio); } - set_new_dnode(&dn, dir, ifolio, NULL, 0); + set_new_dnode(&dn, dir, ientry, NULL, 0); err = f2fs_reserve_block(&dn, 0); if (err) goto out; @@ -441,7 +444,7 @@ static int f2fs_move_inline_dirents(struct inode *dir, struct folio *ifolio, goto out; } - f2fs_folio_wait_writeback(folio, DATA, true, true); + f2fs_folio_wait_writeback(folio, true, true); dentry_blk = folio_address(folio); @@ -449,7 +452,7 @@ static int f2fs_move_inline_dirents(struct inode *dir, struct folio *ifolio, * Start by zeroing the full block, to ensure that all unused space is * zeroed and no uninitialized memory is leaked to disk. */ - memset(dentry_blk, 0, F2FS_BLKSIZE); + memset(dentry_blk, 0, F2FS_BLKSIZE(F2FS_I_SB(dir))); make_dentry_ptr_inline(dir, &src, inline_dentry); make_dentry_ptr_block(dir, &dst, dentry_blk); @@ -464,7 +467,7 @@ static int f2fs_move_inline_dirents(struct inode *dir, struct folio *ifolio, folio_mark_dirty(folio); /* clear inline dir and flag after data writeback */ - f2fs_truncate_inline_inode(dir, ifolio, 0); + f2fs_truncate_inline_inode(dir, ientry, 0); stat_dec_inline_dir(dir); clear_inode_flag(dir, FI_INLINE_DENTRY); @@ -478,8 +481,8 @@ static int f2fs_move_inline_dirents(struct inode *dir, struct folio *ifolio, F2FS_I(dir)->i_inline_xattr_size = 0; f2fs_i_depth_write(dir, 1); - if (i_size_read(dir) < PAGE_SIZE) - f2fs_i_size_write(dir, PAGE_SIZE); + if (i_size_read(dir) < F2FS_BLKSIZE(F2FS_I_SB(dir))) + f2fs_i_size_write(dir, F2FS_BLKSIZE(F2FS_I_SB(dir))); out: f2fs_folio_put(folio, true); return err; @@ -544,8 +547,8 @@ punch_dentry_pages: return err; } -static int f2fs_move_rehashed_dirents(struct inode *dir, struct folio *ifolio, - void *inline_dentry) +static int f2fs_move_rehashed_dirents(struct inode *dir, + struct f2fs_cached_block *ientry, void *inline_dentry) { void *backup_dentry; int err; @@ -553,20 +556,20 @@ static int f2fs_move_rehashed_dirents(struct inode *dir, struct folio *ifolio, backup_dentry = f2fs_kmalloc(F2FS_I_SB(dir), MAX_INLINE_DATA(dir), GFP_F2FS_ZERO); if (!backup_dentry) { - f2fs_folio_put(ifolio, true); + f2fs_put_cache(ientry, true); return -ENOMEM; } memcpy(backup_dentry, inline_dentry, MAX_INLINE_DATA(dir)); - f2fs_truncate_inline_inode(dir, ifolio, 0); + f2fs_truncate_inline_inode(dir, ientry, 0); - folio_unlock(ifolio); + f2fs_unlock_cache(ientry); err = f2fs_add_inline_entries(dir, backup_dentry); if (err) goto recover; - folio_lock(ifolio); + f2fs_lock_cache(ientry); stat_dec_inline_dir(dir); clear_inode_flag(dir, FI_INLINE_DENTRY); @@ -582,31 +585,31 @@ static int f2fs_move_rehashed_dirents(struct inode *dir, struct folio *ifolio, kfree(backup_dentry); return 0; recover: - folio_lock(ifolio); - f2fs_folio_wait_writeback(ifolio, NODE, true, true); + f2fs_lock_cache(ientry); + f2fs_cache_wait_writeback(ientry); memcpy(inline_dentry, backup_dentry, MAX_INLINE_DATA(dir)); f2fs_i_depth_write(dir, 0); f2fs_i_size_write(dir, MAX_INLINE_DATA(dir)); - folio_mark_dirty(ifolio); - f2fs_folio_put(ifolio, true); + f2fs_mark_cache_dirty(ientry); + f2fs_put_cache(ientry, true); kfree(backup_dentry); return err; } -static int do_convert_inline_dir(struct inode *dir, struct folio *ifolio, - void *inline_dentry) +static int do_convert_inline_dir(struct inode *dir, + struct f2fs_cached_block *ientry, void *inline_dentry) { if (!F2FS_I(dir)->i_dir_level) - return f2fs_move_inline_dirents(dir, ifolio, inline_dentry); + return f2fs_move_inline_dirents(dir, ientry, inline_dentry); else - return f2fs_move_rehashed_dirents(dir, ifolio, inline_dentry); + return f2fs_move_rehashed_dirents(dir, ientry, inline_dentry); } int f2fs_try_convert_inline_dir(struct inode *dir, struct dentry *dentry) { struct f2fs_sb_info *sbi = F2FS_I_SB(dir); - struct folio *ifolio; + struct f2fs_cached_block *ientry; struct f2fs_filename fname; struct f2fs_lock_context lc; void *inline_dentry = NULL; @@ -621,22 +624,22 @@ int f2fs_try_convert_inline_dir(struct inode *dir, struct dentry *dentry) if (err) goto out; - ifolio = f2fs_get_inode_folio(sbi, dir->i_ino); - if (IS_ERR(ifolio)) { - err = PTR_ERR(ifolio); + ientry = f2fs_get_inode_cache(sbi, dir->i_ino); + if (IS_ERR(ientry)) { + err = PTR_ERR(ientry); goto out_fname; } - if (f2fs_has_enough_room(dir, ifolio, &fname)) { - f2fs_folio_put(ifolio, true); + if (f2fs_has_enough_room(dir, ientry, &fname)) { + f2fs_put_cache(ientry, true); goto out_fname; } - inline_dentry = inline_data_addr(dir, ifolio); + inline_dentry = inline_data_addr(dir, ientry); - err = do_convert_inline_dir(dir, ifolio, inline_dentry); + err = do_convert_inline_dir(dir, ientry, inline_dentry); if (!err) - f2fs_folio_put(ifolio, true); + f2fs_put_cache(ientry, true); out_fname: f2fs_free_filename(&fname); out: @@ -648,24 +651,24 @@ int f2fs_add_inline_entry(struct inode *dir, const struct f2fs_filename *fname, struct inode *inode, nid_t ino, umode_t mode) { struct f2fs_sb_info *sbi = F2FS_I_SB(dir); - struct folio *ifolio; + struct f2fs_cached_block *ientry; unsigned int bit_pos; void *inline_dentry = NULL; struct f2fs_dentry_ptr d; int slots = GET_DENTRY_SLOTS(fname->disk_name.len); - struct folio *folio = NULL; + struct f2fs_cached_block *nentry = NULL; int err = 0; - ifolio = f2fs_get_inode_folio(sbi, dir->i_ino); - if (IS_ERR(ifolio)) - return PTR_ERR(ifolio); + ientry = f2fs_get_inode_cache(sbi, dir->i_ino); + if (IS_ERR(ientry)) + return PTR_ERR(ientry); - inline_dentry = inline_data_addr(dir, ifolio); + inline_dentry = inline_data_addr(dir, ientry); make_dentry_ptr_inline(dir, &d, inline_dentry); bit_pos = f2fs_room_for_filename(d.bitmap, slots, d.max); if (bit_pos >= d.max) { - err = do_convert_inline_dir(dir, ifolio, inline_dentry); + err = do_convert_inline_dir(dir, ientry, inline_dentry); if (err) return err; err = -EAGAIN; @@ -675,19 +678,19 @@ int f2fs_add_inline_entry(struct inode *dir, const struct f2fs_filename *fname, if (inode) { f2fs_down_write_nested(&F2FS_I(inode)->i_sem, SINGLE_DEPTH_NESTING); - folio = f2fs_init_inode_metadata(inode, dir, fname, ifolio); - if (IS_ERR(folio)) { - err = PTR_ERR(folio); + nentry = f2fs_init_inode_metadata(inode, dir, fname, ientry); + if (IS_ERR(nentry)) { + err = PTR_ERR(nentry); goto fail; } } - f2fs_folio_wait_writeback(ifolio, NODE, true, true); + f2fs_cache_wait_writeback(ientry); f2fs_update_dentry(ino, mode, &d, &fname->disk_name, fname->hash, bit_pos); - folio_mark_dirty(ifolio); + f2fs_mark_cache_dirty(ientry); /* we don't need to mark_inode_dirty now */ if (inode) { @@ -695,9 +698,9 @@ int f2fs_add_inline_entry(struct inode *dir, const struct f2fs_filename *fname, /* synchronize inode page's data from inode cache */ if (is_inode_flag_set(inode, FI_NEW_INODE)) - f2fs_update_inode(inode, folio); + f2fs_update_inode(inode, nentry); - f2fs_folio_put(folio, true); + f2fs_put_cache(nentry, true); } f2fs_update_parent_metadata(dir, inode, 0); @@ -705,12 +708,12 @@ fail: if (inode) f2fs_up_write(&F2FS_I(inode)->i_sem); out: - f2fs_folio_put(ifolio, true); + f2fs_put_cache(ientry, true); return err; } void f2fs_delete_inline_entry(struct f2fs_dir_entry *dentry, - struct folio *folio, struct inode *dir, struct inode *inode) + struct f2fs_cached_block *ientry, struct inode *dir, struct inode *inode) { struct f2fs_dentry_ptr d; void *inline_dentry; @@ -718,18 +721,18 @@ void f2fs_delete_inline_entry(struct f2fs_dir_entry *dentry, unsigned int bit_pos; int i; - folio_lock(folio); - f2fs_folio_wait_writeback(folio, NODE, true, true); + f2fs_lock_cache(ientry); + f2fs_cache_wait_writeback(ientry); - inline_dentry = inline_data_addr(dir, folio); + inline_dentry = inline_data_addr(dir, ientry); make_dentry_ptr_inline(dir, &d, inline_dentry); bit_pos = dentry - d.dentry; for (i = 0; i < slots; i++) __clear_bit_le(bit_pos + i, d.bitmap); - folio_mark_dirty(folio); - f2fs_folio_put(folio, true); + f2fs_mark_cache_dirty(ientry); + f2fs_put_cache(ientry, true); inode_set_mtime_to_ts(dir, inode_set_ctime_current(dir)); f2fs_mark_inode_dirty_sync(dir, true); @@ -741,21 +744,21 @@ void f2fs_delete_inline_entry(struct f2fs_dir_entry *dentry, bool f2fs_empty_inline_dir(struct inode *dir) { struct f2fs_sb_info *sbi = F2FS_I_SB(dir); - struct folio *ifolio; + struct f2fs_cached_block *ientry; unsigned int bit_pos = 2; void *inline_dentry; struct f2fs_dentry_ptr d; - ifolio = f2fs_get_inode_folio(sbi, dir->i_ino); - if (IS_ERR(ifolio)) + ientry = f2fs_get_inode_cache(sbi, dir->i_ino); + if (IS_ERR(ientry)) return false; - inline_dentry = inline_data_addr(dir, ifolio); + inline_dentry = inline_data_addr(dir, ientry); make_dentry_ptr_inline(dir, &d, inline_dentry); bit_pos = find_next_bit_le(d.bitmap, d.max, bit_pos); - f2fs_folio_put(ifolio, true); + f2fs_put_cache(ientry, true); if (bit_pos < d.max) return false; @@ -767,7 +770,7 @@ int f2fs_read_inline_dir(struct file *file, struct dir_context *ctx, struct fscrypt_str *fstr) { struct inode *inode = file_inode(file); - struct folio *ifolio = NULL; + struct f2fs_cached_block *ientry = NULL; struct f2fs_dentry_ptr d; void *inline_dentry = NULL; int err; @@ -777,17 +780,17 @@ int f2fs_read_inline_dir(struct file *file, struct dir_context *ctx, if (ctx->pos == d.max) return 0; - ifolio = f2fs_get_inode_folio(F2FS_I_SB(inode), inode->i_ino); - if (IS_ERR(ifolio)) - return PTR_ERR(ifolio); + ientry = f2fs_get_inode_cache(F2FS_I_SB(inode), inode->i_ino); + if (IS_ERR(ientry)) + return PTR_ERR(ientry); /* * f2fs_readdir was protected by inode.i_rwsem, it is safe to access * ipage without page's lock held. */ - folio_unlock(ifolio); + f2fs_unlock_cache(ientry); - inline_dentry = inline_data_addr(inode, ifolio); + inline_dentry = inline_data_addr(inode, ientry); make_dentry_ptr_inline(inode, &d, inline_dentry); @@ -795,7 +798,7 @@ int f2fs_read_inline_dir(struct file *file, struct dir_context *ctx, if (!err) ctx->pos = d.max; - f2fs_folio_put(ifolio, false); + f2fs_put_cache(ientry, false); return err < 0 ? err : 0; } @@ -806,12 +809,12 @@ int f2fs_inline_data_fiemap(struct inode *inode, __u32 flags = FIEMAP_EXTENT_DATA_INLINE | FIEMAP_EXTENT_NOT_ALIGNED | FIEMAP_EXTENT_LAST; struct node_info ni; - struct folio *ifolio; + struct f2fs_cached_block *ientry; int err = 0; - ifolio = f2fs_get_inode_folio(F2FS_I_SB(inode), inode->i_ino); - if (IS_ERR(ifolio)) - return PTR_ERR(ifolio); + ientry = f2fs_get_inode_cache(F2FS_I_SB(inode), inode->i_ino); + if (IS_ERR(ientry)) + return PTR_ERR(ientry); if ((S_ISREG(inode->i_mode) || S_ISLNK(inode->i_mode)) && !f2fs_has_inline_data(inode)) { @@ -825,13 +828,13 @@ int f2fs_inline_data_fiemap(struct inode *inode, } if (fieinfo->fi_flags & FIEMAP_FLAG_SYNC) { - err = f2fs_write_single_node_folio(ifolio, true, false, FS_NODE_IO); + err = f2fs_write_node_cache(ientry, true, false, FS_NODE_IO); if (err) return err; - ifolio = f2fs_get_inode_folio(F2FS_I_SB(inode), inode->i_ino); - if (IS_ERR(ifolio)) - return PTR_ERR(ifolio); - f2fs_folio_wait_writeback(ifolio, NODE, true, true); + ientry = f2fs_get_inode_cache(F2FS_I_SB(inode), inode->i_ino); + if (IS_ERR(ientry)) + return PTR_ERR(ientry); + f2fs_cache_wait_writeback(ientry); } ilen = min_t(size_t, MAX_INLINE_DATA(inode), i_size_read(inode)); if (start >= ilen) @@ -846,8 +849,8 @@ int f2fs_inline_data_fiemap(struct inode *inode, if (__is_valid_data_blkaddr(ni.blk_addr)) { byteaddr = (__u64)ni.blk_addr << inode->i_sb->s_blocksize_bits; - byteaddr += (char *)inline_data_addr(inode, ifolio) - - (char *)F2FS_INODE(ifolio); + byteaddr += (char *)inline_data_addr(inode, ientry) - + (char *)F2FS_INODE(ientry); } else { f2fs_bug_on(F2FS_I_SB(inode), ni.blk_addr != NEW_ADDR); flags |= FIEMAP_EXTENT_DELALLOC | FIEMAP_EXTENT_UNKNOWN; @@ -855,6 +858,6 @@ int f2fs_inline_data_fiemap(struct inode *inode, err = fiemap_fill_next_extent(fieinfo, start, byteaddr, ilen, flags); trace_f2fs_fiemap(inode, start, byteaddr, ilen, flags, err); out: - f2fs_folio_put(ifolio, true); + f2fs_put_cache(ientry, true); return err; } diff --git a/fs/f2fs/inode.c b/fs/f2fs/inode.c index 96cc0e777567..9164a2b5d9f0 100644 --- a/fs/f2fs/inode.c +++ b/fs/f2fs/inode.c @@ -82,9 +82,9 @@ void f2fs_set_inode_flags(struct inode *inode) S_ENCRYPTED|S_VERITY|S_CASEFOLD); } -static void __get_inode_rdev(struct inode *inode, struct folio *node_folio) +static void __get_inode_rdev(struct inode *inode, struct f2fs_cached_block *ientry) { - __le32 *addr = get_dnode_addr(inode, node_folio); + __le32 *addr = get_dnode_addr(inode, ientry); if (S_ISCHR(inode->i_mode) || S_ISBLK(inode->i_mode) || S_ISFIFO(inode->i_mode) || S_ISSOCK(inode->i_mode)) { @@ -95,9 +95,9 @@ static void __get_inode_rdev(struct inode *inode, struct folio *node_folio) } } -static void __set_inode_rdev(struct inode *inode, struct folio *node_folio) +static void __set_inode_rdev(struct inode *inode, struct f2fs_cached_block *ientry) { - __le32 *addr = get_dnode_addr(inode, node_folio); + __le32 *addr = get_dnode_addr(inode, ientry); if (S_ISCHR(inode->i_mode) || S_ISBLK(inode->i_mode)) { if (old_valid_dev(inode->i_rdev)) { @@ -111,34 +111,33 @@ static void __set_inode_rdev(struct inode *inode, struct folio *node_folio) } } -static void __recover_inline_status(struct inode *inode, struct folio *ifolio) +static void __recover_inline_status(struct inode *inode, struct f2fs_cached_block *ientry) { - void *inline_data = inline_data_addr(inode, ifolio); + void *inline_data = inline_data_addr(inode, ientry); __le32 *start = inline_data; __le32 *end = start + MAX_INLINE_DATA(inode) / sizeof(__le32); while (start < end) { if (*start++) { - f2fs_folio_wait_writeback(ifolio, NODE, true, true); + f2fs_cache_wait_writeback(ientry); set_inode_flag(inode, FI_DATA_EXIST); - set_raw_inline(inode, F2FS_INODE(ifolio)); - folio_mark_dirty(ifolio); + set_raw_inline(inode, F2FS_INODE(ientry)); + f2fs_mark_cache_dirty(ientry); return; } } return; } -static -bool f2fs_enable_inode_chksum(struct f2fs_sb_info *sbi, struct folio *folio) +static bool f2fs_enable_inode_chksum(struct f2fs_sb_info *sbi, struct f2fs_cached_block *entry) { - struct f2fs_inode *ri = &F2FS_NODE(folio)->i; + struct f2fs_inode *ri = F2FS_INODE(entry); if (!f2fs_sb_has_inode_chksum(sbi)) return false; - if (!IS_INODE(folio) || !(ri->i_inline & F2FS_EXTRA_ATTR)) + if (!IS_INODE(sbi, entry) || !(ri->i_inline & F2FS_EXTRA_ATTR)) return false; if (!F2FS_FITS_IN_INODE(ri, le16_to_cpu(ri->i_extra_isize), @@ -148,11 +147,10 @@ bool f2fs_enable_inode_chksum(struct f2fs_sb_info *sbi, struct folio *folio) return true; } -static __u32 f2fs_inode_chksum(struct f2fs_sb_info *sbi, struct folio *folio) +static __u32 f2fs_inode_chksum(struct f2fs_sb_info *sbi, struct f2fs_cached_block *entry) { - struct f2fs_node *node = F2FS_NODE(folio); - struct f2fs_inode *ri = &node->i; - __le32 ino = node->footer.ino; + struct f2fs_inode *ri = F2FS_INODE(entry); + __le32 ino = F2FS_NODE_FOOTER(sbi, entry)->ino; __le32 gen = ri->i_generation; __u32 chksum, chksum_seed; __u32 dummy_cs = 0; @@ -166,11 +164,11 @@ static __u32 f2fs_inode_chksum(struct f2fs_sb_info *sbi, struct folio *folio) chksum = f2fs_chksum(chksum, (__u8 *)&dummy_cs, cs_size); offset += cs_size; chksum = f2fs_chksum(chksum, (__u8 *)ri + offset, - F2FS_BLKSIZE - offset); + F2FS_BLKSIZE(sbi) - offset); return chksum; } -bool f2fs_inode_chksum_verify(struct f2fs_sb_info *sbi, struct folio *folio) +bool f2fs_inode_chksum_verify(struct f2fs_sb_info *sbi, struct f2fs_cached_block *entry) { struct f2fs_inode *ri; __u32 provided, calculated; @@ -178,35 +176,33 @@ bool f2fs_inode_chksum_verify(struct f2fs_sb_info *sbi, struct folio *folio) if (unlikely(is_sbi_flag_set(sbi, SBI_IS_SHUTDOWN))) return true; -#ifdef CONFIG_F2FS_CHECK_FS - if (!f2fs_enable_inode_chksum(sbi, folio)) -#else - if (!f2fs_enable_inode_chksum(sbi, folio) || - folio_test_dirty(folio) || - folio_test_writeback(folio)) -#endif + if (!f2fs_enable_inode_chksum(sbi, entry)) + return true; +#ifndef CONFIG_F2FS_CHECK_FS + if (f2fs_cache_test_dirty(entry) || f2fs_cache_test_writeback(entry)) return true; +#endif - ri = &F2FS_NODE(folio)->i; + ri = F2FS_INODE(entry); provided = le32_to_cpu(ri->i_inode_checksum); - calculated = f2fs_inode_chksum(sbi, folio); + calculated = f2fs_inode_chksum(sbi, entry); if (provided != calculated) - f2fs_warn(sbi, "checksum invalid, nid = %lu, ino_of_node = %x, %x vs. %x", - folio->index, ino_of_node(folio), + f2fs_warn(sbi, "checksum invalid, nid = %lu, ino_of_node = %u, %x vs. %x", + entry->index, ino_of_node(sbi, entry), provided, calculated); return provided == calculated; } -void f2fs_inode_chksum_set(struct f2fs_sb_info *sbi, struct folio *folio) +void f2fs_inode_chksum_set(struct f2fs_sb_info *sbi, struct f2fs_cached_block *entry) { - struct f2fs_inode *ri = &F2FS_NODE(folio)->i; + struct f2fs_inode *ri = F2FS_INODE(entry); - if (!f2fs_enable_inode_chksum(sbi, folio)) + if (!f2fs_enable_inode_chksum(sbi, entry)) return; - ri->i_inode_checksum = cpu_to_le32(f2fs_inode_chksum(sbi, folio)); + ri->i_inode_checksum = cpu_to_le32(f2fs_inode_chksum(sbi, entry)); } static bool sanity_check_compress_inode(struct inode *inode, @@ -222,11 +218,11 @@ static bool sanity_check_compress_inode(struct inode *inode, return false; } if (le64_to_cpu(ri->i_compr_blocks) > - SECTOR_TO_BLOCK(inode->i_blocks)) { + SECTOR_TO_BLOCK(sbi, inode->i_blocks)) { f2fs_warn(sbi, "%s: inode (ino=%llx) has inconsistent i_compr_blocks:%llu, i_blocks:%llu, run fsck to fix", __func__, inode->i_ino, le64_to_cpu(ri->i_compr_blocks), - SECTOR_TO_BLOCK(inode->i_blocks)); + SECTOR_TO_BLOCK(sbi, inode->i_blocks)); return false; } if (ri->i_log_cluster_size < MIN_COMPRESS_LOG_SIZE || @@ -281,28 +277,29 @@ err_level: return false; } -static bool sanity_check_inode(struct inode *inode, struct folio *node_folio) +static bool sanity_check_inode(struct inode *inode, + struct f2fs_cached_block *node_entry) { struct f2fs_sb_info *sbi = F2FS_I_SB(inode); struct f2fs_inode_info *fi = F2FS_I(inode); - struct f2fs_inode *ri = F2FS_INODE(node_folio); + struct f2fs_inode *ri = F2FS_INODE(node_entry); unsigned long long iblocks; - iblocks = le64_to_cpu(F2FS_INODE(node_folio)->i_blocks); + iblocks = le64_to_cpu(ri->i_blocks); if (!iblocks) { f2fs_warn(sbi, "%s: corrupted inode i_blocks i_ino=%llx iblocks=%llu, run fsck to fix.", __func__, inode->i_ino, iblocks); return false; } - if (ino_of_node(node_folio) != nid_of_node(node_folio)) { + if (ino_of_node(sbi, node_entry) != nid_of_node(sbi, node_entry)) { f2fs_warn(sbi, "%s: corrupted inode footer i_ino=%llx, ino,nid: [%u, %u] run fsck to fix.", __func__, inode->i_ino, - ino_of_node(node_folio), nid_of_node(node_folio)); + ino_of_node(sbi, node_entry), nid_of_node(sbi, node_entry)); return false; } - if (ino_of_node(node_folio) == fi->i_xattr_nid) { + if (ino_of_node(sbi, node_entry) == fi->i_xattr_nid) { f2fs_warn(sbi, "%s: corrupted inode i_ino=%llx, xnid=%x, run fsck to fix.", __func__, inode->i_ino, fi->i_xattr_nid); return false; @@ -338,12 +335,13 @@ static bool sanity_check_inode(struct inode *inode, struct folio *node_folio) } if (f2fs_sb_has_flexible_inline_xattr(sbi) && - (fi->i_inline_xattr_size > MAX_INLINE_XATTR_SIZE || + (fi->i_inline_xattr_size > MAX_INLINE_XATTR_SIZE(i_blocksize(inode)) || (f2fs_has_inline_xattr(inode) && fi->i_inline_xattr_size < MIN_INLINE_XATTR_SIZE))) { - f2fs_warn(sbi, "%s: inode (ino=%llx) has corrupted i_inline_xattr_size: %d, min: %zu, max: %lu", + f2fs_warn(sbi, "%s: inode (ino=%llx) has corrupted i_inline_xattr_size: %d, min: %zu, max: %zu", __func__, inode->i_ino, fi->i_inline_xattr_size, - MIN_INLINE_XATTR_SIZE, MAX_INLINE_XATTR_SIZE); + MIN_INLINE_XATTR_SIZE, + (size_t)MAX_INLINE_XATTR_SIZE(i_blocksize(inode))); return false; } @@ -375,7 +373,7 @@ static bool sanity_check_inode(struct inode *inode, struct folio *node_folio) } } - if (f2fs_sanity_check_inline_data(inode, node_folio)) { + if (f2fs_sanity_check_inline_data(inode, node_entry)) { f2fs_warn(sbi, "%s: inode (ino=%llx, mode=%u) should not have inline_data, run fsck to fix", __func__, inode->i_ino, inode->i_mode); return false; @@ -428,7 +426,7 @@ static int do_read_inode(struct inode *inode) { struct f2fs_sb_info *sbi = F2FS_I_SB(inode); struct f2fs_inode_info *fi = F2FS_I(inode); - struct folio *node_folio; + struct f2fs_cached_block *node_entry; struct f2fs_inode *ri; projid_t i_projid; @@ -436,18 +434,19 @@ static int do_read_inode(struct inode *inode) if (f2fs_check_nid_range(sbi, inode->i_ino)) return -EINVAL; - node_folio = f2fs_get_inode_folio(sbi, inode->i_ino); - if (IS_ERR(node_folio)) - return PTR_ERR(node_folio); + node_entry = f2fs_get_inode_cache(sbi, inode->i_ino); + if (IS_ERR(node_entry)) + return PTR_ERR(node_entry); - ri = F2FS_INODE(node_folio); + ri = F2FS_INODE(node_entry); inode->i_mode = le16_to_cpu(ri->i_mode); i_uid_write(inode, le32_to_cpu(ri->i_uid)); i_gid_write(inode, le32_to_cpu(ri->i_gid)); set_nlink(inode, le32_to_cpu(ri->i_links)); inode->i_size = le64_to_cpu(ri->i_size); - inode->i_blocks = SECTOR_FROM_BLOCK(le64_to_cpu(ri->i_blocks) - 1); + inode->i_blocks = SECTOR_FROM_BLOCK(sbi, + le64_to_cpu(ri->i_blocks) - 1); inode_set_atime(inode, le64_to_cpu(ri->i_atime), le32_to_cpu(ri->i_atime_nsec)); @@ -490,8 +489,8 @@ static int do_read_inode(struct inode *inode) fi->i_inline_xattr_size = 0; } - if (!sanity_check_inode(inode, node_folio)) { - f2fs_folio_put(node_folio, true); + if (!sanity_check_inode(inode, node_entry)) { + f2fs_put_cache(node_entry, true); set_sbi_flag(sbi, SBI_NEED_FSCK); f2fs_handle_error(sbi, ERROR_CORRUPTED_INODE); fserror_report_file_metadata(inode, -EFSCORRUPTED, GFP_NOFS); @@ -500,17 +499,17 @@ static int do_read_inode(struct inode *inode) /* check data exist */ if (f2fs_has_inline_data(inode) && !f2fs_exist_data(inode)) - __recover_inline_status(inode, node_folio); + __recover_inline_status(inode, node_entry); /* try to recover cold bit for non-dir inode */ - if (!S_ISDIR(inode->i_mode) && !is_cold_node(node_folio)) { - f2fs_folio_wait_writeback(node_folio, NODE, true, true); - set_cold_node(node_folio, false); - folio_mark_dirty(node_folio); + if (!S_ISDIR(inode->i_mode) && !is_cold_node(sbi, node_entry)) { + f2fs_cache_wait_writeback(node_entry); + set_cold_node(sbi, node_entry, false); + f2fs_mark_cache_dirty(node_entry); } /* get rdev by using inline_info */ - __get_inode_rdev(inode, node_folio); + __get_inode_rdev(inode, node_entry); if (!f2fs_need_inode_block_update(sbi, inode->i_ino)) fi->last_disk_size = inode->i_size; @@ -553,18 +552,18 @@ static int do_read_inode(struct inode *inode) init_idisk_time(inode); - if (!sanity_check_extent_cache(inode, node_folio)) { - f2fs_folio_put(node_folio, true); + if (!sanity_check_extent_cache(inode, node_entry)) { + f2fs_put_cache(node_entry, true); f2fs_handle_error(sbi, ERROR_CORRUPTED_INODE); fserror_report_file_metadata(inode, -EFSCORRUPTED, GFP_NOFS); return -EFSCORRUPTED; } /* Need all the flag bits */ - f2fs_init_read_extent_tree(inode, node_folio); + f2fs_init_read_extent_tree(inode, node_entry); f2fs_init_age_extent_tree(inode); - f2fs_folio_put(node_folio, true); + f2fs_put_cache(node_entry, true); stat_inc_inline_xattr(inode); stat_inc_inline_inode(inode); @@ -575,17 +574,6 @@ static int do_read_inode(struct inode *inode) return 0; } -static bool is_meta_ino(struct f2fs_sb_info *sbi, unsigned int ino) -{ - if (ino == F2FS_NODE_INO(sbi) || ino == F2FS_META_INO(sbi)) - return true; -#ifdef CONFIG_F2FS_FS_COMPRESSION - if (test_opt(sbi, COMPRESS_CACHE) && ino == F2FS_COMPRESS_INO(sbi)) - return true; -#endif - return false; -} - struct inode *f2fs_iget(struct super_block *sb, unsigned long ino) { struct f2fs_sb_info *sbi = F2FS_SB(sb); @@ -597,48 +585,17 @@ struct inode *f2fs_iget(struct super_block *sb, unsigned long ino) return ERR_PTR(-ENOMEM); if (!(inode_state_read_once(inode) & I_NEW)) { - if (is_meta_ino(sbi, ino)) { - f2fs_err(sbi, "inaccessible inode: %lu, run fsck to repair", ino); - set_sbi_flag(sbi, SBI_NEED_FSCK); - ret = -EFSCORRUPTED; - trace_f2fs_iget_exit(inode, ret); - iput(inode); - f2fs_handle_error(sbi, ERROR_CORRUPTED_INODE); - fserror_report_file_metadata(inode, ret, GFP_NOFS); - return ERR_PTR(ret); - } - trace_f2fs_iget(inode); return inode; } - if (is_meta_ino(sbi, ino)) - goto make_now; - ret = do_read_inode(inode); if (ret) goto bad_inode; -make_now: + f2fs_set_inode_flags(inode); - if (ino == F2FS_NODE_INO(sbi)) { - inode->i_mapping->a_ops = &f2fs_node_aops; - mapping_set_gfp_mask(inode->i_mapping, GFP_NOFS); - } else if (ino == F2FS_META_INO(sbi)) { - inode->i_mapping->a_ops = &f2fs_meta_aops; - mapping_set_gfp_mask(inode->i_mapping, GFP_NOFS); - } else if (ino == F2FS_COMPRESS_INO(sbi)) { -#ifdef CONFIG_F2FS_FS_COMPRESSION - inode->i_mapping->a_ops = &f2fs_compress_aops; - /* - * generic_error_remove_folio only truncates pages of regular - * inode - */ - inode->i_mode |= S_IFREG; -#endif - mapping_set_gfp_mask(inode->i_mapping, - GFP_NOFS | __GFP_HIGHMEM | __GFP_MOVABLE); - } else if (S_ISREG(inode->i_mode)) { + if (S_ISREG(inode->i_mode)) { inode->i_op = &f2fs_file_inode_operations; inode->i_fop = &f2fs_file_operations; inode->i_mapping->a_ops = &f2fs_dblock_aops; @@ -694,25 +651,28 @@ retry: return inode; } -void f2fs_update_inode(struct inode *inode, struct folio *node_folio) +void f2fs_update_inode(struct inode *inode, + struct f2fs_cached_block *node_entry) { + struct f2fs_sb_info *sbi = F2FS_I_SB(inode); struct f2fs_inode_info *fi = F2FS_I(inode); struct f2fs_inode *ri; struct extent_tree *et = fi->extent_tree[EX_READ]; - f2fs_folio_wait_writeback(node_folio, NODE, true, true); - folio_mark_dirty(node_folio); + f2fs_cache_wait_writeback(node_entry); + f2fs_mark_cache_dirty(node_entry); f2fs_inode_synced(inode); - ri = F2FS_INODE(node_folio); + ri = F2FS_INODE(node_entry); ri->i_mode = cpu_to_le16(inode->i_mode); ri->i_advise = fi->i_advise; ri->i_uid = cpu_to_le32(i_uid_read(inode)); ri->i_gid = cpu_to_le32(i_gid_read(inode)); ri->i_links = cpu_to_le32(inode->i_nlink); - ri->i_blocks = cpu_to_le64(SECTOR_TO_BLOCK(READ_ONCE(inode->i_blocks)) + 1); + ri->i_blocks = cpu_to_le64(SECTOR_TO_BLOCK(sbi, + READ_ONCE(inode->i_blocks)) + 1); if (!f2fs_is_atomic_file(inode) || is_inode_flag_set(inode, FI_ATOMIC_COMMITTED)) @@ -780,27 +740,27 @@ void f2fs_update_inode(struct inode *inode, struct folio *node_folio) } } - __set_inode_rdev(inode, node_folio); + __set_inode_rdev(inode, node_entry); /* deleted inode */ if (inode->i_nlink == 0) - folio_clear_f2fs_inline(node_folio); + f2fs_cache_clear_inline(node_entry); init_idisk_time(inode); #ifdef CONFIG_F2FS_CHECK_FS - f2fs_inode_chksum_set(F2FS_I_SB(inode), node_folio); + f2fs_inode_chksum_set(F2FS_I_SB(inode), node_entry); #endif } -void f2fs_update_inode_page(struct inode *inode) +void f2fs_update_inode_cache(struct inode *inode) { struct f2fs_sb_info *sbi = F2FS_I_SB(inode); - struct folio *node_folio; + struct f2fs_cached_block *node_entry; int count = 0; retry: - node_folio = f2fs_get_inode_folio(sbi, inode->i_ino); - if (IS_ERR(node_folio)) { - int err = PTR_ERR(node_folio); + node_entry = f2fs_get_inode_cache(sbi, inode->i_ino); + if (IS_ERR(node_entry)) { + int err = PTR_ERR(node_entry); /* The node block was truncated. */ if (err == -ENOENT) @@ -816,18 +776,14 @@ stop_checkpoint: f2fs_stop_checkpoint(sbi, false, STOP_CP_REASON_UPDATE_INODE); return; } - f2fs_update_inode(inode, node_folio); - f2fs_folio_put(node_folio, true); + f2fs_update_inode(inode, node_entry); + f2fs_put_cache(node_entry, true); } int f2fs_write_inode(struct inode *inode, struct writeback_control *wbc) { struct f2fs_sb_info *sbi = F2FS_I_SB(inode); - if (inode->i_ino == F2FS_NODE_INO(sbi) || - inode->i_ino == F2FS_META_INO(sbi)) - return 0; - /* * atime could be updated without dirtying f2fs inode in lazytime mode */ @@ -851,7 +807,7 @@ int f2fs_write_inode(struct inode *inode, struct writeback_control *wbc) * We need to balance fs here to prevent from producing dirty node pages * during the urgent cleaning time when running out of free sections. */ - f2fs_update_inode_page(inode); + f2fs_update_inode_cache(inode); if (wbc && wbc->nr_to_write) f2fs_balance_fs(sbi, true); return 0; @@ -892,7 +848,7 @@ static void f2fs_evict_inode_work(struct work_struct *work) /* * Return true, if we shouldn't go through post_evict_inode. */ -static bool f2fs_pre_evict_inode(struct inode *inode) +static void f2fs_pre_evict_inode(struct inode *inode) { struct f2fs_sb_info *sbi = F2FS_I_SB(inode); struct f2fs_inode_info *fi = F2FS_I(inode); @@ -918,19 +874,12 @@ static bool f2fs_pre_evict_inode(struct inode *inode) test_opt(sbi, COMPRESS_CACHE) && f2fs_compressed_file(inode)) f2fs_invalidate_compress_pages(sbi, inode->i_ino); - if (inode->i_ino == F2FS_NODE_INO(sbi) || - inode->i_ino == F2FS_META_INO(sbi) || - inode->i_ino == F2FS_COMPRESS_INO(sbi)) - return true; - f2fs_bug_on(sbi, get_dirty_pages(inode)); f2fs_remove_dirty_inode(inode); f2fs_remove_donate_inode(inode); if (!IS_DEVICE_ALIASING(inode)) f2fs_destroy_extent_tree(inode); - - return false; } static void f2fs_delete_inode(struct inode *inode) @@ -964,7 +913,7 @@ retry: goto error_check; f2fs_lock_op(sbi, &lc); - err = f2fs_remove_inode_page(inode); + err = f2fs_remove_inode_cache(inode); f2fs_unlock_op(sbi, &lc); if (err == -ENOENT) { @@ -996,13 +945,13 @@ error_check: if (!err) goto unfreeze_out; - f2fs_update_inode_page(inode); + f2fs_update_inode_cache(inode); if (dquot_initialize_needed(inode)) set_sbi_flag(sbi, SBI_QUOTA_NEED_REPAIR); /* - * If both f2fs_truncate() and f2fs_update_inode_page() failed + * If both f2fs_truncate() and f2fs_update_inode_cache() failed * due to fuzzed corrupted inode, call f2fs_inode_synced() to * avoid triggering later f2fs_bug_on(). */ @@ -1046,18 +995,17 @@ static void f2fs_post_evict_inode(struct inode *inode) /* for the case f2fs_new_inode() was failed, .i_ino is zero, skip it */ if (inode->i_ino) - invalidate_mapping_pages(NODE_MAPPING(sbi), inode->i_ino, - inode->i_ino); + f2fs_invalidate_node_cache(sbi, inode->i_ino); if (xnid) - invalidate_mapping_pages(NODE_MAPPING(sbi), xnid, xnid); + f2fs_invalidate_node_cache(sbi, xnid); if (!inode->i_nlink) goto skip_record; if (is_inode_flag_set(inode, FI_APPEND_WRITE)) - record_bits = BIT(APPEND_INO); + record_bits |= BIT(APPEND_INO); if (is_inode_flag_set(inode, FI_UPDATE_WRITE)) - record_bits = BIT(UPDATE_INO); + record_bits |= BIT(UPDATE_INO); if (!record_bits) goto skip_record; @@ -1093,15 +1041,13 @@ skip_record: */ void f2fs_evict_inode(struct inode *inode) { - if (f2fs_pre_evict_inode(inode)) - goto clear_out; + f2fs_pre_evict_inode(inode); if (!inode->i_nlink && !is_bad_inode(inode)) f2fs_delete_inode(inode); f2fs_post_evict_inode(inode); -clear_out: fscrypt_put_encryption_info(inode); clear_inode(inode); } @@ -1124,7 +1070,7 @@ void f2fs_handle_failed_inode(struct inode *inode, * we must call this to avoid inode being remained as dirty, resulting * in a panic when flushing dirty inodes in gdirty_list. */ - f2fs_update_inode_page(inode); + f2fs_update_inode_cache(inode); f2fs_inode_synced(inode); /* don't make bad inode, since it becomes a regular file. */ diff --git a/fs/f2fs/iostat.h b/fs/f2fs/iostat.h index 2025225b5bed..61c6bc8e3119 100644 --- a/fs/f2fs/iostat.h +++ b/fs/f2fs/iostat.h @@ -60,6 +60,13 @@ static inline struct bio_post_read_ctx *get_post_read_ctx(struct bio *bio) return iostat_ctx->post_read_ctx; } +static inline void iostat_set_post_read_ctx(struct bio *bio, void *ctx) +{ + struct bio_iostat_ctx *iostat_ctx = bio->bi_private; + + iostat_ctx->post_read_ctx = ctx; +} + extern void iostat_update_and_unbind_ctx(struct bio *bio); extern void iostat_alloc_and_bind_ctx(struct f2fs_sb_info *sbi, struct bio *bio, struct bio_post_read_ctx *ctx); @@ -81,6 +88,10 @@ static inline struct bio_post_read_ctx *get_post_read_ctx(struct bio *bio) { return bio->bi_private; } +static inline void iostat_set_post_read_ctx(struct bio *bio, void *ctx) +{ + bio->bi_private = ctx; +} static inline int f2fs_init_iostat_processing(void) { return 0; } static inline void f2fs_destroy_iostat_processing(void) {} static inline int f2fs_init_iostat(struct f2fs_sb_info *sbi) { return 0; } diff --git a/fs/f2fs/namei.c b/fs/f2fs/namei.c index ff86ee07290d..ce5d536d892b 100644 --- a/fs/f2fs/namei.c +++ b/fs/f2fs/namei.c @@ -231,7 +231,7 @@ static void set_file_temperature(struct f2fs_sb_info *sbi, struct inode *inode, file_set_hot(inode); } -static struct inode *f2fs_new_inode(struct mnt_idmap *idmap, +static struct inode *f2fs_new_inode(const struct mnt_idmap *idmap, struct inode *dir, umode_t mode, const char *name) { @@ -365,7 +365,7 @@ fail_drop: return ERR_PTR(err); } -static int f2fs_create(struct mnt_idmap *idmap, struct inode *dir, +static int f2fs_create(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *dentry, umode_t mode) { struct f2fs_sb_info *sbi = F2FS_I_SB(dir); @@ -473,12 +473,13 @@ out: struct dentry *f2fs_get_parent(struct dentry *child) { - struct folio *folio; - unsigned long ino = f2fs_inode_by_name(d_inode(child), &dotdot_name, &folio); + void *dentry_block = NULL; + unsigned long ino = f2fs_inode_by_name(d_inode(child), + &dotdot_name, &dentry_block); if (!ino) { - if (IS_ERR(folio)) - return ERR_CAST(folio); + if (IS_ERR(dentry_block)) + return ERR_CAST(dentry_block); return ERR_PTR(-ENOENT); } return d_obtain_alias(f2fs_iget(child->d_sb, ino)); @@ -489,7 +490,7 @@ static struct dentry *f2fs_lookup(struct inode *dir, struct dentry *dentry, { struct inode *inode = NULL; struct f2fs_dir_entry *de; - struct folio *folio; + void *dentry_block = NULL; struct dentry *new; nid_t ino = -1; int err = 0; @@ -507,12 +508,12 @@ static struct dentry *f2fs_lookup(struct inode *dir, struct dentry *dentry, goto out_splice; if (err) goto out; - de = __f2fs_find_entry(dir, &fname, &folio); + de = __f2fs_find_entry(dir, &fname, &dentry_block); f2fs_free_filename(&fname); if (!de) { - if (IS_ERR(folio)) { - err = PTR_ERR(folio); + if (IS_ERR(dentry_block)) { + err = PTR_ERR(dentry_block); goto out; } err = -ENOENT; @@ -520,7 +521,7 @@ static struct dentry *f2fs_lookup(struct inode *dir, struct dentry *dentry, } ino = le32_to_cpu(de->ino); - f2fs_folio_put(folio, false); + f2fs_put_dentry_block(dentry_block, false); inode = f2fs_iget(dir->i_sb, ino); if (IS_ERR(inode)) { @@ -572,7 +573,7 @@ static int __do_unlink(struct inode *dir, struct inode *inode, struct f2fs_sb_info *sbi = F2FS_I_SB(dir); struct f2fs_dir_entry *de; struct f2fs_lock_context lc; - struct folio *folio; + void *dentry_block = NULL; int err; if (IS_DEVICE_ALIASING(inode)) @@ -588,9 +589,9 @@ static int __do_unlink(struct inode *dir, struct inode *inode, if (err) return err; - de = f2fs_find_entry(dir, name, &folio); + de = f2fs_find_entry(dir, name, &dentry_block); if (!de) - return IS_ERR(folio) ? PTR_ERR(folio) : 0; + return IS_ERR(dentry_block) ? PTR_ERR(dentry_block) : 0; if (unlikely(inode->i_nlink == 0)) { f2fs_warn(sbi, "%s: inode (ino=%llx) has zero i_nlink", @@ -610,7 +611,7 @@ static int __do_unlink(struct inode *dir, struct inode *inode, f2fs_unlock_op(sbi, &lc); goto err_out; } - f2fs_delete_entry(de, folio, dir, inode); + f2fs_delete_entry(de, dentry_block, dir, inode); f2fs_unlock_op(sbi, &lc); return 0; @@ -618,7 +619,7 @@ corrupted: err = -EFSCORRUPTED; set_sbi_flag(sbi, SBI_NEED_FSCK); err_out: - f2fs_folio_put(folio, false); + f2fs_put_dentry_block(dentry_block, false); return err; } @@ -662,7 +663,7 @@ static const char *f2fs_get_link(struct dentry *dentry, return link; } -static int f2fs_symlink(struct mnt_idmap *idmap, struct inode *dir, +static int f2fs_symlink(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *dentry, const char *symname) { struct f2fs_sb_info *sbi = F2FS_I_SB(dir); @@ -751,7 +752,7 @@ free_inode: goto out; } -static struct dentry *f2fs_mkdir(struct mnt_idmap *idmap, struct inode *dir, +static struct dentry *f2fs_mkdir(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *dentry, umode_t mode) { struct f2fs_sb_info *sbi = F2FS_I_SB(dir); @@ -810,7 +811,7 @@ static int f2fs_rmdir(struct inode *dir, struct dentry *dentry) return -ENOTEMPTY; } -static int f2fs_mknod(struct mnt_idmap *idmap, struct inode *dir, +static int f2fs_mknod(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *dentry, umode_t mode, dev_t rdev) { struct f2fs_sb_info *sbi = F2FS_I_SB(dir); @@ -857,7 +858,7 @@ out: return err; } -static int __f2fs_tmpfile(struct mnt_idmap *idmap, struct inode *dir, +static int __f2fs_tmpfile(const struct mnt_idmap *idmap, struct inode *dir, struct file *file, umode_t mode, bool is_whiteout, struct inode **new_inode, struct f2fs_filename *fname) { @@ -928,7 +929,7 @@ out: return err; } -static int f2fs_tmpfile(struct mnt_idmap *idmap, struct inode *dir, +static int f2fs_tmpfile(const struct mnt_idmap *idmap, struct inode *dir, struct file *file, umode_t mode) { struct f2fs_sb_info *sbi = F2FS_I_SB(dir); @@ -944,7 +945,7 @@ static int f2fs_tmpfile(struct mnt_idmap *idmap, struct inode *dir, return finish_open_simple(file, err); } -static int f2fs_create_whiteout(struct mnt_idmap *idmap, +static int f2fs_create_whiteout(const struct mnt_idmap *idmap, struct inode *dir, struct inode **whiteout, struct f2fs_filename *fname) { @@ -952,14 +953,14 @@ static int f2fs_create_whiteout(struct mnt_idmap *idmap, true, whiteout, fname); } -int f2fs_get_tmpfile(struct mnt_idmap *idmap, struct inode *dir, +int f2fs_get_tmpfile(const struct mnt_idmap *idmap, struct inode *dir, struct inode **new_inode) { return __f2fs_tmpfile(idmap, dir, NULL, S_IFREG, false, new_inode, NULL); } -static int f2fs_rename(struct mnt_idmap *idmap, struct inode *old_dir, +static int f2fs_rename(const struct mnt_idmap *idmap, struct inode *old_dir, struct dentry *old_dentry, struct inode *new_dir, struct dentry *new_dentry, unsigned int flags) { @@ -967,8 +968,9 @@ static int f2fs_rename(struct mnt_idmap *idmap, struct inode *old_dir, struct inode *old_inode = d_inode(old_dentry); struct inode *new_inode = d_inode(new_dentry); struct inode *whiteout = NULL; - struct folio *old_dir_folio = NULL; - struct folio *old_folio, *new_folio = NULL; + void *old_dir_block = NULL; + void *old_block = NULL; + void *new_block = NULL; struct f2fs_dir_entry *old_dir_entry = NULL; struct f2fs_dir_entry *old_entry; struct f2fs_dir_entry *new_entry; @@ -1032,18 +1034,18 @@ static int f2fs_rename(struct mnt_idmap *idmap, struct inode *old_dir, } err = -ENOENT; - old_entry = f2fs_find_entry(old_dir, &old_dentry->d_name, &old_folio); + old_entry = f2fs_find_entry(old_dir, &old_dentry->d_name, &old_block); if (!old_entry) { - if (IS_ERR(old_folio)) - err = PTR_ERR(old_folio); + if (IS_ERR(old_block)) + err = PTR_ERR(old_block); goto out; } if (old_is_dir && old_dir != new_dir) { - old_dir_entry = f2fs_parent_dir(old_inode, &old_dir_folio); + old_dir_entry = f2fs_parent_dir(old_inode, &old_dir_block); if (!old_dir_entry) { - if (IS_ERR(old_dir_folio)) - err = PTR_ERR(old_dir_folio); + if (IS_ERR(old_dir_block)) + err = PTR_ERR(old_dir_block); goto out_old; } } @@ -1060,10 +1062,10 @@ static int f2fs_rename(struct mnt_idmap *idmap, struct inode *old_dir, err = -ENOENT; new_entry = f2fs_find_entry(new_dir, &new_dentry->d_name, - &new_folio); + &new_block); if (!new_entry) { - if (IS_ERR(new_folio)) - err = PTR_ERR(new_folio); + if (IS_ERR(new_block)) + err = PTR_ERR(new_block); goto out_dir; } @@ -1075,8 +1077,8 @@ static int f2fs_rename(struct mnt_idmap *idmap, struct inode *old_dir, if (err) goto put_out_dir; - f2fs_set_link(new_dir, new_entry, new_folio, old_inode); - new_folio = NULL; + f2fs_set_link(new_dir, new_entry, new_block, old_inode); + new_block = NULL; inode_set_ctime_current(new_inode); f2fs_down_write(&F2FS_I(new_inode)->i_sem); @@ -1115,8 +1117,8 @@ static int f2fs_rename(struct mnt_idmap *idmap, struct inode *old_dir, inode_set_ctime_current(old_inode); f2fs_mark_inode_dirty_sync(old_inode, true); - f2fs_delete_entry(old_entry, old_folio, old_dir, NULL); - old_folio = NULL; + f2fs_delete_entry(old_entry, old_block, old_dir, NULL); + old_block = NULL; if (whiteout) { set_inode_flag(whiteout, FI_INC_LINK); @@ -1134,7 +1136,7 @@ static int f2fs_rename(struct mnt_idmap *idmap, struct inode *old_dir, } if (old_dir_entry) - f2fs_set_link(old_inode, old_dir_entry, old_dir_folio, new_dir); + f2fs_set_link(old_inode, old_dir_entry, old_dir_block, new_dir); if (old_is_dir) f2fs_i_links_write(old_dir, false); @@ -1158,12 +1160,12 @@ static int f2fs_rename(struct mnt_idmap *idmap, struct inode *old_dir, put_out_dir: f2fs_unlock_op(sbi, &lc); - f2fs_folio_put(new_folio, false); + f2fs_put_dentry_block(new_block, false); out_dir: if (old_dir_entry) - f2fs_folio_put(old_dir_folio, false); + f2fs_put_dentry_block(old_dir_block, false); out_old: - f2fs_folio_put(old_folio, false); + f2fs_put_dentry_block(old_block, false); out: iput(whiteout); return err; @@ -1175,8 +1177,10 @@ static int f2fs_cross_rename(struct inode *old_dir, struct dentry *old_dentry, struct f2fs_sb_info *sbi = F2FS_I_SB(old_dir); struct inode *old_inode = d_inode(old_dentry); struct inode *new_inode = d_inode(new_dentry); - struct folio *old_dir_folio, *new_dir_folio; - struct folio *old_folio, *new_folio; + void *old_dir_block = NULL; + void *new_dir_block = NULL; + void *old_block = NULL; + void *new_block = NULL; struct f2fs_dir_entry *old_dir_entry = NULL, *new_dir_entry = NULL; struct f2fs_dir_entry *old_entry, *new_entry; struct f2fs_lock_context lc; @@ -1208,17 +1212,17 @@ static int f2fs_cross_rename(struct inode *old_dir, struct dentry *old_dentry, goto out; err = -ENOENT; - old_entry = f2fs_find_entry(old_dir, &old_dentry->d_name, &old_folio); + old_entry = f2fs_find_entry(old_dir, &old_dentry->d_name, &old_block); if (!old_entry) { - if (IS_ERR(old_folio)) - err = PTR_ERR(old_folio); + if (IS_ERR(old_block)) + err = PTR_ERR(old_block); goto out; } - new_entry = f2fs_find_entry(new_dir, &new_dentry->d_name, &new_folio); + new_entry = f2fs_find_entry(new_dir, &new_dentry->d_name, &new_block); if (!new_entry) { - if (IS_ERR(new_folio)) - err = PTR_ERR(new_folio); + if (IS_ERR(new_block)) + err = PTR_ERR(new_block); goto out_old; } @@ -1226,20 +1230,20 @@ static int f2fs_cross_rename(struct inode *old_dir, struct dentry *old_dentry, if (old_dir != new_dir) { if (S_ISDIR(old_inode->i_mode)) { old_dir_entry = f2fs_parent_dir(old_inode, - &old_dir_folio); + &old_dir_block); if (!old_dir_entry) { - if (IS_ERR(old_dir_folio)) - err = PTR_ERR(old_dir_folio); + if (IS_ERR(old_dir_block)) + err = PTR_ERR(old_dir_block); goto out_new; } } if (S_ISDIR(new_inode->i_mode)) { new_dir_entry = f2fs_parent_dir(new_inode, - &new_dir_folio); + &new_dir_block); if (!new_dir_entry) { - if (IS_ERR(new_dir_folio)) - err = PTR_ERR(new_dir_folio); + if (IS_ERR(new_dir_block)) + err = PTR_ERR(new_dir_block); goto out_old_dir; } } @@ -1266,14 +1270,14 @@ static int f2fs_cross_rename(struct inode *old_dir, struct dentry *old_dentry, /* update ".." directory entry info of old dentry */ if (old_dir_entry) - f2fs_set_link(old_inode, old_dir_entry, old_dir_folio, new_dir); + f2fs_set_link(old_inode, old_dir_entry, old_dir_block, new_dir); /* update ".." directory entry info of new dentry */ if (new_dir_entry) - f2fs_set_link(new_inode, new_dir_entry, new_dir_folio, old_dir); + f2fs_set_link(new_inode, new_dir_entry, new_dir_block, old_dir); /* update directory entry info of old dir inode */ - f2fs_set_link(old_dir, old_entry, old_folio, new_inode); + f2fs_set_link(old_dir, old_entry, old_block, new_inode); f2fs_down_write(&F2FS_I(old_inode)->i_sem); if (!old_dir_entry) @@ -1292,7 +1296,7 @@ static int f2fs_cross_rename(struct inode *old_dir, struct dentry *old_dentry, f2fs_mark_inode_dirty_sync(old_dir, true); /* update directory entry info of new dir inode */ - f2fs_set_link(new_dir, new_entry, new_folio, old_inode); + f2fs_set_link(new_dir, new_entry, new_block, old_inode); f2fs_down_write(&F2FS_I(new_inode)->i_sem); if (!new_dir_entry) @@ -1327,21 +1331,21 @@ static int f2fs_cross_rename(struct inode *old_dir, struct dentry *old_dentry, return 0; out_new_dir: if (new_dir_entry) { - f2fs_folio_put(new_dir_folio, false); + f2fs_put_dentry_block(new_dir_block, false); } out_old_dir: if (old_dir_entry) { - f2fs_folio_put(old_dir_folio, false); + f2fs_put_dentry_block(old_dir_block, false); } out_new: - f2fs_folio_put(new_folio, false); + f2fs_put_dentry_block(new_block, false); out_old: - f2fs_folio_put(old_folio, false); + f2fs_put_dentry_block(old_block, false); out: return err; } -static int f2fs_rename2(struct mnt_idmap *idmap, +static int f2fs_rename2(const struct mnt_idmap *idmap, struct inode *old_dir, struct dentry *old_dentry, struct inode *new_dir, struct dentry *new_dentry, unsigned int flags) @@ -1394,7 +1398,7 @@ static const char *f2fs_encrypted_get_link(struct dentry *dentry, return target; } -static int f2fs_encrypted_symlink_getattr(struct mnt_idmap *idmap, +static int f2fs_encrypted_symlink_getattr(const struct mnt_idmap *idmap, const struct path *path, struct kstat *stat, u32 request_mask, unsigned int query_flags) diff --git a/fs/f2fs/node.c b/fs/f2fs/node.c index 86c2e67e43b6..be14dfa31e53 100644 --- a/fs/f2fs/node.c +++ b/fs/f2fs/node.c @@ -7,12 +7,12 @@ */ #include <linux/fs.h> #include <linux/f2fs_fs.h> -#include <linux/mpage.h> #include <linux/sched/mm.h> #include <linux/blkdev.h> #include <linux/folio_batch.h> #include <linux/swap.h> #include <linux/fserror.h> +#include <linux/freezer.h> #include "f2fs.h" #include "node.h" @@ -85,7 +85,7 @@ bool f2fs_available_free_memory(struct f2fs_sb_info *sbi, int type) } else if (type == DIRTY_DENTS) { if (bdi_wb_dirty_exceeded(sbi->sb->s_bdi)) return false; - mem_size = get_pages(sbi, F2FS_DIRTY_DENTS); + mem_size = get_nr_caches(sbi, F2FS_DIRTY_DENTS); res = mem_size < ((avail_ram * nm_i->ram_thresh / 100) >> 1); } else if (type == INO_ENTRIES) { int i; @@ -109,7 +109,7 @@ bool f2fs_available_free_memory(struct f2fs_sb_info *sbi, int type) mem_size = (atomic_read(&dcc->discard_cmd_cnt) * sizeof(struct discard_cmd)) >> PAGE_SHIFT; res = mem_size < (avail_ram * nm_i->ram_thresh / 100); - } else if (type == COMPRESS_PAGE) { + } else if (type == COMPRESS_BLOCK) { #ifdef CONFIG_F2FS_FS_COMPRESSION unsigned long free_ram = val.freeram; @@ -118,7 +118,7 @@ bool f2fs_available_free_memory(struct f2fs_sb_info *sbi, int type) * exceed threshold, deny caching compress page. */ res = (free_ram > avail_ram * sbi->compress_watermark / 100) && - (COMPRESS_MAPPING(sbi)->nrpages < + (COMPRESS_CACHE(sbi)->num_entries < free_ram * sbi->compress_percent / 100); #else res = false; @@ -130,48 +130,36 @@ bool f2fs_available_free_memory(struct f2fs_sb_info *sbi, int type) return res; } -static void clear_node_folio_dirty(struct folio *folio) +static struct f2fs_cached_block *get_current_nat_cache(struct f2fs_sb_info *sbi, + nid_t nid) { - if (folio_test_dirty(folio)) { - f2fs_clear_page_cache_dirty_tag(folio); - folio_clear_dirty_for_io(folio); - dec_page_count(F2FS_F_SB(folio), F2FS_DIRTY_NODES); - } - folio_clear_uptodate(folio); -} - -static struct folio *get_current_nat_folio(struct f2fs_sb_info *sbi, nid_t nid) -{ - return f2fs_get_meta_folio_retry(sbi, current_nat_addr(sbi, nid)); + return f2fs_get_meta_cache_retry(sbi, current_nat_addr(sbi, nid)); } -static struct folio *get_next_nat_folio(struct f2fs_sb_info *sbi, nid_t nid) +static struct f2fs_cached_block *get_next_nat_cache(struct f2fs_sb_info *sbi, + nid_t nid) { - struct folio *src_folio; - struct folio *dst_folio; + struct f2fs_cached_block *src_entry; + struct f2fs_cached_block *dst_entry; pgoff_t dst_off; - void *src_addr; - void *dst_addr; struct f2fs_nm_info *nm_i = NM_I(sbi); dst_off = next_nat_addr(sbi, current_nat_addr(sbi, nid)); - /* get current nat block page with lock */ - src_folio = get_current_nat_folio(sbi, nid); - if (IS_ERR(src_folio)) - return src_folio; - dst_folio = f2fs_grab_meta_folio(sbi, dst_off); - f2fs_bug_on(sbi, folio_test_dirty(src_folio)); + /* get current nat cached block with lock */ + src_entry = get_current_nat_cache(sbi, nid); + if (IS_ERR(src_entry)) + return src_entry; + dst_entry = f2fs_grab_meta_cache(sbi, dst_off); + f2fs_bug_on(sbi, f2fs_cache_test_dirty(src_entry)); - src_addr = folio_address(src_folio); - dst_addr = folio_address(dst_folio); - memcpy(dst_addr, src_addr, PAGE_SIZE); - folio_mark_dirty(dst_folio); - f2fs_folio_put(src_folio, true); + memcpy(cache_address(dst_entry), cache_address(src_entry), F2FS_BLKSIZE(sbi)); + f2fs_mark_cache_dirty(dst_entry); + f2fs_put_cache(src_entry, true); - set_to_next_nat(nm_i, nid); + set_to_next_nat(sbi, nm_i, nid); - return dst_folio; + return dst_entry; } static struct nat_entry *__alloc_nat_entry(struct f2fs_sb_info *sbi, @@ -254,10 +242,11 @@ static void __del_from_nat_cache(struct f2fs_nm_info *nm_i, struct nat_entry *e) __free_nat_entry(e); } -static struct nat_entry_set *__grab_nat_entry_set(struct f2fs_nm_info *nm_i, +static struct nat_entry_set *__grab_nat_entry_set(struct f2fs_sb_info *sbi, + struct f2fs_nm_info *nm_i, struct nat_entry *ne) { - nid_t set = NAT_BLOCK_OFFSET(ne->ni.nid); + nid_t set = NAT_BLOCK_OFFSET(sbi, ne->ni.nid); struct nat_entry_set *head; head = radix_tree_lookup(&nm_i->nat_set_root, set); @@ -274,14 +263,15 @@ static struct nat_entry_set *__grab_nat_entry_set(struct f2fs_nm_info *nm_i, return head; } -static void __set_nat_cache_dirty(struct f2fs_nm_info *nm_i, +static void __set_nat_cache_dirty(struct f2fs_sb_info *sbi, + struct f2fs_nm_info *nm_i, struct nat_entry *ne, bool init_dirty) { struct nat_entry_set *head; bool new_ne = nat_get_blkaddr(ne) == NEW_ADDR; if (!new_ne) - head = __grab_nat_entry_set(nm_i, ne); + head = __grab_nat_entry_set(sbi, nm_i, ne); /* * update entry_cnt in below condition: @@ -330,9 +320,11 @@ static unsigned int __gang_lookup_nat_set(struct f2fs_nm_info *nm_i, start, nr); } -bool f2fs_in_warm_node_list(struct folio *folio) +bool f2fs_in_warm_node_list(struct f2fs_sb_info *sbi, + struct f2fs_cached_block *entry) { - return is_node_folio(folio) && IS_DNODE(folio) && is_cold_node(folio); + return f2fs_is_node_cache(entry) && + IS_DNODE(sbi, entry) && is_cold_node(sbi, entry); } void f2fs_init_fsync_node_info(struct f2fs_sb_info *sbi) @@ -344,7 +336,7 @@ void f2fs_init_fsync_node_info(struct f2fs_sb_info *sbi) } static unsigned int f2fs_add_fsync_node_entry(struct f2fs_sb_info *sbi, - struct folio *folio) + struct f2fs_cached_block *entry) { struct fsync_node_entry *fn; unsigned long flags; @@ -353,8 +345,8 @@ static unsigned int f2fs_add_fsync_node_entry(struct f2fs_sb_info *sbi, fn = f2fs_kmem_cache_alloc(fsync_node_entry_slab, GFP_NOFS, true, NULL); - folio_get(folio); - fn->folio = folio; + f2fs_cache_get(entry); + fn->entry = entry; INIT_LIST_HEAD(&fn->list); spin_lock_irqsave(&sbi->fsync_node_lock, flags); @@ -367,19 +359,19 @@ static unsigned int f2fs_add_fsync_node_entry(struct f2fs_sb_info *sbi, return seq_id; } -void f2fs_del_fsync_node_entry(struct f2fs_sb_info *sbi, struct folio *folio) +void f2fs_del_fsync_node_entry(struct f2fs_sb_info *sbi, struct f2fs_cached_block *entry) { struct fsync_node_entry *fn; unsigned long flags; spin_lock_irqsave(&sbi->fsync_node_lock, flags); list_for_each_entry(fn, &sbi->fsync_node_list, list) { - if (fn->folio == folio) { + if (fn->entry == entry) { list_del(&fn->list); sbi->fsync_node_num--; spin_unlock_irqrestore(&sbi->fsync_node_lock, flags); + f2fs_put_cache(fn->entry, false); kmem_cache_free(fsync_node_entry_slab, fn); - folio_put(folio); return; } } @@ -527,7 +519,7 @@ static void set_node_addr(struct f2fs_sb_info *sbi, struct node_info *ni, nat_set_blkaddr(e, new_blkaddr); if (!__is_valid_data_blkaddr(new_blkaddr)) set_nat_flag(e, IS_CHECKPOINTED, false); - __set_nat_cache_dirty(nm_i, e, init_dirty); + __set_nat_cache_dirty(sbi, nm_i, e, init_dirty); /* update fsync_mark if its inode nat entry is still alive */ if (ni->nid != ni->ino) @@ -578,9 +570,9 @@ int f2fs_get_node_info(struct f2fs_sb_info *sbi, nid_t nid, struct f2fs_nm_info *nm_i = NM_I(sbi); struct curseg_info *curseg = CURSEG_I(sbi, CURSEG_HOT_DATA); struct f2fs_journal *journal = curseg->journal; - nid_t start_nid = START_NID(nid); + nid_t start_nid = f2fs_start_nid(sbi, nid); struct f2fs_nat_block *nat_blk; - struct folio *folio = NULL; + struct f2fs_cached_block *entry = NULL; struct f2fs_nat_entry ne; struct nat_entry *e; pgoff_t index; @@ -631,18 +623,18 @@ retry: goto sanity_check; } - /* Fill node_info from nat page */ + /* Fill node_info from nat block */ index = current_nat_addr(sbi, nid); f2fs_up_read_trace(&nm_i->nat_tree_lock, &lc); - folio = f2fs_get_meta_folio(sbi, index); - if (IS_ERR(folio)) - return PTR_ERR(folio); + entry = f2fs_get_meta_cache(sbi, index); + if (IS_ERR(entry)) + return PTR_ERR(entry); - nat_blk = folio_address(folio); + nat_blk = cache_address(entry); ne = nat_blk->entries[nid - start_nid]; node_info_from_raw_nat(ni, &ne); - f2fs_folio_put(folio, true); + f2fs_put_cache(entry, true); sanity_check: if (__is_valid_data_blkaddr(ni->blk_addr) && !f2fs_is_valid_blkaddr(sbi, ni->blk_addr, @@ -676,11 +668,11 @@ sanity_check: } /* - * readahead MAX_RA_NODE number of node pages. + * readahead MAX_RA_NODE number of node caches. */ -static void f2fs_ra_node_pages(struct folio *parent, int start, int n) +static void f2fs_ra_node_caches(struct f2fs_cached_block *parent, int start, int n) { - struct f2fs_sb_info *sbi = F2FS_F_SB(parent); + struct f2fs_sb_info *sbi = parent->cache->sbi; struct blk_plug plug; int i, end; nid_t nid; @@ -689,10 +681,10 @@ static void f2fs_ra_node_pages(struct folio *parent, int start, int n) /* Then, try readahead for siblings of the desired node */ end = start + n; - end = min(end, (int)NIDS_PER_BLOCK); + end = min_t(int, end, NIDS_PER_BLOCK(sbi)); for (i = start; i < end; i++) { - nid = get_nid(parent, i, false); - f2fs_ra_node_page(sbi, nid); + nid = get_nid(sbi, parent, i, false); + f2fs_ra_node_cache(sbi, nid); } blk_finish_plug(&plug); @@ -700,9 +692,11 @@ static void f2fs_ra_node_pages(struct folio *parent, int start, int n) pgoff_t f2fs_get_next_page_offset(struct dnode_of_data *dn, pgoff_t pgofs) { + struct f2fs_sb_info *sbi = F2FS_I_SB(dn->inode); const long direct_index = ADDRS_PER_INODE(dn->inode); const long direct_blks = ADDRS_PER_BLOCK(dn->inode); - const long indirect_blks = ADDRS_PER_BLOCK(dn->inode) * NIDS_PER_BLOCK; + const long indirect_blks = ADDRS_PER_BLOCK(dn->inode) * + NIDS_PER_BLOCK(sbi); unsigned int skipped_unit = ADDRS_PER_BLOCK(dn->inode); int cur_level = dn->cur_level; int max_level = dn->max_level; @@ -712,7 +706,7 @@ pgoff_t f2fs_get_next_page_offset(struct dnode_of_data *dn, pgoff_t pgofs) return pgofs + 1; while (max_level-- > cur_level) - skipped_unit *= NIDS_PER_BLOCK; + skipped_unit *= NIDS_PER_BLOCK(sbi); switch (dn->max_level) { case 3: @@ -738,11 +732,13 @@ pgoff_t f2fs_get_next_page_offset(struct dnode_of_data *dn, pgoff_t pgofs) static int get_node_path(struct inode *inode, long block, int offset[4], unsigned int noffset[4]) { + struct f2fs_sb_info *sbi = F2FS_I_SB(inode); const long direct_index = ADDRS_PER_INODE(inode); const long direct_blks = ADDRS_PER_BLOCK(inode); - const long dptrs_per_blk = NIDS_PER_BLOCK; - const long indirect_blks = ADDRS_PER_BLOCK(inode) * NIDS_PER_BLOCK; - const long dindirect_blks = indirect_blks * NIDS_PER_BLOCK; + const long dptrs_per_blk = NIDS_PER_BLOCK(sbi); + const long indirect_blks = ADDRS_PER_BLOCK(inode) * + NIDS_PER_BLOCK(sbi); + const long dindirect_blks = indirect_blks * NIDS_PER_BLOCK(sbi); int n = 0; int level = 0; @@ -754,7 +750,7 @@ static int get_node_path(struct inode *inode, long block, } block -= direct_index; if (block < direct_blks) { - offset[n++] = NODE_DIR1_BLOCK; + offset[n++] = NODE_DIR1_BLOCK(sbi); noffset[n] = 1; offset[n] = block; level = 1; @@ -762,7 +758,7 @@ static int get_node_path(struct inode *inode, long block, } block -= direct_blks; if (block < direct_blks) { - offset[n++] = NODE_DIR2_BLOCK; + offset[n++] = NODE_DIR2_BLOCK(sbi); noffset[n] = 2; offset[n] = block; level = 1; @@ -770,7 +766,7 @@ static int get_node_path(struct inode *inode, long block, } block -= direct_blks; if (block < indirect_blks) { - offset[n++] = NODE_IND1_BLOCK; + offset[n++] = NODE_IND1_BLOCK(sbi); noffset[n] = 3; offset[n++] = block / direct_blks; noffset[n] = 4 + offset[n - 1]; @@ -780,7 +776,7 @@ static int get_node_path(struct inode *inode, long block, } block -= indirect_blks; if (block < indirect_blks) { - offset[n++] = NODE_IND2_BLOCK; + offset[n++] = NODE_IND2_BLOCK(sbi); noffset[n] = 4 + dptrs_per_blk; offset[n++] = block / direct_blks; noffset[n] = 5 + dptrs_per_blk + offset[n - 1]; @@ -790,7 +786,7 @@ static int get_node_path(struct inode *inode, long block, } block -= indirect_blks; if (block < dindirect_blks) { - offset[n++] = NODE_DIND_BLOCK; + offset[n++] = NODE_DIND_BLOCK(sbi); noffset[n] = 5 + (dptrs_per_blk * 2); offset[n++] = block / indirect_blks; noffset[n] = 6 + (dptrs_per_blk * 2) + @@ -809,7 +805,8 @@ got: return level; } -static struct folio *f2fs_get_node_folio_ra(struct folio *parent, int start); +static struct f2fs_cached_block *f2fs_get_node_cache_ra( + struct f2fs_cached_block *parent, int start); /* * Caller should call f2fs_put_dnode(dn). @@ -819,8 +816,8 @@ static struct folio *f2fs_get_node_folio_ra(struct folio *parent, int start); int f2fs_get_dnode_of_data(struct dnode_of_data *dn, pgoff_t index, int mode) { struct f2fs_sb_info *sbi = F2FS_I_SB(dn->inode); - struct folio *nfolio[4]; - struct folio *parent = NULL; + struct f2fs_cached_block *nentry[4]; + struct f2fs_cached_block *parent = NULL; int offset[4]; unsigned int noffset[4]; nid_t nids[4]; @@ -833,26 +830,26 @@ int f2fs_get_dnode_of_data(struct dnode_of_data *dn, pgoff_t index, int mode) nids[0] = dn->inode->i_ino; - if (!dn->inode_folio) { - nfolio[0] = f2fs_get_inode_folio(sbi, nids[0]); - if (IS_ERR(nfolio[0])) - return PTR_ERR(nfolio[0]); + if (!dn->inode_entry) { + nentry[0] = f2fs_get_inode_cache(sbi, nids[0]); + if (IS_ERR(nentry[0])) + return PTR_ERR(nentry[0]); } else { - nfolio[0] = dn->inode_folio; + nentry[0] = dn->inode_entry; } /* if inline_data is set, should not report any block indices */ if (f2fs_has_inline_data(dn->inode) && index) { err = -ENOENT; - f2fs_folio_put(nfolio[0], true); + f2fs_put_cache(nentry[0], true); goto release_out; } - parent = nfolio[0]; + parent = nentry[0]; if (level != 0) - nids[1] = get_nid(parent, offset[0], true); - dn->inode_folio = nfolio[0]; - dn->inode_folio_locked = true; + nids[1] = get_nid(sbi, parent, offset[0], true); + dn->inode_entry = nentry[0]; + dn->inode_entry_locked = true; /* get indirect or direct nodes */ for (i = 1; i <= level; i++) { @@ -865,59 +862,59 @@ int f2fs_get_dnode_of_data(struct dnode_of_data *dn, pgoff_t index, int mode) "ino:%llu, nid:%u, level:%d, offset:%d", dn->inode->i_ino, nids[i], level, offset[level]); set_sbi_flag(sbi, SBI_NEED_FSCK); - goto release_pages; + goto release_caches; } if (!nids[i] && mode == ALLOC_NODE) { /* alloc new node */ if (!f2fs_alloc_nid(sbi, &(nids[i]))) { err = -ENOSPC; - goto release_pages; + goto release_caches; } dn->nid = nids[i]; - nfolio[i] = f2fs_new_node_folio(dn, noffset[i]); - if (IS_ERR(nfolio[i])) { + nentry[i] = f2fs_new_node_cache(dn, noffset[i]); + if (IS_ERR(nentry[i])) { f2fs_alloc_nid_failed(sbi, nids[i]); - err = PTR_ERR(nfolio[i]); - goto release_pages; + err = PTR_ERR(nentry[i]); + goto release_caches; } - set_nid(parent, offset[i - 1], nids[i], i == 1); + set_nid(sbi, parent, offset[i - 1], nids[i], i == 1); f2fs_alloc_nid_done(sbi, nids[i]); done = true; } else if (mode == LOOKUP_NODE_RA && i == level && level > 1) { - nfolio[i] = f2fs_get_node_folio_ra(parent, offset[i - 1]); - if (IS_ERR(nfolio[i])) { - err = PTR_ERR(nfolio[i]); - goto release_pages; + nentry[i] = f2fs_get_node_cache_ra(parent, offset[i - 1]); + if (IS_ERR(nentry[i])) { + err = PTR_ERR(nentry[i]); + goto release_caches; } done = true; } if (i == 1) { - dn->inode_folio_locked = false; - folio_unlock(parent); + dn->inode_entry_locked = false; + f2fs_unlock_cache(parent); } else { - f2fs_folio_put(parent, true); + f2fs_put_cache(parent, true); } if (!done) { - nfolio[i] = f2fs_get_node_folio(sbi, nids[i], + nentry[i] = f2fs_get_node_cache(sbi, nids[i], NODE_TYPE_NON_INODE); - if (IS_ERR(nfolio[i])) { - err = PTR_ERR(nfolio[i]); - f2fs_folio_put(nfolio[0], false); + if (IS_ERR(nentry[i])) { + err = PTR_ERR(nentry[i]); + f2fs_put_cache(nentry[0], false); goto release_out; } } if (i < level) { - parent = nfolio[i]; - nids[i + 1] = get_nid(parent, offset[i], false); + parent = nentry[i]; + nids[i + 1] = get_nid(sbi, parent, offset[i], false); } } dn->nid = nids[level]; dn->ofs_in_node = offset[level]; - dn->node_folio = nfolio[level]; + dn->node_entry = nentry[level]; dn->data_blkaddr = f2fs_data_blkaddr(dn); if (is_inode_flag_set(dn->inode, FI_COMPRESSED_FILE) && @@ -938,9 +935,9 @@ int f2fs_get_dnode_of_data(struct dnode_of_data *dn, pgoff_t index, int mode) if (!c_len) goto out; - blkaddr = data_blkaddr(dn->inode, dn->node_folio, ofs_in_node); + blkaddr = data_blkaddr(dn->inode, dn->node_entry, ofs_in_node); if (blkaddr == COMPRESS_ADDR) - blkaddr = data_blkaddr(dn->inode, dn->node_folio, + blkaddr = data_blkaddr(dn->inode, dn->node_entry, ofs_in_node + 1); f2fs_update_read_extent_tree_range_compressed(dn->inode, @@ -949,13 +946,13 @@ int f2fs_get_dnode_of_data(struct dnode_of_data *dn, pgoff_t index, int mode) out: return 0; -release_pages: - f2fs_folio_put(parent, true); +release_caches: + f2fs_put_cache(parent, true); if (i > 1) - f2fs_folio_put(nfolio[0], false); + f2fs_put_cache(nentry[0], false); release_out: - dn->inode_folio = NULL; - dn->node_folio = NULL; + dn->inode_entry = NULL; + dn->node_entry = NULL; if (err == -ENOENT) { dn->cur_level = i; dn->max_level = level; @@ -996,16 +993,15 @@ static int truncate_node(struct dnode_of_data *dn) f2fs_inode_synced(dn->inode); } - clear_node_folio_dirty(dn->node_folio); + f2fs_drop_cache_dirty(dn->node_entry); set_sbi_flag(sbi, SBI_IS_DIRTY); - index = dn->node_folio->index; - f2fs_folio_put(dn->node_folio, true); + index = dn->node_entry->index; + f2fs_put_cache(dn->node_entry, true); - invalidate_mapping_pages(NODE_MAPPING(sbi), - index, index); + f2fs_invalidate_node_cache(sbi, index); - dn->node_folio = NULL; + dn->node_entry = NULL; trace_f2fs_truncate_node(dn->inode, dn->nid, ni.blk_addr); return 0; @@ -1014,35 +1010,35 @@ static int truncate_node(struct dnode_of_data *dn) static int truncate_dnode(struct dnode_of_data *dn) { struct f2fs_sb_info *sbi = F2FS_I_SB(dn->inode); - struct folio *folio; + struct f2fs_cached_block *entry; int err; if (dn->nid == 0) return 1; /* get direct node */ - folio = f2fs_get_node_folio(sbi, dn->nid, NODE_TYPE_NON_INODE); - if (PTR_ERR(folio) == -ENOENT) + entry = f2fs_get_node_cache(sbi, dn->nid, NODE_TYPE_NON_INODE); + if (PTR_ERR(entry) == -ENOENT) return 1; - else if (IS_ERR(folio)) - return PTR_ERR(folio); + else if (IS_ERR(entry)) + return PTR_ERR(entry); - if (IS_INODE(folio) || ino_of_node(folio) != dn->inode->i_ino) { + if (IS_INODE(sbi, entry) || ino_of_node(sbi, entry) != dn->inode->i_ino) { f2fs_err(sbi, "incorrect node reference, ino: %llu, nid: %u, ino_of_node: %u", - dn->inode->i_ino, dn->nid, ino_of_node(folio)); + dn->inode->i_ino, dn->nid, ino_of_node(sbi, entry)); set_sbi_flag(sbi, SBI_NEED_FSCK); f2fs_handle_error(sbi, ERROR_INVALID_NODE_REFERENCE); - f2fs_folio_put(folio, true); + f2fs_put_cache(entry, true); return -EFSCORRUPTED; } /* Make dnode_of_data for parameter */ - dn->node_folio = folio; + dn->node_entry = entry; dn->ofs_in_node = 0; f2fs_truncate_data_blocks_range(dn, ADDRS_PER_BLOCK(dn->inode)); err = truncate_node(dn); if (err) { - f2fs_folio_put(folio, true); + f2fs_put_cache(entry, true); return err; } @@ -1052,8 +1048,9 @@ static int truncate_dnode(struct dnode_of_data *dn) static int truncate_nodes(struct dnode_of_data *dn, unsigned int nofs, int ofs, int depth) { + struct f2fs_sb_info *sbi = F2FS_I_SB(dn->inode); struct dnode_of_data rdn = *dn; - struct folio *folio; + struct f2fs_cached_block *entry; struct f2fs_node *rn; nid_t child_nid; unsigned int child_nofs; @@ -1061,22 +1058,22 @@ static int truncate_nodes(struct dnode_of_data *dn, unsigned int nofs, int i, ret; if (dn->nid == 0) - return NIDS_PER_BLOCK + 1; + return NIDS_PER_BLOCK(sbi) + 1; trace_f2fs_truncate_nodes_enter(dn->inode, dn->nid, dn->data_blkaddr); - folio = f2fs_get_node_folio(F2FS_I_SB(dn->inode), dn->nid, + entry = f2fs_get_node_cache(F2FS_I_SB(dn->inode), dn->nid, NODE_TYPE_NON_INODE); - if (IS_ERR(folio)) { - trace_f2fs_truncate_nodes_exit(dn->inode, PTR_ERR(folio)); - return PTR_ERR(folio); + if (IS_ERR(entry)) { + trace_f2fs_truncate_nodes_exit(dn->inode, PTR_ERR(entry)); + return PTR_ERR(entry); } - f2fs_ra_node_pages(folio, ofs, NIDS_PER_BLOCK); + f2fs_ra_node_caches(entry, ofs, NIDS_PER_BLOCK(sbi)); - rn = F2FS_NODE(folio); + rn = CACHED_NODE(entry); if (depth < 3) { - for (i = ofs; i < NIDS_PER_BLOCK; i++, freed++) { + for (i = ofs; i < NIDS_PER_BLOCK(sbi); i++, freed++) { child_nid = le32_to_cpu(rn->in.nid[i]); if (child_nid == 0) continue; @@ -1084,21 +1081,21 @@ static int truncate_nodes(struct dnode_of_data *dn, unsigned int nofs, ret = truncate_dnode(&rdn); if (ret < 0) goto out_err; - if (set_nid(folio, i, 0, false)) + if (set_nid(sbi, entry, i, 0, false)) dn->node_changed = true; } } else { - child_nofs = nofs + ofs * (NIDS_PER_BLOCK + 1) + 1; - for (i = ofs; i < NIDS_PER_BLOCK; i++) { + child_nofs = nofs + ofs * (NIDS_PER_BLOCK(sbi) + 1) + 1; + for (i = ofs; i < NIDS_PER_BLOCK(sbi); i++) { child_nid = le32_to_cpu(rn->in.nid[i]); if (child_nid == 0) { - child_nofs += NIDS_PER_BLOCK + 1; + child_nofs += NIDS_PER_BLOCK(sbi) + 1; continue; } rdn.nid = child_nid; ret = truncate_nodes(&rdn, child_nofs, 0, depth - 1); - if (ret == (NIDS_PER_BLOCK + 1)) { - if (set_nid(folio, i, 0, false)) + if (ret == (NIDS_PER_BLOCK(sbi) + 1)) { + if (set_nid(sbi, entry, i, 0, false)) dn->node_changed = true; child_nofs += ret; } else if (ret < 0 && ret != -ENOENT) { @@ -1110,19 +1107,19 @@ static int truncate_nodes(struct dnode_of_data *dn, unsigned int nofs, if (!ofs) { /* remove current indirect node */ - dn->node_folio = folio; + dn->node_entry = entry; ret = truncate_node(dn); if (ret) goto out_err; freed++; } else { - f2fs_folio_put(folio, true); + f2fs_put_cache(entry, true); } trace_f2fs_truncate_nodes_exit(dn->inode, freed); return freed; out_err: - f2fs_folio_put(folio, true); + f2fs_put_cache(entry, true); trace_f2fs_truncate_nodes_exit(dn->inode, ret); return ret; } @@ -1130,60 +1127,62 @@ out_err: static int truncate_partial_nodes(struct dnode_of_data *dn, int *offset, int depth) { - struct folio *folios[2]; + struct f2fs_cached_block *entries[2]; nid_t nid[3]; nid_t child_nid; int err = 0; int i; int idx = depth - 2; - nid[0] = get_nid(dn->inode_folio, offset[0], true); + nid[0] = get_nid(F2FS_I_SB(dn->inode), dn->inode_entry, offset[0], true); if (!nid[0]) return 0; /* get indirect nodes in the path */ for (i = 0; i < idx + 1; i++) { /* reference count'll be increased */ - folios[i] = f2fs_get_node_folio(F2FS_I_SB(dn->inode), nid[i], + entries[i] = f2fs_get_node_cache(F2FS_I_SB(dn->inode), nid[i], NODE_TYPE_NON_INODE); - if (IS_ERR(folios[i])) { - err = PTR_ERR(folios[i]); + if (IS_ERR(entries[i])) { + err = PTR_ERR(entries[i]); idx = i - 1; goto fail; } - nid[i + 1] = get_nid(folios[i], offset[i + 1], false); + nid[i + 1] = get_nid(F2FS_I_SB(dn->inode), entries[i], offset[i + 1], false); } - f2fs_ra_node_pages(folios[idx], offset[idx + 1], NIDS_PER_BLOCK); + f2fs_ra_node_caches(entries[idx], offset[idx + 1], + NIDS_PER_BLOCK(F2FS_I_SB(dn->inode))); /* free direct nodes linked to a partial indirect node */ - for (i = offset[idx + 1]; i < NIDS_PER_BLOCK; i++) { - child_nid = get_nid(folios[idx], i, false); + for (i = offset[idx + 1]; + i < NIDS_PER_BLOCK(F2FS_I_SB(dn->inode)); i++) { + child_nid = get_nid(F2FS_I_SB(dn->inode), entries[idx], i, false); if (!child_nid) continue; dn->nid = child_nid; err = truncate_dnode(dn); if (err < 0) goto fail; - if (set_nid(folios[idx], i, 0, false)) + if (set_nid(F2FS_I_SB(dn->inode), entries[idx], i, 0, false)) dn->node_changed = true; } if (offset[idx + 1] == 0) { - dn->node_folio = folios[idx]; + dn->node_entry = entries[idx]; dn->nid = nid[idx]; err = truncate_node(dn); if (err) goto fail; } else { - f2fs_folio_put(folios[idx], true); + f2fs_put_cache(entries[idx], true); } offset[idx]++; offset[idx + 1] = 0; idx--; fail: for (i = idx; i >= 0; i--) - f2fs_folio_put(folios[i], true); + f2fs_put_cache(entries[i], true); trace_f2fs_truncate_partial_nodes(dn->inode, nid, depth, err); @@ -1191,7 +1190,9 @@ fail: } /* - * All the block addresses of data and nodes should be nullified. + * All the node blocks actually belong to the inode will be released. + * If the level is 0, we will simply truncate the dnode, + * or else we should do dynamic truncate for the node pointers with the depth. */ int f2fs_truncate_inode_blocks(struct inode *inode, pgoff_t from) { @@ -1200,7 +1201,7 @@ int f2fs_truncate_inode_blocks(struct inode *inode, pgoff_t from) int level, offset[4], noffset[4]; unsigned int nofs = 0; struct dnode_of_data dn; - struct folio *folio; + struct f2fs_cached_block *entry; trace_f2fs_truncate_inode_blocks_enter(inode, from); @@ -1217,14 +1218,14 @@ int f2fs_truncate_inode_blocks(struct inode *inode, pgoff_t from) return level; } - folio = f2fs_get_inode_folio(sbi, inode->i_ino); - if (IS_ERR(folio)) { - trace_f2fs_truncate_inode_blocks_exit(inode, PTR_ERR(folio)); - return PTR_ERR(folio); + entry = f2fs_get_inode_cache(sbi, inode->i_ino); + if (IS_ERR(entry)) { + trace_f2fs_truncate_inode_blocks_exit(inode, PTR_ERR(entry)); + return PTR_ERR(entry); } - set_new_dnode(&dn, inode, folio, NULL, 0); - folio_unlock(folio); + set_new_dnode(&dn, inode, entry, NULL, 0); + f2fs_unlock_cache(entry); switch (level) { case 0: @@ -1238,10 +1239,10 @@ int f2fs_truncate_inode_blocks(struct inode *inode, pgoff_t from) err = truncate_partial_nodes(&dn, offset, level); if (err < 0 && err != -ENOENT) goto fail; - nofs += 1 + NIDS_PER_BLOCK; + nofs += 1 + NIDS_PER_BLOCK(sbi); break; case 3: - nofs = 5 + 2 * NIDS_PER_BLOCK; + nofs = 5 + 2 * NIDS_PER_BLOCK(sbi); if (!offset[level - 1]) goto skip_partial; err = truncate_partial_nodes(&dn, offset, level); @@ -1254,28 +1255,21 @@ int f2fs_truncate_inode_blocks(struct inode *inode, pgoff_t from) skip_partial: while (cont) { - dn.nid = get_nid(folio, offset[0], true); - switch (offset[0]) { - case NODE_DIR1_BLOCK: - case NODE_DIR2_BLOCK: + dn.nid = get_nid(sbi, entry, offset[0], true); + if (offset[0] == NODE_DIR1_BLOCK(sbi) || + offset[0] == NODE_DIR2_BLOCK(sbi)) { err = truncate_dnode(&dn); - break; - - case NODE_IND1_BLOCK: - case NODE_IND2_BLOCK: + } else if (offset[0] == NODE_IND1_BLOCK(sbi) || + offset[0] == NODE_IND2_BLOCK(sbi)) { err = truncate_nodes(&dn, nofs, offset[1], 2); - break; - - case NODE_DIND_BLOCK: + } else if (offset[0] == NODE_DIND_BLOCK(sbi)) { err = truncate_nodes(&dn, nofs, offset[1], 3); cont = 0; - break; - - default: + } else { BUG(); } if (err == -ENOENT) { - set_sbi_flag(F2FS_F_SB(folio), SBI_NEED_FSCK); + set_sbi_flag(sbi, SBI_NEED_FSCK); f2fs_handle_error(sbi, ERROR_INVALID_BLKADDR); fserror_report_file_metadata(dn.inode, -EFSCORRUPTED, GFP_NOFS); @@ -1288,42 +1282,42 @@ skip_partial: } if (err < 0) goto fail; - if (offset[1] == 0 && get_nid(folio, offset[0], true)) { - folio_lock(folio); - BUG_ON(!is_node_folio(folio)); - set_nid(folio, offset[0], 0, true); - folio_unlock(folio); + if (offset[1] == 0 && get_nid(sbi, entry, offset[0], true)) { + f2fs_lock_cache(entry); + f2fs_bug_on(sbi, !f2fs_is_node_cache(entry)); + set_nid(sbi, entry, offset[0], 0, true); + f2fs_unlock_cache(entry); } offset[1] = 0; offset[0]++; nofs += err; } fail: - f2fs_folio_put(folio, false); + f2fs_put_cache(entry, false); trace_f2fs_truncate_inode_blocks_exit(inode, err); return err > 0 ? 0 : err; } -/* caller must lock inode page */ +/* caller must lock inode cached block */ int f2fs_truncate_xattr_node(struct inode *inode) { struct f2fs_sb_info *sbi = F2FS_I_SB(inode); nid_t nid = F2FS_I(inode)->i_xattr_nid; struct dnode_of_data dn; - struct folio *nfolio; + struct f2fs_cached_block *nentry; int err; if (!nid) return 0; - nfolio = f2fs_get_xnode_folio(sbi, nid); - if (IS_ERR(nfolio)) - return PTR_ERR(nfolio); + nentry = f2fs_get_xnode_cache(sbi, nid); + if (IS_ERR(nentry)) + return PTR_ERR(nentry); - set_new_dnode(&dn, inode, NULL, nfolio, nid); + set_new_dnode(&dn, inode, NULL, nentry, nid); err = truncate_node(&dn); if (err) { - f2fs_folio_put(nfolio, true); + f2fs_put_cache(nentry, true); return err; } @@ -1336,7 +1330,7 @@ int f2fs_truncate_xattr_node(struct inode *inode) * Caller should grab and release a rwsem by calling f2fs_lock_op() and * f2fs_unlock_op(). */ -int f2fs_remove_inode_page(struct inode *inode) +int f2fs_remove_inode_cache(struct inode *inode) { struct dnode_of_data dn; int err; @@ -1366,12 +1360,12 @@ int f2fs_remove_inode_page(struct inode *inode) if (unlikely(inode->i_blocks != 0 && inode->i_blocks != 8)) { f2fs_warn(F2FS_I_SB(inode), - "f2fs_remove_inode_page: inconsistent i_blocks, ino:%llu, iblocks:%llu", + "f2fs_remove_inode_cache: inconsistent i_blocks, ino:%llu, iblocks:%llu", inode->i_ino, (unsigned long long)inode->i_blocks); set_sbi_flag(F2FS_I_SB(inode), SBI_NEED_FSCK); } - /* will put inode & node pages */ + /* will put inode & node cached blocks */ err = truncate_node(&dn); if (err) { f2fs_put_dnode(&dn); @@ -1380,30 +1374,30 @@ int f2fs_remove_inode_page(struct inode *inode) return 0; } -struct folio *f2fs_new_inode_folio(struct inode *inode) +struct f2fs_cached_block *f2fs_new_inode_cache(struct inode *inode) { struct dnode_of_data dn; - /* allocate inode page for new inode */ + /* allocate inode block for new inode */ set_new_dnode(&dn, inode, NULL, NULL, inode->i_ino); - /* caller should f2fs_folio_put(folio, true); */ - return f2fs_new_node_folio(&dn, 0); + /* caller should f2fs_put_cache(entry, true); */ + return f2fs_new_node_cache(&dn, 0); } -struct folio *f2fs_new_node_folio(struct dnode_of_data *dn, unsigned int ofs) +struct f2fs_cached_block *f2fs_new_node_cache(struct dnode_of_data *dn, unsigned int ofs) { struct f2fs_sb_info *sbi = F2FS_I_SB(dn->inode); struct node_info new_ni; - struct folio *folio; + struct f2fs_cached_block *entry; int err; if (unlikely(is_inode_flag_set(dn->inode, FI_NO_ALLOC))) return ERR_PTR(-EPERM); - folio = f2fs_grab_cache_folio(NODE_MAPPING(sbi), dn->nid, false); - if (IS_ERR(folio)) - return folio; + entry = f2fs_grab_node_cache(sbi, dn->nid); + if (IS_ERR(entry)) + return entry; if (unlikely((err = inc_valid_node_count(sbi, dn->inode, !ofs)))) goto fail; @@ -1419,7 +1413,7 @@ struct folio *f2fs_new_node_folio(struct dnode_of_data *dn, unsigned int ofs) dec_valid_node_count(sbi, dn->inode, !ofs); set_sbi_flag(sbi, SBI_NEED_FSCK); f2fs_warn_ratelimited(sbi, - "f2fs_new_node_folio: inconsistent nat entry, " + "f2fs_new_node_cache: inconsistent nat entry, " "ino:%u, nid:%u, blkaddr:%u, ver:%u, flag:%u", new_ni.ino, new_ni.nid, new_ni.blk_addr, new_ni.version, new_ni.flag); @@ -1434,12 +1428,11 @@ struct folio *f2fs_new_node_folio(struct dnode_of_data *dn, unsigned int ofs) new_ni.version = 0; set_node_addr(sbi, &new_ni, NEW_ADDR, false); - f2fs_folio_wait_writeback(folio, NODE, true, true); - fill_node_footer(folio, dn->nid, dn->inode->i_ino, ofs, true); - set_cold_node(folio, S_ISDIR(dn->inode->i_mode)); - if (!folio_test_uptodate(folio)) - folio_mark_uptodate(folio); - if (folio_mark_dirty(folio)) + f2fs_cache_wait_writeback(entry); + fill_node_footer(sbi, entry, dn->nid, dn->inode->i_ino, ofs, true); + set_cold_node(sbi, entry, S_ISDIR(dn->inode->i_mode)); + f2fs_cache_set_uptodate(entry); + if (f2fs_mark_cache_dirty(entry)) dn->node_changed = true; if (f2fs_has_xattr_block(ofs)) @@ -1447,66 +1440,67 @@ struct folio *f2fs_new_node_folio(struct dnode_of_data *dn, unsigned int ofs) if (ofs == 0) inc_valid_inode_count(sbi); - return folio; + return entry; fail: - clear_node_folio_dirty(folio); - f2fs_folio_put(folio, true); + f2fs_drop_cache_dirty(entry); + f2fs_put_cache(entry, true); return ERR_PTR(err); } /* * Caller should do after getting the following values. - * 0: f2fs_folio_put(folio, false) - * LOCKED_PAGE or error: f2fs_folio_put(folio, true) + * 0: f2fs_put_cache(cache, false) + * LOCKED_CACHE or error: f2fs_put_cache(entry, true) */ -static int read_node_folio(struct folio *folio, blk_opf_t op_flags) +static int read_node_cache(struct f2fs_cached_block *entry, blk_opf_t op_flags) { - struct f2fs_sb_info *sbi = F2FS_F_SB(folio); + struct f2fs_sb_info *sbi = entry->cache->sbi; struct node_info ni; struct f2fs_io_info fio = { .sbi = sbi, .type = NODE, .op = REQ_OP_READ, .op_flags = op_flags, - .folio = folio, .encrypted_page = NULL, + .cache_entry = entry, + .is_cache = 1, }; int err; - if (folio_test_uptodate(folio)) { - if (!f2fs_inode_chksum_verify(sbi, folio)) { - folio_clear_uptodate(folio); + if (f2fs_cache_test_uptodate(entry)) { + if (!f2fs_inode_chksum_verify(sbi, entry)) { + f2fs_cache_clear_uptodate(entry); return -EFSBADCRC; } - return LOCKED_PAGE; + return LOCKED_CACHE; } - err = f2fs_get_node_info(sbi, folio->index, &ni, false); + err = f2fs_get_node_info(sbi, entry->index, &ni, false); if (err) return err; - /* NEW_ADDR can be seen, after cp_error drops some dirty node pages */ + /* NEW_ADDR can be seen, after cp_error drops some dirty node caches */ if (unlikely(ni.blk_addr == NULL_ADDR || ni.blk_addr == NEW_ADDR)) { - folio_clear_uptodate(folio); + f2fs_cache_clear_uptodate(entry); return -ENOENT; } fio.new_blkaddr = fio.old_blkaddr = ni.blk_addr; - err = f2fs_submit_page_bio(&fio); - + err = f2fs_submit_cache_read(&fio); if (!err) - f2fs_update_iostat(sbi, NULL, FS_NODE_READ_IO, F2FS_BLKSIZE); + f2fs_update_iostat(sbi, NULL, FS_NODE_READ_IO, + F2FS_BLKSIZE(sbi)); return err; } /* - * Readahead a node page + * Readahead a node cache */ -void f2fs_ra_node_page(struct f2fs_sb_info *sbi, nid_t nid) +void f2fs_ra_node_cache(struct f2fs_sb_info *sbi, nid_t nid) { - struct folio *afolio; + struct f2fs_cached_block *entry; int err; if (!nid) @@ -1514,29 +1508,29 @@ void f2fs_ra_node_page(struct f2fs_sb_info *sbi, nid_t nid) if (f2fs_check_nid_range(sbi, nid)) return; - afolio = xa_load(&NODE_MAPPING(sbi)->i_pages, nid); - if (afolio) + entry = xa_load(&NODE_CACHE(sbi)->root, nid); + if (entry) return; - afolio = f2fs_grab_cache_folio(NODE_MAPPING(sbi), nid, false); - if (IS_ERR(afolio)) + entry = f2fs_grab_node_cache(sbi, nid); + if (IS_ERR(entry)) return; - err = read_node_folio(afolio, REQ_RAHEAD); - f2fs_folio_put(afolio, err ? true : false); + err = read_node_cache(entry, REQ_RAHEAD); + f2fs_put_cache(entry, err ? true : false); } int f2fs_sanity_check_node_footer(struct f2fs_sb_info *sbi, - struct folio *folio, pgoff_t nid, + struct f2fs_cached_block *entry, pgoff_t nid, enum node_type ntype, bool in_irq) { bool is_inode, is_xnode; - if (unlikely(nid != nid_of_node(folio))) + if (unlikely(nid != nid_of_node(sbi, entry))) goto out_err; - is_inode = IS_INODE(folio); - is_xnode = f2fs_has_xattr_block(ofs_of_node(folio)); + is_inode = IS_INODE(sbi, entry); + is_xnode = f2fs_has_xattr_block(ofs_of_node(sbi, entry)); switch (ntype) { case NODE_TYPE_REGULAR: @@ -1569,20 +1563,20 @@ out_err: set_sbi_flag(sbi, SBI_NEED_FSCK); f2fs_warn_ratelimited(sbi, "inconsistent node block, node_type:%d, nid:%lu, " "node_footer[nid:%u,ino:%u,ofs:%u,cpver:%llu,blkaddr:%u]", - ntype, nid, nid_of_node(folio), ino_of_node(folio), - ofs_of_node(folio), cpver_of_node(folio), - next_blkaddr_of_node(folio)); + ntype, nid, nid_of_node(sbi, entry), ino_of_node(sbi, entry), + ofs_of_node(sbi, entry), cpver_of_node(sbi, entry), + next_blkaddr_of_node(sbi, entry)); f2fs_handle_error(sbi, ERROR_INCONSISTENT_FOOTER); - fserror_report_file_metadata(folio->mapping->host, - -EFSCORRUPTED, in_irq ? GFP_NOWAIT : GFP_NOFS); + fserror_report_metadata(sbi->sb, -EFSCORRUPTED, + in_irq ? GFP_NOWAIT : GFP_NOFS); return -EFSCORRUPTED; } -static struct folio *__get_node_folio(struct f2fs_sb_info *sbi, pgoff_t nid, - struct folio *parent, int start, enum node_type ntype) +static struct f2fs_cached_block *__get_node_cache(struct f2fs_sb_info *sbi, pgoff_t nid, + struct f2fs_cached_block *parent, int start, enum node_type ntype) { - struct folio *folio; + struct f2fs_cached_block *entry; int err; if (!nid) @@ -1590,71 +1584,71 @@ static struct folio *__get_node_folio(struct f2fs_sb_info *sbi, pgoff_t nid, if (f2fs_check_nid_range(sbi, nid)) return ERR_PTR(-EINVAL); repeat: - folio = f2fs_grab_cache_folio(NODE_MAPPING(sbi), nid, false); - if (IS_ERR(folio)) - return folio; + entry = f2fs_grab_node_cache(sbi, nid); + if (IS_ERR(entry)) + return entry; - err = read_node_folio(folio, 0); + err = read_node_cache(entry, 0); if (err < 0) goto out_put_err; - if (err == LOCKED_PAGE) - goto page_hit; + if (err == LOCKED_CACHE) + goto entry_hit; if (parent) - f2fs_ra_node_pages(parent, start + 1, MAX_RA_NODE); + f2fs_ra_node_caches(parent, start + 1, MAX_RA_NODE); - folio_lock(folio); + f2fs_lock_cache(entry); - if (unlikely(!is_node_folio(folio))) { - f2fs_folio_put(folio, true); + if (unlikely(!f2fs_is_node_cache(entry))) { + f2fs_put_cache(entry, true); goto repeat; } - if (unlikely(!folio_test_uptodate(folio))) { + if (unlikely(!f2fs_cache_test_uptodate(entry))) { err = -EIO; goto out_put_err; } - if (!f2fs_inode_chksum_verify(sbi, folio)) { + if (!f2fs_inode_chksum_verify(sbi, entry)) { err = -EFSBADCRC; goto out_err; } -page_hit: - err = f2fs_sanity_check_node_footer(sbi, folio, nid, ntype, false); +entry_hit: + err = f2fs_sanity_check_node_footer(sbi, entry, nid, ntype, false); if (!err) - return folio; + return entry; out_err: - clear_node_folio_dirty(folio); + f2fs_drop_cache_dirty(entry); out_put_err: /* ENOENT comes from read_node_folio which is not an error. */ if (err != -ENOENT) - f2fs_handle_page_eio(sbi, folio, NODE); - f2fs_folio_put(folio, true); + f2fs_handle_page_eio(sbi, entry->index, NODE); + f2fs_put_cache(entry, true); return ERR_PTR(err); } -struct folio *f2fs_get_node_folio(struct f2fs_sb_info *sbi, pgoff_t nid, +struct f2fs_cached_block *f2fs_get_node_cache(struct f2fs_sb_info *sbi, pgoff_t nid, enum node_type node_type) { - return __get_node_folio(sbi, nid, NULL, 0, node_type); + return __get_node_cache(sbi, nid, NULL, 0, node_type); } -struct folio *f2fs_get_inode_folio(struct f2fs_sb_info *sbi, pgoff_t ino) +struct f2fs_cached_block *f2fs_get_inode_cache(struct f2fs_sb_info *sbi, pgoff_t ino) { - return __get_node_folio(sbi, ino, NULL, 0, NODE_TYPE_INODE); + return __get_node_cache(sbi, ino, NULL, 0, NODE_TYPE_INODE); } -struct folio *f2fs_get_xnode_folio(struct f2fs_sb_info *sbi, pgoff_t xnid) +struct f2fs_cached_block *f2fs_get_xnode_cache(struct f2fs_sb_info *sbi, pgoff_t xnid) { - return __get_node_folio(sbi, xnid, NULL, 0, NODE_TYPE_XATTR); + return __get_node_cache(sbi, xnid, NULL, 0, NODE_TYPE_XATTR); } -static struct folio *f2fs_get_node_folio_ra(struct folio *parent, int start) +static struct f2fs_cached_block *f2fs_get_node_cache_ra(struct f2fs_cached_block *parent, int start) { - struct f2fs_sb_info *sbi = F2FS_F_SB(parent); - nid_t nid = get_nid(parent, start, false); + struct f2fs_sb_info *sbi = parent->cache->sbi; + nid_t nid = get_nid(sbi, parent, start, false); - return __get_node_folio(sbi, nid, parent, start, NODE_TYPE_NON_IXNODE); + return __get_node_cache(sbi, nid, parent, start, NODE_TYPE_NON_IXNODE); } static void flush_inline_data(struct f2fs_sb_info *sbi, nid_t ino) @@ -1693,110 +1687,106 @@ iput_out: iput(inode); } -static struct folio *last_fsync_dnode(struct f2fs_sb_info *sbi, nid_t ino) +static struct f2fs_cached_block *last_fsync_dnode(struct f2fs_sb_info *sbi, nid_t ino) { - pgoff_t index; - struct folio_batch fbatch; - struct folio *last_folio = NULL; - int nr_folios; - - folio_batch_init(&fbatch); - index = 0; + pgoff_t index = 0; + struct f2fs_cached_block *entries[F2FS_ONSTACK_CACHES]; + struct f2fs_cached_block *last_entry = NULL; + unsigned int nr; - while ((nr_folios = filemap_get_folios_tag(NODE_MAPPING(sbi), &index, - (pgoff_t)-1, PAGECACHE_TAG_DIRTY, - &fbatch))) { + while ((nr = f2fs_cache_gang_lookup_tag(NODE_CACHE(sbi), + entries, &index, F2FS_ONSTACK_CACHES, + F2FS_CACHE_TAG_DIRTY))) { int i; - for (i = 0; i < nr_folios; i++) { - struct folio *folio = fbatch.folios[i]; + for (i = 0; i < nr; i++) { + struct f2fs_cached_block *entry = entries[i]; if (unlikely(f2fs_cp_error(sbi))) { - f2fs_folio_put(last_folio, false); - folio_batch_release(&fbatch); + f2fs_put_cache(last_entry, false); + f2fs_cache_gang_release(entries, nr); return ERR_PTR(-EIO); } - if (!IS_DNODE(folio) || !is_cold_node(folio)) + if (!IS_DNODE(sbi, entry) || !is_cold_node(sbi, entry)) continue; - if (ino_of_node(folio) != ino) + if (ino_of_node(sbi, entry) != ino) continue; - folio_lock(folio); + f2fs_lock_cache(entry); - if (unlikely(!is_node_folio(folio))) { + if (unlikely(!f2fs_is_node_cache(entry))) { continue_unlock: - folio_unlock(folio); + f2fs_unlock_cache(entry); continue; } - if (ino_of_node(folio) != ino) + if (ino_of_node(sbi, entry) != ino) goto continue_unlock; - if (!folio_test_dirty(folio)) { + if (!f2fs_cache_test_dirty(entry)) { /* someone wrote it for us */ goto continue_unlock; } - if (last_folio) - f2fs_folio_put(last_folio, false); + if (last_entry) + f2fs_put_cache(last_entry, false); - folio_get(folio); - last_folio = folio; - folio_unlock(folio); + f2fs_cache_get(entry); + last_entry = entry; + f2fs_unlock_cache(entry); } - folio_batch_release(&fbatch); + f2fs_cache_gang_release(entries, nr); cond_resched(); } - return last_folio; + return last_entry; } -static bool __write_node_folio(struct folio *folio, bool atomic, bool do_fsync, - bool *submitted, struct writeback_control *wbc, - bool do_balance, enum iostat_type io_type, - unsigned int *seq_id) +static bool __write_node_cache(struct f2fs_cached_block *entry, + bool atomic, bool do_fsync, bool *submitted, + bool sync, bool do_balance, + enum iostat_type io_type, unsigned int *seq_id) { - struct f2fs_sb_info *sbi = F2FS_F_SB(folio); + struct f2fs_sb_info *sbi = entry->cache->sbi; nid_t nid; struct node_info ni; struct f2fs_io_info fio = { .sbi = sbi, - .ino = ino_of_node(folio), + .ino = ino_of_node(sbi, entry), .type = NODE, .op = REQ_OP_WRITE, - .op_flags = wbc_to_write_flags(wbc), - .folio = folio, + .op_flags = sync ? REQ_SYNC : REQ_BACKGROUND, + .cache_entry = entry, .encrypted_page = NULL, .submitted = 0, .io_type = io_type, - .io_wbc = wbc, + .is_cache = true, }; struct f2fs_lock_context lc; unsigned int seq; - trace_f2fs_writepage(folio, NODE); - if (unlikely(f2fs_cp_error(sbi))) { - /* keep node pages in remount-ro mode */ + /* keep node caches in remount-ro mode */ if (F2FS_OPTION(sbi).errors == MOUNT_ERRORS_READONLY) goto redirty_out; - folio_clear_uptodate(folio); - dec_page_count(sbi, F2FS_DIRTY_NODES); - folio_unlock(folio); + f2fs_cache_clear_uptodate(entry); + f2fs_cache_update_tag(entry, F2FS_CACHE_TAG_DIRTY, + F2FS_CACHE_TAG_NONE); + dec_cache_count(entry->cache->sbi, F2FS_DIRTY_NODES); + f2fs_unlock_cache(entry); return true; } if (unlikely(is_sbi_flag_set(sbi, SBI_POR_DOING))) goto redirty_out; - if (!is_sbi_flag_set(sbi, SBI_CP_DISABLED) && - wbc->sync_mode == WB_SYNC_NONE && - IS_DNODE(folio) && is_cold_node(folio)) + if (!is_sbi_flag_set(sbi, SBI_CP_DISABLED) && !sync && + IS_DNODE(sbi, entry) && is_cold_node(sbi, entry)) goto redirty_out; - /* get old block addr of this node page */ - nid = nid_of_node(folio); + /* get old block addr of this node cache */ + nid = nid_of_node(sbi, entry); - if (f2fs_sanity_check_node_footer(sbi, folio, folio->index, + if (f2fs_sanity_check_node_footer(sbi, entry, entry->index, NODE_TYPE_REGULAR, false)) { fserror_report_metadata(sbi->sb, -EFSCORRUPTED, GFP_NOFS); f2fs_stop_checkpoint(sbi, false, STOP_CP_REASON_CORRUPTED_NID); @@ -1808,12 +1798,14 @@ static bool __write_node_folio(struct folio *folio, bool atomic, bool do_fsync, f2fs_down_read_trace(&sbi->node_write, &lc); - /* This page is already truncated */ + /* This cache is already truncated */ if (unlikely(ni.blk_addr == NULL_ADDR)) { - folio_clear_uptodate(folio); - dec_page_count(sbi, F2FS_DIRTY_NODES); + f2fs_cache_clear_uptodate(entry); + f2fs_cache_update_tag(entry, F2FS_CACHE_TAG_DIRTY, + F2FS_CACHE_TAG_NONE); + dec_cache_count(entry->cache->sbi, F2FS_DIRTY_NODES); f2fs_up_read_trace(&sbi->node_write, &lc); - folio_unlock(folio); + f2fs_unlock_cache(entry); return true; } @@ -1827,28 +1819,28 @@ static bool __write_node_folio(struct folio *folio, bool atomic, bool do_fsync, if (atomic && !test_opt(sbi, NOBARRIER)) fio.op_flags |= REQ_PREFLUSH | REQ_FUA; - set_dentry_mark(folio, false); - set_fsync_mark(folio, do_fsync); - if (IS_INODE(folio) && (atomic || is_fsync_dnode(folio))) - set_dentry_mark(folio, - f2fs_need_dentry_mark(sbi, ino_of_node(folio))); + set_dentry_mark(sbi, entry, false); + set_fsync_mark(sbi, entry, do_fsync); + if (IS_INODE(sbi, entry) && (atomic || is_fsync_dnode(sbi, entry))) + set_dentry_mark(sbi, entry, + f2fs_need_dentry_mark(sbi, ino_of_node(sbi, entry))); - /* should add to global list before clearing PAGECACHE status */ - if (f2fs_in_warm_node_list(folio)) { - seq = f2fs_add_fsync_node_entry(sbi, folio); + /* should add to global list before clearing CACHE status */ + if (f2fs_in_warm_node_list(sbi, entry)) { + seq = f2fs_add_fsync_node_entry(sbi, entry); if (seq_id) *seq_id = seq; } - folio_start_writeback(folio); + f2fs_start_cache_writeback(entry); fio.old_blkaddr = ni.blk_addr; - f2fs_do_write_node_page(nid, &fio); - set_node_addr(sbi, &ni, fio.new_blkaddr, is_fsync_dnode(folio)); - dec_page_count(sbi, F2FS_DIRTY_NODES); + f2fs_do_write_node_cache(nid, &fio); + set_node_addr(sbi, &ni, fio.new_blkaddr, is_fsync_dnode(sbi, entry)); + dec_cache_count(sbi, F2FS_DIRTY_NODES); f2fs_up_read_trace(&sbi->node_write, &lc); - folio_unlock(folio); + f2fs_unlock_cache(entry); if (unlikely(f2fs_cp_error(sbi))) { f2fs_submit_merged_write(sbi, NODE); @@ -1862,176 +1854,171 @@ static bool __write_node_folio(struct folio *folio, bool atomic, bool do_fsync, return true; redirty_out: - folio_redirty_for_writepage(wbc, folio); - folio_unlock(folio); + f2fs_cache_set_dirty(entry); + f2fs_unlock_cache(entry); return false; } -int f2fs_write_single_node_folio(struct folio *node_folio, int sync_mode, +int f2fs_write_node_cache(struct f2fs_cached_block *node_entry, int sync_mode, bool mark_dirty, enum iostat_type io_type) { int err = 0; - struct writeback_control wbc = { - .sync_mode = WB_SYNC_ALL, - .nr_to_write = 1, - }; if (!sync_mode) { - /* set page dirty and write it */ - if (!folio_test_writeback(node_folio)) - folio_mark_dirty(node_folio); - goto out_folio; + /* set cache dirty and write it */ + if (!f2fs_cache_test_writeback(node_entry)) + f2fs_mark_cache_dirty(node_entry); + goto out_entry; } - f2fs_folio_wait_writeback(node_folio, NODE, true, true); + f2fs_cache_wait_writeback(node_entry); if (mark_dirty) - folio_mark_dirty(node_folio); - else if (!folio_test_dirty(node_folio)) - goto out_folio; + f2fs_mark_cache_dirty(node_entry); + else if (!f2fs_cache_test_dirty(node_entry)) + goto out_entry; - if (!folio_clear_dirty_for_io(node_folio)) { + if (!f2fs_cache_test_and_clear_dirty(node_entry)) { err = -EAGAIN; - goto out_folio; + goto out_entry; } - if (!__write_node_folio(node_folio, false, false, NULL, - &wbc, false, io_type, NULL)) + if (!__write_node_cache(node_entry, false, false, NULL, + true, false, io_type, NULL)) err = -EAGAIN; - goto release_folio; -out_folio: - folio_unlock(node_folio); -release_folio: - f2fs_folio_put(node_folio, false); + goto release_entry; +out_entry: + f2fs_unlock_cache(node_entry); +release_entry: + f2fs_put_cache(node_entry, false); return err; } -int f2fs_move_node_folio(struct folio *node_folio, int gc_type) +int f2fs_move_node_cache(struct f2fs_cached_block *entry, int gc_type) { - return f2fs_write_single_node_folio(node_folio, gc_type == FG_GC, + return f2fs_write_node_cache(entry, gc_type == FG_GC, true, FS_GC_NODE_IO); } -int f2fs_fsync_node_pages(struct f2fs_sb_info *sbi, struct inode *inode, - struct writeback_control *wbc, bool atomic, - unsigned int *seq_id) +int f2fs_fsync_node_caches(struct f2fs_sb_info *sbi, struct inode *inode, + bool atomic, unsigned int *seq_id) { pgoff_t index; - struct folio_batch fbatch; + struct f2fs_cached_block *entries[F2FS_ONSTACK_CACHES]; int ret = 0; - struct folio *last_folio = NULL; + struct f2fs_cached_block *last_entry = NULL; bool marked = false; nid_t ino = inode->i_ino; - int nr_folios; + int nr; int nwritten = 0; if (atomic) { - last_folio = last_fsync_dnode(sbi, ino); - if (IS_ERR_OR_NULL(last_folio)) - return PTR_ERR_OR_ZERO(last_folio); + last_entry = last_fsync_dnode(sbi, ino); + if (IS_ERR_OR_NULL(last_entry)) + return PTR_ERR_OR_ZERO(last_entry); } retry: - folio_batch_init(&fbatch); index = 0; - while ((nr_folios = filemap_get_folios_tag(NODE_MAPPING(sbi), &index, - (pgoff_t)-1, PAGECACHE_TAG_DIRTY, - &fbatch))) { + while ((nr = f2fs_cache_gang_lookup_tag(NODE_CACHE(sbi), + entries, &index, F2FS_ONSTACK_CACHES, + F2FS_CACHE_TAG_DIRTY))) { int i; - for (i = 0; i < nr_folios; i++) { - struct folio *folio = fbatch.folios[i]; + for (i = 0; i < nr; i++) { + struct f2fs_cached_block *entry = entries[i]; bool submitted = false; bool do_fsync = false; if (unlikely(f2fs_cp_error(sbi))) { - f2fs_folio_put(last_folio, false); - folio_batch_release(&fbatch); + f2fs_put_cache(last_entry, false); + f2fs_cache_gang_release(entries, nr); ret = -EIO; goto out; } - if (!IS_DNODE(folio) || !is_cold_node(folio)) + if (!IS_DNODE(sbi, entry) || !is_cold_node(sbi, entry)) continue; - if (ino_of_node(folio) != ino) + if (ino_of_node(sbi, entry) != ino) continue; - folio_lock(folio); + f2fs_lock_cache(entry); - if (unlikely(!is_node_folio(folio))) { + if (unlikely(!f2fs_is_node_cache(entry))) { continue_unlock: - folio_unlock(folio); + f2fs_unlock_cache(entry); continue; } - if (ino_of_node(folio) != ino) + if (ino_of_node(sbi, entry) != ino) goto continue_unlock; - if (!folio_test_dirty(folio) && folio != last_folio) { + if (!f2fs_cache_test_dirty(entry) && entry != last_entry) { /* someone wrote it for us */ goto continue_unlock; } - f2fs_folio_wait_writeback(folio, NODE, true, true); + f2fs_cache_wait_writeback(entry); - if (!atomic || folio == last_folio) { + if (!atomic || entry == last_entry) { do_fsync = true; percpu_counter_inc(&sbi->rf_node_block_count); - if (IS_INODE(folio)) { + if (IS_INODE(sbi, entry)) { if (is_inode_flag_set(inode, FI_DIRTY_INODE)) - f2fs_update_inode(inode, folio); + f2fs_update_inode(inode, entry); } /* may be written by other thread */ - if (!folio_test_dirty(folio)) - folio_mark_dirty(folio); + if (!f2fs_cache_test_dirty(entry)) + f2fs_mark_cache_dirty(entry); } - if (!folio_clear_dirty_for_io(folio)) + if (!f2fs_cache_test_and_clear_dirty(entry)) goto continue_unlock; - if (!__write_node_folio(folio, atomic && - folio == last_folio, + if (!__write_node_cache(entry, atomic && + entry == last_entry, do_fsync, &submitted, - wbc, true, FS_NODE_IO, + true, true, FS_NODE_IO, seq_id)) { - f2fs_folio_put(last_folio, false); - folio_batch_release(&fbatch); + f2fs_put_cache(last_entry, false); + f2fs_cache_gang_release(entries, nr); ret = -EIO; goto out; } if (submitted) nwritten++; - if (folio == last_folio) { - f2fs_folio_put(folio, false); - folio_batch_release(&fbatch); + if (entry == last_entry) { + f2fs_put_cache(entry, false); + f2fs_cache_gang_release(entries, nr); marked = true; goto out; } } - folio_batch_release(&fbatch); + f2fs_cache_gang_release(entries, nr); cond_resched(); } if (atomic && !marked) { f2fs_debug(sbi, "Retry to write fsync mark: ino=%u, idx=%lx", - ino, last_folio->index); - folio_lock(last_folio); - if (unlikely(!is_node_folio(last_folio))) { - f2fs_folio_put(last_folio, true); + ino, last_entry->index); + f2fs_lock_cache(last_entry); + if (unlikely(!f2fs_is_node_cache(last_entry))) { + f2fs_put_cache(last_entry, true); ret = -EAGAIN; goto out; } - f2fs_folio_wait_writeback(last_folio, NODE, true, true); - folio_mark_dirty(last_folio); - folio_unlock(last_folio); + f2fs_cache_wait_writeback(last_entry); + f2fs_mark_cache_dirty(last_entry); + f2fs_unlock_cache(last_entry); goto retry; } out: if (nwritten) - f2fs_submit_merged_write_cond(sbi, NULL, NULL, ino, NODE); + f2fs_submit_merged_write_cache(sbi, NULL, ino, NODE); return ret; } + static int f2fs_match_ino(struct inode *inode, u64 ino, void *data) { struct f2fs_sb_info *sbi = F2FS_I_SB(inode); @@ -2056,18 +2043,18 @@ static int f2fs_match_ino(struct inode *inode, u64 ino, void *data) return 1; } -static bool flush_dirty_inode(struct folio *folio) +static bool flush_dirty_cache_inode(struct f2fs_cached_block *entry) { - struct f2fs_sb_info *sbi = F2FS_F_SB(folio); + struct f2fs_sb_info *sbi = entry->cache->sbi; struct inode *inode; - nid_t ino = ino_of_node(folio); + nid_t ino = ino_of_node(sbi, entry); inode = find_inode_nowait(sbi->sb, ino, f2fs_match_ino, NULL); if (!inode) return false; - f2fs_update_inode(inode, folio); - folio_unlock(folio); + f2fs_update_inode(inode, entry); + f2fs_unlock_cache(entry); iput(inode); return true; @@ -2076,72 +2063,74 @@ static bool flush_dirty_inode(struct folio *folio) void f2fs_flush_inline_data(struct f2fs_sb_info *sbi) { pgoff_t index = 0; - struct folio_batch fbatch; - int nr_folios; + struct f2fs_cached_block *entries[F2FS_ONSTACK_CACHES]; + unsigned int nr; - folio_batch_init(&fbatch); - - while ((nr_folios = filemap_get_folios_tag(NODE_MAPPING(sbi), &index, - (pgoff_t)-1, PAGECACHE_TAG_DIRTY, - &fbatch))) { + while ((nr = f2fs_cache_gang_lookup_tag(NODE_CACHE(sbi), + entries, &index, F2FS_ONSTACK_CACHES, + F2FS_CACHE_TAG_DIRTY))) { int i; - for (i = 0; i < nr_folios; i++) { - struct folio *folio = fbatch.folios[i]; - - if (!IS_INODE(folio)) - continue; + for (i = 0; i < nr; i++) { + struct f2fs_cached_block *entry = entries[i]; - folio_lock(folio); + f2fs_lock_cache(entry); - if (unlikely(!is_node_folio(folio))) + if (unlikely(!f2fs_is_node_cache(entry))) + goto unlock; + if (!IS_INODE(sbi, entry)) goto unlock; - if (!folio_test_dirty(folio)) + if (!f2fs_cache_test_dirty(entry)) goto unlock; /* flush inline_data, if it's async context. */ - if (folio_test_f2fs_inline(folio)) { - folio_clear_f2fs_inline(folio); - folio_unlock(folio); - flush_inline_data(sbi, ino_of_node(folio)); + if (f2fs_cache_test_inline(entry)) { + nid_t ino = ino_of_node(sbi, entry); + + f2fs_cache_clear_inline(entry); + f2fs_unlock_cache(entry); + flush_inline_data(sbi, ino); continue; } unlock: - folio_unlock(folio); + f2fs_unlock_cache(entry); } - folio_batch_release(&fbatch); + f2fs_cache_gang_release(entries, nr); cond_resched(); } } -int f2fs_sync_node_pages(struct f2fs_sb_info *sbi, - struct writeback_control *wbc, - bool do_balance, enum iostat_type io_type) +int f2fs_writeback_node_caches(struct f2fs_sb_info *sbi, long nr_to_write, + bool sync, bool do_balance, enum iostat_type io_type) { pgoff_t index; - struct folio_batch fbatch; + struct f2fs_cached_block *entries[F2FS_ONSTACK_CACHES]; int step = 0; int nwritten = 0; int ret = 0; - int nr_folios, done = 0; + int nr, done = 0; - folio_batch_init(&fbatch); + trace_f2fs_write_caches(sbi, nr_to_write, 0, NODE); next_step: index = 0; - while (!done && (nr_folios = filemap_get_folios_tag(NODE_MAPPING(sbi), - &index, (pgoff_t)-1, PAGECACHE_TAG_DIRTY, - &fbatch))) { + while (!done && (nr = f2fs_cache_gang_lookup_tag(NODE_CACHE(sbi), + entries, &index, F2FS_ONSTACK_CACHES, + F2FS_CACHE_TAG_DIRTY))) { int i; - for (i = 0; i < nr_folios; i++) { - struct folio *folio = fbatch.folios[i]; + for (i = 0; i < nr; i++) { + struct f2fs_cached_block *entry = entries[i]; bool submitted = false; + if (!sync && unlikely(freezing(current))) { + done = 1; + break; + } + /* give a priority to WB_SYNC threads */ - if (atomic_read(&sbi->wb_sync_req[NODE]) && - wbc->sync_mode == WB_SYNC_NONE) { + if (atomic_read(&sbi->wb_sync_req[NODE]) && !sync) { done = 1; break; } @@ -2152,27 +2141,27 @@ next_step: * 1. dentry dnodes * 2. file dnodes */ - if (step == 0 && IS_DNODE(folio)) + if (step == 0 && IS_DNODE(sbi, entry)) continue; - if (step == 1 && (!IS_DNODE(folio) || - is_cold_node(folio))) + if (step == 1 && (!IS_DNODE(sbi, entry) || + is_cold_node(sbi, entry))) continue; - if (step == 2 && (!IS_DNODE(folio) || - !is_cold_node(folio))) + if (step == 2 && (!IS_DNODE(sbi, entry) || + !is_cold_node(sbi, entry))) continue; lock_node: - if (wbc->sync_mode == WB_SYNC_ALL) - folio_lock(folio); - else if (!folio_trylock(folio)) + if (sync) + f2fs_lock_cache(entry); + else if (!f2fs_trylock_cache(entry)) continue; - if (unlikely(!is_node_folio(folio))) { + if (unlikely(!f2fs_is_node_cache(entry))) { continue_unlock: - folio_unlock(folio); + f2fs_unlock_cache(entry); continue; } - if (!folio_test_dirty(folio)) { + if (!f2fs_cache_test_dirty(entry)) { /* someone wrote it for us */ goto continue_unlock; } @@ -2182,38 +2171,38 @@ continue_unlock: goto write_node; /* flush inline_data */ - if (folio_test_f2fs_inline(folio)) { - folio_clear_f2fs_inline(folio); - folio_unlock(folio); - flush_inline_data(sbi, ino_of_node(folio)); + if (f2fs_cache_test_inline(entry)) { + f2fs_cache_clear_inline(entry); + f2fs_unlock_cache(entry); + flush_inline_data(sbi, ino_of_node(sbi, entry)); goto lock_node; } /* flush dirty inode */ - if (IS_INODE(folio) && flush_dirty_inode(folio)) + if (IS_INODE(sbi, entry) && flush_dirty_cache_inode(entry)) goto lock_node; write_node: - f2fs_folio_wait_writeback(folio, NODE, true, true); + f2fs_cache_wait_writeback(entry); - if (!folio_clear_dirty_for_io(folio)) + if (!f2fs_cache_test_and_clear_dirty(entry)) goto continue_unlock; - if (!__write_node_folio(folio, false, false, &submitted, - wbc, do_balance, io_type, NULL)) { - folio_batch_release(&fbatch); + if (!__write_node_cache(entry, false, false, &submitted, + sync, do_balance, io_type, NULL)) { + f2fs_cache_gang_release(entries, nr); ret = -EIO; goto out; } if (submitted) nwritten++; - if (--wbc->nr_to_write == 0) + if (nwritten >= nr_to_write) break; } - folio_batch_release(&fbatch); + f2fs_cache_gang_release(entries, nr); cond_resched(); - if (wbc->nr_to_write == 0) { + if (nwritten >= nr_to_write) { step = 2; break; } @@ -2221,7 +2210,7 @@ write_node: if (step < 2) { if (!is_sbi_flag_set(sbi, SBI_CP_DISABLED) && - wbc->sync_mode == WB_SYNC_NONE && step == 1) + !sync && step == 1) goto out; step++; goto next_step; @@ -2230,12 +2219,14 @@ out: if (nwritten) f2fs_submit_merged_write(sbi, NODE); + trace_f2fs_write_caches(sbi, nr_to_write, nwritten, NODE); + if (unlikely(f2fs_cp_error(sbi))) return -EIO; return ret; } -int f2fs_wait_on_node_pages_writeback(struct f2fs_sb_info *sbi, +int f2fs_wait_on_node_caches_writeback(struct f2fs_sb_info *sbi, unsigned int seq_id) { struct fsync_node_entry *fn; @@ -2244,7 +2235,7 @@ int f2fs_wait_on_node_pages_writeback(struct f2fs_sb_info *sbi, unsigned int cur_seq_id = 0; while (seq_id && cur_seq_id < seq_id) { - struct folio *folio; + struct f2fs_cached_block *entry; spin_lock_irqsave(&sbi->fsync_node_lock, flags); if (list_empty(head)) { @@ -2257,94 +2248,49 @@ int f2fs_wait_on_node_pages_writeback(struct f2fs_sb_info *sbi, break; } cur_seq_id = fn->seq_id; - folio = fn->folio; - folio_get(folio); + entry = fn->entry; + f2fs_cache_get(entry); spin_unlock_irqrestore(&sbi->fsync_node_lock, flags); - f2fs_folio_wait_writeback(folio, NODE, true, false); + f2fs_lock_cache(entry); + f2fs_cache_wait_writeback(entry); - folio_put(folio); + f2fs_put_cache(entry, true); } - return filemap_check_errors(NODE_MAPPING(sbi)); + return f2fs_cp_error(sbi) ? -EIO : 0; } -static int f2fs_write_node_pages(struct address_space *mapping, - struct writeback_control *wbc) +int f2fs_write_node_caches(struct f2fs_sb_info *sbi) { - struct f2fs_sb_info *sbi = F2FS_M_SB(mapping); struct blk_plug plug; - long diff; + long nr_to_write = LONG_MAX; if (unlikely(is_sbi_flag_set(sbi, SBI_POR_DOING))) - goto skip_write; + return -EAGAIN; /* balancing f2fs's metadata in background */ f2fs_balance_fs_bg(sbi, true); - /* collect a number of dirty node pages and write together */ - if (wbc->sync_mode != WB_SYNC_ALL && - get_pages(sbi, F2FS_DIRTY_NODES) < - nr_pages_to_skip(sbi, NODE)) - goto skip_write; + /* collect a number of dirty node caches and write together */ + if (get_nr_caches(sbi, F2FS_DIRTY_NODES) < + nr_caches_to_skip(sbi, NODE)) + return -EAGAIN; - if (wbc->sync_mode == WB_SYNC_ALL) - atomic_inc(&sbi->wb_sync_req[NODE]); - else if (atomic_read(&sbi->wb_sync_req[NODE])) { + if (atomic_read(&sbi->wb_sync_req[NODE])) { /* to avoid potential deadlock */ if (current->plug) blk_finish_plug(current->plug); - goto skip_write; + return -EAGAIN; } - trace_f2fs_writepages(mapping->host, wbc, NODE); - - diff = nr_pages_to_write(sbi, NODE, wbc); + nr_to_write = adjust_flush_cache_number(sbi, NODE); blk_start_plug(&plug); - f2fs_sync_node_pages(sbi, wbc, true, FS_NODE_IO); + f2fs_writeback_node_caches(sbi, nr_to_write, false, true, FS_NODE_IO); blk_finish_plug(&plug); - wbc->nr_to_write = max((long)0, wbc->nr_to_write - diff); - - if (wbc->sync_mode == WB_SYNC_ALL) - atomic_dec(&sbi->wb_sync_req[NODE]); return 0; - -skip_write: - wbc->pages_skipped += get_pages(sbi, F2FS_DIRTY_NODES); - trace_f2fs_writepages(mapping->host, wbc, NODE); - return 0; -} - -static bool f2fs_dirty_node_folio(struct address_space *mapping, - struct folio *folio) -{ - trace_f2fs_set_page_dirty(folio, NODE); - - if (!folio_test_uptodate(folio)) - folio_mark_uptodate(folio); -#ifdef CONFIG_F2FS_CHECK_FS - if (IS_INODE(folio)) - f2fs_inode_chksum_set(F2FS_M_SB(mapping), folio); -#endif - if (filemap_dirty_folio(mapping, folio)) { - inc_page_count(F2FS_M_SB(mapping), F2FS_DIRTY_NODES); - folio_set_f2fs_reference(folio); - return true; - } - return false; } -/* - * Structure of the f2fs node operations - */ -const struct address_space_operations f2fs_node_aops = { - .writepages = f2fs_write_node_pages, - .dirty_folio = f2fs_dirty_node_folio, - .invalidate_folio = f2fs_invalidate_folio, - .release_folio = f2fs_release_folio, - .migrate_folio = filemap_migrate_folio, -}; - static struct free_nid *__lookup_free_nid_list(struct f2fs_nm_info *nm_i, nid_t n) { @@ -2403,8 +2349,8 @@ static void update_free_nid_bitmap(struct f2fs_sb_info *sbi, nid_t nid, bool set, bool build) { struct f2fs_nm_info *nm_i = NM_I(sbi); - unsigned int nat_ofs = NAT_BLOCK_OFFSET(nid); - unsigned int nid_ofs = nid - START_NID(nid); + unsigned int nat_ofs = NAT_BLOCK_OFFSET(sbi, nid); + unsigned int nid_ofs = nid - f2fs_start_nid(sbi, nid); if (!test_bit_le(nat_ofs, nm_i->nat_block_bitmap)) return; @@ -2461,7 +2407,7 @@ static bool add_free_nid(struct f2fs_sb_info *sbi, * - f2fs_balance_fs_bg * - f2fs_build_free_nids * - __f2fs_build_free_nids - * - scan_nat_page + * - scan_nat_blocks * - add_free_nid * - __lookup_nat_cache * - f2fs_add_link @@ -2519,19 +2465,19 @@ static void remove_free_nid(struct f2fs_sb_info *sbi, nid_t nid) kmem_cache_free(free_nid_slab, i); } -static int scan_nat_page(struct f2fs_sb_info *sbi, +static int scan_nat_block(struct f2fs_sb_info *sbi, struct f2fs_nat_block *nat_blk, nid_t start_nid) { struct f2fs_nm_info *nm_i = NM_I(sbi); block_t blk_addr; - unsigned int nat_ofs = NAT_BLOCK_OFFSET(start_nid); + unsigned int nat_ofs = NAT_BLOCK_OFFSET(sbi, start_nid); int i; __set_bit_le(nat_ofs, nm_i->nat_block_bitmap); - i = start_nid % NAT_ENTRY_PER_BLOCK; + i = start_nid % NAT_ENTRY_PER_BLOCK(sbi); - for (; i < NAT_ENTRY_PER_BLOCK; i++, start_nid++) { + for (; i < NAT_ENTRY_PER_BLOCK(sbi); i++, start_nid++) { if (unlikely(start_nid >= nm_i->max_nid)) break; @@ -2587,16 +2533,16 @@ static void scan_free_nid_bits(struct f2fs_sb_info *sbi) continue; if (!nm_i->free_nid_count[i]) continue; - for (idx = 0; idx < NAT_ENTRY_PER_BLOCK; idx++) { + for (idx = 0; idx < NAT_ENTRY_PER_BLOCK(sbi); idx++) { idx = find_next_bit_le(nm_i->free_nid_bitmap[i], - NAT_ENTRY_PER_BLOCK, idx); - if (idx >= NAT_ENTRY_PER_BLOCK) + NAT_ENTRY_PER_BLOCK(sbi), idx); + if (idx >= NAT_ENTRY_PER_BLOCK(sbi)) break; - nid = i * NAT_ENTRY_PER_BLOCK + idx; + nid = i * NAT_ENTRY_PER_BLOCK(sbi) + idx; add_free_nid(sbi, nid, true, false); - if (nm_i->nid_cnt[FREE_NID] >= MAX_FREE_NIDS) + if (nm_i->nid_cnt[FREE_NID] >= MAX_FREE_NIDS(sbi)) goto out; } } @@ -2610,6 +2556,7 @@ static int __f2fs_build_free_nids(struct f2fs_sb_info *sbi, bool sync, bool mount) { struct f2fs_nm_info *nm_i = NM_I(sbi); + struct f2fs_cached_block *entry = NULL; int i = 0, ret; nid_t nid = nm_i->next_scan_nid; struct f2fs_lock_context lc; @@ -2617,11 +2564,11 @@ static int __f2fs_build_free_nids(struct f2fs_sb_info *sbi, if (unlikely(nid >= nm_i->max_nid)) nid = 0; - if (unlikely(nid % NAT_ENTRY_PER_BLOCK)) - nid = NAT_BLOCK_OFFSET(nid) * NAT_ENTRY_PER_BLOCK; + if (unlikely(nid % NAT_ENTRY_PER_BLOCK(sbi))) + nid = NAT_BLOCK_OFFSET(sbi, nid) * NAT_ENTRY_PER_BLOCK(sbi); /* Enough entries */ - if (nm_i->nid_cnt[FREE_NID] >= NAT_ENTRY_PER_BLOCK) + if (nm_i->nid_cnt[FREE_NID] >= NAT_ENTRY_PER_BLOCK(sbi)) return 0; if (!sync && !f2fs_available_free_memory(sbi, FREE_NIDS)) @@ -2631,27 +2578,27 @@ static int __f2fs_build_free_nids(struct f2fs_sb_info *sbi, /* try to find free nids in free_nid_bitmap */ scan_free_nid_bits(sbi); - if (nm_i->nid_cnt[FREE_NID] >= NAT_ENTRY_PER_BLOCK) + if (nm_i->nid_cnt[FREE_NID] >= NAT_ENTRY_PER_BLOCK(sbi)) return 0; } - /* readahead nat pages to be scanned */ - f2fs_ra_meta_pages(sbi, NAT_BLOCK_OFFSET(nid), FREE_NID_PAGES, + /* readahead nat blocks to be scanned */ + f2fs_ra_meta_caches(sbi, NAT_BLOCK_OFFSET(sbi, nid), FREE_NID_BLOCKS, META_NAT, true); f2fs_down_read_trace(&nm_i->nat_tree_lock, &lc); while (1) { - if (!test_bit_le(NAT_BLOCK_OFFSET(nid), + if (!test_bit_le(NAT_BLOCK_OFFSET(sbi, nid), nm_i->nat_block_bitmap)) { - struct folio *folio = get_current_nat_folio(sbi, nid); + entry = get_current_nat_cache(sbi, nid); - if (IS_ERR(folio)) { - ret = PTR_ERR(folio); + if (IS_ERR(entry)) { + ret = PTR_ERR(entry); } else { - ret = scan_nat_page(sbi, folio_address(folio), + ret = scan_nat_block(sbi, cache_address(entry), nid); - f2fs_folio_put(folio, true); + f2fs_put_cache(entry, true); } if (ret) { @@ -2668,24 +2615,25 @@ static int __f2fs_build_free_nids(struct f2fs_sb_info *sbi, } } - nid += (NAT_ENTRY_PER_BLOCK - (nid % NAT_ENTRY_PER_BLOCK)); + nid += NAT_ENTRY_PER_BLOCK(sbi) - + (nid % NAT_ENTRY_PER_BLOCK(sbi)); if (unlikely(nid >= nm_i->max_nid)) nid = 0; - if (++i >= FREE_NID_PAGES) + if (++i >= FREE_NID_BLOCKS) break; } - /* go to the next free nat pages to find free nids abundantly */ + /* go to the next free nat blocks to find free nids abundantly */ nm_i->next_scan_nid = nid; - /* find free nids from current sum_pages */ + /* find free nids from current sum_blocks */ scan_curseg_cache(sbi); f2fs_up_read_trace(&nm_i->nat_tree_lock, &lc); - f2fs_ra_meta_pages(sbi, NAT_BLOCK_OFFSET(nm_i->next_scan_nid), - nm_i->ra_nid_pages, META_NAT, false); + f2fs_ra_meta_caches(sbi, NAT_BLOCK_OFFSET(sbi, nm_i->next_scan_nid), + nm_i->ra_nid_blocks, META_NAT, false); return 0; } @@ -2750,7 +2698,7 @@ retry: } spin_unlock(&nm_i->nid_list_lock); - /* Let's scan nat pages and its caches to get free nids */ + /* Let's scan nat blocks and its caches to get free nids */ if (!f2fs_build_free_nids(sbi, true, false)) goto retry; return false; @@ -2811,20 +2759,20 @@ int f2fs_try_to_free_nids(struct f2fs_sb_info *sbi, int nr_shrink) struct f2fs_nm_info *nm_i = NM_I(sbi); int nr = nr_shrink; - if (nm_i->nid_cnt[FREE_NID] <= MAX_FREE_NIDS) + if (nm_i->nid_cnt[FREE_NID] <= MAX_FREE_NIDS(sbi)) return 0; if (!mutex_trylock(&nm_i->build_lock)) return 0; - while (nr_shrink && nm_i->nid_cnt[FREE_NID] > MAX_FREE_NIDS) { + while (nr_shrink && nm_i->nid_cnt[FREE_NID] > MAX_FREE_NIDS(sbi)) { struct free_nid *i, *next; unsigned int batch = SHRINK_NID_BATCH_SIZE; spin_lock(&nm_i->nid_list_lock); list_for_each_entry_safe(i, next, &nm_i->free_nid_list, list) { if (!nr_shrink || !batch || - nm_i->nid_cnt[FREE_NID] <= MAX_FREE_NIDS) + nm_i->nid_cnt[FREE_NID] <= MAX_FREE_NIDS(sbi)) break; __remove_free_nid(sbi, i, FREE_NID); kmem_cache_free(free_nid_slab, i); @@ -2839,18 +2787,18 @@ int f2fs_try_to_free_nids(struct f2fs_sb_info *sbi, int nr_shrink) return nr - nr_shrink; } -int f2fs_recover_inline_xattr(struct inode *inode, struct folio *folio) +int f2fs_recover_inline_xattr(struct inode *inode, struct f2fs_cached_block *entry) { void *src_addr, *dst_addr; size_t inline_size; - struct folio *ifolio; + struct f2fs_cached_block *ientry; struct f2fs_inode *ri; - ifolio = f2fs_get_inode_folio(F2FS_I_SB(inode), inode->i_ino); - if (IS_ERR(ifolio)) - return PTR_ERR(ifolio); + ientry = f2fs_get_inode_cache(F2FS_I_SB(inode), inode->i_ino); + if (IS_ERR(ientry)) + return PTR_ERR(ientry); - ri = F2FS_INODE(folio); + ri = &CACHED_NODE(entry)->i; if (ri->i_inline & F2FS_INLINE_XATTR) { if (!f2fs_has_inline_xattr(inode)) { set_inode_flag(inode, FI_INLINE_XATTR); @@ -2864,26 +2812,26 @@ int f2fs_recover_inline_xattr(struct inode *inode, struct folio *folio) goto update_inode; } - dst_addr = inline_xattr_addr(inode, ifolio); - src_addr = inline_xattr_addr(inode, folio); + dst_addr = inline_xattr_addr(inode, ientry); + src_addr = inline_xattr_addr(inode, entry); inline_size = inline_xattr_size(inode); - f2fs_folio_wait_writeback(ifolio, NODE, true, true); + f2fs_cache_wait_writeback(ientry); memcpy(dst_addr, src_addr, inline_size); update_inode: - f2fs_update_inode(inode, ifolio); - f2fs_folio_put(ifolio, true); + f2fs_update_inode(inode, ientry); + f2fs_put_cache(ientry, true); return 0; } -int f2fs_recover_xattr_data(struct inode *inode, struct folio *folio) +int f2fs_recover_xattr_data(struct inode *inode, struct f2fs_cached_block *entry) { struct f2fs_sb_info *sbi = F2FS_I_SB(inode); nid_t prev_xnid = F2FS_I(inode)->i_xattr_nid; nid_t new_xnid; struct dnode_of_data dn; struct node_info ni; - struct folio *xfolio; + struct f2fs_cached_block *xentry; int err; if (!prev_xnid) @@ -2904,32 +2852,33 @@ recover_xnid: return -ENOSPC; set_new_dnode(&dn, inode, NULL, NULL, new_xnid); - xfolio = f2fs_new_node_folio(&dn, XATTR_NODE_OFFSET); - if (IS_ERR(xfolio)) { + xentry = f2fs_new_node_cache(&dn, XATTR_NODE_OFFSET); + if (IS_ERR(xentry)) { f2fs_alloc_nid_failed(sbi, new_xnid); - return PTR_ERR(xfolio); + return PTR_ERR(xentry); } f2fs_alloc_nid_done(sbi, new_xnid); - f2fs_update_inode_page(inode); + f2fs_update_inode_cache(inode); - /* 3: update and set xattr node page dirty */ - if (folio) { - memcpy(F2FS_NODE(xfolio), F2FS_NODE(folio), - VALID_XATTR_BLOCK_SIZE); - folio_mark_dirty(xfolio); + /* 3: update and set xattr node cache dirty */ + if (entry) { + memcpy(CACHED_NODE(xentry), CACHED_NODE(entry), + VALID_XATTR_BLOCK_SIZE(inode)); + f2fs_mark_cache_dirty(xentry); } - f2fs_folio_put(xfolio, true); + f2fs_put_cache(xentry, true); return 0; } -int f2fs_recover_inode_page(struct f2fs_sb_info *sbi, struct folio *folio) +int f2fs_recover_inode_cache(struct f2fs_sb_info *sbi, + struct f2fs_cached_block *entry) { struct f2fs_inode *src, *dst; - nid_t ino = ino_of_node(folio); + nid_t ino = ino_of_node(sbi, entry); struct node_info old_ni, new_ni; - struct folio *ifolio; + struct f2fs_cached_block *ientry; int err; err = f2fs_get_node_info(sbi, ino, &old_ni, false); @@ -2939,8 +2888,8 @@ int f2fs_recover_inode_page(struct f2fs_sb_info *sbi, struct folio *folio) if (unlikely(old_ni.blk_addr != NULL_ADDR)) return -EINVAL; retry: - ifolio = f2fs_grab_cache_folio(NODE_MAPPING(sbi), ino, false); - if (IS_ERR(ifolio)) { + ientry = f2fs_grab_node_cache(sbi, ino); + if (IS_ERR(ientry)) { memalloc_retry_wait(GFP_NOFS); goto retry; } @@ -2948,13 +2897,12 @@ retry: /* Should not use this inode from free nid list */ remove_free_nid(sbi, ino); - if (!folio_test_uptodate(ifolio)) - folio_mark_uptodate(ifolio); - fill_node_footer(ifolio, ino, ino, 0, true); - set_cold_node(ifolio, false); + f2fs_cache_set_uptodate(ientry); + fill_node_footer(sbi, ientry, ino, ino, 0, true); + set_cold_node(sbi, ientry, false); - src = F2FS_INODE(folio); - dst = F2FS_INODE(ifolio); + src = &CACHED_NODE(entry)->i; + dst = F2FS_INODE(ientry); memcpy(dst, src, offsetof(struct f2fs_inode, i_ext)); dst->i_size = 0; @@ -2990,46 +2938,46 @@ retry: WARN_ON(1); set_node_addr(sbi, &new_ni, NEW_ADDR, false); inc_valid_inode_count(sbi); - folio_mark_dirty(ifolio); - f2fs_folio_put(ifolio, true); + f2fs_mark_cache_dirty(ientry); + f2fs_put_cache(ientry, true); return 0; } int f2fs_restore_node_summary(struct f2fs_sb_info *sbi, unsigned int segno, struct f2fs_summary_block *sum) { - struct f2fs_node *rn; struct f2fs_summary *sum_entry; block_t addr; - int i, idx, last_offset, nrpages; + int i, idx, last_offset, nrblocks; /* scan the node segment */ last_offset = BLKS_PER_SEG(sbi); addr = START_BLOCK(sbi, segno); sum_entry = sum_entries(sum); - for (i = 0; i < last_offset; i += nrpages, addr += nrpages) { - nrpages = bio_max_segs(last_offset - i); + for (i = 0; i < last_offset; i += nrblocks, addr += nrblocks) { + nrblocks = bio_max_segs(last_offset - i); - /* readahead node pages */ - f2fs_ra_meta_pages(sbi, addr, nrpages, META_POR, true); + /* readahead node blocks */ + f2fs_ra_meta_caches(sbi, addr, nrblocks, META_POR, true); - for (idx = addr; idx < addr + nrpages; idx++) { - struct folio *folio = f2fs_get_tmp_folio(sbi, idx); + for (idx = addr; idx < addr + nrblocks; idx++) { + struct f2fs_cached_block *entry = + f2fs_get_tmp_cache(sbi, idx); - if (IS_ERR(folio)) - return PTR_ERR(folio); + if (IS_ERR(entry)) + return PTR_ERR(entry); - rn = F2FS_NODE(folio); - sum_entry->nid = rn->footer.nid; + struct node_footer *footer = (struct node_footer *)(cache_address(entry) + + F2FS_BLKSIZE(sbi) - sizeof(struct node_footer)); + sum_entry->nid = footer->nid; sum_entry->version = 0; sum_entry->ofs_in_node = 0; sum_entry++; - f2fs_folio_put(folio, true); + f2fs_put_cache(entry, true); } - invalidate_mapping_pages(META_MAPPING(sbi), addr, - addr + nrpages); + f2fs_truncate_meta_caches(sbi, addr, nrblocks); } return 0; } @@ -3074,7 +3022,7 @@ static void remove_nats_in_journal(struct f2fs_sb_info *sbi) spin_unlock(&nm_i->nid_list_lock); } - __set_nat_cache_dirty(nm_i, ne, init_dirty); + __set_nat_cache_dirty(sbi, nm_i, ne, init_dirty); } update_nats_in_cursum(journal, -i); up_write(&curseg->journal_rwsem); @@ -3102,7 +3050,7 @@ static void __update_nat_bits(struct f2fs_sb_info *sbi, nid_t start_nid, const struct f2fs_nat_block *nat_blk) { struct f2fs_nm_info *nm_i = NM_I(sbi); - unsigned int nat_index = start_nid / NAT_ENTRY_PER_BLOCK; + unsigned int nat_index = start_nid / NAT_ENTRY_PER_BLOCK(sbi); int valid = 0; int i = 0; @@ -3113,7 +3061,7 @@ static void __update_nat_bits(struct f2fs_sb_info *sbi, nid_t start_nid, valid = 1; i = 1; } - for (; i < NAT_ENTRY_PER_BLOCK; i++) { + for (; i < NAT_ENTRY_PER_BLOCK(sbi); i++) { if (le32_to_cpu(nat_blk->entries[i].block_addr) != NULL_ADDR) valid++; } @@ -3124,7 +3072,7 @@ static void __update_nat_bits(struct f2fs_sb_info *sbi, nid_t start_nid, } __clear_bit_le(nat_index, nm_i->empty_nat_bits); - if (valid == NAT_ENTRY_PER_BLOCK) + if (valid == NAT_ENTRY_PER_BLOCK(sbi)) __set_bit_le(nat_index, nm_i->full_nat_bits); else __clear_bit_le(nat_index, nm_i->full_nat_bits); @@ -3135,16 +3083,16 @@ static int __flush_nat_entry_set(struct f2fs_sb_info *sbi, { struct curseg_info *curseg = CURSEG_I(sbi, CURSEG_HOT_DATA); struct f2fs_journal *journal = curseg->journal; - nid_t start_nid = set->set * NAT_ENTRY_PER_BLOCK; + nid_t start_nid = set->set * NAT_ENTRY_PER_BLOCK(sbi); bool to_journal = true; struct f2fs_nat_block *nat_blk; struct nat_entry *ne, *cur; - struct folio *folio = NULL; + struct f2fs_cached_block *entry = NULL; /* * there are two steps to flush nat entries: * #1, flush nat entries to journal in current hot data summary block. - * #2, flush nat entries to nat page. + * #2, flush nat entries to nat block. */ if (enabled_nat_bits(sbi, cpc) || !__has_cursum_space(sbi, journal, set->entry_cnt, NAT_JOURNAL)) @@ -3153,11 +3101,11 @@ static int __flush_nat_entry_set(struct f2fs_sb_info *sbi, if (to_journal) { down_write(&curseg->journal_rwsem); } else { - folio = get_next_nat_folio(sbi, start_nid); - if (IS_ERR(folio)) - return PTR_ERR(folio); + entry = get_next_nat_cache(sbi, start_nid); + if (IS_ERR(entry)) + return PTR_ERR(entry); - nat_blk = folio_address(folio); + nat_blk = cache_address(entry); f2fs_bug_on(sbi, !nat_blk); } @@ -3194,7 +3142,7 @@ static int __flush_nat_entry_set(struct f2fs_sb_info *sbi, up_write(&curseg->journal_rwsem); } else { __update_nat_bits(sbi, start_nid, nat_blk); - f2fs_folio_put(folio, true); + f2fs_put_cache(entry, true); } /* Allow dirty nats by node block allocation in write_begin */ @@ -3266,7 +3214,7 @@ int f2fs_flush_nat_entries(struct f2fs_sb_info *sbi, struct cp_control *cpc) __has_cursum_space(sbi, journal, entry_count, NAT_JOURNAL)) continue; - f2fs_ra_meta_pages(sbi, set->set, 1, META_NAT, true); + f2fs_ra_meta_caches(sbi, set->set, 1, META_NAT, true); } /* flush dirty nats in nat entry set */ list_for_each_entry_safe(set, tmp, &sets, set_list) { @@ -3293,24 +3241,25 @@ static int __get_nat_bitmaps(struct f2fs_sb_info *sbi) if (!enabled_nat_bits(sbi, NULL)) return 0; - nm_i->nat_bits_blocks = F2FS_BLK_ALIGN((nat_bits_bytes << 1) + 8); + nm_i->nat_bits_blocks = F2FS_BLK_ALIGN(sbi, + (nat_bits_bytes << 1) + 8); nm_i->nat_bits = f2fs_kvzalloc(sbi, - F2FS_BLK_TO_BYTES(nm_i->nat_bits_blocks), GFP_KERNEL); + F2FS_BLK_TO_BYTES(sbi, nm_i->nat_bits_blocks), GFP_KERNEL); if (!nm_i->nat_bits) return -ENOMEM; nat_bits_addr = __start_cp_addr(sbi) + BLKS_PER_SEG(sbi) - nm_i->nat_bits_blocks; for (i = 0; i < nm_i->nat_bits_blocks; i++) { - struct folio *folio; + struct f2fs_cached_block *entry; - folio = f2fs_get_meta_folio(sbi, nat_bits_addr++); - if (IS_ERR(folio)) - return PTR_ERR(folio); + entry = f2fs_get_meta_cache(sbi, nat_bits_addr++); + if (IS_ERR(entry)) + return PTR_ERR(entry); - memcpy(nm_i->nat_bits + F2FS_BLK_TO_BYTES(i), - folio_address(folio), F2FS_BLKSIZE); - f2fs_folio_put(folio, true); + memcpy(nm_i->nat_bits + F2FS_BLK_TO_BYTES(sbi, i), + cache_address(entry), F2FS_BLKSIZE(sbi)); + f2fs_put_cache(entry, true); } cp_ver |= (cur_cp_crc(ckpt) << 32); @@ -3342,8 +3291,8 @@ static inline void load_free_nid_bitmap(struct f2fs_sb_info *sbi) __set_bit_le(i, nm_i->nat_block_bitmap); - nid = i * NAT_ENTRY_PER_BLOCK; - last_nid = nid + NAT_ENTRY_PER_BLOCK; + nid = i * NAT_ENTRY_PER_BLOCK(sbi); + last_nid = nid + NAT_ENTRY_PER_BLOCK(sbi); spin_lock(&NM_I(sbi)->nid_list_lock); for (; nid < last_nid; nid++) @@ -3373,7 +3322,7 @@ static int init_node_manager(struct f2fs_sb_info *sbi) /* segment_count_nat includes pair segment so divide to 2. */ nat_segs = le32_to_cpu(sb_raw->segment_count_nat) >> 1; nm_i->nat_blocks = nat_segs << le32_to_cpu(sb_raw->log_blocks_per_seg); - nm_i->max_nid = NAT_ENTRY_PER_BLOCK * nm_i->nat_blocks; + nm_i->max_nid = NAT_ENTRY_PER_BLOCK(sbi) * nm_i->nat_blocks; /* not used nids: 0, node, meta, (and root counted as valid node) */ nm_i->available_nids = nm_i->max_nid - sbi->total_valid_node_count - @@ -3381,7 +3330,7 @@ static int init_node_manager(struct f2fs_sb_info *sbi) nm_i->nid_cnt[FREE_NID] = 0; nm_i->nid_cnt[PREALLOC_NID] = 0; nm_i->ram_thresh = DEF_RAM_THRESHOLD; - nm_i->ra_nid_pages = DEF_RA_NID_PAGES; + nm_i->ra_nid_blocks = DEF_RA_NID_BLOCKS; nm_i->dirty_nats_ratio = DEF_DIRTY_NAT_RATIO_THRESHOLD; nm_i->max_rf_node_blocks = DEF_RF_NODE_BLOCKS; @@ -3436,7 +3385,7 @@ static int init_free_nid_cache(struct f2fs_sb_info *sbi) for (i = 0; i < nm_i->nat_blocks; i++) { nm_i->free_nid_bitmap[i] = f2fs_kvzalloc(sbi, - f2fs_bitmap_size(NAT_ENTRY_PER_BLOCK), GFP_KERNEL); + f2fs_bitmap_size(NAT_ENTRY_PER_BLOCK(sbi)), GFP_KERNEL); if (!nm_i->free_nid_bitmap[i]) return -ENOMEM; } diff --git a/fs/f2fs/node.h b/fs/f2fs/node.h index 5e114f352099..2704a5c6a54d 100644 --- a/fs/f2fs/node.h +++ b/fs/f2fs/node.h @@ -6,19 +6,26 @@ * http://www.samsung.com/ */ /* start node id of a node block dedicated to the given node id */ -#define START_NID(nid) (((nid) / NAT_ENTRY_PER_BLOCK) * NAT_ENTRY_PER_BLOCK) +static inline nid_t f2fs_start_nid(struct f2fs_sb_info *sbi, nid_t nid) +{ + unsigned int entries = NAT_ENTRY_PER_BLOCK(sbi); + + return (nid / entries) * entries; +} /* node block offset on the NAT area dedicated to the given start node id */ -#define NAT_BLOCK_OFFSET(start_nid) ((start_nid) / NAT_ENTRY_PER_BLOCK) +#define NAT_BLOCK_OFFSET(sbi, start_nid) \ + ((start_nid) / NAT_ENTRY_PER_BLOCK(sbi)) -/* # of pages to perform synchronous readahead before building free nids */ -#define FREE_NID_PAGES 8 -#define MAX_FREE_NIDS (NAT_ENTRY_PER_BLOCK * FREE_NID_PAGES) +/* # of blocks to perform synchronous readahead before building free nids */ +#define FREE_NID_BLOCKS 8 +#define MAX_FREE_NIDS(sbi) ((unsigned long)NAT_ENTRY_PER_BLOCK(sbi) * \ + FREE_NID_BLOCKS) /* size of free nid batch when shrinking */ #define SHRINK_NID_BATCH_SIZE 8 -#define DEF_RA_NID_PAGES 0 /* # of nid pages to be readaheaded */ +#define DEF_RA_NID_BLOCKS 0 /* # of nid blocks to be readaheaded */ /* maximum readahead size for node during getting data blocks */ #define MAX_RA_NODE 128 @@ -37,8 +44,8 @@ /* vector size for gang look-up from nat cache that consists of radix tree */ #define NAT_VEC_SIZE 32 -/* return value for read_node_page */ -#define LOCKED_PAGE 1 +/* return value for read_node_cache */ +#define LOCKED_CACHE 1 /* check pinned file's alignment status of physical blocks */ #define FILE_NOT_ALIGNED 1 @@ -150,7 +157,7 @@ enum mem_type { READ_EXTENT_CACHE, /* indicates read extent cache */ AGE_EXTENT_CACHE, /* indicates age extent cache */ DISCARD_CACHE, /* indicates memory of cached discard cmds */ - COMPRESS_PAGE, /* indicates memory of cached compressed pages */ + COMPRESS_BLOCK, /* indicates memory of cached compressed blocks */ BASE_CHECK, /* check kernel status */ }; @@ -208,7 +215,7 @@ static inline pgoff_t current_nat_addr(struct f2fs_sb_info *sbi, nid_t start) * OLD = (segment_off * 512) * 2 + off_in_segment * NEW = 2 * (segment_off * 512 + off_in_segment) - off_in_segment */ - block_off = NAT_BLOCK_OFFSET(start); + block_off = NAT_BLOCK_OFFSET(sbi, start); block_addr = (pgoff_t)(nm_i->nat_blkaddr + (block_off << 1) - @@ -230,9 +237,10 @@ static inline pgoff_t next_nat_addr(struct f2fs_sb_info *sbi, return block_addr + nm_i->nat_blkaddr; } -static inline void set_to_next_nat(struct f2fs_nm_info *nm_i, nid_t start_nid) +static inline void set_to_next_nat(struct f2fs_sb_info *sbi, + struct f2fs_nm_info *nm_i, nid_t start_nid) { - unsigned int block_off = NAT_BLOCK_OFFSET(start_nid); + unsigned int block_off = NAT_BLOCK_OFFSET(sbi, start_nid); f2fs_change_bit(block_off, nm_i->nat_bitmap); #ifdef CONFIG_F2FS_CHECK_FS @@ -240,90 +248,95 @@ static inline void set_to_next_nat(struct f2fs_nm_info *nm_i, nid_t start_nid) #endif } -static inline nid_t ino_of_node(const struct folio *node_folio) +static inline nid_t ino_of_node(struct f2fs_sb_info *sbi, + const struct f2fs_cached_block *entry) { - struct f2fs_node *rn = F2FS_NODE(node_folio); - return le32_to_cpu(rn->footer.ino); + return le32_to_cpu(F2FS_NODE_FOOTER(sbi, entry)->ino); } -static inline nid_t nid_of_node(const struct folio *node_folio) +static inline nid_t nid_of_node(struct f2fs_sb_info *sbi, + const struct f2fs_cached_block *entry) { - struct f2fs_node *rn = F2FS_NODE(node_folio); - return le32_to_cpu(rn->footer.nid); + return le32_to_cpu(F2FS_NODE_FOOTER(sbi, entry)->nid); } -static inline unsigned int ofs_of_node(const struct folio *node_folio) +static inline unsigned int ofs_of_node(struct f2fs_sb_info *sbi, + const struct f2fs_cached_block *entry) { - struct f2fs_node *rn = F2FS_NODE(node_folio); - unsigned flag = le32_to_cpu(rn->footer.flag); + unsigned int flag = le32_to_cpu(F2FS_NODE_FOOTER(sbi, entry)->flag); return flag >> OFFSET_BIT_SHIFT; } -static inline __u64 cpver_of_node(const struct folio *node_folio) +static inline __u64 cpver_of_node(struct f2fs_sb_info *sbi, + const struct f2fs_cached_block *entry) { - struct f2fs_node *rn = F2FS_NODE(node_folio); - return le64_to_cpu(rn->footer.cp_ver); + return le64_to_cpu(F2FS_NODE_FOOTER(sbi, entry)->cp_ver); } -static inline block_t next_blkaddr_of_node(const struct folio *node_folio) +static inline block_t next_blkaddr_of_node(struct f2fs_sb_info *sbi, + const struct f2fs_cached_block *entry) { - struct f2fs_node *rn = F2FS_NODE(node_folio); - return le32_to_cpu(rn->footer.next_blkaddr); + return le32_to_cpu(F2FS_NODE_FOOTER(sbi, entry)->next_blkaddr); } -static inline void fill_node_footer(const struct folio *folio, nid_t nid, +static inline void fill_node_footer(struct f2fs_sb_info *sbi, + struct f2fs_cached_block *entry, nid_t nid, nid_t ino, unsigned int ofs, bool reset) { - struct f2fs_node *rn = F2FS_NODE(folio); + struct f2fs_node *rn = CACHED_NODE(entry); + struct node_footer *footer = F2FS_NODE_FOOTER(sbi, entry); unsigned int old_flag = 0; if (reset) - memset(rn, 0, sizeof(*rn)); + memset(rn, 0, F2FS_BLKSIZE(sbi)); else - old_flag = le32_to_cpu(rn->footer.flag); + old_flag = le32_to_cpu(footer->flag); - rn->footer.nid = cpu_to_le32(nid); - rn->footer.ino = cpu_to_le32(ino); + memset(footer, 0, sizeof(*footer)); + footer->nid = cpu_to_le32(nid); + footer->ino = cpu_to_le32(ino); /* should remain old flag bits such as COLD_BIT_SHIFT */ - rn->footer.flag = cpu_to_le32((ofs << OFFSET_BIT_SHIFT) | + footer->flag = cpu_to_le32((ofs << OFFSET_BIT_SHIFT) | (old_flag & OFFSET_BIT_MASK)); } -static inline void copy_node_footer(const struct folio *dst, - const struct folio *src) +static inline void copy_node_footer(struct f2fs_sb_info *sbi, + struct f2fs_cached_block *dst, + const struct f2fs_cached_block *src) { - struct f2fs_node *src_rn = F2FS_NODE(src); - struct f2fs_node *dst_rn = F2FS_NODE(dst); - memcpy(&dst_rn->footer, &src_rn->footer, sizeof(struct node_footer)); + memcpy(F2FS_NODE_FOOTER(sbi, dst), F2FS_NODE_FOOTER(sbi, src), + sizeof(struct node_footer)); } -static inline void fill_node_footer_blkaddr(struct folio *folio, block_t blkaddr) +static inline void fill_node_footer_blkaddr(struct f2fs_sb_info *sbi, + struct f2fs_cached_block *entry, block_t blkaddr) { - struct f2fs_checkpoint *ckpt = F2FS_CKPT(F2FS_F_SB(folio)); - struct f2fs_node *rn = F2FS_NODE(folio); + struct f2fs_checkpoint *ckpt = F2FS_CKPT(sbi); + struct node_footer *footer = F2FS_NODE_FOOTER(sbi, entry); __u64 cp_ver = cur_cp_version(ckpt); if (__is_set_ckpt_flags(ckpt, CP_CRC_RECOVERY_FLAG)) cp_ver |= (cur_cp_crc(ckpt) << 32); - rn->footer.cp_ver = cpu_to_le64(cp_ver); - rn->footer.next_blkaddr = cpu_to_le32(blkaddr); + footer->cp_ver = cpu_to_le64(cp_ver); + footer->next_blkaddr = cpu_to_le32(blkaddr); } -static inline bool is_recoverable_dnode(const struct folio *folio) +static inline bool is_recoverable_dnode(struct f2fs_sb_info *sbi, + const struct f2fs_cached_block *entry) { - struct f2fs_checkpoint *ckpt = F2FS_CKPT(F2FS_F_SB(folio)); + struct f2fs_checkpoint *ckpt = F2FS_CKPT(sbi); __u64 cp_ver = cur_cp_version(ckpt); /* Don't care crc part, if fsck.f2fs sets it. */ if (__is_set_ckpt_flags(ckpt, CP_NOCRC_RECOVERY_FLAG)) - return (cp_ver << 32) == (cpver_of_node(folio) << 32); + return (cp_ver << 32) == (cpver_of_node(sbi, entry) << 32); if (__is_set_ckpt_flags(ckpt, CP_CRC_RECOVERY_FLAG)) cp_ver |= (cur_cp_crc(ckpt) << 32); - return cp_ver == cpver_of_node(folio); + return cp_ver == cpver_of_node(sbi, entry); } /* @@ -347,43 +360,50 @@ static inline bool is_recoverable_dnode(const struct folio *folio) * `- indirect node ((6 + 2N) + (N - 1)(N + 1)) * `- direct node */ -static inline bool IS_DNODE(const struct folio *node_folio) +static inline bool IS_DNODE(struct f2fs_sb_info *sbi, + const struct f2fs_cached_block *entry) { - unsigned int ofs = ofs_of_node(node_folio); + unsigned int ofs = ofs_of_node(sbi, entry); if (f2fs_has_xattr_block(ofs)) return true; - if (ofs == 3 || ofs == 4 + NIDS_PER_BLOCK || - ofs == 5 + 2 * NIDS_PER_BLOCK) + if (ofs == 3 || ofs == 4 + NIDS_PER_BLOCK(sbi) || + ofs == 5 + 2 * NIDS_PER_BLOCK(sbi)) return false; - if (ofs >= 6 + 2 * NIDS_PER_BLOCK) { - ofs -= 6 + 2 * NIDS_PER_BLOCK; - if (!((long int)ofs % (NIDS_PER_BLOCK + 1))) + if (ofs >= 6 + 2 * NIDS_PER_BLOCK(sbi)) { + ofs -= 6 + 2 * NIDS_PER_BLOCK(sbi); + if (!((long)ofs % (NIDS_PER_BLOCK(sbi) + 1))) return false; } return true; } -static inline int set_nid(struct folio *folio, int off, nid_t nid, bool i) +static inline bool set_nid(struct f2fs_sb_info *sbi, + struct f2fs_cached_block *entry, int off, nid_t nid, bool i) { - struct f2fs_node *rn = F2FS_NODE(folio); + struct f2fs_node *rn = CACHED_NODE(entry); + __le32 *inode_nids = F2FS_INODE_NIDS(sbi, entry); + __le32 *addr = i ? &inode_nids[off - NODE_DIR1_BLOCK(sbi)] : &rn->in.nid[off]; - f2fs_folio_wait_writeback(folio, NODE, true, true); + f2fs_cache_wait_writeback(entry); + if (*addr == cpu_to_le32(nid)) + return false; - if (i) - rn->i.i_nid[off - NODE_DIR1_BLOCK] = cpu_to_le32(nid); - else - rn->in.nid[off] = cpu_to_le32(nid); - return folio_mark_dirty(folio); + *addr = cpu_to_le32(nid); + f2fs_mark_cache_dirty(entry); + return true; } -static inline nid_t get_nid(const struct folio *folio, int off, bool i) +static inline nid_t get_nid(struct f2fs_sb_info *sbi, + const struct f2fs_cached_block *entry, int off, bool i) { - struct f2fs_node *rn = F2FS_NODE(folio); + struct f2fs_node *rn = CACHED_NODE(entry); + const __le32 *inode_nids = F2FS_INODE_NIDS(sbi, entry); + int nid_index = off - NODE_DIR1_BLOCK(sbi); if (i) - return le32_to_cpu(rn->i.i_nid[off - NODE_DIR1_BLOCK]); + return le32_to_cpu(inode_nids[nid_index]); return le32_to_cpu(rn->in.nid[off]); } @@ -394,40 +414,43 @@ static inline nid_t get_nid(const struct folio *folio, int off, bool i) * - Mark cold data pages in page cache */ -static inline int is_node(const struct folio *folio, int type) +static inline int is_node(struct f2fs_sb_info *sbi, + const struct f2fs_cached_block *entry, int type) { - struct f2fs_node *rn = F2FS_NODE(folio); - return le32_to_cpu(rn->footer.flag) & BIT(type); + return le32_to_cpu(F2FS_NODE_FOOTER(sbi, entry)->flag) & BIT(type); } -#define is_cold_node(folio) is_node(folio, COLD_BIT_SHIFT) -#define is_fsync_dnode(folio) is_node(folio, FSYNC_BIT_SHIFT) -#define is_dent_dnode(folio) is_node(folio, DENT_BIT_SHIFT) +#define is_cold_node(sbi, entry) is_node(sbi, entry, COLD_BIT_SHIFT) +#define is_fsync_dnode(sbi, entry) is_node(sbi, entry, FSYNC_BIT_SHIFT) +#define is_dent_dnode(sbi, entry) is_node(sbi, entry, DENT_BIT_SHIFT) -static inline void __set_mark(const struct folio *folio, bool mark, int type) +static inline void __set_mark(struct f2fs_sb_info *sbi, + struct f2fs_cached_block *entry, bool mark, int type) { - struct f2fs_node *rn = F2FS_NODE(folio); - unsigned int flag = le32_to_cpu(rn->footer.flag); + struct node_footer *footer = F2FS_NODE_FOOTER(sbi, entry); + unsigned int flag = le32_to_cpu(footer->flag); if (mark) flag |= BIT(type); else flag &= ~BIT(type); - rn->footer.flag = cpu_to_le32(flag); + footer->flag = cpu_to_le32(flag); } -static inline void set_cold_node(const struct folio *folio, bool is_dir) +static inline void set_cold_node(struct f2fs_sb_info *sbi, + struct f2fs_cached_block *entry, bool is_dir) { - __set_mark(folio, !is_dir, COLD_BIT_SHIFT); + __set_mark(sbi, entry, !is_dir, COLD_BIT_SHIFT); } -static inline void set_mark(struct folio *folio, bool mark, int type) +static inline void set_mark(struct f2fs_sb_info *sbi, + struct f2fs_cached_block *entry, bool mark, int type) { - __set_mark(folio, mark, type); - + __set_mark(sbi, entry, mark, type); #ifdef CONFIG_F2FS_CHECK_FS - f2fs_inode_chksum_set(F2FS_F_SB(folio), folio); + f2fs_inode_chksum_set(sbi, entry); #endif } -#define set_dentry_mark(folio, mark) set_mark(folio, mark, DENT_BIT_SHIFT) -#define set_fsync_mark(folio, mark) set_mark(folio, mark, FSYNC_BIT_SHIFT) + +#define set_dentry_mark(sbi, entry, mark) set_mark(sbi, entry, mark, DENT_BIT_SHIFT) +#define set_fsync_mark(sbi, entry, mark) set_mark(sbi, entry, mark, FSYNC_BIT_SHIFT) diff --git a/fs/f2fs/recovery.c b/fs/f2fs/recovery.c index aaa5227739c8..54b5246609de 100644 --- a/fs/f2fs/recovery.c +++ b/fs/f2fs/recovery.c @@ -182,38 +182,39 @@ static const char *recover_printable_name(struct inode *inode, return raw->i_name; } -static int recover_dentry(struct inode *inode, struct folio *ifolio, +static int recover_dentry(struct inode *inode, struct f2fs_cached_block *entry, struct list_head *dir_list) { - struct f2fs_inode *raw_inode = F2FS_INODE(ifolio); + struct f2fs_inode *raw_inode = &CACHED_NODE(entry)->i; nid_t pino = le32_to_cpu(raw_inode->i_pino); struct f2fs_dir_entry *de; struct f2fs_filename fname; struct qstr usr_fname; - struct folio *folio; + void *dentry_blk = NULL; struct inode *dir, *einode; - struct fsync_inode_entry *entry; + struct fsync_inode_entry *fsync_entry; int err = 0; const char *name; int name_len; - entry = get_fsync_inode(dir_list, pino); - if (!entry) { - entry = add_fsync_inode(F2FS_I_SB(inode), dir_list, + fsync_entry = get_fsync_inode(dir_list, pino); + if (!fsync_entry) { + fsync_entry = add_fsync_inode(F2FS_I_SB(inode), dir_list, pino, false); - if (IS_ERR(entry)) { - dir = ERR_CAST(entry); - err = PTR_ERR(entry); + if (IS_ERR(fsync_entry)) { + dir = ERR_CAST(fsync_entry); + err = PTR_ERR(fsync_entry); goto out; } } - dir = entry->inode; + dir = fsync_entry->inode; err = init_recovered_filename(dir, inode, raw_inode, &fname, &usr_fname); if (err) goto out; retry: - de = __f2fs_find_entry(dir, &fname, &folio); + dentry_blk = NULL; + de = __f2fs_find_entry(dir, &fname, &dentry_blk); if (de && inode->i_ino == le32_to_cpu(de->ino)) goto out_put; @@ -238,11 +239,11 @@ retry: iput(einode); goto out_put; } - f2fs_delete_entry(de, folio, dir, einode); + f2fs_delete_entry(de, dentry_blk, dir, einode); iput(einode); goto retry; - } else if (IS_ERR(folio)) { - err = PTR_ERR(folio); + } else if (IS_ERR(dentry_blk)) { + err = PTR_ERR(dentry_blk); } else { err = f2fs_add_dentry(dir, &fname, inode, inode->i_ino, inode->i_mode); @@ -252,18 +253,18 @@ retry: goto out; out_put: - f2fs_folio_put(folio, false); + f2fs_put_dentry_block(dentry_blk, false); out: name = recover_printable_name(inode, raw_inode, &name_len); f2fs_notice(F2FS_I_SB(inode), "%s: ino = %x, name = %.*s, dir = %llu, err = %d", - __func__, ino_of_node(ifolio), name_len, name, + __func__, ino_of_node(F2FS_I_SB(inode), entry), name_len, name, IS_ERR(dir) ? 0 : dir->i_ino, err); return err; } -static int recover_quota_data(struct inode *inode, struct folio *folio) +static int recover_quota_data(struct inode *inode, struct f2fs_cached_block *entry) { - struct f2fs_inode *raw = F2FS_INODE(folio); + struct f2fs_inode *raw = &CACHED_NODE(entry)->i; struct iattr attr; uid_t i_uid = le32_to_cpu(raw->i_uid); gid_t i_gid = le32_to_cpu(raw->i_gid); @@ -300,9 +301,9 @@ static void recover_inline_flags(struct inode *inode, struct f2fs_inode *ri) clear_inode_flag(inode, FI_DATA_EXIST); } -static int recover_inode(struct inode *inode, struct folio *folio) +static int recover_inode(struct inode *inode, struct f2fs_cached_block *entry) { - struct f2fs_inode *raw = F2FS_INODE(folio); + struct f2fs_inode *raw = &CACHED_NODE(entry)->i; struct f2fs_inode_info *fi = F2FS_I(inode); const char *name; int name_len; @@ -310,7 +311,7 @@ static int recover_inode(struct inode *inode, struct folio *folio) inode->i_mode = le16_to_cpu(raw->i_mode); - err = recover_quota_data(inode, folio); + err = recover_quota_data(inode, entry); if (err) return err; @@ -357,7 +358,7 @@ static int recover_inode(struct inode *inode, struct folio *folio) name = recover_printable_name(inode, raw, &name_len); f2fs_notice(F2FS_I_SB(inode), "%s: ino = %x, name = %.*s, inline = %x", - __func__, ino_of_node(folio), name_len, name, + __func__, ino_of_node(F2FS_I_SB(inode), entry), name_len, name, raw->i_inline); return 0; } @@ -386,30 +387,30 @@ static int sanity_check_node_chain(struct f2fs_sb_info *sbi, block_t blkaddr, return 0; for (i = 0; i < 2; i++) { - struct folio *folio; + struct f2fs_cached_block *entry; if (!f2fs_is_valid_blkaddr(sbi, *blkaddr_fast, META_POR)) { *is_detecting = false; return 0; } - folio = f2fs_get_tmp_folio(sbi, *blkaddr_fast); - if (IS_ERR(folio)) - return PTR_ERR(folio); + entry = f2fs_get_tmp_cache(sbi, *blkaddr_fast); + if (IS_ERR(entry)) + return PTR_ERR(entry); - if (!is_recoverable_dnode(folio)) { - f2fs_folio_put(folio, true); + if (!is_recoverable_dnode(sbi, entry)) { + f2fs_put_cache(entry, true); *is_detecting = false; return 0; } ra_blocks = adjust_por_ra_blocks(sbi, ra_blocks, *blkaddr_fast, - next_blkaddr_of_node(folio)); + next_blkaddr_of_node(sbi, entry)); - *blkaddr_fast = next_blkaddr_of_node(folio); - f2fs_folio_put(folio, true); + *blkaddr_fast = next_blkaddr_of_node(sbi, entry); + f2fs_put_cache(entry, true); - f2fs_ra_meta_pages_cond(sbi, *blkaddr_fast, ra_blocks); + f2fs_ra_meta_caches_cond(sbi, *blkaddr_fast, ra_blocks); } if (*blkaddr_fast == blkaddr) { @@ -434,45 +435,45 @@ static int find_fsync_dnodes(struct f2fs_sb_info *sbi, struct list_head *head, blkaddr_fast = blkaddr; while (1) { - struct fsync_inode_entry *entry; - struct folio *folio; + struct fsync_inode_entry *fsync_entry; + struct f2fs_cached_block *entry; if (!f2fs_is_valid_blkaddr(sbi, blkaddr, META_POR)) return 0; - folio = f2fs_get_tmp_folio(sbi, blkaddr); - if (IS_ERR(folio)) { - err = PTR_ERR(folio); + entry = f2fs_get_tmp_cache(sbi, blkaddr); + if (IS_ERR(entry)) { + err = PTR_ERR(entry); break; } - if (!is_recoverable_dnode(folio)) { - f2fs_folio_put(folio, true); + if (!is_recoverable_dnode(sbi, entry)) { + f2fs_put_cache(entry, true); break; } - if (!is_fsync_dnode(folio)) + if (!is_fsync_dnode(sbi, entry)) goto next; - entry = get_fsync_inode(head, ino_of_node(folio)); - if (!entry) { + fsync_entry = get_fsync_inode(head, ino_of_node(sbi, entry)); + if (!fsync_entry) { bool quota_inode = false; if (!check_only && - IS_INODE(folio) && - is_dent_dnode(folio)) { - err = f2fs_recover_inode_page(sbi, folio); + IS_INODE(sbi, entry) && + is_dent_dnode(sbi, entry)) { + err = f2fs_recover_inode_cache(sbi, entry); if (err) { - f2fs_folio_put(folio, true); + f2fs_put_cache(entry, true); break; } quota_inode = true; } - entry = add_fsync_inode(sbi, head, ino_of_node(folio), + fsync_entry = add_fsync_inode(sbi, head, ino_of_node(sbi, entry), quota_inode); - if (IS_ERR(entry)) { - err = PTR_ERR(entry); + if (IS_ERR(fsync_entry)) { + err = PTR_ERR(fsync_entry); /* * CP | dnode(F) | inode(DF) * For this case, we should not give up now. @@ -482,18 +483,18 @@ static int find_fsync_dnodes(struct f2fs_sb_info *sbi, struct list_head *head, *new_inode = true; goto next; } - f2fs_folio_put(folio, true); + f2fs_put_cache(entry, true); break; } } - entry->blkaddr = blkaddr; + fsync_entry->blkaddr = blkaddr; - if (IS_INODE(folio) && is_dent_dnode(folio)) - entry->last_dentry = blkaddr; + if (IS_INODE(sbi, entry) && is_dent_dnode(sbi, entry)) + fsync_entry->last_dentry = blkaddr; next: /* check next segment */ - blkaddr = next_blkaddr_of_node(folio); - f2fs_folio_put(folio, true); + blkaddr = next_blkaddr_of_node(sbi, entry); + f2fs_put_cache(entry, true); err = sanity_check_node_chain(sbi, blkaddr, &blkaddr_fast, &is_detecting); @@ -519,7 +520,8 @@ static int check_index_in_prev_nodes(struct f2fs_sb_info *sbi, unsigned short blkoff = GET_BLKOFF_FROM_SEG0(sbi, blkaddr); struct f2fs_summary_block *sum_node; struct f2fs_summary sum; - struct folio *sum_folio, *node_folio; + struct f2fs_cached_block *entry = NULL; + struct f2fs_cached_block *node_entry; struct dnode_of_data tdn = *dn; nid_t ino, nid; struct inode *inode; @@ -541,18 +543,18 @@ static int check_index_in_prev_nodes(struct f2fs_sb_info *sbi, } } - sum_folio = f2fs_get_sum_folio(sbi, segno); - if (IS_ERR(sum_folio)) - return PTR_ERR(sum_folio); - sum_node = SUM_BLK_PAGE_ADDR(sbi, sum_folio, segno); + entry = f2fs_get_sum_cache(sbi, segno); + if (IS_ERR(entry)) + return PTR_ERR(entry); + sum_node = SUM_BLK_ENTRY_ADDR(sbi, entry, segno); sum = sum_entries(sum_node)[blkoff]; - f2fs_folio_put(sum_folio, true); + f2fs_put_cache(entry, true); got_it: /* Use the locked dnode page and inode */ nid = le32_to_cpu(sum.nid); ofs_in_node = le16_to_cpu(sum.ofs_in_node); - max_addrs = ADDRS_PER_PAGE(dn->node_folio, dn->inode); + max_addrs = ADDRS_PER_PAGE(dn->node_entry, dn->inode); if (ofs_in_node >= max_addrs) { f2fs_err(sbi, "Inconsistent ofs_in_node:%u in summary, ino:%llu, nid:%u, max:%u", ofs_in_node, dn->inode->i_ino, nid, max_addrs); @@ -562,9 +564,9 @@ got_it: if (dn->inode->i_ino == nid) { tdn.nid = nid; - if (!dn->inode_folio_locked) - folio_lock(dn->inode_folio); - tdn.node_folio = dn->inode_folio; + if (!dn->inode_entry_locked) + f2fs_lock_cache(dn->inode_entry); + tdn.node_entry = dn->inode_entry; tdn.ofs_in_node = ofs_in_node; goto truncate_out; } else if (dn->nid == nid) { @@ -573,13 +575,13 @@ got_it: } /* Get the node page */ - node_folio = f2fs_get_node_folio(sbi, nid, NODE_TYPE_REGULAR); - if (IS_ERR(node_folio)) - return PTR_ERR(node_folio); + node_entry = f2fs_get_node_cache(sbi, nid, NODE_TYPE_REGULAR); + if (IS_ERR(node_entry)) + return PTR_ERR(node_entry); - offset = ofs_of_node(node_folio); - ino = ino_of_node(node_folio); - f2fs_folio_put(node_folio, true); + offset = ofs_of_node(sbi, node_entry); + ino = ino_of_node(sbi, node_entry); + f2fs_put_cache(node_entry, true); if (ino != dn->inode->i_ino) { int ret; @@ -605,8 +607,8 @@ got_it: * if inode page is locked, unlock temporarily, but its reference * count keeps alive. */ - if (ino == dn->inode->i_ino && dn->inode_folio_locked) - folio_unlock(dn->inode_folio); + if (ino == dn->inode->i_ino && dn->inode_entry_locked) + f2fs_unlock_cache(dn->inode_entry); set_new_dnode(&tdn, inode, NULL, NULL, 0); if (f2fs_get_dnode_of_data(&tdn, bidx, LOOKUP_NODE)) @@ -619,15 +621,15 @@ got_it: out: if (ino != dn->inode->i_ino) iput(inode); - else if (dn->inode_folio_locked) - folio_lock(dn->inode_folio); + else if (dn->inode_entry_locked) + f2fs_lock_cache(dn->inode_entry); return 0; truncate_out: if (f2fs_data_blkaddr(&tdn) == blkaddr) f2fs_truncate_data_blocks_range(&tdn, 1); - if (dn->inode->i_ino == nid && !dn->inode_folio_locked) - folio_unlock(dn->inode_folio); + if (dn->inode->i_ino == nid && !dn->inode_entry_locked) + f2fs_unlock_cache(dn->inode_entry); return 0; } @@ -645,7 +647,7 @@ static int f2fs_reserve_new_block_retry(struct dnode_of_data *dn) } static int do_recover_data(struct f2fs_sb_info *sbi, struct inode *inode, - struct folio *folio) + struct f2fs_cached_block *entry) { struct dnode_of_data dn; struct node_info ni; @@ -653,19 +655,19 @@ static int do_recover_data(struct f2fs_sb_info *sbi, struct inode *inode, int err = 0, recovered = 0; /* step 1: recover xattr */ - if (IS_INODE(folio)) { - err = f2fs_recover_inline_xattr(inode, folio); + if (IS_INODE(sbi, entry)) { + err = f2fs_recover_inline_xattr(inode, entry); if (err) goto out; - } else if (f2fs_has_xattr_block(ofs_of_node(folio))) { - err = f2fs_recover_xattr_data(inode, folio); + } else if (f2fs_has_xattr_block(ofs_of_node(sbi, entry))) { + err = f2fs_recover_xattr_data(inode, entry); if (!err) recovered++; goto out; } /* step 2: recover inline data */ - err = f2fs_recover_inline_data(inode, folio); + err = f2fs_recover_inline_data(inode, entry); if (err) { if (err == 1) err = 0; @@ -673,8 +675,8 @@ static int do_recover_data(struct f2fs_sb_info *sbi, struct inode *inode, } /* step 3: recover data indices */ - start = f2fs_start_bidx_of_node(ofs_of_node(folio), inode); - end = start + ADDRS_PER_PAGE(folio, inode); + start = f2fs_start_bidx_of_node(ofs_of_node(sbi, entry), inode); + end = start + addrs_per_page(inode, IS_INODE(sbi, entry)); set_new_dnode(&dn, inode, NULL, NULL, 0); retry_dn: @@ -687,18 +689,18 @@ retry_dn: goto out; } - f2fs_folio_wait_writeback(dn.node_folio, NODE, true, true); + f2fs_cache_wait_writeback(dn.node_entry); err = f2fs_get_node_info(sbi, dn.nid, &ni, false); if (err) goto err; - f2fs_bug_on(sbi, ni.ino != ino_of_node(folio)); + f2fs_bug_on(sbi, ni.ino != ino_of_node(sbi, entry)); - if (ofs_of_node(dn.node_folio) != ofs_of_node(folio)) { + if (ofs_of_node(sbi, dn.node_entry) != ofs_of_node(sbi, entry)) { f2fs_warn(sbi, "Inconsistent ofs_of_node, ino:%llu, ofs:%u, %u", - inode->i_ino, ofs_of_node(dn.node_folio), - ofs_of_node(folio)); + inode->i_ino, ofs_of_node(sbi, dn.node_entry), + ofs_of_node(sbi, entry)); err = -EFSCORRUPTED; f2fs_handle_error(sbi, ERROR_INCONSISTENT_FOOTER); fserror_report_file_metadata(dn.inode, err, GFP_NOFS); @@ -709,7 +711,7 @@ retry_dn: block_t src, dest; src = f2fs_data_blkaddr(&dn); - dest = data_blkaddr(dn.inode, folio, dn.ofs_in_node); + dest = data_blkaddr(dn.inode, entry, dn.ofs_in_node); if (__is_valid_data_blkaddr(src) && !f2fs_is_valid_blkaddr(sbi, src, META_POR)) { @@ -734,9 +736,9 @@ retry_dn: } if (!file_keep_isize(inode) && - (i_size_read(inode) <= ((loff_t)index << PAGE_SHIFT))) + (i_size_read(inode) <= F2FS_BLK_TO_BYTES(sbi, index))) f2fs_i_size_write(inode, - (loff_t)(index + 1) << PAGE_SHIFT); + F2FS_BLK_TO_BYTES(sbi, index + 1)); /* * dest is reserved block, invalidate src block @@ -784,16 +786,16 @@ retry_prev: } } - copy_node_footer(dn.node_folio, folio); - fill_node_footer(dn.node_folio, dn.nid, ni.ino, - ofs_of_node(folio), false); - folio_mark_dirty(dn.node_folio); + copy_node_footer(sbi, dn.node_entry, entry); + fill_node_footer(sbi, dn.node_entry, dn.nid, ni.ino, + ofs_of_node(sbi, entry), false); + f2fs_mark_cache_dirty(dn.node_entry); err: f2fs_put_dnode(&dn); out: f2fs_notice(sbi, "recover_data: ino = %llx, nid = %x (i_size: %s), " "range (%u, %u), recovered = %d, err = %d", - inode->i_ino, nid_of_node(folio), + inode->i_ino, nid_of_node(sbi, entry), file_keep_isize(inode) ? "keep" : "recover", start, end, recovered, err); return err; @@ -820,26 +822,26 @@ static int recover_data(struct f2fs_sb_info *sbi, struct list_head *inode_list, blkaddr = NEXT_FREE_BLKADDR(sbi, curseg); while (1) { - struct fsync_inode_entry *entry; - struct folio *folio; + struct fsync_inode_entry *fsync_entry; + struct f2fs_cached_block *entry; if (!f2fs_is_valid_blkaddr(sbi, blkaddr, META_POR)) break; - folio = f2fs_get_tmp_folio(sbi, blkaddr); - if (IS_ERR(folio)) { - err = PTR_ERR(folio); + entry = f2fs_get_tmp_cache(sbi, blkaddr); + if (IS_ERR(entry)) { + err = PTR_ERR(entry); break; } - if (!is_recoverable_dnode(folio)) { - f2fs_folio_put(folio, true); + if (!is_recoverable_dnode(sbi, entry)) { + f2fs_put_cache(entry, true); break; } recoverable_dnode++; - entry = get_fsync_inode(inode_list, ino_of_node(folio)); - if (!entry) + fsync_entry = get_fsync_inode(inode_list, ino_of_node(sbi, entry)); + if (!fsync_entry) goto next; fsynced_dnode++; /* @@ -847,40 +849,40 @@ static int recover_data(struct f2fs_sb_info *sbi, struct list_head *inode_list, * In this case, we can lose the latest inode(x). * So, call recover_inode for the inode update. */ - if (IS_INODE(folio)) { - err = recover_inode(entry->inode, folio); + if (IS_INODE(sbi, entry)) { + err = recover_inode(fsync_entry->inode, entry); if (err) { - f2fs_folio_put(folio, true); + f2fs_put_cache(entry, true); break; } recovered_inode++; } - if (entry->last_dentry == blkaddr) { - err = recover_dentry(entry->inode, folio, dir_list); + if (fsync_entry->last_dentry == blkaddr) { + err = recover_dentry(fsync_entry->inode, entry, dir_list); if (err) { - f2fs_folio_put(folio, true); + f2fs_put_cache(entry, true); break; } recovered_dentry++; } - err = do_recover_data(sbi, entry->inode, folio); + err = do_recover_data(sbi, fsync_entry->inode, entry); if (err) { - f2fs_folio_put(folio, true); + f2fs_put_cache(entry, true); break; } recovered_dnode++; - if (entry->blkaddr == blkaddr) - list_move_tail(&entry->list, tmp_inode_list); + if (fsync_entry->blkaddr == blkaddr) + list_move_tail(&fsync_entry->list, tmp_inode_list); next: ra_blocks = adjust_por_ra_blocks(sbi, ra_blocks, blkaddr, - next_blkaddr_of_node(folio)); + next_blkaddr_of_node(sbi, entry)); /* check next segment */ - blkaddr = next_blkaddr_of_node(folio); - f2fs_folio_put(folio, true); + blkaddr = next_blkaddr_of_node(sbi, entry); + f2fs_put_cache(entry, true); - f2fs_ra_meta_pages_cond(sbi, blkaddr, ra_blocks); + f2fs_ra_meta_caches_cond(sbi, blkaddr, ra_blocks); total_dnode++; } if (!err) @@ -937,12 +939,11 @@ skip: destroy_fsync_dnodes(&tmp_inode_list, err); /* truncate meta pages to be used by the recovery */ - truncate_inode_pages_range(META_MAPPING(sbi), - (loff_t)MAIN_BLKADDR(sbi) << PAGE_SHIFT, -1); - + f2fs_truncate_meta_caches(sbi, MAIN_BLKADDR(sbi), + MAX_BLKADDR(sbi) - MAIN_BLKADDR(sbi)); if (err) { - truncate_inode_pages_final(NODE_MAPPING(sbi)); - truncate_inode_pages_final(META_MAPPING(sbi)); + f2fs_truncate_node_caches(sbi, 0, ULONG_MAX); + f2fs_truncate_meta_caches(sbi, 0, ULONG_MAX); } /* diff --git a/fs/f2fs/segment.c b/fs/f2fs/segment.c index 8c156e1fd37d..b09ffe916642 100644 --- a/fs/f2fs/segment.c +++ b/fs/f2fs/segment.c @@ -335,8 +335,8 @@ static int __f2fs_commit_atomic_write(struct inode *inode) goto next; } - blen = min((pgoff_t)ADDRS_PER_PAGE(dn.node_folio, cow_inode), - len); + blen = min((pgoff_t)addrs_per_page(cow_inode, + IS_INODE(F2FS_I_SB(cow_inode), dn.node_entry)), len); index = off; for (i = 0; i < blen; i++, dn.ofs_in_node++, index++) { blkaddr = f2fs_data_blkaddr(&dn); @@ -479,11 +479,11 @@ void f2fs_balance_fs(struct f2fs_sb_info *sbi, bool need) static inline bool excess_dirty_threshold(struct f2fs_sb_info *sbi) { int factor = f2fs_rwsem_is_locked(&sbi->cp_rwsem) ? 3 : 2; - unsigned int dents = get_pages(sbi, F2FS_DIRTY_DENTS); - unsigned int qdata = get_pages(sbi, F2FS_DIRTY_QDATA); - unsigned int nodes = get_pages(sbi, F2FS_DIRTY_NODES); - unsigned int meta = get_pages(sbi, F2FS_DIRTY_META); - unsigned int imeta = get_pages(sbi, F2FS_DIRTY_IMETA); + unsigned int dents = get_nr_caches(sbi, F2FS_DIRTY_DENTS); + unsigned int qdata = get_nr_caches(sbi, F2FS_DIRTY_QDATA); + unsigned int nodes = get_nr_caches(sbi, F2FS_DIRTY_NODES); + unsigned int meta = get_nr_caches(sbi, F2FS_DIRTY_META); + unsigned int imeta = get_nr_caches(sbi, F2FS_DIRTY_IMETA); unsigned int threshold = SEGS_TO_BLKS(sbi, (factor * DEFAULT_DIRTY_THRESHOLD)); unsigned int global_threshold = threshold * 3 / 2; @@ -512,10 +512,10 @@ void f2fs_balance_fs_bg(struct f2fs_sb_info *sbi, bool from_bg) /* check the # of cached NAT entries */ if (!f2fs_available_free_memory(sbi, NAT_ENTRIES)) - f2fs_try_to_free_nats(sbi, NAT_ENTRY_PER_BLOCK); + f2fs_try_to_free_nats(sbi, NAT_ENTRY_PER_BLOCK(sbi)); if (!f2fs_available_free_memory(sbi, FREE_NIDS)) - f2fs_try_to_free_nids(sbi, MAX_FREE_NIDS); + f2fs_try_to_free_nids(sbi, MAX_FREE_NIDS(sbi)); else f2fs_build_free_nids(sbi, false, false); @@ -1310,13 +1310,14 @@ static void __submit_zone_reset_cmd(struct f2fs_sb_info *sbi, /* sanity check on discard range */ __check_sit_bitmap(sbi, dc->di.lstart, dc->di.lstart + dc->di.len); - bio->bi_iter.bi_sector = SECTOR_FROM_BLOCK(dc->di.start); + bio->bi_iter.bi_sector = SECTOR_FROM_BLOCK(sbi, dc->di.start); bio->bi_private = dc; bio->bi_end_io = f2fs_submit_discard_endio; submit_bio(bio); atomic_inc(&dcc->issued_discard); - f2fs_update_iostat(sbi, NULL, FS_ZONE_RESET_IO, dc->di.len * F2FS_BLKSIZE); + f2fs_update_iostat(sbi, NULL, FS_ZONE_RESET_IO, + dc->di.len * F2FS_BLKSIZE(sbi)); } #endif @@ -1327,7 +1328,7 @@ static int __submit_discard_cmd(struct f2fs_sb_info *sbi, { struct block_device *bdev = dc->bdev; unsigned int max_discard_blocks = - SECTOR_TO_BLOCK(bdev_max_discard_sectors(bdev)); + SECTOR_TO_BLOCK(sbi, bdev_max_discard_sectors(bdev)); struct discard_cmd_control *dcc = SM_I(sbi)->dcc_info; struct list_head *wait_list = (dpolicy->type == DPOLICY_FSTRIM) ? &(dcc->fstrim_list) : &(dcc->wait_list); @@ -1389,8 +1390,8 @@ static int __submit_discard_cmd(struct f2fs_sb_info *sbi, dc->di.len += len; - __blkdev_issue_discard(bdev, SECTOR_FROM_BLOCK(start), - SECTOR_FROM_BLOCK(len), GFP_NOFS, &bio); + __blkdev_issue_discard(bdev, SECTOR_FROM_BLOCK(sbi, start), + SECTOR_FROM_BLOCK(sbi, len), GFP_NOFS, &bio); f2fs_bug_on(sbi, !bio); /* @@ -1419,7 +1420,8 @@ static int __submit_discard_cmd(struct f2fs_sb_info *sbi, atomic_inc(&dcc->issued_discard); - f2fs_update_iostat(sbi, NULL, FS_DISCARD_IO, len * F2FS_BLKSIZE); + f2fs_update_iostat(sbi, NULL, FS_DISCARD_IO, + len * F2FS_BLKSIZE(sbi)); lstart += len; start += len; @@ -1518,7 +1520,7 @@ static void __update_discard_tree_range(struct f2fs_sb_info *sbi, struct discard_info di = {0}; struct rb_node **insert_p = NULL, *insert_parent = NULL; unsigned int max_discard_blocks = - SECTOR_TO_BLOCK(bdev_max_discard_sectors(bdev)); + SECTOR_TO_BLOCK(sbi, bdev_max_discard_sectors(bdev)); block_t end = lstart + len; dc = __lookup_discard_cmd_ret(&dcc->root, lstart, @@ -1871,6 +1873,60 @@ static unsigned int __wait_all_discard_cmd(struct f2fs_sb_info *sbi, return discard_blks; } +void f2fs_drop_discard_cmd_range(struct f2fs_sb_info *sbi, + block_t start, block_t len) +{ + struct discard_cmd_control *dcc = SM_I(sbi)->dcc_info; + struct discard_cmd *prev_dc = NULL, *next_dc = NULL; + struct rb_node **insert_p = NULL, *insert_parent = NULL; + struct discard_cmd *dc, *wait_dc; + u64 cur = start; + u64 end = (u64)start + len; + int count; + + if (!f2fs_realtime_discard_enable(sbi)) + return; + +next: + count = 0; + wait_dc = NULL; + + mutex_lock(&dcc->cmd_lock); + while (cur < end) { + dc = __lookup_discard_cmd_ret(&dcc->root, cur, + &prev_dc, &next_dc, &insert_p, &insert_parent); + if (!dc) + dc = next_dc; + + if (!dc || (u64)dc->di.lstart >= end) + break; + + if (dc->state == D_PREP) { + cur = (u64)dc->di.lstart + dc->di.len; + __remove_discard_cmd(sbi, dc); + if (++count >= MAX_DISCARD_DROP_COUNT) + break; + continue; + } + + dc->ref++; + cur = (u64)dc->di.lstart + dc->di.len; + wait_dc = dc; + break; + } + mutex_unlock(&dcc->cmd_lock); + + if (wait_dc) { + __wait_one_discard_bio(sbi, wait_dc); + goto next; + } + + if (count >= MAX_DISCARD_DROP_COUNT) { + cond_resched(); + goto next; + } +} + /* This should be covered by global mutex, &sit_i->sentry_lock */ static void f2fs_wait_discard_bio(struct f2fs_sb_info *sbi, block_t blkaddr) { @@ -2041,8 +2097,8 @@ static int __f2fs_issue_discard_zone(struct f2fs_sb_info *sbi, /* For sequential zones, reset the zone write pointer */ if (f2fs_blkz_is_seq(sbi, devi, blkstart)) { - sector = SECTOR_FROM_BLOCK(blkstart); - nr_sects = SECTOR_FROM_BLOCK(blklen); + sector = SECTOR_FROM_BLOCK(sbi, blkstart); + nr_sects = SECTOR_FROM_BLOCK(sbi, blklen); div64_u64_rem(sector, bdev_zone_sectors(bdev), &remainder); if (remainder || nr_sects != bdev_zone_sectors(bdev)) { @@ -2749,12 +2805,12 @@ static unsigned short f2fs_curseg_valid_blocks(struct f2fs_sb_info *sbi, int typ } /* - * Calculate the number of current summary pages for writing + * Calculate the number of current summary blocks for writing */ -int f2fs_npages_for_summary_flush(struct f2fs_sb_info *sbi, bool for_ra) +int f2fs_nblocks_for_summary_flush(struct f2fs_sb_info *sbi, bool for_ra) { int valid_sum_count = 0; - int i, sum_in_page; + int i, sum_in_block; for (i = CURSEG_HOT_DATA; i <= CURSEG_COLD_DATA; i++) { if (sbi->ckpt->alloc_type[i] != SSR && for_ra) @@ -2764,73 +2820,72 @@ int f2fs_npages_for_summary_flush(struct f2fs_sb_info *sbi, bool for_ra) valid_sum_count += f2fs_curseg_valid_blocks(sbi, i); } - sum_in_page = (sbi->blocksize - 2 * sbi->sum_journal_size - + sum_in_block = (sbi->blocksize - 2 * sbi->sum_journal_size - SUM_FOOTER_SIZE) / SUMMARY_SIZE; - if (valid_sum_count <= sum_in_page) + if (valid_sum_count <= sum_in_block) return 1; - else if ((valid_sum_count - sum_in_page) <= + else if ((valid_sum_count - sum_in_block) <= (sbi->blocksize - SUM_FOOTER_SIZE) / SUMMARY_SIZE) return 2; return 3; } -/* - * Caller should put this summary folio - */ -struct folio *f2fs_get_sum_folio(struct f2fs_sb_info *sbi, unsigned int segno) + +struct f2fs_cached_block *f2fs_get_sum_cache(struct f2fs_sb_info *sbi, + unsigned int segno) { if (unlikely(f2fs_cp_error(sbi))) return ERR_PTR(-EIO); - return f2fs_get_meta_folio_retry(sbi, GET_SUM_BLOCK(sbi, segno)); + return f2fs_get_meta_cache_retry(sbi, GET_SUM_BLOCK(sbi, segno)); } -void f2fs_update_meta_page(struct f2fs_sb_info *sbi, +void f2fs_update_meta_block(struct f2fs_sb_info *sbi, void *src, block_t blk_addr) { - struct folio *folio; + struct f2fs_cached_block *entry; if (!f2fs_sb_has_packed_ssa(sbi)) - folio = f2fs_grab_meta_folio(sbi, blk_addr); + entry = f2fs_grab_meta_cache(sbi, blk_addr); else - folio = f2fs_get_meta_folio_retry(sbi, blk_addr); + entry = f2fs_get_meta_cache_retry(sbi, blk_addr); - if (IS_ERR(folio)) + if (IS_ERR(entry)) return; - memcpy(folio_address(folio), src, PAGE_SIZE); - folio_mark_dirty(folio); - f2fs_folio_put(folio, true); + memcpy(cache_address(entry), src, F2FS_BLKSIZE(sbi)); + f2fs_mark_cache_dirty(entry); + f2fs_put_cache(entry, true); } -static void write_sum_page(struct f2fs_sb_info *sbi, +static void write_sum_block(struct f2fs_sb_info *sbi, struct f2fs_summary_block *sum_blk, unsigned int segno) { - struct folio *folio; + struct f2fs_cached_block *entry; if (!f2fs_sb_has_packed_ssa(sbi)) - return f2fs_update_meta_page(sbi, (void *)sum_blk, + return f2fs_update_meta_block(sbi, (void *)sum_blk, GET_SUM_BLOCK(sbi, segno)); - folio = f2fs_get_sum_folio(sbi, segno); - if (IS_ERR(folio)) + entry = f2fs_get_sum_cache(sbi, segno); + if (IS_ERR(entry)) return; - memcpy(SUM_BLK_PAGE_ADDR(sbi, folio, segno), sum_blk, + memcpy(SUM_BLK_ENTRY_ADDR(sbi, entry, segno), sum_blk, sbi->sum_blocksize); - folio_mark_dirty(folio); - f2fs_folio_put(folio, true); + f2fs_mark_cache_dirty(entry); + f2fs_put_cache(entry, true); } -static void write_current_sum_page(struct f2fs_sb_info *sbi, +static void write_current_sum_block(struct f2fs_sb_info *sbi, int type, block_t blk_addr) { struct curseg_info *curseg = CURSEG_I(sbi, type); - struct folio *folio = f2fs_grab_meta_folio(sbi, blk_addr); + struct f2fs_cached_block *entry = f2fs_grab_meta_cache(sbi, blk_addr); struct f2fs_summary_block *src = curseg->sum_blk; struct f2fs_summary_block *dst; - dst = folio_address(folio); - memset(dst, 0, PAGE_SIZE); + dst = cache_address(entry); + memset(dst, 0, sbi->blocksize); mutex_lock(&curseg->curseg_mutex); @@ -2843,8 +2898,8 @@ static void write_current_sum_page(struct f2fs_sb_info *sbi, mutex_unlock(&curseg->curseg_mutex); - folio_mark_dirty(folio); - f2fs_folio_put(folio, true); + f2fs_mark_cache_dirty(entry); + f2fs_put_cache(entry, true); } static int is_next_segment_free(struct f2fs_sb_info *sbi, @@ -3104,7 +3159,7 @@ static int new_curseg(struct f2fs_sb_info *sbi, int type, bool new_sec) int ret; if (curseg->inited) - write_sum_page(sbi, curseg->sum_blk, segno); + write_sum_block(sbi, curseg->sum_blk, segno); segno = __get_next_segno(sbi, type); ret = get_new_segment(sbi, &segno, new_sec, pinning); @@ -3160,10 +3215,10 @@ static int change_curseg(struct f2fs_sb_info *sbi, int type) struct curseg_info *curseg = CURSEG_I(sbi, type); unsigned int new_segno = curseg->next_segno; struct f2fs_summary_block *sum_node; - struct folio *sum_folio; + struct f2fs_cached_block *entry = NULL; if (curseg->inited) - write_sum_page(sbi, curseg->sum_blk, curseg->segno); + write_sum_block(sbi, curseg->sum_blk, curseg->segno); __set_test_and_inuse(sbi, new_segno); @@ -3176,15 +3231,15 @@ static int change_curseg(struct f2fs_sb_info *sbi, int type) curseg->alloc_type = SSR; curseg->next_blkoff = __next_free_blkoff(sbi, curseg->segno, 0); - sum_folio = f2fs_get_sum_folio(sbi, new_segno); - if (IS_ERR(sum_folio)) { - /* GC won't be able to use stale summary pages by cp_error */ + entry = f2fs_get_sum_cache(sbi, new_segno); + if (IS_ERR(entry)) { + /* GC won't be able to use stale summary blocks by cp_error */ memset(curseg->sum_blk, 0, sbi->sum_entry_size); - return PTR_ERR(sum_folio); + return PTR_ERR(entry); } - sum_node = SUM_BLK_PAGE_ADDR(sbi, sum_folio, new_segno); + sum_node = SUM_BLK_ENTRY_ADDR(sbi, entry, new_segno); memcpy(curseg->sum_blk, sum_node, sbi->sum_entry_size); - f2fs_folio_put(sum_folio, true); + f2fs_put_cache(entry, true); return 0; } @@ -3271,7 +3326,7 @@ static void __f2fs_save_inmem_curseg(struct f2fs_sb_info *sbi, int type) goto out; if (get_valid_blocks(sbi, curseg->segno, false)) { - write_sum_page(sbi, curseg->sum_blk, curseg->segno); + write_sum_block(sbi, curseg->sum_blk, curseg->segno); } else { mutex_lock(&DIRTY_I(sbi)->seglist_lock); __set_test_and_free(sbi, curseg->segno, true); @@ -3592,8 +3647,8 @@ skip: int f2fs_trim_fs(struct f2fs_sb_info *sbi, struct fstrim_range *range) { - __u64 start = F2FS_BYTES_TO_BLK(range->start); - __u64 end = start + F2FS_BYTES_TO_BLK(range->len) - 1; + __u64 start = F2FS_BYTES_TO_BLK(sbi, range->start); + __u64 end = start + F2FS_BYTES_TO_BLK(sbi, range->len) - 1; unsigned int start_segno, end_segno; block_t start_block, end_block; struct cp_control cpc; @@ -3624,7 +3679,8 @@ int f2fs_trim_fs(struct f2fs_sb_info *sbi, struct fstrim_range *range) } cpc.reason = CP_DISCARD; - cpc.trim_minlen = max_t(__u64, 1, F2FS_BYTES_TO_BLK(range->minlen)); + cpc.trim_minlen = max_t(__u64, 1, + F2FS_BYTES_TO_BLK(sbi, range->minlen)); cpc.trim_start = start_segno; cpc.trim_end = end_segno; @@ -3658,7 +3714,7 @@ int f2fs_trim_fs(struct f2fs_sb_info *sbi, struct fstrim_range *range) start_block, end_block); out: if (!err) - range->len = F2FS_BLK_TO_BYTES(trimmed); + range->len = F2FS_BLK_TO_BYTES(sbi, trimmed); return err; } @@ -3770,7 +3826,8 @@ static int __get_segment_type_4(struct f2fs_io_info *fio) else return CURSEG_COLD_DATA; } else { - if (IS_DNODE(fio->folio) && is_cold_node(fio->folio)) + f2fs_bug_on(fio->sbi, !fio->is_cache); + if (IS_DNODE(fio->sbi, fio->cache_entry) && is_cold_node(fio->sbi, fio->cache_entry)) return CURSEG_WARM_NODE; else return CURSEG_COLD_NODE; @@ -3828,9 +3885,9 @@ static int __get_segment_type_6(struct f2fs_io_info *fio) return f2fs_rw_hint_to_seg_type(F2FS_I_SB(inode), inode->i_write_hint); } else { - if (IS_DNODE(fio->folio)) - return is_cold_node(fio->folio) ? CURSEG_WARM_NODE : - CURSEG_HOT_NODE; + f2fs_bug_on(fio->sbi, !fio->is_cache); + if (IS_DNODE(fio->sbi, fio->cache_entry)) + return is_cold_node(fio->sbi, fio->cache_entry) ? CURSEG_WARM_NODE : CURSEG_HOT_NODE; return CURSEG_COLD_NODE; } } @@ -3897,7 +3954,7 @@ static void f2fs_randomize_chunk(struct f2fs_sb_info *sbi, get_random_u32_inclusive(1, sbi->max_fragment_hole); } -int f2fs_allocate_data_block(struct f2fs_sb_info *sbi, struct folio *folio, +int f2fs_allocate_data_block(struct f2fs_sb_info *sbi, block_t old_blkaddr, block_t *new_blkaddr, struct f2fs_summary *sum, int type, struct f2fs_io_info *fio) @@ -3966,7 +4023,7 @@ int f2fs_allocate_data_block(struct f2fs_sb_info *sbi, struct folio *folio, if (segment_full) { if (type == CURSEG_COLD_DATA_PINNED && !((curseg->segno + 1) % sbi->segs_per_sec)) { - write_sum_page(sbi, curseg->sum_blk, curseg->segno); + write_sum_block(sbi, curseg->sum_blk, curseg->segno); reset_curseg_fields(curseg); goto skip_new_segment; } @@ -4005,10 +4062,10 @@ skip_new_segment: up_write(&sit_i->sentry_lock); - if (folio && IS_NODESEG(curseg->seg_type)) { - fill_node_footer_blkaddr(folio, NEXT_FREE_BLKADDR(sbi, curseg)); - - f2fs_inode_chksum_set(sbi, folio); + if (fio && fio->is_cache && IS_NODESEG(curseg->seg_type)) { + fill_node_footer_blkaddr(sbi, fio->cache_entry, + NEXT_FREE_BLKADDR(sbi, curseg)); + f2fs_inode_chksum_set(sbi, fio->cache_entry); } if (fio) { @@ -4084,7 +4141,7 @@ static int log_type_to_seg_type(enum log_type type) return seg_type; } -static void do_write_page(struct f2fs_summary *sum, struct f2fs_io_info *fio) +static void do_write_block(struct f2fs_summary *sum, struct f2fs_io_info *fio) { struct folio *folio = fio->folio; enum log_type type = __get_segment_type(fio); @@ -4096,16 +4153,24 @@ static void do_write_page(struct f2fs_summary *sum, struct f2fs_io_info *fio) if (keep_order) f2fs_down_read(&fio->sbi->io_order_lock); - err = f2fs_allocate_data_block(fio->sbi, folio, fio->old_blkaddr, + err = f2fs_allocate_data_block(fio->sbi, fio->old_blkaddr, &fio->new_blkaddr, sum, type, fio); if (unlikely(err)) { - f2fs_err_ratelimited(fio->sbi, - "%s Failed to allocate data block, ino:%u, index:%lu, type:%d, old_blkaddr:0x%x, new_blkaddr:0x%x, err:%d", - __func__, fio->ino, folio->index, type, - fio->old_blkaddr, fio->new_blkaddr, err); - folio_end_writeback(folio); - if (f2fs_in_warm_node_list(folio)) - f2fs_del_fsync_node_entry(fio->sbi, folio); + if (fio->is_cache) { + f2fs_err_ratelimited(fio->sbi, + "%s Failed to allocate data block, ino:%u, index:%lu, type:%d, old_blkaddr:0x%x, new_blkaddr:0x%x, err:%d", + __func__, fio->ino, fio->cache_entry->index, type, + fio->old_blkaddr, fio->new_blkaddr, err); + if (f2fs_in_warm_node_list(fio->sbi, fio->cache_entry)) + f2fs_del_fsync_node_entry(fio->sbi, fio->cache_entry); + f2fs_end_cache_writeback(fio->cache_entry); + } else { + f2fs_err_ratelimited(fio->sbi, + "%s Failed to allocate data block, ino:%u, index:%lu, type:%d, old_blkaddr:0x%x, new_blkaddr:0x%x, err:%d", + __func__, fio->ino, folio->index, type, + fio->old_blkaddr, fio->new_blkaddr, err); + folio_end_writeback(folio); + } f2fs_bug_on(fio->sbi, !is_set_ckpt_flags(fio->sbi, CP_ERROR_FLAG)); goto out; @@ -4118,7 +4183,10 @@ static void do_write_page(struct f2fs_summary *sum, struct f2fs_io_info *fio) f2fs_invalidate_internal_cache(fio->sbi, fio->old_blkaddr, 1); /* writeout dirty page into bdev */ - f2fs_submit_page_write(fio); + if (fio->is_cache) + f2fs_submit_cache_write(fio); + else + f2fs_submit_page_write(fio); f2fs_update_device_state(fio->sbi, fio->ino, fio->new_blkaddr, 1); out: @@ -4126,8 +4194,9 @@ out: f2fs_up_read(&fio->sbi->io_order_lock); } -void f2fs_do_write_meta_page(struct f2fs_sb_info *sbi, struct folio *folio, - enum iostat_type io_type) +void f2fs_do_write_meta_cache(struct f2fs_sb_info *sbi, + struct f2fs_cached_block *entry, + enum iostat_type io_type) { struct f2fs_io_info fio = { .sbi = sbi, @@ -4135,31 +4204,36 @@ void f2fs_do_write_meta_page(struct f2fs_sb_info *sbi, struct folio *folio, .temp = HOT, .op = REQ_OP_WRITE, .op_flags = REQ_SYNC | REQ_META | REQ_PRIO, - .old_blkaddr = folio->index, - .new_blkaddr = folio->index, - .folio = folio, + .old_blkaddr = entry->index, + .new_blkaddr = entry->index, .encrypted_page = NULL, .in_list = 0, + .cache_entry = entry, + .is_cache = 1, }; - if (unlikely(folio->index >= MAIN_BLKADDR(sbi))) + if (unlikely(entry->index >= MAIN_BLKADDR(sbi))) fio.op_flags &= ~REQ_META; - folio_start_writeback(folio); - f2fs_submit_page_write(&fio); + f2fs_start_cache_writeback(entry); + f2fs_submit_cache_write(&fio); - stat_inc_meta_count(sbi, folio->index); - f2fs_update_iostat(sbi, NULL, io_type, F2FS_BLKSIZE); + stat_inc_meta_count(sbi, entry->index); + f2fs_update_iostat(sbi, NULL, io_type, F2FS_BLKSIZE(sbi)); } -void f2fs_do_write_node_page(unsigned int nid, struct f2fs_io_info *fio) +void f2fs_do_write_node_cache(unsigned int nid, struct f2fs_io_info *fio) { struct f2fs_summary sum; + if (fio->is_cache) + trace_f2fs_write_cache(fio->cache_entry, NODE); + set_summary(&sum, nid, 0, 0); - do_write_page(&sum, fio); + do_write_block(&sum, fio); - f2fs_update_iostat(fio->sbi, NULL, fio->io_type, F2FS_BLKSIZE); + f2fs_update_iostat(fio->sbi, NULL, fio->io_type, + F2FS_BLKSIZE(fio->sbi)); } void f2fs_outplace_write_data(struct dnode_of_data *dn, @@ -4172,10 +4246,11 @@ void f2fs_outplace_write_data(struct dnode_of_data *dn, if (fio->io_type == FS_DATA_IO || fio->io_type == FS_CP_DATA_IO) f2fs_update_age_extent_cache(dn); set_summary(&sum, dn->nid, dn->ofs_in_node, fio->version); - do_write_page(&sum, fio); + do_write_block(&sum, fio); f2fs_update_data_blkaddr(dn, fio->new_blkaddr); - f2fs_update_iostat(sbi, dn->inode, fio->io_type, F2FS_BLKSIZE); + f2fs_update_iostat(sbi, dn->inode, fio->io_type, + F2FS_BLKSIZE(sbi)); } int f2fs_inplace_write_data(struct f2fs_io_info *fio) @@ -4205,7 +4280,7 @@ int f2fs_inplace_write_data(struct f2fs_io_info *fio) } if (fio->meta_gc) - f2fs_truncate_meta_inode_pages(sbi, fio->new_blkaddr, 1); + f2fs_truncate_meta_caches(sbi, fio->new_blkaddr, 1); stat_inc_inplace_blocks(fio->sbi); @@ -4217,7 +4292,7 @@ int f2fs_inplace_write_data(struct f2fs_io_info *fio) f2fs_update_device_state(fio->sbi, fio->ino, fio->new_blkaddr, 1); f2fs_update_iostat(fio->sbi, fio_inode(fio), - fio->io_type, F2FS_BLKSIZE); + fio->io_type, F2FS_BLKSIZE(fio->sbi)); } return err; @@ -4349,14 +4424,13 @@ void f2fs_replace_block(struct f2fs_sb_info *sbi, struct dnode_of_data *dn, f2fs_update_data_blkaddr(dn, new_addr); } -void f2fs_folio_wait_writeback(struct folio *folio, enum page_type type, - bool ordered, bool locked) +void f2fs_folio_wait_writeback(struct folio *folio, bool ordered, bool locked) { if (folio_test_writeback(folio)) { struct f2fs_sb_info *sbi = F2FS_F_SB(folio); /* submit cached LFS IO */ - f2fs_submit_merged_write_folio(sbi, folio, type); + f2fs_submit_merged_write_folio(sbi, folio); /* submit cached IPU IO */ f2fs_submit_merged_ipu_write(sbi, NULL, folio); if (ordered) { @@ -4371,7 +4445,7 @@ void f2fs_folio_wait_writeback(struct folio *folio, enum page_type type, void f2fs_wait_on_block_writeback(struct inode *inode, block_t blkaddr) { struct f2fs_sb_info *sbi = F2FS_I_SB(inode); - struct folio *cfolio; + struct f2fs_cached_block *entry; if (!f2fs_meta_inode_gc_required(inode)) return; @@ -4379,11 +4453,12 @@ void f2fs_wait_on_block_writeback(struct inode *inode, block_t blkaddr) if (!__is_valid_data_blkaddr(blkaddr)) return; - cfolio = filemap_lock_folio(META_MAPPING(sbi), blkaddr); - if (!IS_ERR(cfolio)) { - f2fs_folio_wait_writeback(cfolio, DATA, true, true); - f2fs_folio_put(cfolio, true); - } + entry = f2fs_find_cache(META_CACHE(sbi), blkaddr, 0); + if (IS_ERR(entry)) + return; + f2fs_lock_cache(entry); + f2fs_cache_wait_writeback_cond(entry, DATA); + f2fs_put_cache(entry, true); } void f2fs_wait_on_block_writeback_range(struct inode *inode, block_t blkaddr, @@ -4398,7 +4473,7 @@ void f2fs_wait_on_block_writeback_range(struct inode *inode, block_t blkaddr, for (i = 0; i < len; i++) f2fs_wait_on_block_writeback(inode, blkaddr + i); - f2fs_truncate_meta_inode_pages(sbi, blkaddr, len); + f2fs_truncate_meta_caches(sbi, blkaddr, len); } static int read_compacted_summaries(struct f2fs_sb_info *sbi) @@ -4406,16 +4481,16 @@ static int read_compacted_summaries(struct f2fs_sb_info *sbi) struct f2fs_checkpoint *ckpt = F2FS_CKPT(sbi); struct curseg_info *seg_i; unsigned char *kaddr; - struct folio *folio; + struct f2fs_cached_block *entry; block_t start; int i, j, offset; start = start_sum_block(sbi); - folio = f2fs_get_meta_folio(sbi, start++); - if (IS_ERR(folio)) - return PTR_ERR(folio); - kaddr = folio_address(folio); + entry = f2fs_get_meta_cache(sbi, start++); + if (IS_ERR(entry)) + return PTR_ERR(entry); + kaddr = cache_address(entry); /* Step 1: restore nat cache */ seg_i = CURSEG_I(sbi, CURSEG_HOT_DATA); @@ -4452,16 +4527,16 @@ static int read_compacted_summaries(struct f2fs_sb_info *sbi) SUM_FOOTER_SIZE) continue; - f2fs_folio_put(folio, true); + f2fs_put_cache(entry, true); - folio = f2fs_get_meta_folio(sbi, start++); - if (IS_ERR(folio)) - return PTR_ERR(folio); - kaddr = folio_address(folio); + entry = f2fs_get_meta_cache(sbi, start++); + if (IS_ERR(entry)) + return PTR_ERR(entry); + kaddr = cache_address(entry); offset = 0; } } - f2fs_folio_put(folio, true); + f2fs_put_cache(entry, true); return 0; } @@ -4470,7 +4545,7 @@ static int read_normal_summaries(struct f2fs_sb_info *sbi, int type) struct f2fs_checkpoint *ckpt = F2FS_CKPT(sbi); struct f2fs_summary_block *sum; struct curseg_info *curseg; - struct folio *new; + struct f2fs_cached_block *entry; unsigned short blk_off; unsigned int segno = 0; block_t blk_addr = 0; @@ -4497,10 +4572,10 @@ static int read_normal_summaries(struct f2fs_sb_info *sbi, int type) blk_addr = GET_SUM_BLOCK(sbi, segno); } - new = f2fs_get_meta_folio(sbi, blk_addr); - if (IS_ERR(new)) - return PTR_ERR(new); - sum = folio_address(new); + entry = f2fs_get_meta_cache(sbi, blk_addr); + if (IS_ERR(entry)) + return PTR_ERR(entry); + sum = cache_address(entry); if (IS_NODESEG(type)) { if (__exist_node_summaries(sbi)) { @@ -4537,7 +4612,7 @@ static int read_normal_summaries(struct f2fs_sb_info *sbi, int type) curseg->next_blkoff = blk_off; mutex_unlock(&curseg->curseg_mutex); out: - f2fs_folio_put(new, true); + f2fs_put_cache(entry, true); return err; } @@ -4549,11 +4624,11 @@ static int restore_curseg_summaries(struct f2fs_sb_info *sbi) int err; if (is_set_ckpt_flags(sbi, CP_COMPACT_SUM_FLAG)) { - int npages = f2fs_npages_for_summary_flush(sbi, true); + int nblocks = f2fs_nblocks_for_summary_flush(sbi, true); - if (npages >= 2) - f2fs_ra_meta_pages(sbi, start_sum_block(sbi), npages, - META_CP, true); + if (nblocks >= 2) + f2fs_ra_meta_caches(sbi, start_sum_block(sbi), + nblocks, META_CP, true); /* restore for compacted data summary */ err = read_compacted_summaries(sbi); @@ -4563,7 +4638,7 @@ static int restore_curseg_summaries(struct f2fs_sb_info *sbi) } if (__exist_node_summaries(sbi)) - f2fs_ra_meta_pages(sbi, + f2fs_ra_meta_caches(sbi, sum_blk_addr(sbi, NR_CURSEG_PERSIST_TYPE, type), NR_CURSEG_PERSIST_TYPE - type, META_CP, true); @@ -4586,16 +4661,16 @@ static int restore_curseg_summaries(struct f2fs_sb_info *sbi) static void write_compacted_summaries(struct f2fs_sb_info *sbi, block_t blkaddr) { - struct folio *folio; + struct f2fs_cached_block *entry = NULL; unsigned char *kaddr; struct f2fs_summary *summary; struct curseg_info *seg_i; int written_size = 0; int i, j; - folio = f2fs_grab_meta_folio(sbi, blkaddr++); - kaddr = folio_address(folio); - memset(kaddr, 0, PAGE_SIZE); + entry = f2fs_grab_meta_cache(sbi, blkaddr++); + kaddr = cache_address(entry); + memset(kaddr, 0, sbi->blocksize); /* Step 1: write nat cache */ seg_i = CURSEG_I(sbi, CURSEG_HOT_DATA); @@ -4611,10 +4686,10 @@ static void write_compacted_summaries(struct f2fs_sb_info *sbi, block_t blkaddr) for (i = CURSEG_HOT_DATA; i <= CURSEG_COLD_DATA; i++) { seg_i = CURSEG_I(sbi, i); for (j = 0; j < f2fs_curseg_valid_blocks(sbi, i); j++) { - if (!folio) { - folio = f2fs_grab_meta_folio(sbi, blkaddr++); - kaddr = folio_address(folio); - memset(kaddr, 0, PAGE_SIZE); + if (!entry) { + entry = f2fs_grab_meta_cache(sbi, blkaddr++); + kaddr = cache_address(entry); + memset(kaddr, 0, sbi->blocksize); written_size = 0; } summary = (struct f2fs_summary *)(kaddr + written_size); @@ -4625,14 +4700,14 @@ static void write_compacted_summaries(struct f2fs_sb_info *sbi, block_t blkaddr) SUM_FOOTER_SIZE) continue; - folio_mark_dirty(folio); - f2fs_folio_put(folio, true); - folio = NULL; + f2fs_mark_cache_dirty(entry); + f2fs_put_cache(entry, true); + entry = NULL; } } - if (folio) { - folio_mark_dirty(folio); - f2fs_folio_put(folio, true); + if (entry) { + f2fs_mark_cache_dirty(entry); + f2fs_put_cache(entry, true); } } @@ -4647,7 +4722,7 @@ static void write_normal_summaries(struct f2fs_sb_info *sbi, end = type + NR_CURSEG_NODE_TYPE; for (i = type; i < end; i++) - write_current_sum_page(sbi, i, blkaddr + (i - type)); + write_current_sum_block(sbi, i, blkaddr + (i - type)); } void f2fs_write_data_summaries(struct f2fs_sb_info *sbi, block_t start_blk) @@ -4686,29 +4761,29 @@ int f2fs_lookup_journal_in_cursum(struct f2fs_sb_info *sbi, return -1; } -static struct folio *get_current_sit_folio(struct f2fs_sb_info *sbi, +static struct f2fs_cached_block *get_current_sit_cache(struct f2fs_sb_info *sbi, unsigned int segno) { - return f2fs_get_meta_folio(sbi, current_sit_addr(sbi, segno)); + return f2fs_get_meta_cache(sbi, current_sit_addr(sbi, segno)); } -static struct folio *get_next_sit_folio(struct f2fs_sb_info *sbi, +static struct f2fs_cached_block *get_next_sit_cache(struct f2fs_sb_info *sbi, unsigned int start) { struct sit_info *sit_i = SIT_I(sbi); - struct folio *folio; + struct f2fs_cached_block *entry; pgoff_t src_off, dst_off; src_off = current_sit_addr(sbi, start); dst_off = next_sit_addr(sbi, src_off); - folio = f2fs_grab_meta_folio(sbi, dst_off); - seg_info_to_sit_folio(sbi, folio, start); + entry = f2fs_grab_meta_cache(sbi, dst_off); + seg_info_to_sit_block(sbi, entry, start); - folio_mark_dirty(folio); - set_to_next_sit(sit_i, start); + f2fs_mark_cache_dirty(entry); + set_to_next_sit(sbi, sit_i, start); - return folio; + return entry; } static struct sit_entry_set *grab_sit_entry_set(void) @@ -4745,10 +4820,11 @@ static void adjust_sit_entry_set(struct sit_entry_set *ses, list_move_tail(&ses->set_list, head); } -static void add_sit_entry(unsigned int segno, struct list_head *head) +static void add_sit_entry(struct f2fs_sb_info *sbi, unsigned int segno, + struct list_head *head) { struct sit_entry_set *ses; - unsigned int start_segno = START_SEGNO(segno); + unsigned int start_segno = f2fs_start_segno(sbi, segno); list_for_each_entry(ses, head, set_list) { if (ses->start_segno == start_segno) { @@ -4773,7 +4849,7 @@ static void add_sits_in_set(struct f2fs_sb_info *sbi) unsigned int segno; for_each_set_bit(segno, bitmap, MAIN_SEGS(sbi)) - add_sit_entry(segno, set_list); + add_sit_entry(sbi, segno, set_list); } static void remove_sits_in_journal(struct f2fs_sb_info *sbi) @@ -4791,7 +4867,7 @@ static void remove_sits_in_journal(struct f2fs_sb_info *sbi) dirtied = __mark_sit_entry_dirty(sbi, segno); if (!dirtied) - add_sit_entry(segno, &SM_I(sbi)->sit_entry_set); + add_sit_entry(sbi, segno, &SM_I(sbi)->sit_entry_set); } update_sits_in_cursum(journal, -i); up_write(&curseg->journal_rwsem); @@ -4838,10 +4914,11 @@ void f2fs_flush_sit_entries(struct f2fs_sb_info *sbi, struct cp_control *cpc) * #2, flush sit entries to sit page. */ list_for_each_entry_safe(ses, tmp, head, set_list) { - struct folio *folio = NULL; + struct f2fs_cached_block *entry = NULL; struct f2fs_sit_block *raw_sit = NULL; + unsigned int start_segno = ses->start_segno; - unsigned int end = min(start_segno + SIT_ENTRY_PER_BLOCK, + unsigned int end = min(start_segno + SIT_ENTRY_PER_BLOCK(sbi), (unsigned long)MAIN_SEGS(sbi)); unsigned int segno = start_segno; @@ -4853,8 +4930,8 @@ void f2fs_flush_sit_entries(struct f2fs_sb_info *sbi, struct cp_control *cpc) if (to_journal) { down_write(&curseg->journal_rwsem); } else { - folio = get_next_sit_folio(sbi, start_segno); - raw_sit = folio_address(folio); + entry = get_next_sit_cache(sbi, start_segno); + raw_sit = cache_address(entry); } /* flush dirty sit entries in region of current sit set */ @@ -4899,7 +4976,7 @@ void f2fs_flush_sit_entries(struct f2fs_sb_info *sbi, struct cp_control *cpc) if (to_journal) up_write(&curseg->journal_rwsem); else - f2fs_folio_put(folio, true); + f2fs_put_cache(entry, true); f2fs_bug_on(sbi, ses->entry_cnt); release_sit_entry_set(ses); @@ -5006,7 +5083,7 @@ static int build_sit_info(struct f2fs_sb_info *sbi) sit_i->written_valid_blocks = 0; sit_i->bitmap_size = sit_bitmap_size; sit_i->dirty_sentries = 0; - sit_i->sents_per_block = SIT_ENTRY_PER_BLOCK; + sit_i->sents_per_block = SIT_ENTRY_PER_BLOCK(sbi); sit_i->elapsed_time = le64_to_cpu(sbi->ckpt->elapsed_time); sit_i->mounted_time = ktime_get_boottime_seconds(); init_rwsem(&sit_i->sentry_lock); @@ -5090,7 +5167,7 @@ static int build_sit_entries(struct f2fs_sb_info *sbi) block_t sit_valid_blocks[2] = {0, 0}; do { - readed = f2fs_ra_meta_pages(sbi, start_blk, BIO_MAX_VECS, + readed = f2fs_ra_meta_caches(sbi, start_blk, BIO_MAX_VECS, META_SIT, true); start = start_blk * sit_i->sents_per_block; @@ -5098,15 +5175,15 @@ static int build_sit_entries(struct f2fs_sb_info *sbi) for (; start < end && start < MAIN_SEGS(sbi); start++) { struct f2fs_sit_block *sit_blk; - struct folio *folio; + struct f2fs_cached_block *entry; se = &sit_i->sentries[start]; - folio = get_current_sit_folio(sbi, start); - if (IS_ERR(folio)) - return PTR_ERR(folio); - sit_blk = folio_address(folio); + entry = get_current_sit_cache(sbi, start); + if (IS_ERR(entry)) + return PTR_ERR(entry); + sit_blk = cache_address(entry); sit = sit_blk->entries[SIT_ENTRY_OFFSET(sit_i, start)]; - f2fs_folio_put(folio, true); + f2fs_put_cache(entry, true); err = check_block_count(sbi, start, &sit); if (err) diff --git a/fs/f2fs/segment.h b/fs/f2fs/segment.h index 5949aa5200ac..526764ba31ae 100644 --- a/fs/f2fs/segment.h +++ b/fs/f2fs/segment.h @@ -96,25 +96,30 @@ static inline void sanity_check_seg_type(struct f2fs_sb_info *sbi, #define GET_SUM_BLKOFF(sbi, segno) (segno % (sbi)->sums_per_block) #define SUM_BLK_PAGE_ADDR(sbi, folio, segno) \ (folio_address(folio) + GET_SUM_BLKOFF(sbi, segno) * (sbi)->sum_blocksize) +#define SUM_BLK_ENTRY_ADDR(sbi, entry, segno) \ + (cache_address(entry) + GET_SUM_BLKOFF(sbi, segno) * (sbi)->sum_blocksize) #define GET_SUM_TYPE(footer) ((footer)->entry_type) #define SET_SUM_TYPE(footer, type) ((footer)->entry_type = (type)) #define SIT_ENTRY_OFFSET(sit_i, segno) \ ((segno) % (sit_i)->sents_per_block) -#define SIT_BLOCK_OFFSET(segno) \ - ((segno) / SIT_ENTRY_PER_BLOCK) -#define START_SEGNO(segno) \ - (SIT_BLOCK_OFFSET(segno) * SIT_ENTRY_PER_BLOCK) +#define SIT_BLOCK_OFFSET(sbi, segno) \ + ((segno) / SIT_ENTRY_PER_BLOCK(sbi)) +static inline unsigned int +f2fs_start_segno(struct f2fs_sb_info *sbi, unsigned int segno) +{ + return SIT_BLOCK_OFFSET(sbi, segno) * SIT_ENTRY_PER_BLOCK(sbi); +} #define SIT_BLK_CNT(sbi) \ - DIV_ROUND_UP(MAIN_SEGS(sbi), SIT_ENTRY_PER_BLOCK) + DIV_ROUND_UP(MAIN_SEGS(sbi), SIT_ENTRY_PER_BLOCK(sbi)) #define f2fs_bitmap_size(nr) \ (BITS_TO_LONGS(nr) * sizeof(unsigned long)) -#define SECTOR_FROM_BLOCK(blk_addr) \ - (((sector_t)blk_addr) << F2FS_LOG_SECTORS_PER_BLOCK) -#define SECTOR_TO_BLOCK(sectors) \ - ((sectors) >> F2FS_LOG_SECTORS_PER_BLOCK) +#define SECTOR_FROM_BLOCK(sbi, blk_addr) \ + (((sector_t)blk_addr) << F2FS_LOG_SECTORS_PER_BLOCK(sbi)) +#define SECTOR_TO_BLOCK(sbi, sectors) \ + ((sectors) >> F2FS_LOG_SECTORS_PER_BLOCK(sbi)) /* * In the victim_sel_policy->alloc_mode, there are three block allocation modes. @@ -418,18 +423,18 @@ static inline void __seg_info_to_raw_sit(struct seg_entry *se, rs->mtime = cpu_to_le64(se->mtime); } -static inline void seg_info_to_sit_folio(struct f2fs_sb_info *sbi, - struct folio *folio, unsigned int start) +static inline void seg_info_to_sit_block(struct f2fs_sb_info *sbi, + struct f2fs_cached_block *entry, unsigned int start) { struct f2fs_sit_block *raw_sit; struct seg_entry *se; struct f2fs_sit_entry *rs; - unsigned int end = min(start + SIT_ENTRY_PER_BLOCK, + unsigned int end = min(start + SIT_ENTRY_PER_BLOCK(sbi), (unsigned long)MAIN_SEGS(sbi)); int i; - raw_sit = folio_address(folio); - memset(raw_sit, 0, PAGE_SIZE); + raw_sit = cache_address(entry); + memset(raw_sit, 0, F2FS_BLKSIZE(sbi)); for (i = 0; i < end - start; i++) { rs = &raw_sit->entries[i]; se = get_seg_entry(sbi, start + i); @@ -651,15 +656,15 @@ static inline void get_additional_blocks_required(struct f2fs_sb_info *sbi, */ static inline int __get_secs_required(struct f2fs_sb_info *sbi) { - unsigned int total_node_blocks = get_pages(sbi, F2FS_DIRTY_NODES) + - get_pages(sbi, F2FS_DIRTY_DENTS) + - get_pages(sbi, F2FS_DIRTY_IMETA); - unsigned int total_dent_blocks = get_pages(sbi, F2FS_DIRTY_DENTS); + unsigned int total_node_blocks = get_nr_caches(sbi, F2FS_DIRTY_NODES) + + get_nr_caches(sbi, F2FS_DIRTY_DENTS) + + get_nr_caches(sbi, F2FS_DIRTY_IMETA); + unsigned int total_dent_blocks = get_nr_caches(sbi, F2FS_DIRTY_DENTS); unsigned int total_data_blocks = 0; bool separate_dent = true; if (f2fs_lfs_mode(sbi)) - total_data_blocks = get_pages(sbi, F2FS_DIRTY_DATA); + total_data_blocks = get_nr_caches(sbi, F2FS_DIRTY_DATA); /* * When active_logs != 4, dentry blocks and data blocks can be @@ -869,7 +874,7 @@ static inline pgoff_t current_sit_addr(struct f2fs_sb_info *sbi, unsigned int start) { struct sit_info *sit_i = SIT_I(sbi); - unsigned int offset = SIT_BLOCK_OFFSET(start); + unsigned int offset = SIT_BLOCK_OFFSET(sbi, start); block_t blk_addr = sit_i->sit_base_addr + offset; f2fs_bug_on(sbi, !valid_main_segno(sbi, start)); @@ -894,9 +899,10 @@ static inline pgoff_t next_sit_addr(struct f2fs_sb_info *sbi, return block_addr + sit_i->sit_base_addr; } -static inline void set_to_next_sit(struct sit_info *sit_i, unsigned int start) +static inline void set_to_next_sit(struct f2fs_sb_info *sbi, + struct sit_info *sit_i, unsigned int start) { - unsigned int block_off = SIT_BLOCK_OFFSET(start); + unsigned int block_off = SIT_BLOCK_OFFSET(sbi, start); f2fs_change_bit(block_off, sit_i->sit_bitmap); } @@ -971,13 +977,13 @@ static inline bool sec_usage_check(struct f2fs_sb_info *sbi, unsigned int secno) } /* - * It is very important to gather dirty pages and write at once, so that we can + * It is very important to gather dirty blocks and write at once, so that we can * submit a big bio without interfering other data writes. - * By default, 512 pages for directory data, - * 512 pages (2MB) * 8 for nodes, and - * 256 pages * 8 for meta are set. + * By default, 512 blocks for directory data, + * 512 blocks (2MB) * 8 for nodes, and + * 256 blocks * 8 for meta are set. */ -static inline int nr_pages_to_skip(struct f2fs_sb_info *sbi, int type) +static inline int nr_caches_to_skip(struct f2fs_sb_info *sbi, int type) { if (bdi_wb_dirty_exceeded(sbi->sb->s_bdi)) return 0; @@ -993,23 +999,24 @@ static inline int nr_pages_to_skip(struct f2fs_sb_info *sbi, int type) } /* - * When writing pages, it'd better align nr_to_write for segment size. + * When writing cache asynchronously, align nr_to_write to BIO_MAX_VECS. */ -static inline long nr_pages_to_write(struct f2fs_sb_info *sbi, int type, - struct writeback_control *wbc) +static inline long adjust_flush_cache_number(struct f2fs_sb_info *sbi, int type) { - long nr_to_write, desired; - - if (wbc->sync_mode != WB_SYNC_NONE) + long nr_to_write; + + switch (type) { + case META: + nr_to_write = BIO_MAX_VECS; + break; + case NODE: + nr_to_write = BIO_MAX_VECS << 1; + break; + default: + f2fs_bug_on(sbi, 1); return 0; - - nr_to_write = wbc->nr_to_write; - desired = BIO_MAX_VECS; - if (type == NODE) - desired <<= 1; - - wbc->nr_to_write = desired; - return desired - nr_to_write; + } + return nr_to_write; } static inline void wake_up_discard_thread(struct f2fs_sb_info *sbi, bool force) diff --git a/fs/f2fs/shrinker.c b/fs/f2fs/shrinker.c index 4f6bf5926de4..29f488531492 100644 --- a/fs/f2fs/shrinker.c +++ b/fs/f2fs/shrinker.c @@ -23,7 +23,7 @@ static unsigned long __count_nat_entries(struct f2fs_sb_info *sbi) static unsigned long __count_free_nids(struct f2fs_sb_info *sbi) { - long count = NM_I(sbi)->nid_cnt[FREE_NID] - MAX_FREE_NIDS; + long count = NM_I(sbi)->nid_cnt[FREE_NID] - MAX_FREE_NIDS(sbi); return count > 0 ? count : 0; } @@ -37,6 +37,13 @@ static unsigned long __count_extent_cache(struct f2fs_sb_info *sbi, atomic_read(&eti->total_ext_node); } +static unsigned long __count_cache(struct f2fs_sb_info *sbi) +{ + return META_CACHE(sbi)->num_entries + + NODE_CACHE(sbi)->num_entries + + COMPRESS_CACHE(sbi)->num_entries; +} + unsigned long f2fs_shrink_count(struct shrinker *shrink, struct shrink_control *sc) { @@ -68,6 +75,9 @@ unsigned long f2fs_shrink_count(struct shrinker *shrink, /* count free nids cache entries */ count += __count_free_nids(sbi); + /* count generic cache entries */ + count += __count_cache(sbi); + spin_lock(&f2fs_list_lock); p = p->next; mutex_unlock(&sbi->umount_mutex); @@ -120,6 +130,10 @@ unsigned long f2fs_shrink_scan(struct shrinker *shrink, if (freed < nr) freed += f2fs_try_to_free_nids(sbi, nr - freed); + /* shrink generic cache entries */ + if (freed < nr) + freed += f2fs_shrink_cache(sbi, nr - freed); + spin_lock(&f2fs_list_lock); p = p->next; list_move_tail(&sbi->s_list, &f2fs_list); diff --git a/fs/f2fs/super.c b/fs/f2fs/super.c index 1314b6ccced9..294f6f2c28a5 100644 --- a/fs/f2fs/super.c +++ b/fs/f2fs/super.c @@ -855,9 +855,11 @@ static int f2fs_parse_param(struct fs_context *fc, struct fs_parameter *param) break; case Opt_inline_xattr_size: if (result.int_32 < MIN_INLINE_XATTR_SIZE || - result.int_32 > MAX_INLINE_XATTR_SIZE) { + result.int_32 > + MAX_INLINE_XATTR_SIZE(F2FS_MAX_BLKSIZE)) { f2fs_err(NULL, "inline xattr size is out of range: %u ~ %u", - (u32)MIN_INLINE_XATTR_SIZE, (u32)MAX_INLINE_XATTR_SIZE); + (u32)MIN_INLINE_XATTR_SIZE, + (u32)MAX_INLINE_XATTR_SIZE(F2FS_MAX_BLKSIZE)); return -EINVAL; } ctx_set_opt(ctx, F2FS_MOUNT_INLINE_XATTR_SIZE); @@ -1596,6 +1598,8 @@ static int f2fs_check_opt_consistency(struct fs_context *fc, } if (ctx_test_opt(ctx, F2FS_MOUNT_INLINE_XATTR_SIZE)) { + int min_size, max_size; + if (!f2fs_sb_has_extra_attr(sbi) || !f2fs_sb_has_flexible_inline_xattr(sbi)) { f2fs_err(sbi, "extra_attr or flexible_inline_xattr feature is off"); @@ -1605,6 +1609,15 @@ static int f2fs_check_opt_consistency(struct fs_context *fc, f2fs_err(sbi, "inline_xattr_size option should be set with inline_xattr option"); return -EINVAL; } + min_size = MIN_INLINE_XATTR_SIZE; + max_size = MAX_INLINE_XATTR_SIZE(F2FS_BLKSIZE(sbi)); + + if (F2FS_OPTION(sbi).inline_xattr_size < min_size || + F2FS_OPTION(sbi).inline_xattr_size > max_size) { + f2fs_err(sbi, "inline xattr size is out of range: %d ~ %d", + min_size, max_size); + return -EINVAL; + } } if (ctx_test_opt(ctx, F2FS_MOUNT_ATGC) && @@ -1859,22 +1872,9 @@ static struct inode *f2fs_alloc_inode(struct super_block *sb) static int f2fs_drop_inode(struct inode *inode) { - struct f2fs_sb_info *sbi = F2FS_I_SB(inode); int ret; /* - * during filesystem shutdown, if checkpoint is disabled, - * drop useless meta/node dirty pages. - */ - if (unlikely(is_sbi_flag_set(sbi, SBI_CP_DISABLED))) { - if (inode->i_ino == F2FS_NODE_INO(sbi) || - inode->i_ino == F2FS_META_INO(sbi)) { - trace_f2fs_drop_inode(inode, 1); - return 1; - } - } - - /* * This is to avoid a deadlock condition like below. * writeback_single_inode(inode) * - f2fs_write_data_page @@ -1894,7 +1894,7 @@ static int f2fs_drop_inode(struct inode *inode) f2fs_i_size_write(inode, 0); f2fs_submit_merged_write_cond(F2FS_I_SB(inode), - inode, NULL, 0, DATA); + inode, NULL); truncate_inode_pages_final(inode->i_mapping); if (F2FS_HAS_BLOCKS(inode)) @@ -1930,7 +1930,7 @@ int f2fs_inode_dirtied(struct inode *inode, bool sync) if (sync && list_empty(&F2FS_I(inode)->gdirty_list)) { list_add_tail(&F2FS_I(inode)->gdirty_list, &sbi->inode_list[DIRTY_META]); - inc_page_count(sbi, F2FS_DIRTY_IMETA); + inc_cache_count(sbi, F2FS_DIRTY_IMETA); } spin_unlock(&sbi->inode_lock[DIRTY_META]); @@ -1953,7 +1953,7 @@ void f2fs_inode_synced(struct inode *inode) } if (!list_empty(&F2FS_I(inode)->gdirty_list)) { list_del_init(&F2FS_I(inode)->gdirty_list); - dec_page_count(sbi, F2FS_DIRTY_IMETA); + dec_cache_count(sbi, F2FS_DIRTY_IMETA); } clear_inode_flag(inode, FI_DIRTY_INODE); clear_inode_flag(inode, FI_AUTO_RECOVER); @@ -1968,12 +1968,6 @@ void f2fs_inode_synced(struct inode *inode) */ static void f2fs_dirty_inode(struct inode *inode, int flags) { - struct f2fs_sb_info *sbi = F2FS_I_SB(inode); - - if (inode->i_ino == F2FS_NODE_INO(sbi) || - inode->i_ino == F2FS_META_INO(sbi)) - return; - if (is_inode_flag_set(inode, FI_AUTO_RECOVER)) clear_inode_flag(inode, FI_AUTO_RECOVER); @@ -2026,6 +2020,7 @@ static void f2fs_put_super(struct super_block *sb) * flush all issued checkpoints and stop checkpoint issue thread. * after then, all checkpoints should be done by each process context. */ + f2fs_stop_cache_wb_thread(sbi); f2fs_stop_ckpt_thread(sbi); /* @@ -2064,30 +2059,27 @@ static void f2fs_put_super(struct super_block *sb) /* our cp_error case, we can wait for any writeback page */ f2fs_flush_merged_writes(sbi); - f2fs_wait_on_all_pages(sbi, F2FS_WB_CP_DATA); + f2fs_sync_dirty_data(sbi, F2FS_WB_CP_DATA); - if (err || f2fs_cp_error(sbi)) { - truncate_inode_pages_final(NODE_MAPPING(sbi)); - truncate_inode_pages_final(META_MAPPING(sbi)); + if (err || f2fs_cp_error(sbi) || + unlikely(is_sbi_flag_set(sbi, SBI_CP_DISABLED))) { + f2fs_truncate_node_caches(sbi, 0, ULONG_MAX); + f2fs_truncate_meta_caches(sbi, 0, ULONG_MAX); } f2fs_bug_on(sbi, sbi->fsync_node_num); - f2fs_destroy_compress_inode(sbi); - - iput(sbi->node_inode); - sbi->node_inode = NULL; - - iput(sbi->meta_inode); - sbi->meta_inode = NULL; + f2fs_destroy_cache(COMPRESS_CACHE(sbi)); + f2fs_destroy_cache(NODE_CACHE(sbi)); + f2fs_destroy_cache(META_CACHE(sbi)); /* Should check the page counts after dropping all node/meta pages */ for (i = 0; i < NR_COUNT_TYPE; i++) { - if (!get_pages(sbi, i)) + if (!get_nr_caches(sbi, i)) continue; f2fs_err(sbi, "detect filesystem reference count leak during " "umount, type: %d, count: %lld, err: %d, cp_err: %d", - i, get_pages(sbi, i), err, f2fs_cp_error(sbi)); + i, get_nr_caches(sbi, i), err, f2fs_cp_error(sbi)); f2fs_bug_on(sbi, 1); } @@ -2701,9 +2693,9 @@ static int f2fs_disable_checkpoint(struct f2fs_sb_info *sbi) f2fs_info(sbi, "%s: call sync_filesystem() to persist meta: %lld, node: %lld, data: %lld", __func__, - get_pages(sbi, F2FS_DIRTY_META), - get_pages(sbi, F2FS_DIRTY_NODES), - get_pages(sbi, F2FS_DIRTY_DATA)); + get_nr_caches(sbi, F2FS_DIRTY_META), + get_nr_caches(sbi, F2FS_DIRTY_NODES), + get_nr_caches(sbi, F2FS_DIRTY_DATA)); ret = sync_filesystem(sbi->sb); if (ret || err) { @@ -2720,9 +2712,9 @@ static int f2fs_disable_checkpoint(struct f2fs_sb_info *sbi) skip_gc: f2fs_info(sbi, "%s: call f2fs_write_checkpoint(), meta: %lld, node: %lld, data: %lld", __func__, - get_pages(sbi, F2FS_DIRTY_META), - get_pages(sbi, F2FS_DIRTY_NODES), - get_pages(sbi, F2FS_DIRTY_DATA)); + get_nr_caches(sbi, F2FS_DIRTY_META), + get_nr_caches(sbi, F2FS_DIRTY_NODES), + get_nr_caches(sbi, F2FS_DIRTY_DATA)); f2fs_down_write_trace(&sbi->gc_lock, &lc); cpc.reason = CP_PAUSE; @@ -2754,9 +2746,9 @@ static int f2fs_enable_checkpoint(struct f2fs_sb_info *sbi) long long skipped_write, dirty_data; f2fs_info(sbi, "f2fs_enable_checkpoint() starts, meta: %lld, node: %lld, data: %lld", - get_pages(sbi, F2FS_DIRTY_META), - get_pages(sbi, F2FS_DIRTY_NODES), - get_pages(sbi, F2FS_DIRTY_DATA)); + get_nr_caches(sbi, F2FS_DIRTY_META), + get_nr_caches(sbi, F2FS_DIRTY_NODES), + get_nr_caches(sbi, F2FS_DIRTY_DATA)); start = ktime_get(); @@ -2764,26 +2756,26 @@ static int f2fs_enable_checkpoint(struct f2fs_sb_info *sbi) /* we should flush all the data to keep data consistency */ do { - skipped_write = get_pages(sbi, F2FS_SKIPPED_WRITE); - dirty_data = get_pages(sbi, F2FS_DIRTY_DATA); + skipped_write = get_nr_caches(sbi, F2FS_SKIPPED_WRITE); + dirty_data = get_nr_caches(sbi, F2FS_DIRTY_DATA); sync_inodes_sb(sbi->sb); f2fs_io_schedule_timeout(DEFAULT_SCHEDULE_TIMEOUT); f2fs_info(sbi, "sync_inode_sb done, dirty_data: %lld, %lld, " "skipped write: %lld, %lld, retry: %d", - get_pages(sbi, F2FS_DIRTY_DATA), + get_nr_caches(sbi, F2FS_DIRTY_DATA), dirty_data, - get_pages(sbi, F2FS_SKIPPED_WRITE), + get_nr_caches(sbi, F2FS_SKIPPED_WRITE), skipped_write, retry); /* * sync_inodes_sb() has retry logic, so let's check dirty_data * in prior to skipped_write in case there is no dirty data. */ - if (!get_pages(sbi, F2FS_DIRTY_DATA)) + if (!get_nr_caches(sbi, F2FS_DIRTY_DATA)) break; - if (get_pages(sbi, F2FS_SKIPPED_WRITE) == skipped_write) + if (get_nr_caches(sbi, F2FS_SKIPPED_WRITE) == skipped_write) break; } while (retry--); @@ -2791,14 +2783,14 @@ static int f2fs_enable_checkpoint(struct f2fs_sb_info *sbi) writeback = ktime_get(); - if (unlikely(get_pages(sbi, F2FS_DIRTY_DATA) || - get_pages(sbi, F2FS_SKIPPED_WRITE))) + if (unlikely(get_nr_caches(sbi, F2FS_DIRTY_DATA) || + get_nr_caches(sbi, F2FS_SKIPPED_WRITE))) f2fs_warn(sbi, "checkpoint=enable unwritten data: %lld, skipped data: %lld, retry: %d", - get_pages(sbi, F2FS_DIRTY_DATA), - get_pages(sbi, F2FS_SKIPPED_WRITE), retry); + get_nr_caches(sbi, F2FS_DIRTY_DATA), + get_nr_caches(sbi, F2FS_SKIPPED_WRITE), retry); - if (get_pages(sbi, F2FS_SKIPPED_WRITE)) - atomic_set(&sbi->nr_pages[F2FS_SKIPPED_WRITE], 0); + if (get_nr_caches(sbi, F2FS_SKIPPED_WRITE)) + atomic_set(&sbi->nr_caches[F2FS_SKIPPED_WRITE], 0); f2fs_down_write_trace(&sbi->gc_lock, &lc); f2fs_dirty_to_prefree(sbi); @@ -2830,6 +2822,7 @@ static int __f2fs_remount(struct fs_context *fc, struct super_block *sb) unsigned int flags = fc->sb_flags; int err; bool need_restart_gc = false, need_stop_gc = false; + bool need_restart_wb = false, need_stop_wb = false; bool need_restart_flush = false, need_stop_flush = false; bool need_restart_discard = false, need_stop_discard = false; bool need_enable_checkpoint = false, need_disable_checkpoint = false; @@ -2864,7 +2857,8 @@ static int __f2fs_remount(struct fs_context *fc, struct super_block *sb) if (!org_mount_opt.s_qf_names[i]) { for (j = 0; j < i; j++) kfree(org_mount_opt.s_qf_names[j]); - return -ENOMEM; + err = -ENOMEM; + goto restore_holder; } } else { org_mount_opt.s_qf_names[i] = NULL; @@ -2989,13 +2983,25 @@ static int __f2fs_remount(struct fs_context *fc, struct super_block *sb) } if (flags & SB_RDONLY) { + if (sbi->cache_thread.cache_wb_task) { + f2fs_stop_cache_wb_thread(sbi); + need_restart_wb = true; + } + } else if (!sbi->cache_thread.cache_wb_task) { + err = f2fs_start_cache_wb_thread(sbi); + if (err) + goto restore_gc; + need_stop_wb = true; + } + + if (flags & SB_RDONLY) { sync_inodes_sb(sb); set_sbi_flag(sbi, SBI_IS_DIRTY); set_sbi_flag(sbi, SBI_IS_CLOSE); err = f2fs_sync_fs(sb, 1); if (err) - goto restore_gc; + goto restore_wb; clear_sbi_flag(sbi, SBI_IS_CLOSE); } @@ -3010,7 +3016,7 @@ static int __f2fs_remount(struct fs_context *fc, struct super_block *sb) } else { err = f2fs_create_flush_cmd_control(sbi); if (err) - goto restore_gc; + goto restore_wb; need_stop_flush = true; } @@ -3107,6 +3113,13 @@ restore_flush: clear_opt(sbi, FLUSH_MERGE); f2fs_destroy_flush_cmd_control(sbi, false); } +restore_wb: + if (need_restart_wb) { + if (f2fs_start_cache_wb_thread(sbi)) + f2fs_warn(sbi, "background cache writeback thread has stopped"); + } else if (need_stop_wb) { + f2fs_stop_cache_wb_thread(sbi); + } restore_gc: if (need_restart_gc) { if (f2fs_start_gc_thread(sbi)) @@ -3125,6 +3138,7 @@ restore_opts: sbi->mount_opt = org_mount_opt; sb->s_flags = old_sb_flags; +restore_holder: sbi->umount_lock_holder = NULL; return err; } @@ -3871,13 +3885,13 @@ static const struct export_operations f2fs_export_ops = { .get_parent = f2fs_get_parent, }; -loff_t max_file_blocks(struct inode *inode) +loff_t max_file_blocks(struct f2fs_sb_info *sbi, struct inode *inode) { loff_t result = 0; loff_t leaf_count; /* - * note: previously, result is equal to (DEF_ADDRS_PER_INODE - + * note: previously, result is equal to (DEF_ADDRS_PER_INODE(sbi) - * DEFAULT_INLINE_XATTR_ADDRS), but now f2fs try to reserve more * space in inode.i_addr, it will be more safe to reassign * result as zero. @@ -3886,17 +3900,17 @@ loff_t max_file_blocks(struct inode *inode) if (inode && f2fs_compressed_file(inode)) leaf_count = ADDRS_PER_BLOCK(inode); else - leaf_count = DEF_ADDRS_PER_BLOCK; + leaf_count = DEF_ADDRS_PER_BLOCK(sbi); /* two direct node blocks */ result += (leaf_count * 2); /* two indirect node blocks */ - leaf_count *= NIDS_PER_BLOCK; + leaf_count *= NIDS_PER_BLOCK(sbi); result += (leaf_count * 2); /* one double indirect node block */ - leaf_count *= NIDS_PER_BLOCK; + leaf_count *= NIDS_PER_BLOCK(sbi); result += leaf_count; /* @@ -3905,7 +3919,8 @@ loff_t max_file_blocks(struct inode *inode) * fit within U32_MAX + 1 data units. */ - result = umin(result, F2FS_BYTES_TO_BLK(((loff_t)U32_MAX + 1) * 4096)); + result = umin(result, F2FS_BYTES_TO_BLK(sbi, + ((loff_t)U32_MAX + 1) * 4096)); return result; } @@ -3931,7 +3946,7 @@ static int __f2fs_commit_super(struct f2fs_sb_info *sbi, struct folio *folio, bio = bio_alloc(sbi->sb->s_bdev, 1, opf, GFP_NOFS); /* it doesn't need to set crypto context for superblock update */ - bio->bi_iter.bi_sector = SECTOR_FROM_BLOCK(folio->index); + bio->bi_iter.bi_sector = SECTOR_FROM_BLOCK(sbi, folio->index); if (!bio_add_folio(bio, folio, folio_size(folio), 0)) f2fs_bug_on(sbi, 1); @@ -4065,10 +4080,10 @@ static int sanity_check_raw_super(struct f2fs_sb_info *sbi, } /* only support block_size equals to PAGE_SIZE */ - if (le32_to_cpu(raw_super->log_blocksize) != F2FS_BLKSIZE_BITS) { + if (le32_to_cpu(raw_super->log_blocksize) != PAGE_SHIFT) { f2fs_info(sbi, "Invalid log_blocksize (%u), supports only %u", le32_to_cpu(raw_super->log_blocksize), - F2FS_BLKSIZE_BITS); + PAGE_SHIFT); return -EFSCORRUPTED; } @@ -4329,7 +4344,7 @@ skip_cross: return 1; } - sit_blk_cnt = DIV_ROUND_UP(main_segs, SIT_ENTRY_PER_BLOCK); + sit_blk_cnt = DIV_ROUND_UP(main_segs, SIT_ENTRY_PER_BLOCK(sbi)); if (sit_bitmap_size * 8 < sit_blk_cnt) { f2fs_err(sbi, "Wrong bitmap size: sit: %u, sit_blk_cnt:%u", sit_bitmap_size, sit_blk_cnt); @@ -4357,7 +4372,7 @@ skip_cross: nat_blocks = nat_segs << log_blocks_per_seg; nat_bits_bytes = nat_blocks / BITS_PER_BYTE; - nat_bits_blocks = F2FS_BLK_ALIGN((nat_bits_bytes << 1) + 8); + nat_bits_blocks = F2FS_BLK_ALIGN(sbi, (nat_bits_bytes << 1) + 8); if (__is_set_ckpt_flags(ckpt, CP_NAT_BITS_FLAG) && (cp_payload + F2FS_CP_PACKS + NR_CURSEG_PERSIST_TYPE + nat_bits_blocks >= blocks_per_seg)) { @@ -4382,6 +4397,22 @@ static void init_sb_info(struct f2fs_sb_info *sbi) le32_to_cpu(raw_super->log_sectors_per_block); sbi->log_blocksize = le32_to_cpu(raw_super->log_blocksize); sbi->blocksize = BIT(sbi->log_blocksize); + sbi->nat_entries_per_block = sbi->blocksize / + sizeof(struct f2fs_nat_entry); + sbi->addrs_per_inode = F2FS_DEF_ADDRS_PER_INODE(sbi->blocksize); + sbi->addrs_per_block = (sbi->blocksize - + sizeof(struct node_footer)) / sizeof(__le32); + sbi->nids_per_block = sbi->addrs_per_block; + sbi->sit_entries_per_block = sbi->blocksize / + sizeof(struct f2fs_sit_entry); + sbi->orphans_per_block = (sbi->blocksize - + sizeof(struct f2fs_orphan_footer)) / sizeof(__le32); + sbi->dentries_per_block = (BITS_PER_BYTE * sbi->blocksize) / + ((SIZE_OF_DIR_ENTRY + F2FS_SLOT_LEN) * BITS_PER_BYTE + 1); + sbi->dentry_bitmap_size = DIV_ROUND_UP(sbi->dentries_per_block, + BITS_PER_BYTE); + sbi->dentry_reserved_size = sbi->blocksize - sbi->dentry_bitmap_size - + (SIZE_OF_DIR_ENTRY + F2FS_SLOT_LEN) * sbi->dentries_per_block; sbi->log_blocks_per_seg = le32_to_cpu(raw_super->log_blocks_per_seg); sbi->blocks_per_seg = BIT(sbi->log_blocks_per_seg); sbi->segs_per_sec = le32_to_cpu(raw_super->segs_per_sec); @@ -4389,12 +4420,10 @@ static void init_sb_info(struct f2fs_sb_info *sbi) sbi->total_sections = le32_to_cpu(raw_super->section_count); sbi->total_node_count = SEGS_TO_BLKS(sbi, ((le32_to_cpu(raw_super->segment_count_nat) / 2) * - NAT_ENTRY_PER_BLOCK)); + NAT_ENTRY_PER_BLOCK(sbi))); sbi->allocate_section_hint = le32_to_cpu(raw_super->section_count); sbi->allocate_section_policy = ALLOCATE_FORWARD_NOHINT; F2FS_ROOT_INO(sbi) = le32_to_cpu(raw_super->root_ino); - F2FS_NODE_INO(sbi) = le32_to_cpu(raw_super->node_ino); - F2FS_META_INO(sbi) = le32_to_cpu(raw_super->meta_ino); sbi->cur_victim_sec = NULL_SECNO; sbi->gc_mode = GC_NORMAL; sbi->next_victim_seg[BG_GC] = NULL_SEGNO; @@ -4436,7 +4465,7 @@ static void init_sb_info(struct f2fs_sb_info *sbi) clear_sbi_flag(sbi, SBI_NEED_FSCK); for (i = 0; i < NR_COUNT_TYPE; i++) - atomic_set(&sbi->nr_pages[i], 0); + atomic_set(&sbi->nr_caches[i], 0); for (i = 0; i < META; i++) atomic_set(&sbi->wb_sync_req[i], 0); @@ -4490,7 +4519,7 @@ static int f2fs_report_zone_cb(struct blk_zone *zone, unsigned int idx, { struct f2fs_report_zones_args *rz_args = data; block_t unusable_blocks = (zone->len - zone->capacity) >> - F2FS_LOG_SECTORS_PER_BLOCK; + F2FS_LOG_SECTORS_PER_BLOCK(rz_args->sbi); if (zone->type == BLK_ZONE_TYPE_CONVENTIONAL) return 0; @@ -4533,10 +4562,10 @@ static int init_blkz_info(struct f2fs_sb_info *sbi, int devi) zone_sectors = bdev_zone_sectors(bdev); if (sbi->blocks_per_blkz && sbi->blocks_per_blkz != - SECTOR_TO_BLOCK(zone_sectors)) + SECTOR_TO_BLOCK(sbi, zone_sectors)) return -EINVAL; - sbi->blocks_per_blkz = SECTOR_TO_BLOCK(zone_sectors); - FDEV(devi).nr_blkz = div_u64(SECTOR_TO_BLOCK(nr_sectors), + sbi->blocks_per_blkz = SECTOR_TO_BLOCK(sbi, zone_sectors); + FDEV(devi).nr_blkz = div_u64(SECTOR_TO_BLOCK(sbi, nr_sectors), sbi->blocks_per_blkz); if (nr_sectors & (zone_sectors - 1)) FDEV(devi).nr_blkz++; @@ -5045,7 +5074,7 @@ static void f2fs_restore_device_alias(struct f2fs_sb_info *sbi) { struct inode *root = d_inode(sbi->sb->s_root); struct f2fs_dir_entry *de; - struct folio *folio; + void *dentry_block = NULL; int i; if (!f2fs_sb_has_device_alias(sbi)) @@ -5060,7 +5089,7 @@ static void f2fs_restore_device_alias(struct f2fs_sb_info *sbi) qstr.name = name; qstr.len = strlen(name); - de = f2fs_find_entry(root, &qstr, &folio); + de = f2fs_find_entry(root, &qstr, &dentry_block); if (!de) continue; @@ -5070,7 +5099,7 @@ static void f2fs_restore_device_alias(struct f2fs_sb_info *sbi) FDEV(i).has_alias = true; iput(inode); } - f2fs_folio_put(folio, 0); + f2fs_put_dentry_block(dentry_block, false); } } @@ -5125,12 +5154,6 @@ try_onemore: } mutex_init(&sbi->flush_lock); - /* set a block size */ - if (unlikely(!sb_set_blocksize(sb, F2FS_BLKSIZE))) { - f2fs_err(sbi, "unable to set blocksize"); - goto free_sbi; - } - err = read_raw_super_block(sbi, &raw_super, &valid_super_block, &recovery); if (err) @@ -5138,6 +5161,14 @@ try_onemore: sb->s_fs_info = sbi; sbi->raw_super = raw_super; + init_sb_info(sbi); + + /* set a block size */ + if (unlikely(!sb_set_blocksize(sb, sbi->blocksize))) { + f2fs_err(sbi, "unable to set blocksize %u", sbi->blocksize); + err = -EINVAL; + goto free_sb_buf; + } sbi->max_atc_write_bio_size = UINT_MAX; INIT_WORK(&sbi->s_error_work, f2fs_record_error_work); @@ -5161,7 +5192,7 @@ try_onemore: if (err) goto free_options; - sb->s_maxbytes = max_file_blocks(NULL) << + sb->s_maxbytes = max_file_blocks(sbi, NULL) << le32_to_cpu(raw_super->log_blocksize); sb->s_max_links = F2FS_LINK_MAX; @@ -5217,8 +5248,6 @@ try_onemore: if (err) goto free_bio_info; - init_sb_info(sbi); - err = f2fs_init_iostat(sbi); if (err) goto free_bio_info; @@ -5231,18 +5260,14 @@ try_onemore: if (err) goto free_percpu; - /* get an inode for meta space */ - sbi->meta_inode = f2fs_iget(sb, F2FS_META_INO(sbi)); - if (IS_ERR(sbi->meta_inode)) { - f2fs_err(sbi, "Failed to read F2FS meta data inode"); - err = PTR_ERR(sbi->meta_inode); - goto free_page_array_cache; - } + f2fs_init_cache(sbi, META_CACHE(sbi), F2FS_META_CACHE); + f2fs_init_cache(sbi, NODE_CACHE(sbi), F2FS_NODE_CACHE); + f2fs_init_cache(sbi, COMPRESS_CACHE(sbi), F2FS_COMPRESS_CACHE); err = f2fs_get_valid_checkpoint(sbi); if (err) { f2fs_err(sbi, "Failed to get valid F2FS checkpoint"); - goto free_meta_inode; + goto free_compress_cache; } if (__is_set_ckpt_flags(F2FS_CKPT(sbi), CP_QUOTA_NEED_FSCK_FLAG)) @@ -5288,6 +5313,8 @@ try_onemore: f2fs_init_fsync_node_info(sbi); + f2fs_init_compress_cache_context(sbi); + /* setup checkpoint request control and start checkpoint issue thread */ f2fs_init_ckpt_req_control(sbi); if (!f2fs_readonly(sb) && !test_opt(sbi, DISABLE_CHECKPOINT) && @@ -5339,42 +5366,30 @@ try_onemore: if (err) goto free_nm; - /* get an inode for node space */ - sbi->node_inode = f2fs_iget(sb, F2FS_NODE_INO(sbi)); - if (IS_ERR(sbi->node_inode)) { - f2fs_err(sbi, "Failed to read node inode"); - err = PTR_ERR(sbi->node_inode); - goto free_stats; - } - /* read root inode and dentry */ root = f2fs_iget(sb, F2FS_ROOT_INO(sbi)); if (IS_ERR(root)) { f2fs_err(sbi, "Failed to read root inode"); err = PTR_ERR(root); - goto free_node_inode; + goto free_ino_entry; } if (!S_ISDIR(root->i_mode) || !root->i_blocks || !root->i_size || !root->i_nlink) { iput(root); err = -EINVAL; - goto free_node_inode; + goto free_ino_entry; } generic_set_sb_d_ops(sb); sb->s_root = d_make_root(root); /* allocate root dentry */ if (!sb->s_root) { err = -ENOMEM; - goto free_node_inode; + goto free_ino_entry; } - err = f2fs_init_compress_inode(sbi); - if (err) - goto free_root_inode; - err = f2fs_register_sysfs(sbi); if (err) - goto free_compress_inode; + goto free_root_inode; sbi->umount_lock_holder = current; #ifdef CONFIG_QUOTA @@ -5491,6 +5506,12 @@ reset_checkpoint: goto sync_free_meta; } + if (!f2fs_readonly(sb)) { + err = f2fs_start_cache_wb_thread(sbi); + if (err) + goto stop_gc_thread; + } + /* recover broken superblock */ if (recovery) { err = f2fs_commit_super(sbi, true); @@ -5513,6 +5534,8 @@ reset_checkpoint: sbi->umount_lock_holder = NULL; return 0; +stop_gc_thread: + f2fs_stop_gc_thread(sbi); sync_free_meta: /* safe to flush all the data */ sync_filesystem(sbi->sb); @@ -5528,23 +5551,18 @@ free_meta: * Some dirty meta pages can be produced by f2fs_recover_orphan_inodes() * failed by EIO. Then, iput(node_inode) can trigger balance_fs_bg() * followed by f2fs_write_checkpoint() through f2fs_write_node_pages(), which - * falls into an infinite loop in f2fs_sync_meta_pages(). + * falls into an infinite loop in f2fs_sync_meta_caches(). */ - truncate_inode_pages_final(META_MAPPING(sbi)); + f2fs_truncate_meta_caches(sbi, 0, ULONG_MAX); /* evict some inodes being cached by GC */ evict_inodes(sb); f2fs_unregister_sysfs(sbi); -free_compress_inode: - f2fs_destroy_compress_inode(sbi); free_root_inode: dput(sb->s_root); sb->s_root = NULL; -free_node_inode: +free_ino_entry: f2fs_release_ino_entry(sbi, true); - truncate_inode_pages_final(NODE_MAPPING(sbi)); - iput(sbi->node_inode); - sbi->node_inode = NULL; -free_stats: + f2fs_truncate_node_caches(sbi, 0, ULONG_MAX); f2fs_destroy_stats(sbi); free_nm: /* stop discard thread before destroying node manager */ @@ -5560,11 +5578,10 @@ stop_ckpt_thread: free_devices: destroy_device_list(sbi); kvfree(sbi->ckpt); -free_meta_inode: - make_bad_inode(sbi->meta_inode); - iput(sbi->meta_inode); - sbi->meta_inode = NULL; -free_page_array_cache: +free_compress_cache: + f2fs_destroy_cache(COMPRESS_CACHE(sbi)); + f2fs_destroy_cache(NODE_CACHE(sbi)); + f2fs_destroy_cache(META_CACHE(sbi)); f2fs_destroy_page_array_cache(sbi); free_percpu: destroy_percpu_info(sbi); @@ -5653,7 +5670,8 @@ static void kill_f2fs_super(struct super_block *sb) * compress inode cache. */ if (test_opt(sbi, COMPRESS_CACHE)) - truncate_inode_pages_final(COMPRESS_MAPPING(sbi)); + f2fs_invalidate_compress_pages_range(sbi, + 0, UINT_MAX); #endif if (is_sbi_flag_set(sbi, SBI_IS_DIRTY) || diff --git a/fs/f2fs/sysfs.c b/fs/f2fs/sysfs.c index 811e350a1430..9749da70089a 100644 --- a/fs/f2fs/sysfs.c +++ b/fs/f2fs/sysfs.c @@ -40,6 +40,7 @@ enum { RESERVED_BLOCKS, /* struct f2fs_sb_info */ CPRC_INFO, /* struct ckpt_req_control */ ATGC_INFO, /* struct atgc_management */ + WB_THREAD, /* struct f2fs_cache_kthread */ }; static const char *gc_mode_names[MAX_GC_MODE] = { @@ -98,6 +99,8 @@ static unsigned char *__struct_ptr(struct f2fs_sb_info *sbi, int struct_type) return (unsigned char *)&sbi->cprc_info; else if (struct_type == ATGC_INFO) return (unsigned char *)&sbi->am; + else if (struct_type == WB_THREAD) + return (unsigned char *)&sbi->cache_thread; return NULL; } @@ -995,6 +998,13 @@ out: return count; } + if (!strcmp(a->attr.name, "cache_wb_interval")) { + if (t < MIN_DIRTY_CACHE_TIMEOUT || t > MAX_DIRTY_CACHE_TIMEOUT) + return -EINVAL; + sbi->cache_thread.cache_wb_interval = t; + return count; + } + __sbi_store_value(a, sbi, ptr + a->offset, t); return count; @@ -1008,7 +1018,8 @@ static ssize_t f2fs_sbi_store(struct f2fs_attr *a, bool gc_entry = (!strcmp(a->attr.name, "gc_urgent") || a->struct_type == GC_THREAD); bool thread_entry = !strcmp(a->attr.name, "ckpt_thread_ioprio") || - !strcmp(a->attr.name, "critical_task_priority"); + !strcmp(a->attr.name, "critical_task_priority") || + !strcmp(a->attr.name, "cache_wb_interval"); if (gc_entry || thread_entry) { if (!down_read_trylock(&sbi->sb->s_umount)) @@ -1218,6 +1229,9 @@ static struct f2fs_attr f2fs_attr_##name = __ATTR(name, 0444, name##_show, NULL) #define ATGC_INFO_RW_ATTR(name, elname) \ F2FS_RW_ATTR(ATGC_INFO, atgc_management, name, elname) +#define WB_THREAD_RW_ATTR(name, elname) \ + F2FS_RW_ATTR(WB_THREAD, f2fs_cache_kthread, name, elname) + /* GC_THREAD ATTR */ GC_THREAD_RW_ATTR(gc_urgent_sleep_time, urgent_sleep_time); GC_THREAD_RW_ATTR(gc_min_sleep_time, min_sleep_time); @@ -1254,7 +1268,7 @@ DCC_INFO_GENERAL_RW_ATTR(discard_io_aware); /* NM_INFO ATTR */ NM_INFO_RW_ATTR(max_roll_forward_node_blocks, max_rf_node_blocks); NM_INFO_GENERAL_RW_ATTR(ram_thresh); -NM_INFO_GENERAL_RW_ATTR(ra_nid_pages); +NM_INFO_RW_ATTR(ra_nid_pages, ra_nid_blocks); NM_INFO_GENERAL_RW_ATTR(dirty_nats_ratio); /* F2FS_SBI ATTR */ @@ -1347,6 +1361,9 @@ ATGC_INFO_RW_ATTR(atgc_candidate_count, max_candidate_count); ATGC_INFO_RW_ATTR(atgc_age_weight, age_weight); ATGC_INFO_RW_ATTR(atgc_age_threshold, age_threshold); +/* WB_THREAD ATTR */ +WB_THREAD_RW_ATTR(cache_wb_interval, cache_wb_interval); + F2FS_GENERAL_RO_ATTR(dirty_segments); F2FS_GENERAL_RO_ATTR(free_segments); F2FS_GENERAL_RO_ATTR(ovp_segments); @@ -1533,6 +1550,7 @@ static struct attribute *f2fs_attrs[] = { ATTR_LIST(lock_duration_priority), ATTR_LIST(adjust_lock_priority), ATTR_LIST(critical_task_priority), + ATTR_LIST(cache_wb_interval), NULL, }; ATTRIBUTE_GROUPS(f2fs); @@ -1871,8 +1889,8 @@ static int __maybe_unused disk_map_seq_show(struct seq_file *seq, struct f2fs_sb_info *sbi = F2FS_SB(sb); int i; - seq_printf(seq, "Address Layout : %5luB Block address (# of Segments)\n", - F2FS_BLKSIZE); + seq_printf(seq, "Address Layout : %5uB Block address (# of Segments)\n", + F2FS_BLKSIZE(sbi)); seq_printf(seq, " SB : %12s\n", "0/1024B"); seq_printf(seq, " seg0_blkaddr : 0x%010x\n", SEG0_BLKADDR(sbi)); seq_printf(seq, " Checkpoint : 0x%010x (%10d)\n", @@ -1889,13 +1907,13 @@ static int __maybe_unused disk_map_seq_show(struct seq_file *seq, seq_printf(seq, " Main : 0x%010x (%10d)\n", SM_I(sbi)->main_blkaddr, le32_to_cpu(F2FS_RAW_SUPER(sbi)->segment_count_main)); - seq_printf(seq, " Block size : %12lu KB\n", F2FS_BLKSIZE >> 10); + seq_printf(seq, " Block size : %12u KB\n", F2FS_BLKSIZE(sbi) >> 10); seq_printf(seq, " Segment size : %12d MB\n", - (BLKS_PER_SEG(sbi) << (F2FS_BLKSIZE_BITS - 10)) >> 10); + (BLKS_PER_SEG(sbi) << (F2FS_BLKSIZE_BITS(sbi) - 10)) >> 10); seq_printf(seq, " Segs/Sections : %12d\n", SEGS_PER_SEC(sbi)); seq_printf(seq, " Section size : %12d MB\n", - (BLKS_PER_SEC(sbi) << (F2FS_BLKSIZE_BITS - 10)) >> 10); + (BLKS_PER_SEC(sbi) << (F2FS_BLKSIZE_BITS(sbi) - 10)) >> 10); seq_printf(seq, " # of Sections : %12d\n", le32_to_cpu(F2FS_RAW_SUPER(sbi)->section_count)); diff --git a/fs/f2fs/verity.c b/fs/f2fs/verity.c index 39f482515445..ee838f880c4a 100644 --- a/fs/f2fs/verity.c +++ b/fs/f2fs/verity.c @@ -75,7 +75,8 @@ static int pagecache_write(struct inode *inode, const void *buf, size_t count, struct address_space *mapping = inode->i_mapping; const struct address_space_operations *aops = mapping->a_ops; - if (pos + count > F2FS_BLK_TO_BYTES(max_file_blocks(inode))) + if (pos + count > F2FS_BLK_TO_BYTES(F2FS_I_SB(inode), + max_file_blocks(F2FS_I_SB(inode), inode))) return -EFBIG; while (count) { @@ -239,7 +240,8 @@ static int f2fs_get_verity_descriptor(struct inode *inode, void *buf, /* Get the descriptor */ if (pos + size < pos || - pos + size > F2FS_BLK_TO_BYTES(max_file_blocks(inode)) || + pos + size > F2FS_BLK_TO_BYTES(F2FS_I_SB(inode), + max_file_blocks(F2FS_I_SB(inode), inode)) || pos < f2fs_verity_metadata_pos(inode) || size > INT_MAX) { f2fs_warn(F2FS_I_SB(inode), "invalid verity xattr"); f2fs_handle_error(F2FS_I_SB(inode), diff --git a/fs/f2fs/xattr.c b/fs/f2fs/xattr.c index 6728d1488cad..0ed879fc3076 100644 --- a/fs/f2fs/xattr.c +++ b/fs/f2fs/xattr.c @@ -67,7 +67,7 @@ static int f2fs_xattr_generic_get(const struct xattr_handler *handler, } static int f2fs_xattr_generic_set(const struct xattr_handler *handler, - struct mnt_idmap *idmap, + const struct mnt_idmap *idmap, struct dentry *unused, struct inode *inode, const char *name, const void *value, size_t size, int flags) @@ -111,7 +111,7 @@ static int f2fs_xattr_advise_get(const struct xattr_handler *handler, } static int f2fs_xattr_advise_set(const struct xattr_handler *handler, - struct mnt_idmap *idmap, + const struct mnt_idmap *idmap, struct dentry *unused, struct inode *inode, const char *name, const void *value, size_t size, int flags) @@ -138,7 +138,7 @@ static int f2fs_xattr_advise_set(const struct xattr_handler *handler, #ifdef CONFIG_F2FS_FS_SECURITY static int f2fs_initxattrs(struct inode *inode, const struct xattr *xattr_array, - void *folio) + void *fs_data) { const struct xattr *xattr; int err = 0; @@ -146,7 +146,7 @@ static int f2fs_initxattrs(struct inode *inode, const struct xattr *xattr_array, for (xattr = xattr_array; xattr->name != NULL; xattr++) { err = f2fs_setxattr(inode, F2FS_XATTR_INDEX_SECURITY, xattr->name, xattr->value, - xattr->value_len, folio, 0); + xattr->value_len, fs_data, 0); if (err < 0) break; } @@ -154,10 +154,10 @@ static int f2fs_initxattrs(struct inode *inode, const struct xattr *xattr_array, } int f2fs_init_security(struct inode *inode, struct inode *dir, - const struct qstr *qstr, struct folio *ifolio) + const struct qstr *qstr, struct f2fs_cached_block *ientry) { return security_inode_init_security(inode, dir, qstr, - f2fs_initxattrs, ifolio); + f2fs_initxattrs, ientry); } #endif @@ -273,25 +273,25 @@ static struct f2fs_xattr_entry *__find_inline_xattr(struct inode *inode, return entry; } -static int read_inline_xattr(struct inode *inode, struct folio *ifolio, +static int read_inline_xattr(struct inode *inode, struct f2fs_cached_block *ientry, void *txattr_addr) { struct f2fs_sb_info *sbi = F2FS_I_SB(inode); unsigned int inline_size = inline_xattr_size(inode); - struct folio *folio = NULL; + struct f2fs_cached_block *in_entry = NULL; void *inline_addr; - if (ifolio) { - inline_addr = inline_xattr_addr(inode, ifolio); + if (ientry) { + inline_addr = inline_xattr_addr(inode, ientry); } else { - folio = f2fs_get_inode_folio(sbi, inode->i_ino); - if (IS_ERR(folio)) - return PTR_ERR(folio); + in_entry = f2fs_get_inode_cache(sbi, inode->i_ino); + if (IS_ERR(in_entry)) + return PTR_ERR(in_entry); - inline_addr = inline_xattr_addr(inode, folio); + inline_addr = inline_xattr_addr(inode, in_entry); } memcpy(txattr_addr, inline_addr, inline_size); - f2fs_folio_put(folio, true); + f2fs_put_cache(in_entry, true); return 0; } @@ -301,22 +301,23 @@ static int read_xattr_block(struct inode *inode, void *txattr_addr) struct f2fs_sb_info *sbi = F2FS_I_SB(inode); nid_t xnid = F2FS_I(inode)->i_xattr_nid; unsigned int inline_size = inline_xattr_size(inode); - struct folio *xfolio; + struct f2fs_cached_block *xentry; void *xattr_addr; /* The inode already has an extended attribute block. */ - xfolio = f2fs_get_xnode_folio(sbi, xnid); - if (IS_ERR(xfolio)) - return PTR_ERR(xfolio); + xentry = f2fs_get_xnode_cache(sbi, xnid); + if (IS_ERR(xentry)) + return PTR_ERR(xentry); - xattr_addr = folio_address(xfolio); - memcpy(txattr_addr + inline_size, xattr_addr, VALID_XATTR_BLOCK_SIZE); - f2fs_folio_put(xfolio, true); + xattr_addr = cache_address(xentry); + memcpy(txattr_addr + inline_size, xattr_addr, + VALID_XATTR_BLOCK_SIZE(inode)); + f2fs_put_cache(xentry, true); return 0; } -static int lookup_all_xattrs(struct inode *inode, struct folio *ifolio, +static int lookup_all_xattrs(struct inode *inode, struct f2fs_cached_block *ientry, unsigned int index, unsigned int len, const char *name, struct f2fs_xattr_entry **xe, void **base_addr, int *base_size, @@ -340,7 +341,7 @@ static int lookup_all_xattrs(struct inode *inode, struct folio *ifolio, /* read from inline xattr */ if (inline_size) { - err = read_inline_xattr(inode, ifolio, txattr_addr); + err = read_inline_xattr(inode, ientry, txattr_addr); if (err) goto out; @@ -388,12 +389,12 @@ out: return err; } -static int read_all_xattrs(struct inode *inode, struct folio *ifolio, +static int read_all_xattrs(struct inode *inode, struct f2fs_cached_block *ientry, void **base_addr) { struct f2fs_xattr_header *header; nid_t xnid = F2FS_I(inode)->i_xattr_nid; - unsigned int size = VALID_XATTR_BLOCK_SIZE; + unsigned int size = VALID_XATTR_BLOCK_SIZE(inode); unsigned int inline_size = inline_xattr_size(inode); void *txattr_addr; int err; @@ -405,7 +406,7 @@ static int read_all_xattrs(struct inode *inode, struct folio *ifolio, /* read from inline xattr */ if (inline_size) { - err = read_inline_xattr(inode, ifolio, txattr_addr); + err = read_inline_xattr(inode, ientry, txattr_addr); if (err) goto fail; } @@ -432,14 +433,14 @@ fail: } static inline int write_all_xattrs(struct inode *inode, __u32 hsize, - void *txattr_addr, struct folio *ifolio) + void *txattr_addr, struct f2fs_cached_block *ientry) { struct f2fs_sb_info *sbi = F2FS_I_SB(inode); size_t inline_size = inline_xattr_size(inode); - struct folio *in_folio = NULL; + struct f2fs_cached_block *in_entry = NULL; void *xattr_addr; void *inline_addr = NULL; - struct folio *xfolio; + struct f2fs_cached_block *xentry; nid_t new_nid = 0; int err = 0; @@ -449,75 +450,75 @@ static inline int write_all_xattrs(struct inode *inode, __u32 hsize, /* write to inline xattr */ if (inline_size) { - if (ifolio) { - inline_addr = inline_xattr_addr(inode, ifolio); + if (ientry) { + inline_addr = inline_xattr_addr(inode, ientry); } else { - in_folio = f2fs_get_inode_folio(sbi, inode->i_ino); - if (IS_ERR(in_folio)) { + in_entry = f2fs_get_inode_cache(sbi, inode->i_ino); + if (IS_ERR(in_entry)) { f2fs_alloc_nid_failed(sbi, new_nid); - return PTR_ERR(in_folio); + return PTR_ERR(in_entry); } - inline_addr = inline_xattr_addr(inode, in_folio); + inline_addr = inline_xattr_addr(inode, in_entry); } - f2fs_folio_wait_writeback(ifolio ? ifolio : in_folio, - NODE, true, true); + f2fs_cache_wait_writeback(ientry ? ientry : in_entry); /* no need to use xattr node block */ if (hsize <= inline_size) { err = f2fs_truncate_xattr_node(inode); f2fs_alloc_nid_failed(sbi, new_nid); if (err) { - f2fs_folio_put(in_folio, true); + f2fs_put_cache(in_entry, true); return err; } memcpy(inline_addr, txattr_addr, inline_size); - folio_mark_dirty(ifolio ? ifolio : in_folio); + f2fs_mark_cache_dirty(ientry ? ientry : in_entry); goto in_page_out; } } /* write to xattr node block */ if (F2FS_I(inode)->i_xattr_nid) { - xfolio = f2fs_get_xnode_folio(sbi, F2FS_I(inode)->i_xattr_nid); - if (IS_ERR(xfolio)) { - err = PTR_ERR(xfolio); + xentry = f2fs_get_xnode_cache(sbi, F2FS_I(inode)->i_xattr_nid); + if (IS_ERR(xentry)) { + err = PTR_ERR(xentry); f2fs_alloc_nid_failed(sbi, new_nid); goto in_page_out; } f2fs_bug_on(sbi, new_nid); - f2fs_folio_wait_writeback(xfolio, NODE, true, true); + f2fs_cache_wait_writeback(xentry); } else { struct dnode_of_data dn; set_new_dnode(&dn, inode, NULL, NULL, new_nid); - xfolio = f2fs_new_node_folio(&dn, XATTR_NODE_OFFSET); - if (IS_ERR(xfolio)) { - err = PTR_ERR(xfolio); + xentry = f2fs_new_node_cache(&dn, XATTR_NODE_OFFSET); + if (IS_ERR(xentry)) { + err = PTR_ERR(xentry); f2fs_alloc_nid_failed(sbi, new_nid); goto in_page_out; } f2fs_alloc_nid_done(sbi, new_nid); } - xattr_addr = folio_address(xfolio); + xattr_addr = cache_address(xentry); if (inline_size) memcpy(inline_addr, txattr_addr, inline_size); - memcpy(xattr_addr, txattr_addr + inline_size, VALID_XATTR_BLOCK_SIZE); + memcpy(xattr_addr, txattr_addr + inline_size, + VALID_XATTR_BLOCK_SIZE(inode)); if (inline_size) - folio_mark_dirty(ifolio ? ifolio : in_folio); - folio_mark_dirty(xfolio); + f2fs_mark_cache_dirty(ientry ? ientry : in_entry); + f2fs_mark_cache_dirty(xentry); - f2fs_folio_put(xfolio, true); + f2fs_put_cache(xentry, true); in_page_out: - f2fs_folio_put(in_folio, true); + f2fs_put_cache(in_entry, true); return err; } int f2fs_getxattr(struct inode *inode, int index, const char *name, - void *buffer, size_t buffer_size, struct folio *ifolio) + void *buffer, size_t buffer_size, struct f2fs_cached_block *ientry) { - struct f2fs_xattr_entry *entry = NULL; + struct f2fs_xattr_entry *xe = NULL; int error; unsigned int size, len; void *base_addr = NULL; @@ -531,16 +532,16 @@ int f2fs_getxattr(struct inode *inode, int index, const char *name, if (len > F2FS_NAME_LEN) return -ERANGE; - if (!ifolio) + if (!ientry) f2fs_down_read(&F2FS_I(inode)->i_xattr_sem); - error = lookup_all_xattrs(inode, ifolio, index, len, name, - &entry, &base_addr, &base_size, &is_inline); - if (!ifolio) + error = lookup_all_xattrs(inode, ientry, index, len, name, + &xe, &base_addr, &base_size, &is_inline); + if (!ientry) f2fs_up_read(&F2FS_I(inode)->i_xattr_sem); if (error) return error; - size = le16_to_cpu(entry->e_value_size); + size = le16_to_cpu(xe->e_value_size); if (buffer && size > buffer_size) { error = -ERANGE; @@ -548,7 +549,7 @@ int f2fs_getxattr(struct inode *inode, int index, const char *name, } if (buffer) { - char *pval = entry->e_name + entry->e_name_len; + char *pval = xe->e_name + xe->e_name_len; if (base_size - (pval - (char *)base_addr) < size) { error = -ERANGE; @@ -632,7 +633,7 @@ static bool f2fs_xattr_value_same(struct f2fs_xattr_entry *entry, static int __f2fs_setxattr(struct inode *inode, int index, const char *name, const void *value, size_t size, - struct folio *ifolio, int flags) + struct f2fs_cached_block *ientry, int flags) { struct f2fs_sb_info *sbi = F2FS_I_SB(inode); struct f2fs_xattr_entry *here, *last; @@ -656,7 +657,7 @@ static int __f2fs_setxattr(struct inode *inode, int index, if (size > MAX_VALUE_LEN(inode)) return -E2BIG; retry: - error = read_all_xattrs(inode, ifolio, &base_addr); + error = read_all_xattrs(inode, ientry, &base_addr); if (error) return error; @@ -773,7 +774,7 @@ retry: *(u32 *)((u8 *)last + newsize) = 0; } - error = write_all_xattrs(inode, new_hsize, base_addr, ifolio); + error = write_all_xattrs(inode, new_hsize, base_addr, ientry); if (error) goto exit; @@ -807,7 +808,7 @@ exit: int f2fs_setxattr(struct inode *inode, int index, const char *name, const void *value, size_t size, - struct folio *ifolio, int flags) + struct f2fs_cached_block *ientry, int flags) { struct f2fs_sb_info *sbi = F2FS_I_SB(inode); struct f2fs_lock_context lc; @@ -823,9 +824,9 @@ int f2fs_setxattr(struct inode *inode, int index, const char *name, return err; /* this case is only from f2fs_init_inode_metadata */ - if (ifolio) + if (ientry) return __f2fs_setxattr(inode, index, name, value, - size, ifolio, flags); + size, ientry, flags); f2fs_balance_fs(sbi, true); f2fs_lock_op(sbi, &lc); @@ -848,4 +849,4 @@ int __init f2fs_init_xattr_cache(void) void f2fs_destroy_xattr_cache(void) { kmem_cache_destroy(inline_xattr_slab); -}
\ No newline at end of file +} diff --git a/fs/f2fs/xattr.h b/fs/f2fs/xattr.h index bce3d93e4755..870524b667ed 100644 --- a/fs/f2fs/xattr.h +++ b/fs/f2fs/xattr.h @@ -71,21 +71,22 @@ struct f2fs_xattr_entry { for (entry = XATTR_FIRST_ENTRY(addr);\ !IS_XATTR_LAST_ENTRY(entry);\ entry = XATTR_NEXT_ENTRY(entry)) -#define VALID_XATTR_BLOCK_SIZE (PAGE_SIZE - sizeof(struct node_footer)) +#define VALID_XATTR_BLOCK_SIZE(i) (i_blocksize(i) - \ + sizeof(struct node_footer)) #define XATTR_PADDING_SIZE (sizeof(__u32)) #define XATTR_SIZE(i) ((F2FS_I(i)->i_xattr_nid ? \ - VALID_XATTR_BLOCK_SIZE : 0) + \ + VALID_XATTR_BLOCK_SIZE(i) : 0) + \ (inline_xattr_size(i))) #define MIN_OFFSET(i) XATTR_ALIGN(inline_xattr_size(i) + \ - VALID_XATTR_BLOCK_SIZE) + VALID_XATTR_BLOCK_SIZE(i)) #define MAX_VALUE_LEN(i) (MIN_OFFSET(i) - \ sizeof(struct f2fs_xattr_header) - \ sizeof(struct f2fs_xattr_entry)) #define MIN_INLINE_XATTR_SIZE (sizeof(struct f2fs_xattr_header) / sizeof(__le32)) -#define MAX_INLINE_XATTR_SIZE \ - (DEF_ADDRS_PER_INODE - \ +#define MAX_INLINE_XATTR_SIZE(blocksize) \ + (F2FS_DEF_ADDRS_PER_INODE(blocksize) - \ F2FS_TOTAL_EXTRA_ATTR_SIZE / sizeof(__le32) - \ DEF_INLINE_RESERVED_SIZE - \ MIN_INLINE_DENTRY_SIZE / sizeof(__le32)) @@ -130,9 +131,9 @@ extern const struct xattr_handler f2fs_xattr_security_handler; extern const struct xattr_handler * const f2fs_xattr_handlers[]; int f2fs_setxattr(struct inode *, int, const char *, const void *, - size_t, struct folio *, int); + size_t, struct f2fs_cached_block *, int); int f2fs_getxattr(struct inode *, int, const char *, void *, - size_t, struct folio *); + size_t, struct f2fs_cached_block *); ssize_t f2fs_listxattr(struct dentry *, char *, size_t); int __init f2fs_init_xattr_cache(void); void f2fs_destroy_xattr_cache(void); @@ -142,13 +143,13 @@ void f2fs_destroy_xattr_cache(void); #define f2fs_listxattr NULL static inline int f2fs_setxattr(struct inode *inode, int index, const char *name, const void *value, size_t size, - struct folio *folio, int flags) + struct f2fs_cached_block *ientry, int flags) { return -EOPNOTSUPP; } static inline int f2fs_getxattr(struct inode *inode, int index, const char *name, void *buffer, - size_t buffer_size, struct folio *dfolio) + size_t buffer_size, struct f2fs_cached_block *ientry) { return -EOPNOTSUPP; } @@ -158,10 +159,10 @@ static inline void f2fs_destroy_xattr_cache(void) { } #ifdef CONFIG_F2FS_FS_SECURITY int f2fs_init_security(struct inode *, struct inode *, - const struct qstr *, struct folio *); + const struct qstr *, struct f2fs_cached_block *); #else static inline int f2fs_init_security(struct inode *inode, struct inode *dir, - const struct qstr *qstr, struct folio *ifolio) + const struct qstr *qstr, struct f2fs_cached_block *ientry) { return 0; } diff --git a/fs/failfs.c b/fs/failfs.c index 66a36da3d236..437cdc981c2d 100644 --- a/fs/failfs.c +++ b/fs/failfs.c @@ -22,7 +22,7 @@ bool failfs_mnt(const struct vfsmount *mnt) return mnt->mnt_sb == failfs_root_path.mnt->mnt_sb; } -static int failfs_permission(struct mnt_idmap *idmap, struct inode *inode, +static int failfs_permission(const struct mnt_idmap *idmap, struct inode *inode, int mask) { return -EOPNOTSUPP; @@ -35,7 +35,7 @@ static struct dentry *failfs_lookup(struct inode *dir, struct dentry *dentry, return ERR_PTR(-EOPNOTSUPP); } -static int failfs_getattr(struct mnt_idmap *idmap, const struct path *path, +static int failfs_getattr(const struct mnt_idmap *idmap, const struct path *path, struct kstat *stat, u32 request_mask, unsigned int query_flags) { diff --git a/fs/fat/fat.h b/fs/fat/fat.h index fbd207c55859..a4a8958fcfc7 100644 --- a/fs/fat/fat.h +++ b/fs/fat/fat.h @@ -404,10 +404,10 @@ extern long fat_generic_ioctl(struct file *filp, unsigned int cmd, unsigned long arg); extern const struct file_operations fat_file_operations; extern const struct inode_operations fat_file_inode_operations; -extern int fat_setattr(struct mnt_idmap *idmap, struct dentry *dentry, +extern int fat_setattr(const struct mnt_idmap *idmap, struct dentry *dentry, struct iattr *attr); extern void fat_truncate_blocks(struct inode *inode, loff_t offset); -extern int fat_getattr(struct mnt_idmap *idmap, +extern int fat_getattr(const struct mnt_idmap *idmap, const struct path *path, struct kstat *stat, u32 request_mask, unsigned int flags); int fat_fileattr_get(struct dentry *dentry, struct file_kattr *fa); diff --git a/fs/fat/file.c b/fs/fat/file.c index 6c475c53334c..19bb6b3cab12 100644 --- a/fs/fat/file.c +++ b/fs/fat/file.c @@ -433,7 +433,7 @@ int fat_fileattr_get(struct dentry *dentry, struct file_kattr *fa) } EXPORT_SYMBOL_GPL(fat_fileattr_get); -int fat_getattr(struct mnt_idmap *idmap, const struct path *path, +int fat_getattr(const struct mnt_idmap *idmap, const struct path *path, struct kstat *stat, u32 request_mask, unsigned int flags) { struct inode *inode = d_inode(path->dentry); @@ -494,7 +494,7 @@ static int fat_sanitize_mode(const struct msdos_sb_info *sbi, return 0; } -static int fat_allow_set_time(struct mnt_idmap *idmap, +static int fat_allow_set_time(const struct mnt_idmap *idmap, struct msdos_sb_info *sbi, struct inode *inode) { umode_t allow_utime = sbi->options.allow_utime; @@ -515,7 +515,7 @@ static int fat_allow_set_time(struct mnt_idmap *idmap, /* valid file mode bits */ #define FAT_VALID_MODE (S_IFREG | S_IFDIR | S_IRWXUGO) -int fat_setattr(struct mnt_idmap *idmap, struct dentry *dentry, +int fat_setattr(const struct mnt_idmap *idmap, struct dentry *dentry, struct iattr *attr) { struct msdos_sb_info *sbi = MSDOS_SB(dentry->d_sb); diff --git a/fs/fat/misc.c b/fs/fat/misc.c index c44296756eae..9a76ff8c60f0 100644 --- a/fs/fat/misc.c +++ b/fs/fat/misc.c @@ -358,7 +358,7 @@ int fat_sync_bhs(struct buffer_head **bhs, int nr_bhs) for (i = 0; i < nr_bhs; i++) { wait_on_buffer(bhs[i]); - if (!err && !buffer_uptodate(bhs[i])) + if (!err && buffer_write_io_error(bhs[i])) err = -EIO; } return err; diff --git a/fs/fat/namei_msdos.c b/fs/fat/namei_msdos.c index d46d1a3851f2..dde4215616f9 100644 --- a/fs/fat/namei_msdos.c +++ b/fs/fat/namei_msdos.c @@ -263,7 +263,7 @@ static int msdos_add_entry(struct inode *dir, const unsigned char *name, } /***** Create a file */ -static int msdos_create(struct mnt_idmap *idmap, struct inode *dir, +static int msdos_create(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *dentry, umode_t mode) { struct super_block *sb = dir->i_sb; @@ -345,7 +345,7 @@ out: } /***** Make a directory */ -static struct dentry *msdos_mkdir(struct mnt_idmap *idmap, struct inode *dir, +static struct dentry *msdos_mkdir(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *dentry, umode_t mode) { struct super_block *sb = dir->i_sb; @@ -600,7 +600,7 @@ error_inode: } /***** Rename, a wrapper for rename_same_dir & rename_diff_dir */ -static int msdos_rename(struct mnt_idmap *idmap, +static int msdos_rename(const struct mnt_idmap *idmap, struct inode *old_dir, struct dentry *old_dentry, struct inode *new_dir, struct dentry *new_dentry, unsigned int flags) diff --git a/fs/fat/namei_vfat.c b/fs/fat/namei_vfat.c index da3e89c0b16a..3dc063ba0a73 100644 --- a/fs/fat/namei_vfat.c +++ b/fs/fat/namei_vfat.c @@ -753,7 +753,7 @@ error: return ERR_PTR(err); } -static int vfat_create(struct mnt_idmap *idmap, struct inode *dir, +static int vfat_create(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *dentry, umode_t mode) { struct super_block *sb = dir->i_sb; @@ -846,7 +846,7 @@ out: return err; } -static struct dentry *vfat_mkdir(struct mnt_idmap *idmap, struct inode *dir, +static struct dentry *vfat_mkdir(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *dentry, umode_t mode) { struct super_block *sb = dir->i_sb; @@ -1160,7 +1160,7 @@ error_exchange: goto out; } -static int vfat_rename2(struct mnt_idmap *idmap, struct inode *old_dir, +static int vfat_rename2(const struct mnt_idmap *idmap, struct inode *old_dir, struct dentry *old_dentry, struct inode *new_dir, struct dentry *new_dentry, unsigned int flags) { diff --git a/fs/fhandle.c b/fs/fhandle.c index f8829231e3d7..2aa55b8a878a 100644 --- a/fs/fhandle.c +++ b/fs/fhandle.c @@ -201,7 +201,7 @@ static int vfs_dentry_acceptable(void *context, struct dentry *dentry) struct handle_to_path_ctx *ctx = context; struct user_namespace *user_ns = current_user_ns(); struct dentry *d, *root = ctx->root.dentry; - struct mnt_idmap *idmap = mnt_idmap(ctx->root.mnt); + const struct mnt_idmap *idmap = mnt_idmap(ctx->root.mnt); int retval = 0; if (!root) diff --git a/fs/file.c b/fs/file.c index 628ca07dc4b1..88b7a5340815 100644 --- a/fs/file.c +++ b/fs/file.c @@ -352,24 +352,78 @@ static inline bool fd_is_open(unsigned int fd, const struct fdtable *fdt) return test_bit(fd, fdt->open_fds); } +/* Bits of [range->from, range->to] that fall into word @i of a bitmap. */ +static unsigned long fd_range_word(struct fd_range *range, unsigned int i) +{ + unsigned int first = i * BITS_PER_LONG; + unsigned int last = first + BITS_PER_LONG - 1; + + if (range->to < first || range->from > last) + return 0; + return GENMASK(min(range->to, last) - first, + max(range->from, first) - first); +} + +/* Bits of word @i that dup_fd() leaves behind and __range_close() closes. */ +static unsigned long dup_fd_dropped_word(struct fdtable *fdt, unsigned int i, + struct fd_range *range) +{ + unsigned long dropped; + + if (!range) + return 0; + dropped = fd_range_word(range, i); + if (range->flags & FD_RANGE_EXCEPT) + dropped = ~dropped; + if (range->flags & FD_RANGE_CLOEXEC_ONLY) + dropped &= fdt->close_on_exec[i]; + return dropped; +} + /* * Note that a sane fdtable size always has to be a multiple of * BITS_PER_LONG, since we have bitmaps that are sized by this. * - * punch_hole is optional - when close_range() is asked to unshare - * and close, we don't need to copy descriptors in that range, so - * a smaller cloned descriptor table might suffice if the last - * currently opened descriptor falls into that range. + * range is optional. When close_range() is asked to unshare dup_fd() + * will leave any files behind according to the range and its flags. The + * cloned table only has to reach the last open descriptor that is + * carried over. */ -static unsigned int sane_fdtable_size(struct fdtable *fdt, struct fd_range *punch_hole) +static unsigned int sane_fdtable_size(struct fdtable *fdt, struct fd_range *range) { unsigned int last = find_last_bit(fdt->open_fds, fdt->max_fds); + unsigned int i; if (last == fdt->max_fds) return NR_OPEN_DEFAULT; - if (punch_hole && punch_hole->to >= last && punch_hole->from <= last) { - last = find_last_bit(fdt->open_fds, punch_hole->from); - if (last == punch_hole->from) + if (!range) + return ALIGN(last + 1, BITS_PER_LONG); + + if (range->flags & FD_RANGE_CLOEXEC_ONLY) { + /* The close-on-exec bits decide what is dropped, walk the words. */ + i = last / BITS_PER_LONG + 1; + while (i--) { + unsigned long dropped = dup_fd_dropped_word(fdt, i, range); + + if (fdt->open_fds[i] & ~dropped) + return (i + 1) * BITS_PER_LONG; + } + return NR_OPEN_DEFAULT; + } + + if (range->flags & FD_RANGE_EXCEPT) { + /* Only the range is carried over. */ + if (last > range->to) { + last = find_last_bit(fdt->open_fds, range->to + 1); + if (last > range->to) + return NR_OPEN_DEFAULT; + } + if (last < range->from) + return NR_OPEN_DEFAULT; + } else if (last >= range->from && last <= range->to) { + /* The last open descriptor goes, the kept ones sit below the range. */ + last = find_last_bit(fdt->open_fds, range->from); + if (last == range->from) return NR_OPEN_DEFAULT; } return ALIGN(last + 1, BITS_PER_LONG); @@ -378,13 +432,14 @@ static unsigned int sane_fdtable_size(struct fdtable *fdt, struct fd_range *punc /* * Allocate a new descriptor table and copy contents from the passed in * instance. Returns a pointer to cloned table on success, ERR_PTR() - * on failure. For 'punch_hole' see sane_fdtable_size(). + * on failure. For 'range' see sane_fdtable_size(). */ -struct files_struct *dup_fd(struct files_struct *oldf, struct fd_range *punch_hole) +struct files_struct *dup_fd(struct files_struct *oldf, struct fd_range *range) { struct files_struct *newf; struct file **old_fds, **new_fds; - unsigned int open_files, i; + unsigned int open_files, fd; + unsigned long dropped = 0; struct fdtable *old_fdt, *new_fdt; newf = kmem_cache_alloc(files_cachep, GFP_KERNEL); @@ -406,7 +461,7 @@ struct files_struct *dup_fd(struct files_struct *oldf, struct fd_range *punch_ho spin_lock(&oldf->file_lock); old_fdt = files_fdtable(oldf); - open_files = sane_fdtable_size(old_fdt, punch_hole); + open_files = sane_fdtable_size(old_fdt, range); /* * Check whether we need to allocate a larger fd array and fd set. @@ -430,7 +485,7 @@ struct files_struct *dup_fd(struct files_struct *oldf, struct fd_range *punch_ho */ spin_lock(&oldf->file_lock); old_fdt = files_fdtable(oldf); - open_files = sane_fdtable_size(old_fdt, punch_hole); + open_files = sane_fdtable_size(old_fdt, range); } copy_fd_bitmaps(new_fdt, old_fdt, open_files / BITS_PER_LONG); @@ -451,13 +506,18 @@ struct files_struct *dup_fd(struct files_struct *oldf, struct fd_range *punch_ho * * Instead of trying to placate userspace racing with itself, we * ref the file if we see it and mark the fd slot as unused otherwise. + * Descriptors dup_fd() is asked to leave behind get the same treatment. */ - for (i = open_files; i != 0; i--) { + for (fd = 0; fd < open_files; fd++) { struct file *f = rcu_dereference_raw(*old_fds++); - if (f) { + + if (!(fd % BITS_PER_LONG)) + dropped = dup_fd_dropped_word(old_fdt, fd / BITS_PER_LONG, range); + if (f && !(dropped & BIT_MASK(fd))) { get_file(f); } else { - __clear_open_fd(open_files - i, new_fdt); + f = NULL; + __clear_open_fd(fd, new_fdt); } rcu_assign_pointer(*new_fds++, f); } @@ -471,7 +531,25 @@ struct files_struct *dup_fd(struct files_struct *oldf, struct fd_range *punch_ho return newf; } -static struct fdtable *close_files(struct files_struct * files) +/* + * Unshare file descriptor table if it is being shared + */ +int unshare_fd(unsigned long unshare_flags, struct files_struct **new_fdp) +{ + struct files_struct *fd = current->files; + + if ((unshare_flags & CLONE_FILES) && + (fd && atomic_read(&fd->count) > 1)) { + fd = dup_fd(fd, NULL); + if (IS_ERR(fd)) + return PTR_ERR(fd); + *new_fdp = fd; + } + + return 0; +} + +static struct fdtable *close_files(struct files_struct *files) { /* * It is safe to dereference the fd table without RCU or @@ -479,24 +557,21 @@ static struct fdtable *close_files(struct files_struct * files) * files structure. */ struct fdtable *fdt = rcu_dereference_raw(files->fdt); - unsigned int i, j = 0; + unsigned int j = fdt->max_fds / BITS_PER_LONG; + + /* Highest fd first, the order the deferred puts ran in. */ + while (j--) { + unsigned long set = fdt->open_fds[j]; - for (;;) { - unsigned long set; - i = j * BITS_PER_LONG; - if (i >= fdt->max_fds) - break; - set = fdt->open_fds[j++]; while (set) { - if (set & 1) { - struct file *file = fdt->fd[i]; - if (file) { - filp_close(file, files); - cond_resched(); - } + unsigned int bit = __fls(set); + struct file *file = fdt->fd[j * BITS_PER_LONG + bit]; + + set ^= 1UL << bit; + if (file) { + filp_close_sync(file, files); + cond_resched(); } - i++; - set >>= 1; } } @@ -515,16 +590,18 @@ void put_files_struct(struct files_struct *files) } } -void exit_files(struct task_struct *tsk) +/* Install @files on @tsk, consuming the reference, and put the old table. */ +void switch_files_struct(struct task_struct *tsk, struct files_struct *files) { - struct files_struct * files = tsk->files; + scoped_guard(task_lock, tsk) + swap(tsk->files, files); + put_files_struct(files); +} - if (files) { - task_lock(tsk); - tsk->files = NULL; - task_unlock(tsk); - put_files_struct(files); - } +void exit_files(struct task_struct *tsk) +{ + if (tsk->files) + switch_files_struct(tsk, NULL); } struct files_struct init_files = { @@ -732,16 +809,13 @@ struct file *file_close_fd_locked(struct files_struct *files, unsigned fd) int close_fd(unsigned fd) { - struct files_struct *files = current->files; struct file *file; - spin_lock(&files->file_lock); - file = file_close_fd_locked(files, fd); - spin_unlock(&files->file_lock); + file = file_close_fd(fd); if (!file) return -EBADF; - return filp_close(file, files); + return filp_close(file, current->files); } EXPORT_SYMBOL(close_fd); @@ -759,38 +833,90 @@ static inline unsigned last_fd(struct fdtable *fdt) } static inline void __range_cloexec(struct files_struct *cur_fds, - unsigned int fd, unsigned int max_fd) + struct fd_range *range) { struct fdtable *fdt; + unsigned int last; - /* make sure we're using the correct maximum value */ spin_lock(&cur_fds->file_lock); fdt = files_fdtable(cur_fds); - max_fd = min(last_fd(fdt), max_fd); - if (fd <= max_fd) - bitmap_set(fdt->close_on_exec, fd, max_fd - fd + 1); + /* make sure we're using the correct maximum value */ + last = last_fd(fdt); + if (!(range->flags & FD_RANGE_EXCEPT)) { + if (range->from <= last) + bitmap_set(fdt->close_on_exec, range->from, + min(range->to, last) - range->from + 1); + } else { + if (range->from > 0) + bitmap_set(fdt->close_on_exec, 0, + min(range->from - 1, last) + 1); + if (range->to < last) + bitmap_set(fdt->close_on_exec, range->to + 1, + last - range->to); + } spin_unlock(&cur_fds->file_lock); } -static inline void __range_close(struct files_struct *files, unsigned int fd, - unsigned int max_fd) +/* Highest open descriptor below @n that @range selects, or @n. */ +static inline unsigned int last_fd_to_close(struct fdtable *fdt, unsigned int n, + struct fd_range *range) +{ + unsigned int i, lo = 0; + + if (!(range->flags & FD_RANGE_EXCEPT)) + lo = range->from / BITS_PER_LONG; + for (i = n ? (n - 1) / BITS_PER_LONG + 1 : 0; i-- > lo; ) { + unsigned long set = fdt->open_fds[i]; + + if (!set) { + /* Skip the empty stretch at find_last_bit() speed. */ + unsigned int last = find_last_bit(fdt->open_fds, i * BITS_PER_LONG); + + if (last >= i * BITS_PER_LONG) + break; + i = last / BITS_PER_LONG + 1; + continue; + } + /* Hop below the kept window in one step. */ + if ((range->flags & FD_RANGE_EXCEPT) && + i * BITS_PER_LONG >= range->from && + i * BITS_PER_LONG + BITS_PER_LONG - 1 <= range->to) { + if (!range->from) + break; + i = (range->from - 1) / BITS_PER_LONG + 1; + continue; + } + set &= dup_fd_dropped_word(fdt, i, range); + if (i == (n - 1) / BITS_PER_LONG) + set &= BITMAP_LAST_WORD_MASK(n); + if (set) + return i * BITS_PER_LONG + __fls(set); + } + return n; +} + +static inline void __range_close(struct files_struct *files, + struct fd_range *range) { struct file *file; struct fdtable *fdt; - unsigned n; + unsigned int fd, n; spin_lock(&files->file_lock); fdt = files_fdtable(files); - n = last_fd(fdt); - max_fd = min(max_fd, n); + if (range->flags & FD_RANGE_EXCEPT) + /* Outside of the range means the whole table. */ + n = fdt->max_fds; + else + n = min(range->to, last_fd(fdt)) + 1; - for (fd = find_next_bit(fdt->open_fds, max_fd + 1, fd); - fd <= max_fd; - fd = find_next_bit(fdt->open_fds, max_fd + 1, fd + 1)) { + /* Highest fd first, see close_files(). */ + while ((fd = last_fd_to_close(fdt, n, range)) < n) { + n = fd; file = file_close_fd_locked(files, fd); if (file) { spin_unlock(&files->file_lock); - filp_close(file, files); + filp_close_sync(file, files); cond_resched(); spin_lock(&files->file_lock); fdt = files_fdtable(files); @@ -814,21 +940,43 @@ static inline void __range_close(struct files_struct *files, unsigned int fd, * This closes a range of file descriptors. All file descriptors * from @fd up to and including @max_fd are closed. * Currently, errors to close a given file descriptor are ignored. + * + * With CLOSE_RANGE_EXCEPT the range names what to leave alone instead: + * every open file descriptor outside of [@fd, @max_fd] is closed, or + * marked close-on-exec with CLOSE_RANGE_CLOEXEC. + * + * With CLOSE_RANGE_CLOEXEC_ONLY only file descriptors that have + * close-on-exec set are closed. Together with CLOSE_RANGE_EXCEPT the + * range names the close-on-exec file descriptors to keep. To keep none + * of them, name a range that cannot hold an open file descriptor, e.g. + * close_range(~0U, ~0U, ...). */ SYSCALL_DEFINE3(close_range, unsigned int, fd, unsigned int, max_fd, unsigned int, flags) { struct task_struct *me = current; struct files_struct *cur_fds = me->files, *fds = NULL; + struct fd_range range = {fd, max_fd}; + + if (flags & ~(CLOSE_RANGE_UNSHARE | CLOSE_RANGE_CLOEXEC | + CLOSE_RANGE_EXCEPT | CLOSE_RANGE_CLOEXEC_ONLY)) + return -EINVAL; - if (flags & ~(CLOSE_RANGE_UNSHARE | CLOSE_RANGE_CLOEXEC)) + /* One marks close-on-exec, the other closes what is marked. */ + if (hweight32(flags & (CLOSE_RANGE_CLOEXEC | + CLOSE_RANGE_CLOEXEC_ONLY)) > 1) return -EINVAL; if (fd > max_fd) return -EINVAL; + if (flags & CLOSE_RANGE_EXCEPT) + range.flags |= FD_RANGE_EXCEPT; + if (flags & CLOSE_RANGE_CLOEXEC_ONLY) + range.flags |= FD_RANGE_CLOEXEC_ONLY; + if ((flags & CLOSE_RANGE_UNSHARE) && atomic_read(&cur_fds->count) > 1) { - struct fd_range range = {fd, max_fd}, *punch_hole = ⦥ + struct fd_range *drop = ⦥ /* * If the caller requested all fds to be made cloexec we always @@ -836,9 +984,9 @@ SYSCALL_DEFINE3(close_range, unsigned int, fd, unsigned int, max_fd, * use them. */ if (flags & CLOSE_RANGE_CLOEXEC) - punch_hole = NULL; + drop = NULL; - fds = dup_fd(cur_fds, punch_hole); + fds = dup_fd(cur_fds, drop); if (IS_ERR(fds)) return PTR_ERR(fds); /* @@ -848,20 +996,19 @@ SYSCALL_DEFINE3(close_range, unsigned int, fd, unsigned int, max_fd, swap(cur_fds, fds); } - if (flags & CLOSE_RANGE_CLOEXEC) - __range_cloexec(cur_fds, fd, max_fd); - else - __range_close(cur_fds, fd, max_fd); + if (flags & CLOSE_RANGE_CLOEXEC) { + __range_cloexec(cur_fds, &range); + } else if (!fds) { + /* If we unshared, dup_fd() already left behind what we'd close. */ + __range_close(cur_fds, &range); + } if (fds) { /* * We're done closing the files we were supposed to. Time to install * the new file descriptor table and drop the old one. */ - task_lock(me); - me->files = cur_fds; - task_unlock(me); - put_files_struct(fds); + switch_files_struct(me, cur_fds); } return 0; @@ -887,34 +1034,36 @@ struct file *file_close_fd(unsigned int fd) return file; } -void do_close_on_exec(struct files_struct *files) +void close_cloexec_files(struct files_struct *files) { unsigned i; struct fdtable *fdt; /* exec unshares first */ spin_lock(&files->file_lock); - for (i = 0; ; i++) { + fdt = files_fdtable(files); + /* Highest fd first, see close_files(). */ + for (i = fdt->max_fds / BITS_PER_LONG; i--; ) { unsigned long set; - unsigned fd = i * BITS_PER_LONG; + fdt = files_fdtable(files); - if (fd >= fdt->max_fds) - break; set = fdt->close_on_exec[i]; if (!set) continue; fdt->close_on_exec[i] = 0; - for ( ; set ; fd++, set >>= 1) { + while (set) { + unsigned int bit = __fls(set); + unsigned fd = i * BITS_PER_LONG + bit; struct file *file; - if (!(set & 1)) - continue; + + set ^= 1UL << bit; file = fdt->fd[fd]; if (!file) continue; rcu_assign_pointer(fdt->fd[fd], NULL); __put_unused_fd(files, fd); spin_unlock(&files->file_lock); - filp_close(file, files); + filp_close_sync(file, files); cond_resched(); spin_lock(&files->file_lock); } @@ -1391,17 +1540,17 @@ int receive_fd(struct file *file, int __user *ufd, unsigned int o_flags) return error; FD_PREPARE(fdf, o_flags, file); - if (fdf.err) - return fdf.err; + if (fdf->fd < 0) + return fdf->fd; get_file(file); if (ufd) { - error = put_user(fd_prepare_fd(fdf), ufd); + error = put_user(fdf->fd, ufd); if (error) return error; } - __receive_sock(fd_prepare_file(fdf)); + __receive_sock(fdf->file); return fd_publish(fdf); } EXPORT_SYMBOL_GPL(receive_fd); @@ -1529,3 +1678,7 @@ int iterate_fd(struct files_struct *files, unsigned n, return res; } EXPORT_SYMBOL(iterate_fd); + +#ifdef CONFIG_FDTABLE_KUNIT_TEST +#include "tests/fdtable_kunit.c" +#endif diff --git a/fs/file_attr.c b/fs/file_attr.c index bfb00d256dd5..81af4364e33a 100644 --- a/fs/file_attr.c +++ b/fs/file_attr.c @@ -265,7 +265,7 @@ static int fileattr_set_prepare(struct inode *inode, * * Return: 0 on success, or a negative error on failure. */ -int vfs_fileattr_set(struct mnt_idmap *idmap, struct dentry *dentry, +int vfs_fileattr_set(const struct mnt_idmap *idmap, struct dentry *dentry, struct file_kattr *fa) { struct inode *inode = d_inode(dentry); @@ -323,7 +323,7 @@ int ioctl_getflags(struct file *file, unsigned int __user *argp) int ioctl_setflags(struct file *file, unsigned int __user *argp) { - struct mnt_idmap *idmap = file_mnt_idmap(file); + const struct mnt_idmap *idmap = file_mnt_idmap(file); struct dentry *dentry = file->f_path.dentry; struct file_kattr fa = {}; unsigned int flags; @@ -355,7 +355,7 @@ int ioctl_fsgetxattr(struct file *file, void __user *argp) int ioctl_fssetxattr(struct file *file, void __user *argp) { - struct mnt_idmap *idmap = file_mnt_idmap(file); + const struct mnt_idmap *idmap = file_mnt_idmap(file); struct dentry *dentry = file->f_path.dentry; struct file_kattr fa = {}; int err; diff --git a/fs/fs-writeback.c b/fs/fs-writeback.c index ea3eb40bf828..58a6403780aa 100644 --- a/fs/fs-writeback.c +++ b/fs/fs-writeback.c @@ -233,7 +233,7 @@ void wb_wait_for_completion(struct wb_completion *done) * Parameters for foreign inode detection, see wbc_detach_inode() to see * how they're used. * - * These paramters are inherently heuristical as the detection target + * These parameters are inherently heuristical as the detection target * itself is fuzzy. All we want to do is detaching an inode from the * current owner if it's being written to by some other cgroups too much. * @@ -248,7 +248,7 @@ void wb_wait_for_completion(struct wb_completion *done) * to 16 slots. To avoid tiny writes from swinging the decision too much, * writes smaller than 1/8 of avg size are ignored. */ -#define WB_FRN_TIME_SHIFT 13 /* 1s = 2^13, upto 8 secs w/ 16bit */ +#define WB_FRN_TIME_SHIFT 13 /* 1s = 2^13, up to 8 secs w/ 16bit */ #define WB_FRN_TIME_AVG_SHIFT 3 /* avg = avg * 7/8 + new * 1/8 */ #define WB_FRN_TIME_CUT_DIV 8 /* ignore rounds < avg / 8 */ #define WB_FRN_TIME_PERIOD (2 * (1 << WB_FRN_TIME_SHIFT)) /* 2s */ @@ -259,7 +259,7 @@ void wb_wait_for_completion(struct wb_completion *done) #define WB_FRN_HIST_THR_SLOTS (WB_FRN_HIST_SLOTS / 2) /* if foreign slots >= 8, switch */ #define WB_FRN_HIST_MAX_SLOTS (WB_FRN_HIST_THR_SLOTS / 2 + 1) - /* one round can affect upto 5 slots */ + /* one round can affect up to 5 slots */ #define WB_FRN_MAX_IN_FLIGHT 1024 /* don't queue too many concurrently */ /* @@ -1181,7 +1181,7 @@ int cgroup_writeback_by_id(u64 bdi_id, int memcg_id, struct cgroup_subsys_state *memcg_css; struct bdi_writeback *wb; struct wb_writeback_work *work; - unsigned long dirty; + long dirty; int ret; /* lookup bdi and memcg */ @@ -1210,16 +1210,13 @@ int cgroup_writeback_by_id(u64 bdi_id, int memcg_id, } /* - * The caller is attempting to write out most of - * the currently dirty pages. Let's take the current dirty page - * count and inflate it by 25% which should be large enough to - * flush out most dirty pages while avoiding getting livelocked by - * concurrent dirtiers. - * - * BTW the memcg stats are flushed periodically and this is best-effort - * estimation, so some potential error is ok. + * The caller is attempting to write out most of the target wb's + * currently dirty pages. Size the work from the wb's reclaimable pages + * and inflate the count by 25%, which should be large enough to flush + * out most dirty pages while avoiding getting livelocked by concurrent + * dirtiers. */ - dirty = memcg_page_state(mem_cgroup_from_css(memcg_css), NR_FILE_DIRTY); + dirty = wb_stat_sum(wb, WB_RECLAIMABLE); dirty = dirty * 10 / 8; /* issue the writeback work */ diff --git a/fs/fs_pin.c b/fs/fs_pin.c index 47ef3c71ce90..1a508f2167e0 100644 --- a/fs/fs_pin.c +++ b/fs/fs_pin.c @@ -1,5 +1,6 @@ // SPDX-License-Identifier: GPL-2.0 #include <linux/fs.h> +#include <linux/rculist.h> #include <linux/sched.h> #include <linux/slab.h> #include "internal.h" @@ -22,8 +23,8 @@ void pin_remove(struct fs_pin *pin) void pin_insert(struct fs_pin *pin, struct vfsmount *m) { spin_lock(&pin_lock); - hlist_add_head(&pin->s_list, &m->mnt_sb->s_pins); - hlist_add_head(&pin->m_list, &real_mount(m)->mnt_pins); + hlist_add_head_rcu(&pin->s_list, &m->mnt_sb->s_pins); + hlist_add_head_rcu(&pin->m_list, &real_mount(m)->mnt_pins); spin_unlock(&pin_lock); } @@ -73,7 +74,7 @@ void mnt_pin_kill(struct mount *m) while (1) { struct hlist_node *p; rcu_read_lock(); - p = READ_ONCE(m->mnt_pins.first); + p = rcu_dereference(hlist_first_rcu(&m->mnt_pins)); if (!p) { rcu_read_unlock(); break; @@ -87,7 +88,7 @@ void group_pin_kill(struct hlist_head *p) while (1) { struct hlist_node *q; rcu_read_lock(); - q = READ_ONCE(p->first); + q = rcu_dereference(hlist_first_rcu(p)); if (!q) { rcu_read_unlock(); break; diff --git a/fs/fuse/Kconfig b/fs/fuse/Kconfig index 3a4ae632c94a..8c41838b1c3d 100644 --- a/fs/fuse/Kconfig +++ b/fs/fuse/Kconfig @@ -40,9 +40,9 @@ config VIRTIO_FS If you want to share files between guests or with the host, answer Y or M. -config FUSE_DAX +config FUSE_VDAX bool "Virtio Filesystem Direct Host Memory Access support" - default y + default FUSE_DAX select INTERVAL_TREE depends on VIRTIO_FS depends on FS_DAX @@ -54,6 +54,10 @@ config FUSE_DAX If you want to allow mounting a Virtio Filesystem with the "dax" option, answer Y. +config FUSE_DAX + bool + transitional + config FUSE_PASSTHROUGH bool "FUSE passthrough operations support" default y diff --git a/fs/fuse/Makefile b/fs/fuse/Makefile index 245e67852b03..5858feafa916 100644 --- a/fs/fuse/Makefile +++ b/fs/fuse/Makefile @@ -14,7 +14,7 @@ fuse-y := trace.o # put trace.o first so we see ftrace errors sooner fuse-y += dev.o dir.o file.o inode.o control.o xattr.o acl.o readdir.o ioctl.o req_timeout.o req.o fuse-y += poll.o notify.o fuse-y += iomode.o -fuse-$(CONFIG_FUSE_DAX) += dax.o +fuse-$(CONFIG_FUSE_VDAX) += dax.o fuse-$(CONFIG_FUSE_PASSTHROUGH) += passthrough.o backing.o fuse-$(CONFIG_SYSCTL) += sysctl.o fuse-$(CONFIG_FUSE_IO_URING) += dev_uring.o diff --git a/fs/fuse/acl.c b/fs/fuse/acl.c index 31fb50e16aed..738abed9a816 100644 --- a/fs/fuse/acl.c +++ b/fs/fuse/acl.c @@ -62,7 +62,7 @@ static inline bool fuse_no_acl(const struct fuse_conn *fc, return !fc->posix_acl && (i_user_ns(inode) != &init_user_ns); } -struct posix_acl *fuse_get_acl(struct mnt_idmap *idmap, +struct posix_acl *fuse_get_acl(const struct mnt_idmap *idmap, struct dentry *dentry, int type) { struct inode *inode = d_inode(dentry); @@ -90,7 +90,7 @@ struct posix_acl *fuse_get_inode_acl(struct inode *inode, int type, bool rcu) return __fuse_get_acl(fc, inode, type, rcu); } -int fuse_set_acl(struct mnt_idmap *idmap, struct dentry *dentry, +int fuse_set_acl(const struct mnt_idmap *idmap, struct dentry *dentry, struct posix_acl *acl, int type) { struct inode *inode = d_inode(dentry); diff --git a/fs/fuse/backing.c b/fs/fuse/backing.c index 472b6afa7dff..433fa3098d71 100644 --- a/fs/fuse/backing.c +++ b/fs/fuse/backing.c @@ -10,7 +10,7 @@ #include <linux/file.h> -struct fuse_backing *fuse_backing_get(struct fuse_backing *fb) +static struct fuse_backing *fuse_backing_get(struct fuse_backing *fb) { if (fb && refcount_inc_not_zero(&fb->count)) return fb; diff --git a/fs/fuse/dax.c b/fs/fuse/dax.c index a5994f1c637d..a2bb9c041a4c 100644 --- a/fs/fuse/dax.c +++ b/fs/fuse/dax.c @@ -1,6 +1,6 @@ // SPDX-License-Identifier: GPL-2.0 /* - * dax: direct host memory access + * dax: direct host memory access for virtiofs * Copyright (C) 2020 Red Hat, Inc. */ @@ -59,7 +59,7 @@ struct fuse_dax_mapping { }; /* Per-inode dax map */ -struct fuse_inode_dax { +struct fuse_inode_vdax { /* Semaphore to protect modifications to the dmap tree */ struct rw_semaphore sem; @@ -68,7 +68,7 @@ struct fuse_inode_dax { unsigned long nr; }; -struct fuse_conn_dax { +struct fuse_conn_vdax { /* DAX device */ struct dax_device *dev; @@ -102,10 +102,10 @@ node_to_dmap(struct interval_tree_node *node) } static struct fuse_dax_mapping * -alloc_dax_mapping_reclaim(struct fuse_conn_dax *fcd, struct inode *inode); +alloc_dax_mapping_reclaim(struct fuse_conn_vdax *fcd, struct inode *inode); static void -__kick_dmap_free_worker(struct fuse_conn_dax *fcd, unsigned long delay_ms) +__kick_dmap_free_worker(struct fuse_conn_vdax *fcd, unsigned long delay_ms) { unsigned long free_threshold; @@ -117,7 +117,7 @@ __kick_dmap_free_worker(struct fuse_conn_dax *fcd, unsigned long delay_ms) msecs_to_jiffies(delay_ms)); } -static void kick_dmap_free_worker(struct fuse_conn_dax *fcd, +static void kick_dmap_free_worker(struct fuse_conn_vdax *fcd, unsigned long delay_ms) { spin_lock(&fcd->lock); @@ -125,7 +125,7 @@ static void kick_dmap_free_worker(struct fuse_conn_dax *fcd, spin_unlock(&fcd->lock); } -static struct fuse_dax_mapping *alloc_dax_mapping(struct fuse_conn_dax *fcd) +static struct fuse_dax_mapping *alloc_dax_mapping(struct fuse_conn_vdax *fcd) { struct fuse_dax_mapping *dmap; @@ -144,7 +144,7 @@ static struct fuse_dax_mapping *alloc_dax_mapping(struct fuse_conn_dax *fcd) } /* This assumes fcd->lock is held */ -static void __dmap_remove_busy_list(struct fuse_conn_dax *fcd, +static void __dmap_remove_busy_list(struct fuse_conn_vdax *fcd, struct fuse_dax_mapping *dmap) { list_del_init(&dmap->busy_list); @@ -152,7 +152,7 @@ static void __dmap_remove_busy_list(struct fuse_conn_dax *fcd, fcd->nr_busy_ranges--; } -static void dmap_remove_busy_list(struct fuse_conn_dax *fcd, +static void dmap_remove_busy_list(struct fuse_conn_vdax *fcd, struct fuse_dax_mapping *dmap) { spin_lock(&fcd->lock); @@ -161,7 +161,7 @@ static void dmap_remove_busy_list(struct fuse_conn_dax *fcd, } /* This assumes fcd->lock is held */ -static void __dmap_add_to_free_pool(struct fuse_conn_dax *fcd, +static void __dmap_add_to_free_pool(struct fuse_conn_vdax *fcd, struct fuse_dax_mapping *dmap) { list_add_tail(&dmap->list, &fcd->free_ranges); @@ -169,7 +169,7 @@ static void __dmap_add_to_free_pool(struct fuse_conn_dax *fcd, wake_up(&fcd->range_waitq); } -static void dmap_add_to_free_pool(struct fuse_conn_dax *fcd, +static void dmap_add_to_free_pool(struct fuse_conn_vdax *fcd, struct fuse_dax_mapping *dmap) { /* Return fuse_dax_mapping to free list */ @@ -183,7 +183,7 @@ static int fuse_setup_one_mapping(struct inode *inode, unsigned long start_idx, bool upgrade) { struct fuse_mount *fm = get_fuse_mount(inode); - struct fuse_conn_dax *fcd = fm->fc->dax; + struct fuse_conn_vdax *fcd = fm->fc->vdax; struct fuse_inode *fi = get_fuse_inode(inode); struct fuse_setupmapping_in inarg; loff_t offset = start_idx << FUSE_DAX_SHIFT; @@ -218,9 +218,9 @@ static int fuse_setup_one_mapping(struct inode *inode, unsigned long start_idx, */ dmap->inode = inode; dmap->itn.start = dmap->itn.last = start_idx; - /* Protected by fi->dax->sem */ - interval_tree_insert(&dmap->itn, &fi->dax->tree); - fi->dax->nr++; + /* Protected by fi->vdax->sem */ + interval_tree_insert(&dmap->itn, &fi->vdax->tree); + fi->vdax->nr++; spin_lock(&fcd->lock); list_add_tail(&dmap->busy_list, &fcd->busy_ranges); fcd->nr_busy_ranges++; @@ -288,7 +288,7 @@ out: * Cleanup dmap entry and add back to free list. This should be called with * fcd->lock held. */ -static void dmap_reinit_add_to_free_pool(struct fuse_conn_dax *fcd, +static void dmap_reinit_add_to_free_pool(struct fuse_conn_vdax *fcd, struct fuse_dax_mapping *dmap) { pr_debug("fuse: freeing memory range start_idx=0x%lx end_idx=0x%lx window_offset=0x%llx length=0x%llx\n", @@ -306,7 +306,7 @@ static void dmap_reinit_add_to_free_pool(struct fuse_conn_dax *fcd, * called from evict_inode() path where we know all dmap entries can be * reclaimed. */ -static void inode_reclaim_dmap_range(struct fuse_conn_dax *fcd, +static void inode_reclaim_dmap_range(struct fuse_conn_vdax *fcd, struct inode *inode, loff_t start, loff_t end) { @@ -319,14 +319,14 @@ static void inode_reclaim_dmap_range(struct fuse_conn_dax *fcd, struct interval_tree_node *node; while (1) { - node = interval_tree_iter_first(&fi->dax->tree, start_idx, + node = interval_tree_iter_first(&fi->vdax->tree, start_idx, end_idx); if (!node) break; dmap = node_to_dmap(node); /* inode is going away. There should not be any users of dmap */ WARN_ON(refcount_read(&dmap->refcnt) > 1); - interval_tree_remove(&dmap->itn, &fi->dax->tree); + interval_tree_remove(&dmap->itn, &fi->vdax->tree); num++; list_add(&dmap->list, &to_remove); } @@ -335,8 +335,8 @@ static void inode_reclaim_dmap_range(struct fuse_conn_dax *fcd, if (list_empty(&to_remove)) return; - WARN_ON(fi->dax->nr < num); - fi->dax->nr -= num; + WARN_ON(fi->vdax->nr < num); + fi->vdax->nr -= num; err = dmap_removemapping_list(inode, num, &to_remove); if (err && err != -ENOTCONN) { pr_warn("Failed to removemappings. start=0x%llx end=0x%llx\n", @@ -367,11 +367,11 @@ static int dmap_removemapping_one(struct inode *inode, /* * It is called from evict_inode() and by that time inode is going away. So - * this function does not take any locks like fi->dax->sem for traversing + * this function does not take any locks like fi->vdax->sem for traversing * that fuse inode interval tree. If that lock is taken then lock validator * complains of deadlock situation w.r.t fs_reclaim lock. */ -void fuse_dax_inode_cleanup(struct inode *inode) +void fuse_vdax_inode_cleanup(struct inode *inode) { struct fuse_conn *fc = get_fuse_conn(inode); struct fuse_inode *fi = get_fuse_inode(inode); @@ -381,8 +381,8 @@ void fuse_dax_inode_cleanup(struct inode *inode) * before we arrive here. So we should not have to worry about any * pages/exception entries still associated with inode. */ - inode_reclaim_dmap_range(fc->dax, inode, 0, -1); - WARN_ON(fi->dax->nr); + inode_reclaim_dmap_range(fc->vdax, inode, 0, -1); + WARN_ON(fi->vdax->nr); } static void fuse_fill_iomap_hole(struct iomap *iomap, loff_t length) @@ -414,7 +414,7 @@ static void fuse_fill_iomap(struct inode *inode, loff_t pos, loff_t length, iomap->type = IOMAP_MAPPED; /* * increace refcnt so that reclaim code knows this dmap is in - * use. This assumes fi->dax->sem mutex is held either + * use. This assumes fi->vdax->sem mutex is held either * shared/exclusive. */ refcount_inc(&dmap->refcnt); @@ -434,7 +434,7 @@ static int fuse_setup_new_dax_mapping(struct inode *inode, loff_t pos, { struct fuse_inode *fi = get_fuse_inode(inode); struct fuse_conn *fc = get_fuse_conn(inode); - struct fuse_conn_dax *fcd = fc->dax; + struct fuse_conn_vdax *fcd = fc->vdax; struct fuse_dax_mapping *dmap, *alloc_dmap = NULL; int ret; bool writable = flags & IOMAP_WRITE; @@ -469,17 +469,17 @@ static int fuse_setup_new_dax_mapping(struct inode *inode, loff_t pos, * Take write lock so that only one caller can try to setup mapping * and other waits. */ - down_write(&fi->dax->sem); + down_write(&fi->vdax->sem); /* * We dropped lock. Check again if somebody else setup * mapping already. */ - node = interval_tree_iter_first(&fi->dax->tree, start_idx, start_idx); + node = interval_tree_iter_first(&fi->vdax->tree, start_idx, start_idx); if (node) { dmap = node_to_dmap(node); fuse_fill_iomap(inode, pos, length, iomap, dmap, flags); dmap_add_to_free_pool(fcd, alloc_dmap); - up_write(&fi->dax->sem); + up_write(&fi->vdax->sem); return 0; } @@ -488,11 +488,11 @@ static int fuse_setup_new_dax_mapping(struct inode *inode, loff_t pos, writable, false); if (ret < 0) { dmap_add_to_free_pool(fcd, alloc_dmap); - up_write(&fi->dax->sem); + up_write(&fi->vdax->sem); return ret; } fuse_fill_iomap(inode, pos, length, iomap, alloc_dmap, flags); - up_write(&fi->dax->sem); + up_write(&fi->vdax->sem); return 0; } @@ -510,14 +510,14 @@ static int fuse_upgrade_dax_mapping(struct inode *inode, loff_t pos, * Take exclusive lock so that only one caller can try to setup * mapping and others wait. */ - down_write(&fi->dax->sem); - node = interval_tree_iter_first(&fi->dax->tree, idx, idx); + down_write(&fi->vdax->sem); + node = interval_tree_iter_first(&fi->vdax->tree, idx, idx); /* We are holding either inode lock or invalidate_lock, and that should * ensure that dmap can't be truncated. We are holding a reference * on dmap and that should make sure it can't be reclaimed. So dmap * should still be there in tree despite the fact we dropped and - * re-acquired the fi->dax->sem lock. + * re-acquired the fi->vdax->sem lock. */ ret = -EIO; if (WARN_ON(!node)) @@ -526,7 +526,7 @@ static int fuse_upgrade_dax_mapping(struct inode *inode, loff_t pos, dmap = node_to_dmap(node); /* We took an extra reference on dmap to make sure its not reclaimd. - * Now we hold fi->dax->sem lock and that reference is not needed + * Now we hold fi->vdax->sem lock and that reference is not needed * anymore. Drop it. */ if (refcount_dec_and_test(&dmap->refcnt)) { @@ -551,7 +551,7 @@ static int fuse_upgrade_dax_mapping(struct inode *inode, loff_t pos, out_fill_iomap: fuse_fill_iomap(inode, pos, length, iomap, dmap, flags); out_err: - up_write(&fi->dax->sem); + up_write(&fi->vdax->sem); return ret; } @@ -576,7 +576,7 @@ static int fuse_iomap_begin(struct inode *inode, loff_t pos, loff_t length, iomap->offset = pos; iomap->flags = 0; iomap->bdev = NULL; - iomap->dax_dev = fc->dax->dev; + iomap->dax_dev = fc->vdax->dev; /* * Both read/write and mmap path can race here. So we need something @@ -585,33 +585,33 @@ static int fuse_iomap_begin(struct inode *inode, loff_t pos, loff_t length, * For now, use a semaphore for this. It probably needs to be * optimized later. */ - down_read(&fi->dax->sem); - node = interval_tree_iter_first(&fi->dax->tree, start_idx, start_idx); + down_read(&fi->vdax->sem); + node = interval_tree_iter_first(&fi->vdax->tree, start_idx, start_idx); if (node) { dmap = node_to_dmap(node); if (writable && !dmap->writable) { /* Upgrade read-only mapping to read-write. This will - * require exclusive fi->dax->sem lock as we don't want + * require exclusive fi->vdax->sem lock as we don't want * two threads to be trying to this simultaneously * for same dmap. So drop shared lock and acquire * exclusive lock. * - * Before dropping fi->dax->sem lock, take reference + * Before dropping fi->vdax->sem lock, take reference * on dmap so that its not freed by range reclaim. */ refcount_inc(&dmap->refcnt); - up_read(&fi->dax->sem); + up_read(&fi->vdax->sem); pr_debug("%s: Upgrading mapping at offset 0x%llx length 0x%llx\n", __func__, pos, length); return fuse_upgrade_dax_mapping(inode, pos, length, flags, iomap); } else { fuse_fill_iomap(inode, pos, length, iomap, dmap, flags); - up_read(&fi->dax->sem); + up_read(&fi->vdax->sem); return 0; } } else { - up_read(&fi->dax->sem); + up_read(&fi->vdax->sem); pr_debug("%s: no mapping at offset 0x%llx length 0x%llx\n", __func__, pos, length); if (pos >= i_size_read(inode)) @@ -668,14 +668,14 @@ static void fuse_wait_dax_page(struct inode *inode) } /* Should be called with mapping->invalidate_lock held exclusively. */ -int fuse_dax_break_layouts(struct inode *inode, u64 dmap_start, +int fuse_vdax_break_layouts(struct inode *inode, u64 dmap_start, u64 dmap_end) { return dax_break_layout(inode, dmap_start, dmap_end, fuse_wait_dax_page); } -ssize_t fuse_dax_read_iter(struct kiocb *iocb, struct iov_iter *to) +ssize_t fuse_vdax_read_iter(struct kiocb *iocb, struct iov_iter *to) { struct inode *inode = file_inode(iocb->ki_filp); ssize_t ret; @@ -715,7 +715,7 @@ static ssize_t fuse_dax_direct_write(struct kiocb *iocb, struct iov_iter *from) return ret; } -ssize_t fuse_dax_write_iter(struct kiocb *iocb, struct iov_iter *from) +ssize_t fuse_vdax_write_iter(struct kiocb *iocb, struct iov_iter *from) { struct inode *inode = file_inode(iocb->ki_filp); ssize_t ret; @@ -761,7 +761,7 @@ static vm_fault_t __fuse_dax_fault(struct vm_fault *vmf, unsigned int order, unsigned long pfn; int error = 0; struct fuse_conn *fc = get_fuse_conn(inode); - struct fuse_conn_dax *fcd = fc->dax; + struct fuse_conn_vdax *fcd = fc->vdax; bool retry = false; if (write) @@ -822,11 +822,12 @@ static const struct vm_operations_struct fuse_dax_vm_ops = { .pfn_mkwrite = fuse_dax_pfn_mkwrite, }; -int fuse_dax_mmap(struct file *file, struct vm_area_struct *vma) +int fuse_vdax_mmap(struct file *file, struct vm_area_struct *vma) { file_accessed(file); vma->vm_ops = &fuse_dax_vm_ops; vma_set_flags(vma, VMA_HUGEPAGE_BIT); + vma->vm_page_prot = pgprot_decrypted(vma->vm_page_prot); return 0; } @@ -869,8 +870,8 @@ static int reclaim_one_dmap_locked(struct inode *inode, return ret; /* Remove dax mapping from inode interval tree now */ - interval_tree_remove(&dmap->itn, &fi->dax->tree); - fi->dax->nr--; + interval_tree_remove(&dmap->itn, &fi->vdax->tree); + fi->vdax->nr--; /* It is possible that umount/shutdown has killed the fuse connection * and worker thread is trying to reclaim memory in parallel. Don't @@ -885,7 +886,7 @@ static int reclaim_one_dmap_locked(struct inode *inode, } /* Find first mapped dmap for an inode and return file offset. Caller needs - * to hold fi->dax->sem lock either shared or exclusive. + * to hold fi->vdax->sem lock either shared or exclusive. */ static struct fuse_dax_mapping *inode_lookup_first_dmap(struct inode *inode) { @@ -893,7 +894,7 @@ static struct fuse_dax_mapping *inode_lookup_first_dmap(struct inode *inode) struct fuse_dax_mapping *dmap; struct interval_tree_node *node; - for (node = interval_tree_iter_first(&fi->dax->tree, 0, -1); node; + for (node = interval_tree_iter_first(&fi->vdax->tree, 0, -1); node; node = interval_tree_iter_next(node, 0, -1)) { dmap = node_to_dmap(node); /* still in use. */ @@ -911,7 +912,7 @@ static struct fuse_dax_mapping *inode_lookup_first_dmap(struct inode *inode) * it back to free pool. */ static struct fuse_dax_mapping * -inode_inline_reclaim_one_dmap(struct fuse_conn_dax *fcd, struct inode *inode, +inode_inline_reclaim_one_dmap(struct fuse_conn_vdax *fcd, struct inode *inode, bool *retry) { struct fuse_inode *fi = get_fuse_inode(inode); @@ -924,14 +925,14 @@ inode_inline_reclaim_one_dmap(struct fuse_conn_dax *fcd, struct inode *inode, filemap_invalidate_lock(inode->i_mapping); /* Lookup a dmap and corresponding file offset to reclaim. */ - down_read(&fi->dax->sem); + down_read(&fi->vdax->sem); dmap = inode_lookup_first_dmap(inode); if (dmap) { start_idx = dmap->itn.start; dmap_start = start_idx << FUSE_DAX_SHIFT; dmap_end = dmap_start + FUSE_DAX_SZ - 1; } - up_read(&fi->dax->sem); + up_read(&fi->vdax->sem); if (!dmap) goto out_mmap_sem; @@ -939,16 +940,16 @@ inode_inline_reclaim_one_dmap(struct fuse_conn_dax *fcd, struct inode *inode, * Make sure there are no references to inode pages using * get_user_pages() */ - ret = fuse_dax_break_layouts(inode, dmap_start, dmap_end); + ret = fuse_vdax_break_layouts(inode, dmap_start, dmap_end); if (ret) { - pr_debug("fuse: fuse_dax_break_layouts() failed. err=%d\n", + pr_debug("fuse: fuse_vdax_break_layouts() failed. err=%d\n", ret); dmap = ERR_PTR(ret); goto out_mmap_sem; } - down_write(&fi->dax->sem); - node = interval_tree_iter_first(&fi->dax->tree, start_idx, start_idx); + down_write(&fi->vdax->sem); + node = interval_tree_iter_first(&fi->vdax->tree, start_idx, start_idx); /* Range already got reclaimed by somebody else */ if (!node) { if (retry) @@ -980,14 +981,14 @@ inode_inline_reclaim_one_dmap(struct fuse_conn_dax *fcd, struct inode *inode, __func__, inode, dmap->window_offset, dmap->length); out_write_dmap_sem: - up_write(&fi->dax->sem); + up_write(&fi->vdax->sem); out_mmap_sem: filemap_invalidate_unlock(inode->i_mapping); return dmap; } static struct fuse_dax_mapping * -alloc_dax_mapping_reclaim(struct fuse_conn_dax *fcd, struct inode *inode) +alloc_dax_mapping_reclaim(struct fuse_conn_vdax *fcd, struct inode *inode) { struct fuse_dax_mapping *dmap; struct fuse_inode *fi = get_fuse_inode(inode); @@ -1014,18 +1015,18 @@ alloc_dax_mapping_reclaim(struct fuse_conn_dax *fcd, struct inode *inode) * if a deadlock is possible if we sleep with * mapping->invalidate_lock held and worker to free memory * can't make progress due to unavailability of - * mapping->invalidate_lock. So sleep only if fi->dax->nr=0 + * mapping->invalidate_lock. So sleep only if fi->vdax->nr=0 */ if (retry) continue; /* * There are no mappings which can be reclaimed. Wait for one. - * We are not holding fi->dax->sem. So it is possible + * We are not holding fi->vdax->sem. So it is possible * that range gets added now. But as we are not holding * mapping->invalidate_lock, worker should still be able to * free up a range and wake us up. */ - if (!fi->dax->nr && !(fcd->nr_free_ranges > 0)) { + if (!fi->vdax->nr && !(fcd->nr_free_ranges > 0)) { if (wait_event_killable_exclusive(fcd->range_waitq, (fcd->nr_free_ranges > 0))) { return ERR_PTR(-EINTR); @@ -1034,7 +1035,7 @@ alloc_dax_mapping_reclaim(struct fuse_conn_dax *fcd, struct inode *inode) } } -static int lookup_and_reclaim_dmap_locked(struct fuse_conn_dax *fcd, +static int lookup_and_reclaim_dmap_locked(struct fuse_conn_vdax *fcd, struct inode *inode, unsigned long start_idx) { @@ -1044,7 +1045,7 @@ static int lookup_and_reclaim_dmap_locked(struct fuse_conn_dax *fcd, struct interval_tree_node *node; /* Find fuse dax mapping at file offset inode. */ - node = interval_tree_iter_first(&fi->dax->tree, start_idx, start_idx); + node = interval_tree_iter_first(&fi->vdax->tree, start_idx, start_idx); /* Range already got cleaned up by somebody else */ if (!node) @@ -1070,10 +1071,10 @@ static int lookup_and_reclaim_dmap_locked(struct fuse_conn_dax *fcd, * Free a range of memory. * Locking: * 1. Take mapping->invalidate_lock to block dax faults. - * 2. Take fi->dax->sem to protect interval tree and also to make sure + * 2. Take fi->vdax->sem to protect interval tree and also to make sure * read/write can not reuse a dmap which we might be freeing. */ -static int lookup_and_reclaim_dmap(struct fuse_conn_dax *fcd, +static int lookup_and_reclaim_dmap(struct fuse_conn_vdax *fcd, struct inode *inode, unsigned long start_idx, unsigned long end_idx) @@ -1084,22 +1085,22 @@ static int lookup_and_reclaim_dmap(struct fuse_conn_dax *fcd, loff_t dmap_end = (dmap_start + FUSE_DAX_SZ) - 1; filemap_invalidate_lock(inode->i_mapping); - ret = fuse_dax_break_layouts(inode, dmap_start, dmap_end); + ret = fuse_vdax_break_layouts(inode, dmap_start, dmap_end); if (ret) { - pr_debug("virtio_fs: fuse_dax_break_layouts() failed. err=%d\n", + pr_debug("virtio_fs: fuse_vdax_break_layouts() failed. err=%d\n", ret); goto out_mmap_sem; } - down_write(&fi->dax->sem); + down_write(&fi->vdax->sem); ret = lookup_and_reclaim_dmap_locked(fcd, inode, start_idx); - up_write(&fi->dax->sem); + up_write(&fi->vdax->sem); out_mmap_sem: filemap_invalidate_unlock(inode->i_mapping); return ret; } -static int try_to_free_dmap_chunks(struct fuse_conn_dax *fcd, +static int try_to_free_dmap_chunks(struct fuse_conn_vdax *fcd, unsigned long nr_to_free) { struct fuse_dax_mapping *dmap, *pos, *temp; @@ -1160,7 +1161,7 @@ static int try_to_free_dmap_chunks(struct fuse_conn_dax *fcd, static void fuse_dax_free_mem_worker(struct work_struct *work) { int ret; - struct fuse_conn_dax *fcd = container_of(work, struct fuse_conn_dax, + struct fuse_conn_vdax *fcd = container_of(work, struct fuse_conn_vdax, free_work.work); ret = try_to_free_dmap_chunks(fcd, FUSE_DAX_RECLAIM_CHUNK); if (ret) { @@ -1185,16 +1186,16 @@ static void fuse_free_dax_mem_ranges(struct list_head *mem_list) } } -void fuse_dax_conn_free(struct fuse_conn *fc) +void fuse_vdax_conn_free(struct fuse_conn *fc) { - if (fc->dax) { - fuse_free_dax_mem_ranges(&fc->dax->free_ranges); - kfree(fc->dax); - fc->dax = NULL; + if (fc->vdax) { + fuse_free_dax_mem_ranges(&fc->vdax->free_ranges); + kfree(fc->vdax); + fc->vdax = NULL; } } -static int fuse_dax_mem_range_init(struct fuse_conn_dax *fcd) +static int fuse_dax_mem_range_init(struct fuse_conn_vdax *fcd) { long nr_pages, nr_ranges; struct fuse_dax_mapping *range; @@ -1246,13 +1247,13 @@ out_err: return ret; } -int fuse_dax_conn_alloc(struct fuse_conn *fc, enum fuse_dax_mode dax_mode, +int fuse_vdax_conn_alloc(struct fuse_conn *fc, enum fuse_vdax_mode dax_mode, struct dax_device *dax_dev) { - struct fuse_conn_dax *fcd; + struct fuse_conn_vdax *fcd; int err; - fc->dax_mode = dax_mode; + fc->vdax_mode = dax_mode; if (!dax_dev) return 0; @@ -1269,22 +1270,22 @@ int fuse_dax_conn_alloc(struct fuse_conn *fc, enum fuse_dax_mode dax_mode, return err; } - fc->dax = fcd; + fc->vdax = fcd; return 0; } -bool fuse_dax_inode_alloc(struct super_block *sb, struct fuse_inode *fi) +bool fuse_vdax_inode_alloc(struct super_block *sb, struct fuse_inode *fi) { struct fuse_conn *fc = get_fuse_conn_super(sb); - fi->dax = NULL; - if (fc->dax) { - fi->dax = kzalloc_obj(*fi->dax, GFP_KERNEL_ACCOUNT); - if (!fi->dax) + fi->vdax = NULL; + if (fc->vdax) { + fi->vdax = kzalloc_obj(*fi->vdax, GFP_KERNEL_ACCOUNT); + if (!fi->vdax) return false; - init_rwsem(&fi->dax->sem); - fi->dax->tree = RB_ROOT_CACHED; + init_rwsem(&fi->vdax->sem); + fi->vdax->tree = RB_ROOT_CACHED; } return true; @@ -1298,26 +1299,26 @@ static const struct address_space_operations fuse_dax_file_aops = { static bool fuse_should_enable_dax(struct inode *inode, unsigned int flags) { struct fuse_conn *fc = get_fuse_conn(inode); - enum fuse_dax_mode dax_mode = fc->dax_mode; + enum fuse_vdax_mode dax_mode = fc->vdax_mode; - if (dax_mode == FUSE_DAX_NEVER) + if (dax_mode == FUSE_VDAX_NEVER) return false; /* - * fc->dax may be NULL in 'inode' mode when filesystem device doesn't + * fc->vdax may be NULL in 'inode' mode when filesystem device doesn't * support DAX, in which case it will silently fallback to 'never' mode. */ - if (!fc->dax) + if (!fc->vdax) return false; - if (dax_mode == FUSE_DAX_ALWAYS) + if (dax_mode == FUSE_VDAX_ALWAYS) return true; /* dax_mode is FUSE_DAX_INODE* */ - return fc->inode_dax && (flags & FUSE_ATTR_DAX); + return fc->inode_vdax && (flags & FUSE_ATTR_DAX); } -void fuse_dax_inode_init(struct inode *inode, unsigned int flags) +void fuse_vdax_inode_init(struct inode *inode, unsigned int flags) { if (!fuse_should_enable_dax(inode, flags)) return; @@ -1326,18 +1327,18 @@ void fuse_dax_inode_init(struct inode *inode, unsigned int flags) inode->i_data.a_ops = &fuse_dax_file_aops; } -void fuse_dax_dontcache(struct inode *inode, unsigned int flags) +void fuse_vdax_dontcache(struct inode *inode, unsigned int flags) { struct fuse_conn *fc = get_fuse_conn(inode); - if (fuse_is_inode_dax_mode(fc->dax_mode) && + if (fuse_is_inode_vdax_mode(fc->vdax_mode) && ((bool) IS_DAX(inode) != (bool) (flags & FUSE_ATTR_DAX))) d_mark_dontcache(inode); } -bool fuse_dax_check_alignment(struct fuse_conn *fc, unsigned int map_alignment) +bool fuse_vdax_check_alignment(struct fuse_conn *fc, unsigned int map_alignment) { - if (fc->dax && (map_alignment > FUSE_DAX_SHIFT)) { + if (fc->vdax && (map_alignment > FUSE_DAX_SHIFT)) { pr_warn("FUSE: map_alignment %u incompatible with dax mem range size %u\n", map_alignment, FUSE_DAX_SZ); return false; @@ -1345,12 +1346,12 @@ bool fuse_dax_check_alignment(struct fuse_conn *fc, unsigned int map_alignment) return true; } -void fuse_dax_cancel_work(struct fuse_conn *fc) +void fuse_vdax_cancel_work(struct fuse_conn *fc) { - struct fuse_conn_dax *fcd = fc->dax; + struct fuse_conn_vdax *fcd = fc->vdax; if (fcd) cancel_delayed_work_sync(&fcd->free_work); } -EXPORT_SYMBOL_GPL(fuse_dax_cancel_work); +EXPORT_SYMBOL_GPL(fuse_vdax_cancel_work); diff --git a/fs/fuse/dev.c b/fs/fuse/dev.c index 4fec31fc0b84..2d7ee5498f1c 100644 --- a/fs/fuse/dev.c +++ b/fs/fuse/dev.c @@ -1061,7 +1061,8 @@ static int fuse_copy_fill(struct fuse_copy_state *cs) err = iov_iter_get_pages2(cs->iter, &page, PAGE_SIZE, 1, &off); if (err < 0) return err; - BUG_ON(!err); + if (!err) + return -EIO; cs->len = err; cs->offset = off; cs->pg = page; @@ -1252,31 +1253,10 @@ static int fuse_ref_folio(struct fuse_copy_state *cs, struct folio *folio, * done atomically */ int fuse_copy_folio(struct fuse_copy_state *cs, struct folio **foliop, - unsigned offset, unsigned count, int zeroing) + unsigned offset, unsigned count) { int err; struct folio *folio = *foliop; - size_t size; - - if (folio) { - size = folio_size(folio); - if (zeroing && count < size) { - /* - * When the copy is skipped the folio already holds the - * payload, so only the bytes outside [offset, offset + - * count) may be zeroed. - * - * Otherwise, the whole folio is cleared first so that a - * failed copy leaves zeros rather than stale folio - * contents. - */ - if (cs->skip_folio_copy) - folio_zero_segments(folio, 0, offset, - offset + count, size); - else - folio_zero_range(folio, 0, size); - } - } while (!cs->skip_folio_copy && count) { if (cs->write && cs->pipebufs && folio) { @@ -1293,7 +1273,7 @@ int fuse_copy_folio(struct fuse_copy_state *cs, struct folio **foliop, } } else if (!cs->len) { if (cs->move_folios && folio && - offset == 0 && count == size) { + offset == 0 && count == folio_size(folio)) { err = fuse_try_move_folio(cs, foliop); if (err <= 0) return err; @@ -1309,7 +1289,8 @@ int fuse_copy_folio(struct fuse_copy_state *cs, struct folio **foliop, unsigned int copy = count; unsigned int bytes_copied; - if (folio_test_highmem(folio) && count > PAGE_SIZE - offset_in_page(offset)) + if (folio_test_partial_kmap(folio) && + count > PAGE_SIZE - offset_in_page(offset)) copy = PAGE_SIZE - offset_in_page(offset); bytes_copied = fuse_copy_do(cs, &buf, ©); @@ -1334,10 +1315,25 @@ static int fuse_copy_folios(struct fuse_copy_state *cs, unsigned nbytes, for (i = 0; i < ap->num_folios && (nbytes || zeroing); i++) { int err; + struct folio *folio = ap->folios[i]; unsigned int offset = ap->descs[i].offset; - unsigned int count = min(nbytes, ap->descs[i].length); + unsigned int length = ap->descs[i].length; + unsigned int count = min(nbytes, length); + + /* + * The reply may be shorter than what was asked for. The full + * descs[i].length is reported as read, so the tail bytes the + * server did not send are about to be marked uptodate and need + * to be zeroed. + * + * Only [offset, offset + length) can be touched since the + * rest of the folio can hold blocks that are already uptodate + * or dirty, and clearing those would lose data. + */ + if (folio && zeroing && count < length) + folio_zero_range(folio, offset + count, length - count); - err = fuse_copy_folio(cs, &ap->folios[i], offset, count, zeroing); + err = fuse_copy_folio(cs, &ap->folios[i], offset, count); if (err) return err; diff --git a/fs/fuse/dev.h b/fs/fuse/dev.h index 8d25378c0918..f6c47ae0395b 100644 --- a/fs/fuse/dev.h +++ b/fs/fuse/dev.h @@ -90,7 +90,7 @@ int fuse_backing_close(struct fuse_conn *fc, int backing_id); int fuse_copy_one(struct fuse_copy_state *cs, void *val, unsigned size); int fuse_copy_folio(struct fuse_copy_state *cs, struct folio **foliop, - unsigned offset, unsigned count, int zeroing); + unsigned offset, unsigned count); void fuse_copy_finish(struct fuse_copy_state *cs); #ifdef CONFIG_FUSE_IO_URING diff --git a/fs/fuse/dev_uring.c b/fs/fuse/dev_uring.c index c6dd420c4034..460776cd69cd 100644 --- a/fs/fuse/dev_uring.c +++ b/fs/fuse/dev_uring.c @@ -893,6 +893,12 @@ static int fuse_uring_args_to_ring(struct fuse_req *req, fuse_copy_finish(&cs); if (err) { pr_info_ratelimited("%s fuse_copy_args failed\n", __func__); + /* + * fuse-uring does not use pipe buffers, so -EIO here can only + * mean the in-args did not fit ent->payload + */ + if (err == -EIO && args->opcode == FUSE_SETXATTR) + err = -E2BIG; return err; } diff --git a/fs/fuse/dir.c b/fs/fuse/dir.c index e49b4e874b15..d11e2aabbfd5 100644 --- a/fs/fuse/dir.c +++ b/fs/fuse/dir.c @@ -751,7 +751,7 @@ static u32 fuse_ext_size(size_t size) /* * This adds just a single supplementary group that matches the parent's group. */ -static int get_create_supp_group(struct mnt_idmap *idmap, +static int get_create_supp_group(const struct mnt_idmap *idmap, struct inode *dir, struct fuse_in_arg *ext) { @@ -782,7 +782,7 @@ static int get_create_supp_group(struct mnt_idmap *idmap, return 0; } -static int get_create_ext(struct mnt_idmap *idmap, +static int get_create_ext(const struct mnt_idmap *idmap, struct fuse_args *args, struct inode *dir, struct dentry *dentry, umode_t mode) @@ -820,7 +820,7 @@ static void free_ext_value(struct fuse_args *args) * If the filesystem doesn't support this, then fall back to separate * 'mknod' + 'open' requests. */ -static int fuse_create_open(struct mnt_idmap *idmap, struct inode *dir, +static int fuse_create_open(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *entry, struct file *file, unsigned int flags, umode_t mode, u32 opcode) { @@ -934,14 +934,14 @@ out_err: return err; } -static int fuse_mknod(struct mnt_idmap *, struct inode *, struct dentry *, +static int fuse_mknod(const struct mnt_idmap *, struct inode *, struct dentry *, umode_t, dev_t); static int fuse_atomic_open(struct inode *dir, struct dentry *entry, struct file *file, unsigned flags, umode_t mode) { int err; - struct mnt_idmap *idmap = file_mnt_idmap(file); + const struct mnt_idmap *idmap = file_mnt_idmap(file); struct fuse_conn *fc = get_fuse_conn(dir); if (fuse_is_bad(dir)) @@ -980,7 +980,7 @@ mknod: /* * Code shared between mknod, mkdir, symlink and link */ -static struct dentry *create_new_entry(struct mnt_idmap *idmap, struct fuse_mount *fm, +static struct dentry *create_new_entry(const struct mnt_idmap *idmap, struct fuse_mount *fm, struct fuse_args *args, struct inode *dir, struct dentry *entry, umode_t mode) { @@ -1053,7 +1053,7 @@ static struct dentry *create_new_entry(struct mnt_idmap *idmap, struct fuse_moun return ERR_PTR(err); } -static int create_new_nondir(struct mnt_idmap *idmap, struct fuse_mount *fm, +static int create_new_nondir(const struct mnt_idmap *idmap, struct fuse_mount *fm, struct fuse_args *args, struct inode *dir, struct dentry *entry, umode_t mode) { @@ -1069,7 +1069,7 @@ static int create_new_nondir(struct mnt_idmap *idmap, struct fuse_mount *fm, return PTR_ERR(create_new_entry(idmap, fm, args, dir, entry, mode)); } -static int fuse_mknod(struct mnt_idmap *idmap, struct inode *dir, +static int fuse_mknod(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *entry, umode_t mode, dev_t rdev) { struct fuse_mknod_in inarg; @@ -1092,13 +1092,13 @@ static int fuse_mknod(struct mnt_idmap *idmap, struct inode *dir, return create_new_nondir(idmap, fm, &args, dir, entry, mode); } -static int fuse_create(struct mnt_idmap *idmap, struct inode *dir, +static int fuse_create(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *entry, umode_t mode) { return fuse_mknod(idmap, dir, entry, mode, 0); } -static int fuse_tmpfile(struct mnt_idmap *idmap, struct inode *dir, +static int fuse_tmpfile(const struct mnt_idmap *idmap, struct inode *dir, struct file *file, umode_t mode) { struct fuse_conn *fc = get_fuse_conn(dir); @@ -1116,7 +1116,7 @@ static int fuse_tmpfile(struct mnt_idmap *idmap, struct inode *dir, return err; } -static struct dentry *fuse_mkdir(struct mnt_idmap *idmap, struct inode *dir, +static struct dentry *fuse_mkdir(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *entry, umode_t mode) { struct fuse_mkdir_in inarg; @@ -1146,7 +1146,7 @@ static struct dentry *fuse_mkdir(struct mnt_idmap *idmap, struct inode *dir, return create_new_entry(idmap, fm, &args, dir, entry, S_IFDIR); } -static int fuse_symlink(struct mnt_idmap *idmap, struct inode *dir, +static int fuse_symlink(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *entry, const char *link) { struct fuse_mount *fm = get_fuse_mount(dir); @@ -1256,9 +1256,10 @@ static int fuse_rmdir(struct inode *dir, struct dentry *entry) return err; } -static int fuse_rename_common(struct mnt_idmap *idmap, struct inode *olddir, struct dentry *oldent, - struct inode *newdir, struct dentry *newent, - unsigned int flags, int opcode, size_t argsize) +static int fuse_rename_common(const struct mnt_idmap *idmap, struct inode *olddir, + struct dentry *oldent, struct inode *newdir, + struct dentry *newent, unsigned int flags, + int opcode, size_t argsize) { int err; struct fuse_rename2_in inarg; @@ -1306,7 +1307,7 @@ static int fuse_rename_common(struct mnt_idmap *idmap, struct inode *olddir, str return err; } -static int fuse_rename2(struct mnt_idmap *idmap, struct inode *olddir, +static int fuse_rename2(const struct mnt_idmap *idmap, struct inode *olddir, struct dentry *oldent, struct inode *newdir, struct dentry *newent, unsigned int flags) { @@ -1375,7 +1376,7 @@ out: return err; } -static void fuse_fillattr(struct mnt_idmap *idmap, struct inode *inode, +static void fuse_fillattr(const struct mnt_idmap *idmap, struct inode *inode, struct fuse_attr *attr, struct kstat *stat) { unsigned int blkbits; @@ -1429,7 +1430,7 @@ static void fuse_statx_to_attr(struct fuse_statx *sx, struct fuse_attr *attr) attr->blksize = sx->blksize; } -static int fuse_do_statx(struct mnt_idmap *idmap, struct inode *inode, +static int fuse_do_statx(const struct mnt_idmap *idmap, struct inode *inode, struct file *file, struct kstat *stat) { int err; @@ -1490,7 +1491,7 @@ static int fuse_do_statx(struct mnt_idmap *idmap, struct inode *inode, return 0; } -static int fuse_do_getattr(struct mnt_idmap *idmap, struct inode *inode, +static int fuse_do_getattr(const struct mnt_idmap *idmap, struct inode *inode, struct kstat *stat, struct file *file) { int err; @@ -1536,7 +1537,7 @@ static int fuse_do_getattr(struct mnt_idmap *idmap, struct inode *inode, return err; } -static int fuse_update_get_attr(struct mnt_idmap *idmap, struct inode *inode, +static int fuse_update_get_attr(const struct mnt_idmap *idmap, struct inode *inode, struct file *file, struct kstat *stat, u32 request_mask, unsigned int flags) { @@ -1762,7 +1763,7 @@ static int fuse_perm_getattr(struct inode *inode, int mask) * access request is sent. Execute permission is still checked * locally based on file mode. */ -static int fuse_permission(struct mnt_idmap *idmap, +static int fuse_permission(const struct mnt_idmap *idmap, struct inode *inode, int mask) { struct fuse_conn *fc = get_fuse_conn(inode); @@ -2000,7 +2001,7 @@ static bool update_mtime(unsigned ivalid, bool trust_local_mtime) return true; } -static void iattr_to_fattr(struct mnt_idmap *idmap, struct fuse_conn *fc, +static void iattr_to_fattr(const struct mnt_idmap *idmap, struct fuse_conn *fc, struct iattr *iattr, struct fuse_setattr_in *arg, bool trust_local_cmtime) { @@ -2142,7 +2143,7 @@ int fuse_flush_times(struct inode *inode, struct fuse_file *ff) * vmtruncate() doesn't allow for this case, so do the rlimit checking * and the actual truncation by hand. */ -int fuse_do_setattr(struct mnt_idmap *idmap, struct dentry *dentry, +int fuse_do_setattr(const struct mnt_idmap *idmap, struct dentry *dentry, struct iattr *attr, struct file *file) { struct inode *inode = d_inode(dentry); @@ -2174,10 +2175,10 @@ int fuse_do_setattr(struct mnt_idmap *idmap, struct dentry *dentry, is_truncate = true; } - if (FUSE_IS_DAX(inode) && is_truncate) { + if (FUSE_IS_VDAX(inode) && is_truncate) { filemap_invalidate_lock(mapping); fault_blocked = true; - err = fuse_dax_break_layouts(inode, 0, -1); + err = fuse_vdax_break_layouts(inode, 0, -1); if (err) goto unlock; } @@ -2323,7 +2324,7 @@ unlock: return err; } -static int fuse_setattr(struct mnt_idmap *idmap, struct dentry *entry, +static int fuse_setattr(const struct mnt_idmap *idmap, struct dentry *entry, struct iattr *attr) { struct inode *inode = d_inode(entry); @@ -2386,7 +2387,7 @@ static int fuse_setattr(struct mnt_idmap *idmap, struct dentry *entry, return ret; } -static int fuse_getattr(struct mnt_idmap *idmap, +static int fuse_getattr(const struct mnt_idmap *idmap, const struct path *path, struct kstat *stat, u32 request_mask, unsigned int flags) { diff --git a/fs/fuse/file.c b/fs/fuse/file.c index 9a36d0329e22..3d209e2b71ba 100644 --- a/fs/fuse/file.c +++ b/fs/fuse/file.c @@ -119,7 +119,7 @@ static void fuse_file_put(struct fuse_file *ff, bool sync) * DAX inodes may need to issue a number of synchronous * request for clearing the mappings. */ - if (ra && ra->inode && FUSE_IS_DAX(ra->inode)) + if (ra && ra->inode && FUSE_IS_VDAX(ra->inode)) args->may_block = true; args->end = fuse_release_end; if (fuse_simple_background(ff->fm, args, @@ -256,7 +256,7 @@ static int fuse_open(struct inode *inode, struct file *file) int err; bool is_truncate = (file->f_flags & O_TRUNC) && fc->atomic_o_trunc; bool is_wb_truncate = is_truncate && fc->writeback_cache; - bool dax_truncate = is_truncate && FUSE_IS_DAX(inode); + bool vdax_truncate = is_truncate && FUSE_IS_VDAX(inode); if (fuse_is_bad(inode)) return -EIO; @@ -265,17 +265,17 @@ static int fuse_open(struct inode *inode, struct file *file) if (err) return err; - if (is_wb_truncate || dax_truncate) + if (is_wb_truncate || vdax_truncate) inode_lock(inode); - if (dax_truncate) { + if (vdax_truncate) { filemap_invalidate_lock(inode->i_mapping); - err = fuse_dax_break_layouts(inode, 0, -1); + err = fuse_vdax_break_layouts(inode, 0, -1); if (err) goto out_unlock; } - if (is_wb_truncate || dax_truncate) + if (is_wb_truncate || vdax_truncate) fuse_set_nowrite(inode); err = fuse_do_open(fm, get_node_id(inode), file, false); @@ -288,7 +288,7 @@ static int fuse_open(struct inode *inode, struct file *file) fuse_truncate_update_attr(inode, file); } - if (is_wb_truncate || dax_truncate) + if (is_wb_truncate || vdax_truncate) fuse_release_nowrite(inode); if (!err) { if (is_truncate) @@ -297,9 +297,9 @@ static int fuse_open(struct inode *inode, struct file *file) invalidate_inode_pages2(inode->i_mapping); } out_unlock: - if (dax_truncate) + if (vdax_truncate) filemap_invalidate_unlock(inode->i_mapping); - if (is_wb_truncate || dax_truncate) + if (is_wb_truncate || vdax_truncate) inode_unlock(inode); return err; @@ -311,7 +311,7 @@ static void fuse_prepare_release(struct fuse_inode *fi, struct fuse_file *ff, struct fuse_conn *fc = ff->fm->fc; struct fuse_release_args *ra = &ff->args->release_args; - if (fuse_file_passthrough(ff)) + if (fuse_is_passthrough(ff)) fuse_passthrough_release(ff, fuse_inode_backing(fi)); /* Inode is NULL on error path of fuse_create_open() */ @@ -689,7 +689,7 @@ static void fuse_aio_complete(struct fuse_io_priv *io, int err, ssize_t pos) struct address_space *mapping = io->iocb->ki_filp->f_mapping; ssize_t res = fuse_get_res_by_io(io); - if (res >= 0) { + if (res >= 0 && io->write) { struct fuse_conn *fc = get_fuse_conn(inode); struct fuse_inode *fi = get_fuse_inode(inode); @@ -864,18 +864,29 @@ static int fuse_do_readfolio(struct file *file, struct folio *folio, attr_ver = fuse_get_attr_version(fm->fc); - /* Don't overflow end offset */ - if (pos + (desc.length - 1) == LLONG_MAX) - desc.length--; + /* + * Don't overflow end offset. + * + * Ask the server for len - 1 bytes. desc.length still holds the full + * length. When the reply comes back, it will be one byte shorter than + * desc.length and fuse_copy_folios() will zero that last byte. + * + * For this reason, desc.length must not be decremented too. The caller + * reports the full length to iomap_finish_folio_read(), which marks + * every block it covers uptodate. Shortening the descriptor would + * suppress zeroing and leave the last byte holding stale data. + */ + if (pos + (len - 1) == LLONG_MAX) + len--; - fuse_read_args_fill(&ia, file, pos, desc.length, FUSE_READ); + fuse_read_args_fill(&ia, file, pos, len, FUSE_READ); res = fuse_simple_request(fm, &ia.ap.args); if (res < 0) return res; /* * Short read means EOF. If file size is larger, truncate it */ - if (res < desc.length) + if (res < len) fuse_short_read(inode, attr_ver, res, &ia.ap); return 0; @@ -1068,11 +1079,20 @@ static void fuse_send_readpages(struct fuse_io_args *ia, struct file *file, ap->args.page_zeroing = true; ap->args.page_replace = true; - /* Don't overflow end offset */ - if (pos + (count - 1) == LLONG_MAX) { + /* + * Don't overflow end offset. + * + * Ask the server for count - 1 bytes. The reply is then one byte + * shorter than what the descriptor lengths add up to, so + * fuse_copy_folios() zeroes the last byte when it walks the folios. + * + * ap->descs[] must not be decremented here. It is what + * fuse_readpages_end() reports back to iomap_finish_folio_read(), and + * iomap has already accounted the full descriptor length, so shortening + * it would leave ifs->read_bytes_pending nonzero and the folio locked. + */ + if (pos + (count - 1) == LLONG_MAX) count--; - ap->descs[ap->num_folios - 1].length--; - } WARN_ON((loff_t) (pos + count) < 0); fuse_read_args_fill(ia, file, pos, count, FUSE_READ); @@ -1487,7 +1507,7 @@ static const struct iomap_write_ops fuse_iomap_write_ops = { static ssize_t fuse_cache_write_iter(struct kiocb *iocb, struct iov_iter *from) { struct file *file = iocb->ki_filp; - struct mnt_idmap *idmap = file_mnt_idmap(file); + const struct mnt_idmap *idmap = file_mnt_idmap(file); struct address_space *mapping = file->f_mapping; ssize_t written = 0; struct inode *inode = mapping->host; @@ -1835,13 +1855,13 @@ static ssize_t fuse_file_read_iter(struct kiocb *iocb, struct iov_iter *to) if (fuse_is_bad(inode)) return -EIO; - if (FUSE_IS_DAX(inode)) - return fuse_dax_read_iter(iocb, to); + if (FUSE_IS_VDAX(inode)) + return fuse_vdax_read_iter(iocb, to); /* FOPEN_DIRECT_IO overrides FOPEN_PASSTHROUGH */ if (ff->open_flags & FOPEN_DIRECT_IO) return fuse_direct_read_iter(iocb, to); - else if (fuse_file_passthrough(ff)) + else if (fuse_is_passthrough(ff)) return fuse_passthrough_read_iter(iocb, to); else return fuse_cache_read_iter(iocb, to); @@ -1856,13 +1876,13 @@ static ssize_t fuse_file_write_iter(struct kiocb *iocb, struct iov_iter *from) if (fuse_is_bad(inode)) return -EIO; - if (FUSE_IS_DAX(inode)) - return fuse_dax_write_iter(iocb, from); + if (FUSE_IS_VDAX(inode)) + return fuse_vdax_write_iter(iocb, from); /* FOPEN_DIRECT_IO overrides FOPEN_PASSTHROUGH */ if (ff->open_flags & FOPEN_DIRECT_IO) return fuse_direct_write_iter(iocb, from); - else if (fuse_file_passthrough(ff)) + else if (fuse_is_passthrough(ff)) return fuse_passthrough_write_iter(iocb, from); else return fuse_cache_write_iter(iocb, from); @@ -1875,7 +1895,10 @@ static ssize_t fuse_splice_read(struct file *in, loff_t *ppos, struct fuse_file *ff = in->private_data; /* FOPEN_DIRECT_IO overrides FOPEN_PASSTHROUGH */ - if (fuse_file_passthrough(ff) && !(ff->open_flags & FOPEN_DIRECT_IO)) + + if (ff->open_flags & FOPEN_DIRECT_IO) + return copy_splice_read(in, ppos, pipe, len, flags); + else if (fuse_is_passthrough(ff)) return fuse_passthrough_splice_read(in, ppos, pipe, len, flags); else return filemap_splice_read(in, ppos, pipe, len, flags); @@ -1887,7 +1910,7 @@ static ssize_t fuse_splice_write(struct pipe_inode_info *pipe, struct file *out, struct fuse_file *ff = out->private_data; /* FOPEN_DIRECT_IO overrides FOPEN_PASSTHROUGH */ - if (fuse_file_passthrough(ff) && !(ff->open_flags & FOPEN_DIRECT_IO)) + if (fuse_is_passthrough(ff) && !(ff->open_flags & FOPEN_DIRECT_IO)) return fuse_passthrough_splice_write(pipe, out, ppos, len, flags); else return iter_file_splice_write(pipe, out, ppos, len, flags); @@ -2394,15 +2417,15 @@ static int fuse_file_mmap(struct file *file, struct vm_area_struct *vma) int rc; /* DAX mmap is superior to direct_io mmap */ - if (FUSE_IS_DAX(inode)) - return fuse_dax_mmap(file, vma); + if (FUSE_IS_VDAX(inode)) + return fuse_vdax_mmap(file, vma); /* * If inode is in passthrough io mode, because it has some file open * in passthrough mode, either mmap to backing file or fail mmap, * because mixing cached mmap and passthrough io mode is not allowed. */ - if (fuse_file_passthrough(ff)) + if (fuse_is_passthrough(ff)) return fuse_passthrough_mmap(file, vma); else if (fuse_inode_backing(get_fuse_inode(inode))) return -ENODEV; @@ -2844,7 +2867,7 @@ static long fuse_file_fallocate(struct file *file, int mode, loff_t offset, .mode = mode }; int err; - bool block_faults = FUSE_IS_DAX(inode) && + bool block_faults = FUSE_IS_VDAX(inode) && (!(mode & FALLOC_FL_KEEP_SIZE) || (mode & (FALLOC_FL_PUNCH_HOLE | FALLOC_FL_ZERO_RANGE))); @@ -2858,7 +2881,7 @@ static long fuse_file_fallocate(struct file *file, int mode, loff_t offset, inode_lock(inode); if (block_faults) { filemap_invalidate_lock(inode->i_mapping); - err = fuse_dax_break_layouts(inode, 0, -1); + err = fuse_vdax_break_layouts(inode, 0, -1); if (err) goto out; } @@ -2978,14 +3001,16 @@ static ssize_t __fuse_copy_file_range(struct file *file_in, loff_t pos_in, /* * Write out dirty pages in the destination file before sending the COPY - * request to userspace. After the request is completed, truncate off - * pages (including partial ones) from the cache that have been copied, - * since these contain stale data at that point. + * request to userspace. After the request is completed, drop the + * folios covering the copied range from the cache, since these contain + * stale data at that point. * - * This should be mostly correct, but if the COPY writes to partial - * pages (at the start or end) and the parts not covered by the COPY are + * This should be mostly correct, but if the COPY writes to a partial + * folio (at the start or end) and the parts not covered by the COPY are * written through a memory map after calling fuse_writeback_range(), - * then these partial page modifications will be lost on truncation. + * then that folio is laundered before it is dropped, so the memory map + * modifications are written back over the range the server just copied + * into and the copied data is lost. * * It is unlikely that someone would rely on such mixed style * modifications. Yet this does give less guarantees than if the @@ -3038,9 +3063,10 @@ fallback: goto out; } - truncate_inode_pages_range(inode_out->i_mapping, - ALIGN_DOWN(pos_out, PAGE_SIZE), - ALIGN(pos_out + bytes_copied, PAGE_SIZE) - 1); + if (bytes_copied) + invalidate_inode_pages2_range(inode_out->i_mapping, + pos_out >> PAGE_SHIFT, + (pos_out + bytes_copied - 1) >> PAGE_SHIFT); file_update_time(file_out); fuse_write_update_attr(inode_out, pos_out + bytes_copied, bytes_copied); @@ -3113,6 +3139,7 @@ void fuse_init_file_inode(struct inode *inode, unsigned int flags) { struct fuse_inode *fi = get_fuse_inode(inode); struct fuse_conn *fc = get_fuse_conn(inode); + unsigned int max_folio_pages; inode->i_fop = &fuse_file_operations; inode->i_data.a_ops = &fuse_file_aops; @@ -3126,6 +3153,20 @@ void fuse_init_file_inode(struct inode *inode, unsigned int flags) init_waitqueue_head(&fi->page_waitq); init_waitqueue_head(&fi->direct_io_waitq); - if (IS_ENABLED(CONFIG_FUSE_DAX)) - fuse_dax_inode_init(inode, flags); + if (IS_ENABLED(CONFIG_FUSE_VDAX)) + fuse_vdax_inode_init(inode, flags); + + if (FUSE_IS_VDAX(inode)) + return; + + /* + * A folio is never split across requests so cap the order so that one + * always fits in a single request. + */ + max_folio_pages = min3(fc->max_write >> PAGE_SHIFT, + fc->max_read >> PAGE_SHIFT, fc->max_pages); + + if (max_folio_pages) + mapping_set_folio_order_range(inode->i_mapping, 0, + ilog2(max_folio_pages)); } diff --git a/fs/fuse/fuse_i.h b/fs/fuse/fuse_i.h index c8d4c5f3af7e..87e2bd9d4bb1 100644 --- a/fs/fuse/fuse_i.h +++ b/fs/fuse/fuse_i.h @@ -218,11 +218,11 @@ struct fuse_inode { /** @lock: Lock to protect write-related fields */ spinlock_t lock; -#ifdef CONFIG_FUSE_DAX +#ifdef CONFIG_FUSE_VDAX /** - * @dax: Dax specific inode data + * @vdax: Virtiofs DAX specific inode data */ - struct fuse_inode_dax *dax; + struct fuse_inode_vdax *vdax; #endif /** @submount_lookup: Submount specific lookup tracking */ struct fuse_submount_lookup *submount_lookup; @@ -364,16 +364,16 @@ struct fuse_io_priv { .iocb = i, \ } -enum fuse_dax_mode { - FUSE_DAX_INODE_DEFAULT, /* default */ - FUSE_DAX_ALWAYS, /* "-o dax=always" */ - FUSE_DAX_NEVER, /* "-o dax=never" */ - FUSE_DAX_INODE_USER, /* "-o dax=inode" */ +enum fuse_vdax_mode { + FUSE_VDAX_INODE_DEFAULT, /* default */ + FUSE_VDAX_ALWAYS, /* "-o dax=always" */ + FUSE_VDAX_NEVER, /* "-o dax=never" */ + FUSE_VDAX_INODE_USER, /* "-o dax=inode" */ }; -static inline bool fuse_is_inode_dax_mode(enum fuse_dax_mode mode) +static inline bool fuse_is_inode_vdax_mode(enum fuse_vdax_mode mode) { - return mode == FUSE_DAX_INODE_DEFAULT || mode == FUSE_DAX_INODE_USER; + return mode == FUSE_VDAX_INODE_DEFAULT || mode == FUSE_VDAX_INODE_USER; } struct fuse_fs_context { @@ -391,13 +391,14 @@ struct fuse_fs_context { bool no_control:1; bool no_force_umount:1; bool legacy_opts_show:1; - enum fuse_dax_mode dax_mode; + bool syncfs_capable:1; + enum fuse_vdax_mode vdax_mode; unsigned int max_read; unsigned int blksize; const char *subtype; - /* DAX device, may be NULL */ - struct dax_device *dax_dev; + /* Virtiofs DAX device, may be NULL */ + struct dax_device *vdax_dev; }; struct fuse_sync_bucket { @@ -675,6 +676,14 @@ struct fuse_conn { /** @sync_fs: Propagate syncfs() to server */ unsigned int sync_fs:1; + /** + * @syncfs_capable: the privilege required to honor FUSE_HAS_SYNCFS was + * present when /dev/fuse was opened (CAP_SYS_ADMIN in the initial user + * namespace), i.e. the same privilege that mounting virtiofs/fuseblk + * requires. + */ + unsigned int syncfs_capable:1; + /** @init_security: Initialize security xattrs when creating a new inode */ unsigned int init_security:1; @@ -684,8 +693,8 @@ struct fuse_conn { */ unsigned int create_supp_group:1; - /** @inode_dax: Does the filesystem support per inode DAX? */ - unsigned int inode_dax:1; + /** @inode_vdax: Does the filesystem support per inode virtiofs DAX? */ + unsigned int inode_vdax:1; /** @no_tmpfile: Is tmpfile not implemented by fs? */ unsigned int no_tmpfile:1; @@ -744,12 +753,12 @@ struct fuse_conn { */ struct rw_semaphore killsb; -#ifdef CONFIG_FUSE_DAX - /** @dax_mode: Dax mode */ - enum fuse_dax_mode dax_mode; +#ifdef CONFIG_FUSE_VDAX + /** @vdax_mode: Virtiofs DAX mode */ + enum fuse_vdax_mode vdax_mode; - /** @dax: Dax specific conn data, non-NULL if DAX is enabled */ - struct fuse_conn_dax *dax; + /** @dax: Dax specific conn data, non-NULL if virtiofs DAX is enabled */ + struct fuse_conn_vdax *vdax; #endif /** @mounts: List of filesystems using this connection */ @@ -1003,7 +1012,7 @@ void __exit fuse_ctl_cleanup(void); /* * Simple request sending that does request allocation and freeing */ -ssize_t __fuse_simple_request(struct mnt_idmap *idmap, +ssize_t __fuse_simple_request(const struct mnt_idmap *idmap, struct fuse_mount *fm, struct fuse_args *args); @@ -1012,7 +1021,7 @@ static inline ssize_t fuse_simple_request(struct fuse_mount *fm, struct fuse_arg return __fuse_simple_request(&invalid_mnt_idmap, fm, args); } -static inline ssize_t fuse_simple_idmap_request(struct mnt_idmap *idmap, +static inline ssize_t fuse_simple_idmap_request(const struct mnt_idmap *idmap, struct fuse_mount *fm, struct fuse_args *args) { @@ -1189,7 +1198,7 @@ bool fuse_write_update_attr(struct inode *inode, loff_t pos, ssize_t written); int fuse_flush_times(struct inode *inode, struct fuse_file *ff); int fuse_write_inode(struct inode *inode, struct writeback_control *wbc); -int fuse_do_setattr(struct mnt_idmap *idmap, struct dentry *dentry, +int fuse_do_setattr(const struct mnt_idmap *idmap, struct dentry *dentry, struct iattr *attr, struct file *file); void fuse_unlock_inode(struct inode *inode, bool locked); @@ -1205,9 +1214,9 @@ extern const struct xattr_handler * const fuse_xattr_handlers[]; struct posix_acl; struct posix_acl *fuse_get_inode_acl(struct inode *inode, int type, bool rcu); -struct posix_acl *fuse_get_acl(struct mnt_idmap *idmap, +struct posix_acl *fuse_get_acl(const struct mnt_idmap *idmap, struct dentry *dentry, int type); -int fuse_set_acl(struct mnt_idmap *, struct dentry *dentry, +int fuse_set_acl(const struct mnt_idmap *, struct dentry *dentry, struct posix_acl *acl, int type); /* readdir.c */ @@ -1217,28 +1226,28 @@ void fuse_free_conn(struct fuse_conn *fc); /* dax.c */ -#define FUSE_IS_DAX(inode) (IS_ENABLED(CONFIG_FUSE_DAX) && IS_DAX(inode)) - -ssize_t fuse_dax_read_iter(struct kiocb *iocb, struct iov_iter *to); -ssize_t fuse_dax_write_iter(struct kiocb *iocb, struct iov_iter *from); -int fuse_dax_mmap(struct file *file, struct vm_area_struct *vma); -int fuse_dax_break_layouts(struct inode *inode, u64 dmap_start, u64 dmap_end); -int fuse_dax_conn_alloc(struct fuse_conn *fc, enum fuse_dax_mode mode, - struct dax_device *dax_dev); -void fuse_dax_conn_free(struct fuse_conn *fc); -bool fuse_dax_inode_alloc(struct super_block *sb, struct fuse_inode *fi); -void fuse_dax_inode_init(struct inode *inode, unsigned int flags); -void fuse_dax_inode_cleanup(struct inode *inode); -void fuse_dax_dontcache(struct inode *inode, unsigned int flags); -bool fuse_dax_check_alignment(struct fuse_conn *fc, unsigned int map_alignment); -void fuse_dax_cancel_work(struct fuse_conn *fc); +#define FUSE_IS_VDAX(inode) (IS_ENABLED(CONFIG_FUSE_VDAX) && IS_DAX(inode)) + +ssize_t fuse_vdax_read_iter(struct kiocb *iocb, struct iov_iter *to); +ssize_t fuse_vdax_write_iter(struct kiocb *iocb, struct iov_iter *from); +int fuse_vdax_mmap(struct file *file, struct vm_area_struct *vma); +int fuse_vdax_break_layouts(struct inode *inode, u64 dmap_start, u64 dmap_end); +int fuse_vdax_conn_alloc(struct fuse_conn *fc, enum fuse_vdax_mode mode, + struct dax_device *vdax_dev); +void fuse_vdax_conn_free(struct fuse_conn *fc); +bool fuse_vdax_inode_alloc(struct super_block *sb, struct fuse_inode *fi); +void fuse_vdax_inode_init(struct inode *inode, unsigned int flags); +void fuse_vdax_inode_cleanup(struct inode *inode); +void fuse_vdax_dontcache(struct inode *inode, unsigned int flags); +bool fuse_vdax_check_alignment(struct fuse_conn *fc, unsigned int map_alignment); +void fuse_vdax_cancel_work(struct fuse_conn *fc); /* ioctl.c */ long fuse_file_ioctl(struct file *file, unsigned int cmd, unsigned long arg); long fuse_file_compat_ioctl(struct file *file, unsigned int cmd, unsigned long arg); int fuse_fileattr_get(struct dentry *dentry, struct file_kattr *fa); -int fuse_fileattr_set(struct mnt_idmap *idmap, +int fuse_fileattr_set(const struct mnt_idmap *idmap, struct dentry *dentry, struct file_kattr *fa); /* iomode.c */ @@ -1258,24 +1267,13 @@ void fuse_file_release(struct inode *inode, struct fuse_file *ff, /* backing.c */ #ifdef CONFIG_FUSE_PASSTHROUGH -struct fuse_backing *fuse_backing_get(struct fuse_backing *fb); void fuse_backing_put(struct fuse_backing *fb); struct fuse_backing *fuse_backing_lookup(struct fuse_conn *fc, int backing_id); #else -static inline struct fuse_backing *fuse_backing_get(struct fuse_backing *fb) -{ - return NULL; -} - static inline void fuse_backing_put(struct fuse_backing *fb) { } -static inline struct fuse_backing *fuse_backing_lookup(struct fuse_conn *fc, - int backing_id) -{ - return NULL; -} #endif void fuse_backing_files_init(struct fuse_conn *fc); @@ -1304,13 +1302,9 @@ static inline struct fuse_backing *fuse_inode_backing_set(struct fuse_inode *fi, struct fuse_backing *fuse_passthrough_open(struct file *file, int backing_id); void fuse_passthrough_release(struct fuse_file *ff, struct fuse_backing *fb); -static inline struct file *fuse_file_passthrough(struct fuse_file *ff) +static inline bool fuse_is_passthrough(struct fuse_file *ff) { -#ifdef CONFIG_FUSE_PASSTHROUGH - return ff->passthrough; -#else - return NULL; -#endif + return IS_ENABLED(CONFIG_FUSE_PASSTHROUGH) && (ff->open_flags & FOPEN_PASSTHROUGH); } ssize_t fuse_passthrough_read_iter(struct kiocb *iocb, struct iov_iter *iter); diff --git a/fs/fuse/inode.c b/fs/fuse/inode.c index e9552be3637b..cbb10e19e7e8 100644 --- a/fs/fuse/inode.c +++ b/fs/fuse/inode.c @@ -100,7 +100,7 @@ static struct inode *fuse_alloc_inode(struct super_block *sb) if (!fi->forget) goto out_free; - if (IS_ENABLED(CONFIG_FUSE_DAX) && !fuse_dax_inode_alloc(sb, fi)) + if (IS_ENABLED(CONFIG_FUSE_VDAX) && !fuse_vdax_inode_alloc(sb, fi)) goto out_free_forget; if (IS_ENABLED(CONFIG_FUSE_PASSTHROUGH)) @@ -121,8 +121,8 @@ static void fuse_free_inode(struct inode *inode) mutex_destroy(&fi->mutex); kfree(fi->forget); -#ifdef CONFIG_FUSE_DAX - kfree(fi->dax); +#ifdef CONFIG_FUSE_VDAX + kfree(fi->vdax); #endif if (IS_ENABLED(CONFIG_FUSE_PASSTHROUGH)) fuse_backing_put(fuse_inode_backing(fi)); @@ -148,7 +148,7 @@ static void fuse_evict_inode(struct inode *inode) /* Will write inode on close/munmap and in all other dirtiers */ WARN_ON(inode_state_read_once(inode) & I_DIRTY_INODE); - if (FUSE_IS_DAX(inode)) + if (FUSE_IS_VDAX(inode)) dax_break_layout_final(inode); truncate_inode_pages_final(&inode->i_data); @@ -156,8 +156,8 @@ static void fuse_evict_inode(struct inode *inode) if (inode->i_sb->s_flags & SB_ACTIVE) { struct fuse_conn *fc = get_fuse_conn(inode); - if (FUSE_IS_DAX(inode)) - fuse_dax_inode_cleanup(inode); + if (FUSE_IS_VDAX(inode)) + fuse_vdax_inode_cleanup(inode); if (fi->nlookup) { fuse_chan_queue_forget(fc->chan, fi->forget, fi->nodeid, fi->nlookup); @@ -385,8 +385,8 @@ static void fuse_change_attributes_i(struct inode *inode, struct fuse_attr *attr invalidate_inode_pages2(inode->i_mapping); } - if (IS_ENABLED(CONFIG_FUSE_DAX)) - fuse_dax_dontcache(inode, attr->flags); + if (IS_ENABLED(CONFIG_FUSE_VDAX)) + fuse_vdax_dontcache(inode, attr->flags); } void fuse_change_attributes(struct inode *inode, struct fuse_attr *attr, @@ -803,6 +803,15 @@ static int fuse_opt_fd(struct fs_context *fsc, struct file *file) if (file->f_cred->user_ns != fsc->user_ns) return invalfc(fsc, "wrong user namespace for fuse device"); + /* + * Record whether the server opened /dev/fuse with CAP_SYS_ADMIN in the + * initial user namespace -- the same privilege that mounting virtiofs + * or fuseblk requires. Only such servers are trusted to receive + * FUSE_SYNCFS (see fuse_syncfs_enable()). + */ + ctx->syncfs_capable = file_ns_capable(file, &init_user_ns, + CAP_SYS_ADMIN); + ctx->fud = fuse_dev_grab(file); return 0; @@ -944,12 +953,12 @@ static int fuse_show_options(struct seq_file *m, struct dentry *root) if (sb->s_bdev && sb->s_blocksize != FUSE_DEFAULT_BLKSIZE) seq_printf(m, ",blksize=%lu", sb->s_blocksize); } -#ifdef CONFIG_FUSE_DAX - if (fc->dax_mode == FUSE_DAX_ALWAYS) +#ifdef CONFIG_FUSE_VDAX + if (fc->vdax_mode == FUSE_VDAX_ALWAYS) seq_puts(m, ",dax=always"); - else if (fc->dax_mode == FUSE_DAX_NEVER) + else if (fc->vdax_mode == FUSE_VDAX_NEVER) seq_puts(m, ",dax=never"); - else if (fc->dax_mode == FUSE_DAX_INODE_USER) + else if (fc->vdax_mode == FUSE_VDAX_INODE_USER) seq_puts(m, ",dax=inode"); #endif @@ -1007,8 +1016,8 @@ void fuse_conn_put(struct fuse_conn *fc) if (!refcount_dec_and_test(&fc->count)) return; - if (IS_ENABLED(CONFIG_FUSE_DAX)) - fuse_dax_conn_free(fc); + if (IS_ENABLED(CONFIG_FUSE_VDAX)) + fuse_vdax_conn_free(fc); cancel_work_sync(&fc->epoch_work); fuse_chan_release(fc->chan); put_pid_ns(fc->pid_ns); @@ -1269,6 +1278,16 @@ struct fuse_init_args { struct fuse_mount *fm; }; +/* + * A server can stall syncfs()/sync(), so only honor FUSE_HAS_SYNCFS for + * servers that opened /dev/fuse with CAP_SYS_ADMIN in the initial user + * namespace -- the same privilege required to mount virtiofs or fuseblk. + */ +static bool fuse_syncfs_enable(struct fuse_conn *fc, u64 flags) +{ + return (flags & FUSE_HAS_SYNCFS) && fc->syncfs_capable; +} + static void process_init_reply(struct fuse_args *args, int error) { struct fuse_init_args *ia = container_of(args, typeof(*ia), args); @@ -1354,13 +1373,13 @@ static void process_init_reply(struct fuse_args *args, int error) if (fc->max_pages > 1) fc->name_max = FUSE_NAME_MAX; } - if (IS_ENABLED(CONFIG_FUSE_DAX)) { + if (IS_ENABLED(CONFIG_FUSE_VDAX)) { if (flags & FUSE_MAP_ALIGNMENT && - !fuse_dax_check_alignment(fc, arg->map_alignment)) { + !fuse_vdax_check_alignment(fc, arg->map_alignment)) { ok = false; } if (flags & FUSE_HAS_INODE_DAX) - fc->inode_dax = 1; + fc->inode_vdax = 1; } if (flags & FUSE_HANDLE_KILLPRIV_V2) { fc->handle_killpriv_v2 = 1; @@ -1410,6 +1429,9 @@ static void process_init_reply(struct fuse_args *args, int error) if (flags & FUSE_REQUEST_TIMEOUT) timeout = arg->request_timeout; + + if (fuse_syncfs_enable(fc, flags)) + fc->sync_fs = 1; } else { ra_pages = fc->max_read / PAGE_SIZE; fc->no_lock = 1; @@ -1469,16 +1491,21 @@ static struct fuse_init_args *fuse_new_init(struct fuse_mount *fm) FUSE_HAS_EXPIRE_ONLY | FUSE_DIRECT_IO_ALLOW_MMAP | FUSE_NO_EXPORT_SUPPORT | FUSE_HAS_RESEND | FUSE_ALLOW_IDMAP | FUSE_REQUEST_TIMEOUT; -#ifdef CONFIG_FUSE_DAX - if (fm->fc->dax) +#ifdef CONFIG_FUSE_VDAX + if (fm->fc->vdax) flags |= FUSE_MAP_ALIGNMENT; - if (fuse_is_inode_dax_mode(fm->fc->dax_mode)) + if (fuse_is_inode_vdax_mode(fm->fc->vdax_mode)) flags |= FUSE_HAS_INODE_DAX; #endif if (fm->fc->auto_submounts) flags |= FUSE_SUBMOUNTS; if (IS_ENABLED(CONFIG_FUSE_PASSTHROUGH)) flags |= FUSE_PASSTHROUGH; + /* Only offered to sufficiently privileged servers; see + * fuse_syncfs_enable(). + */ + if (fm->fc->syncfs_capable) + flags |= FUSE_HAS_SYNCFS; if (fuse_uring_enabled()) flags |= FUSE_OVER_IO_URING | FUSE_HAS_IO_URING_BUFPOOL; @@ -1751,8 +1778,8 @@ int fuse_fill_super_common(struct super_block *sb, struct fuse_fs_context *ctx) sb->s_subtype = ctx->subtype; ctx->subtype = NULL; - if (IS_ENABLED(CONFIG_FUSE_DAX)) { - err = fuse_dax_conn_alloc(fc, ctx->dax_mode, ctx->dax_dev); + if (IS_ENABLED(CONFIG_FUSE_VDAX)) { + err = fuse_vdax_conn_alloc(fc, ctx->vdax_mode, ctx->vdax_dev); if (err) goto err; } @@ -1761,7 +1788,7 @@ int fuse_fill_super_common(struct super_block *sb, struct fuse_fs_context *ctx) fm->sb = sb; err = fuse_bdi_init(fc, sb); if (err) - goto err_free_dax; + goto err_free_vdax; /* Handle umasking inside the fuse code */ if (sb->s_flags & SB_POSIXACL) @@ -1770,6 +1797,7 @@ int fuse_fill_super_common(struct super_block *sb, struct fuse_fs_context *ctx) fc->default_permissions = ctx->default_permissions; fc->allow_other = ctx->allow_other; + fc->syncfs_capable = ctx->syncfs_capable; fc->user_id = ctx->user_id; fc->group_id = ctx->group_id; fc->legacy_opts_show = ctx->legacy_opts_show; @@ -1783,7 +1811,7 @@ int fuse_fill_super_common(struct super_block *sb, struct fuse_fs_context *ctx) set_default_d_op(sb, &fuse_dentry_operations); root_dentry = d_make_root(root); if (!root_dentry) - goto err_free_dax; + goto err_free_vdax; mutex_lock(&fuse_mutex); err = -EINVAL; @@ -1809,9 +1837,9 @@ int fuse_fill_super_common(struct super_block *sb, struct fuse_fs_context *ctx) err_unlock: mutex_unlock(&fuse_mutex); dput(root_dentry); - err_free_dax: - if (IS_ENABLED(CONFIG_FUSE_DAX)) - fuse_dax_conn_free(fc); + err_free_vdax: + if (IS_ENABLED(CONFIG_FUSE_VDAX)) + fuse_vdax_conn_free(fc); err: return err; } diff --git a/fs/fuse/ioctl.c b/fs/fuse/ioctl.c index 3614ea603913..ce1807704da6 100644 --- a/fs/fuse/ioctl.c +++ b/fs/fuse/ioctl.c @@ -130,9 +130,6 @@ static int fuse_setup_measure_verity(unsigned long arg, struct iovec *iov) if (copy_from_user(&digest_size, &uarg->digest_size, sizeof(digest_size))) return -EFAULT; - if (digest_size > SIZE_MAX - sizeof(struct fsverity_digest)) - return -EINVAL; - iov->iov_len = sizeof(struct fsverity_digest) + digest_size; return 0; @@ -540,7 +537,7 @@ cleanup: return err; } -int fuse_fileattr_set(struct mnt_idmap *idmap, +int fuse_fileattr_set(const struct mnt_idmap *idmap, struct dentry *dentry, struct file_kattr *fa) { struct inode *inode = d_inode(dentry); diff --git a/fs/fuse/iomode.c b/fs/fuse/iomode.c index 3728933188f3..79637c09e883 100644 --- a/fs/fuse/iomode.c +++ b/fs/fuse/iomode.c @@ -200,10 +200,10 @@ int fuse_file_io_open(struct file *file, struct inode *inode) int err; /* - * io modes are not relevant with DAX and with server that does not + * io modes are not relevant with virtiofs DAX and with server that does not * implement open. */ - if (FUSE_IS_DAX(inode) || !ff->args) + if (FUSE_IS_VDAX(inode) || !ff->args) return 0; /* diff --git a/fs/fuse/notify.c b/fs/fuse/notify.c index 1ba763705d91..c0428e03d138 100644 --- a/fs/fuse/notify.c +++ b/fs/fuse/notify.c @@ -155,7 +155,7 @@ static int fuse_notify_store(struct fuse_conn *fc, unsigned int size, nodeid = outarg.nodeid; pos = outarg.offset; - num = min(outarg.size, MAX_LFS_FILESIZE - pos); + num = umin(outarg.size, MAX_LFS_FILESIZE - pos); down_read(&fc->killsb); @@ -190,7 +190,7 @@ static int fuse_notify_store(struct fuse_conn *fc, unsigned int size, folio_offset = offset_in_folio(folio, pos); nr_bytes = min(num, folio_size(folio) - folio_offset); - err = fuse_copy_folio(cs, &folio, folio_offset, nr_bytes, 0); + err = fuse_copy_folio(cs, &folio, folio_offset, nr_bytes); if (!folio_test_uptodate(folio) && !err && folio_offset == 0 && (nr_bytes == folio_size(folio) || file_size == end)) { folio_zero_segment(folio, nr_bytes, folio_size(folio)); diff --git a/fs/fuse/passthrough.c b/fs/fuse/passthrough.c index f2d08ac2459b..b43d3e0f7081 100644 --- a/fs/fuse/passthrough.c +++ b/fs/fuse/passthrough.c @@ -11,6 +11,10 @@ #include <linux/backing-file.h> #include <linux/splice.h> +static inline struct file *fuse_file_passthrough(struct fuse_file *ff) +{ + return ff->passthrough; +} static void fuse_file_accessed(struct file *file) { struct inode *inode = file_inode(file); @@ -187,10 +191,15 @@ out: void fuse_passthrough_release(struct fuse_file *ff, struct fuse_backing *fb) { + struct file *backing_file = fuse_file_passthrough(ff); + pr_debug("%s: fb=0x%p, backing_file=0x%p\n", __func__, - fb, ff->passthrough); + fb, backing_file); + + if (!backing_file) + return; - fput(ff->passthrough); + fput(backing_file); ff->passthrough = NULL; put_cred(ff->cred); ff->cred = NULL; diff --git a/fs/fuse/req.c b/fs/fuse/req.c index a01ee743d31e..6274a23af445 100644 --- a/fs/fuse/req.c +++ b/fs/fuse/req.c @@ -3,19 +3,20 @@ #include "dev.h" #include "fuse_i.h" -static int fuse_fill_creds(struct fuse_mount *fm, struct fuse_args *args, struct mnt_idmap *idmap) +static int fuse_fill_creds(struct fuse_mount *fm, struct fuse_args *args, + const struct mnt_idmap *idmap) { struct fuse_conn *fc = fm->fc; bool no_idmap = !fm->sb || (fm->sb->s_iflags & SB_I_NOIDMAP); kuid_t fsuid = mapped_fsuid(idmap, fc->user_ns); kgid_t fsgid = mapped_fsgid(idmap, fc->user_ns); + if (args->nocreds) + return 0; + args->pid = pid_nr_ns(task_pid(current), fc->pid_ns); if (args->force) { - if (args->nocreds) - return 0; - if (no_idmap) { args->uid = from_kuid_munged(fc->user_ns, current_fsuid()); args->gid = from_kgid_munged(fc->user_ns, current_fsgid()); @@ -26,7 +27,6 @@ static int fuse_fill_creds(struct fuse_mount *fm, struct fuse_args *args, struct return 0; } - WARN_ON(args->nocreds); /* * Keep the old behavior when idmappings support was not * declared by a FUSE server. @@ -49,7 +49,8 @@ static int fuse_fill_creds(struct fuse_mount *fm, struct fuse_args *args, struct return 0; } -static int fuse_req_prep(struct fuse_mount *fm, struct fuse_args *args, struct mnt_idmap *idmap) +static int fuse_req_prep(struct fuse_mount *fm, struct fuse_args *args, + const struct mnt_idmap *idmap) { if (!args->force && fm->fc->conn_error) return -ECONNREFUSED; @@ -57,7 +58,7 @@ static int fuse_req_prep(struct fuse_mount *fm, struct fuse_args *args, struct m return fuse_fill_creds(fm, args, idmap); } -ssize_t __fuse_simple_request(struct mnt_idmap *idmap, struct fuse_mount *fm, +ssize_t __fuse_simple_request(const struct mnt_idmap *idmap, struct fuse_mount *fm, struct fuse_args *args) { struct fuse_conn *fc = fm->fc; diff --git a/fs/fuse/virtio_fs.c b/fs/fuse/virtio_fs.c index f15e516ebcb5..4f334766b8c3 100644 --- a/fs/fuse/virtio_fs.c +++ b/fs/fuse/virtio_fs.c @@ -102,9 +102,9 @@ static int virtio_fs_enqueue_req(struct virtio_fs_vq *fsvq, gfp_t gfp); static const struct constant_table dax_param_enums[] = { - {"always", FUSE_DAX_ALWAYS }, - {"never", FUSE_DAX_NEVER }, - {"inode", FUSE_DAX_INODE_USER }, + {"always", FUSE_VDAX_ALWAYS }, + {"never", FUSE_VDAX_NEVER }, + {"inode", FUSE_VDAX_INODE_USER }, {} }; @@ -132,10 +132,10 @@ static int virtio_fs_parse_param(struct fs_context *fsc, switch (opt) { case OPT_DAX: - ctx->dax_mode = FUSE_DAX_ALWAYS; + ctx->vdax_mode = FUSE_VDAX_ALWAYS; break; case OPT_DAX_ENUM: - ctx->dax_mode = result.uint_32; + ctx->vdax_mode = result.uint_32; break; default: return -EINVAL; @@ -798,9 +798,11 @@ static void virtio_fs_request_complete(struct fuse_req *req, for (i = 0; i < ap->num_folios; i++) { thislen = ap->descs[i].length; if (len < thislen) { - WARN_ON(ap->descs[i].offset); + unsigned int offset = ap->descs[i].offset; + folio = ap->folios[i]; - folio_zero_segment(folio, len, thislen); + folio_zero_segment(folio, offset + len, + offset + thislen); len = 0; } else { len -= thislen; @@ -1079,7 +1081,7 @@ static int virtio_fs_setup_dax(struct virtio_device *vdev, struct virtio_fs *fs) struct dev_pagemap *pgmap; bool have_cache; - if (!IS_ENABLED(CONFIG_FUSE_DAX)) + if (!IS_ENABLED(CONFIG_FUSE_VDAX)) return 0; dax_dev = alloc_dax(fs, &virtio_fs_dax_ops); @@ -1592,14 +1594,14 @@ static int virtio_fs_fill_super(struct super_block *sb, struct fs_context *fsc) goto err_free_fuse_devs; } - if (ctx->dax_mode != FUSE_DAX_NEVER) { - if (ctx->dax_mode == FUSE_DAX_ALWAYS && !fs->dax_dev) { + if (ctx->vdax_mode != FUSE_VDAX_NEVER) { + if (ctx->vdax_mode == FUSE_VDAX_ALWAYS && !fs->dax_dev) { err = -EINVAL; pr_err("virtio-fs: dax can't be enabled as filesystem" " device does not support it.\n"); goto err_free_fuse_devs; } - ctx->dax_dev = fs->dax_dev; + ctx->vdax_dev = fs->dax_dev; } err = fuse_fill_super_common(sb, ctx); if (err < 0) @@ -1633,8 +1635,8 @@ static void virtio_fs_conn_destroy(struct fuse_mount *fm) /* Stop dax worker. Soon evict_inodes() will be called which * will free all memory ranges belonging to all inodes. */ - if (IS_ENABLED(CONFIG_FUSE_DAX)) - fuse_dax_cancel_work(fc); + if (IS_ENABLED(CONFIG_FUSE_VDAX)) + fuse_vdax_cancel_work(fc); /* Stop forget queue. Soon destroy will be sent */ spin_lock(&fsvq->lock); diff --git a/fs/fuse/xattr.c b/fs/fuse/xattr.c index cab2685acc65..53e3c5e6fff0 100644 --- a/fs/fuse/xattr.c +++ b/fs/fuse/xattr.c @@ -188,7 +188,7 @@ static int fuse_xattr_get(const struct xattr_handler *handler, } static int fuse_xattr_set(const struct xattr_handler *handler, - struct mnt_idmap *idmap, + const struct mnt_idmap *idmap, struct dentry *dentry, struct inode *inode, const char *name, const void *value, size_t size, int flags) diff --git a/fs/gfs2/acl.c b/fs/gfs2/acl.c index a5b60778b91c..f1c6a5e392b4 100644 --- a/fs/gfs2/acl.c +++ b/fs/gfs2/acl.c @@ -59,7 +59,7 @@ static struct posix_acl *__gfs2_get_acl(struct inode *inode, int type) struct posix_acl *gfs2_get_acl(struct inode *inode, int type, bool rcu) { - struct gfs2_inode *ip = GFS2_I(inode); + struct gfs2_glock *gl = gfs2_inode_glock(inode); struct gfs2_holder gh; bool need_unlock = false; struct posix_acl *acl; @@ -67,8 +67,8 @@ struct posix_acl *gfs2_get_acl(struct inode *inode, int type, bool rcu) if (rcu) return ERR_PTR(-ECHILD); - if (!gfs2_glock_is_locked_by_me(ip->i_gl)) { - int ret = gfs2_glock_nq_init(ip->i_gl, LM_ST_SHARED, + if (!gfs2_glock_is_locked_by_me(gl)) { + int ret = gfs2_glock_nq_init(gl, LM_ST_SHARED, LM_FLAG_ANY, &gh); if (ret) return ERR_PTR(ret); @@ -102,10 +102,11 @@ out: return error; } -int gfs2_set_acl(struct mnt_idmap *idmap, struct dentry *dentry, +int gfs2_set_acl(const struct mnt_idmap *idmap, struct dentry *dentry, struct posix_acl *acl, int type) { struct inode *inode = d_inode(dentry); + struct gfs2_glock *gl = gfs2_inode_glock(inode); struct gfs2_inode *ip = GFS2_I(inode); struct gfs2_holder gh; bool need_unlock = false; @@ -119,8 +120,8 @@ int gfs2_set_acl(struct mnt_idmap *idmap, struct dentry *dentry, if (ret) return ret; - if (!gfs2_glock_is_locked_by_me(ip->i_gl)) { - ret = gfs2_glock_nq_init(ip->i_gl, LM_ST_EXCLUSIVE, 0, &gh); + if (!gfs2_glock_is_locked_by_me(gl)) { + ret = gfs2_glock_nq_init(gl, LM_ST_EXCLUSIVE, 0, &gh); if (ret) goto out; need_unlock = true; diff --git a/fs/gfs2/acl.h b/fs/gfs2/acl.h index 82f5b09c04e6..d19d41755936 100644 --- a/fs/gfs2/acl.h +++ b/fs/gfs2/acl.h @@ -13,7 +13,7 @@ struct posix_acl *gfs2_get_acl(struct inode *inode, int type, bool rcu); int __gfs2_set_acl(struct inode *inode, struct posix_acl *acl, int type); -int gfs2_set_acl(struct mnt_idmap *idmap, struct dentry *dentry, +int gfs2_set_acl(const struct mnt_idmap *idmap, struct dentry *dentry, struct posix_acl *acl, int type); #endif /* __ACL_DOT_H__ */ diff --git a/fs/gfs2/aops.c b/fs/gfs2/aops.c index 0a7b8076af3a..ee11d494e47c 100644 --- a/fs/gfs2/aops.c +++ b/fs/gfs2/aops.c @@ -102,7 +102,7 @@ static int __gfs2_jdata_write_folio(struct folio *folio, struct writeback_control *wbc) { struct inode *inode = folio->mapping->host; - struct gfs2_inode *ip = GFS2_I(inode); + struct gfs2_glock *gl = gfs2_inode_glock(inode); if (folio_test_checked(folio)) { folio_clear_checked(folio); @@ -111,7 +111,7 @@ static int __gfs2_jdata_write_folio(struct folio *folio, inode->i_sb->s_blocksize, BIT(BH_Dirty)|BIT(BH_Uptodate)); } - gfs2_trans_add_databufs(ip->i_gl, folio, 0, folio_size(folio)); + gfs2_trans_add_databufs(gl, folio, 0, folio_size(folio)); } return gfs2_write_jdata_folio(folio, wbc); } @@ -126,13 +126,13 @@ static int __gfs2_jdata_write_folio(struct folio *folio, int gfs2_jdata_writeback(struct address_space *mapping, struct writeback_control *wbc) { struct inode *inode = mapping->host; - struct gfs2_inode *ip = GFS2_I(inode); + struct gfs2_glock *gl = gfs2_inode_glock(inode); struct gfs2_sbd *sdp = GFS2_SB(mapping->host); struct folio *folio = NULL; int error; BUG_ON(current->journal_info); - if (gfs2_assert_withdraw(sdp, ip->i_gl->gl_state == LM_ST_EXCLUSIVE)) + if (gfs2_assert_withdraw(sdp, gl->gl_state == LM_ST_EXCLUSIVE)) return 0; while ((folio = writeback_iter(mapping, wbc, folio, &error))) { @@ -362,14 +362,14 @@ retry: static int gfs2_jdata_writepages(struct address_space *mapping, struct writeback_control *wbc) { - struct gfs2_inode *ip = GFS2_I(mapping->host); + struct gfs2_glock *gl = gfs2_inode_glock(mapping->host); struct gfs2_sbd *sdp = GFS2_SB(mapping->host); int ret; ret = gfs2_write_cache_jdata(mapping, wbc); if (ret == 0 && wbc->sync_mode == WB_SYNC_ALL) { - gfs2_log_flush(sdp, ip->i_gl, GFS2_LOG_HEAD_FLUSH_NORMAL | - GFS2_LFC_JDATA_WPAGES); + gfs2_log_flush(sdp, gl, GFS2_LOG_HEAD_FLUSH_NORMAL | + GFS2_LFC_JDATA_WPAGES); ret = gfs2_write_cache_jdata(mapping, wbc); } return ret; @@ -561,12 +561,13 @@ static bool gfs2_jdata_dirty_folio(struct address_space *mapping, static sector_t gfs2_bmap(struct address_space *mapping, sector_t lblock) { + struct gfs2_glock *gl = gfs2_inode_glock(mapping->host); struct gfs2_inode *ip = GFS2_I(mapping->host); struct gfs2_holder i_gh; sector_t dblock = 0; int error; - error = gfs2_glock_nq_init(ip->i_gl, LM_ST_SHARED, LM_FLAG_ANY, &i_gh); + error = gfs2_glock_nq_init(gl, LM_ST_SHARED, LM_FLAG_ANY, &i_gh); if (error) return 0; diff --git a/fs/gfs2/bmap.c b/fs/gfs2/bmap.c index 73c626971163..264170f91c5d 100644 --- a/fs/gfs2/bmap.c +++ b/fs/gfs2/bmap.c @@ -55,6 +55,7 @@ static int gfs2_unstuffer_folio(struct gfs2_inode *ip, struct buffer_head *dibh, u64 block, struct folio *folio) { struct inode *inode = &ip->i_inode; + struct gfs2_glock *gl = gfs2_inode_glock(inode); if (!folio_test_uptodate(folio)) { void *kaddr = kmap_local_folio(folio, 0); @@ -78,7 +79,7 @@ static int gfs2_unstuffer_folio(struct gfs2_inode *ip, struct buffer_head *dibh, map_bh(bh, inode->i_sb, block); set_buffer_uptodate(bh); - gfs2_trans_add_data(ip->i_gl, bh); + gfs2_trans_add_data(gl, bh); } else { folio_mark_dirty(folio); gfs2_ordered_add_inode(ip); @@ -89,6 +90,8 @@ static int gfs2_unstuffer_folio(struct gfs2_inode *ip, struct buffer_head *dibh, static int __gfs2_unstuff_inode(struct gfs2_inode *ip, struct folio *folio) { + struct inode *inode = &ip->i_inode; + struct gfs2_glock *gl = gfs2_inode_glock(inode); struct buffer_head *bh, *dibh; struct gfs2_dinode *di; u64 block = 0; @@ -99,7 +102,7 @@ static int __gfs2_unstuff_inode(struct gfs2_inode *ip, struct folio *folio) if (error) return error; - if (i_size_read(&ip->i_inode)) { + if (i_size_read(inode)) { /* Get a free block, fill it with the stuffed data, and write it out to disk */ @@ -108,7 +111,7 @@ static int __gfs2_unstuff_inode(struct gfs2_inode *ip, struct folio *folio) if (error) goto out_brelse; if (isdir) { - gfs2_trans_remove_revoke(GFS2_SB(&ip->i_inode), block, 1); + gfs2_trans_remove_revoke(GFS2_SB(inode), block, 1); error = gfs2_dir_get_new_buffer(ip, block, &bh); if (error) goto out_brelse; @@ -124,14 +127,14 @@ static int __gfs2_unstuff_inode(struct gfs2_inode *ip, struct folio *folio) /* Set up the pointer to the new block */ - gfs2_trans_add_meta(ip->i_gl, dibh); + gfs2_trans_add_meta(gl, dibh); di = (struct gfs2_dinode *)dibh->b_data; gfs2_buffer_clear_tail(dibh, sizeof(struct gfs2_dinode)); - if (i_size_read(&ip->i_inode)) { + if (i_size_read(inode)) { *(__be64 *)(di + 1) = cpu_to_be64(block); - gfs2_add_inode_blocks(&ip->i_inode, 1); - di->di_blocks = cpu_to_be64(gfs2_get_inode_blocks(&ip->i_inode)); + gfs2_add_inode_blocks(inode, 1); + di->di_blocks = cpu_to_be64(gfs2_get_inode_blocks(inode)); } ip->i_height = 1; @@ -662,6 +665,7 @@ enum alloc_state { static int __gfs2_iomap_alloc(struct inode *inode, struct iomap *iomap, struct metapath *mp) { + struct gfs2_glock *gl = gfs2_inode_glock(inode); struct gfs2_inode *ip = GFS2_I(inode); struct gfs2_sbd *sdp = GFS2_SB(inode); struct buffer_head *dibh = metapath_dibh(mp); @@ -678,7 +682,7 @@ static int __gfs2_iomap_alloc(struct inode *inode, struct iomap *iomap, BUG_ON(dibh == NULL); BUG_ON(dblks < 1); - gfs2_trans_add_meta(ip->i_gl, dibh); + gfs2_trans_add_meta(gl, dibh); down_write(&ip->i_rw_mutex); @@ -722,7 +726,7 @@ static int __gfs2_iomap_alloc(struct inode *inode, struct iomap *iomap, } for (; i - 1 < mp->mp_fheight - ip->i_height && n > 0; i++, n--) - gfs2_indirect_init(mp, ip->i_gl, i, 0, bn++); + gfs2_indirect_init(mp, gl, i, 0, bn++); if (i - 1 == mp->mp_fheight - ip->i_height) { i--; gfs2_buffer_copy_tail(mp->mp_bh[i], @@ -748,9 +752,9 @@ static int __gfs2_iomap_alloc(struct inode *inode, struct iomap *iomap, fallthrough; /* To branching from existing tree */ case ALLOC_GROW_DEPTH: if (i > 1 && i < mp->mp_fheight) - gfs2_trans_add_meta(ip->i_gl, mp->mp_bh[i-1]); + gfs2_trans_add_meta(gl, mp->mp_bh[i-1]); for (; i < mp->mp_fheight && n > 0; i++, n--) - gfs2_indirect_init(mp, ip->i_gl, i, + gfs2_indirect_init(mp, gl, i, mp->mp_list[i-1], bn++); if (i == mp->mp_fheight) state = ALLOC_DATA; @@ -760,7 +764,7 @@ static int __gfs2_iomap_alloc(struct inode *inode, struct iomap *iomap, case ALLOC_DATA: BUG_ON(n > dblks); BUG_ON(mp->mp_bh[end_of_metadata] == NULL); - gfs2_trans_add_meta(ip->i_gl, mp->mp_bh[end_of_metadata]); + gfs2_trans_add_meta(gl, mp->mp_bh[end_of_metadata]); dblks = n; ptr = metapointer(end_of_metadata, mp); iomap->addr = bn << inode->i_blkbits; @@ -989,12 +993,12 @@ static void gfs2_iomap_put_folio(struct inode *inode, loff_t pos, unsigned copied, struct folio *folio) { struct gfs2_trans *tr = current->journal_info; + struct gfs2_glock *gl = gfs2_inode_glock(inode); struct gfs2_inode *ip = GFS2_I(inode); struct gfs2_sbd *sdp = GFS2_SB(inode); if (gfs2_is_jdata(ip) && !gfs2_is_stuffed(ip)) - gfs2_trans_add_databufs(ip->i_gl, folio, - offset_in_folio(folio, pos), + gfs2_trans_add_databufs(gl, folio, offset_in_folio(folio, pos), copied); folio_unlock(folio); @@ -1150,6 +1154,7 @@ out_unlock: static int gfs2_iomap_end(struct inode *inode, loff_t pos, loff_t length, ssize_t written, unsigned flags, struct iomap *iomap) { + struct gfs2_glock *gl = gfs2_inode_glock(inode); struct gfs2_inode *ip = GFS2_I(inode); struct gfs2_sbd *sdp = GFS2_SB(inode); @@ -1196,7 +1201,7 @@ static int gfs2_iomap_end(struct inode *inode, loff_t pos, loff_t length, if (iomap->flags & IOMAP_F_SIZE_CHANGED) mark_inode_dirty(inode); - set_bit(GLF_DIRTY, &ip->i_gl->gl_flags); + set_bit(GLF_DIRTY, &gl->gl_flags); return 0; } @@ -1212,7 +1217,7 @@ const struct iomap_ops gfs2_iomap_ops = { * @inode: The inode * @lblock: The logical block number * @bh_map: The bh to be mapped - * @create: True if its ok to alloc blocks to satify the request + * @create: True if its ok to alloc blocks to satisfy the request * * The size of the requested mapping is defined in bh_map->b_size. * @@ -1387,6 +1392,7 @@ static int gfs2_journaled_truncate(struct inode *inode, u64 oldsize, u64 newsize static int trunc_start(struct inode *inode, u64 newsize) { + struct gfs2_glock *gl = gfs2_inode_glock(inode); struct gfs2_inode *ip = GFS2_I(inode); struct gfs2_sbd *sdp = GFS2_SB(inode); struct buffer_head *dibh = NULL; @@ -1415,7 +1421,7 @@ static int trunc_start(struct inode *inode, u64 newsize) if (error) goto out; - gfs2_trans_add_meta(ip->i_gl, dibh); + gfs2_trans_add_meta(gl, dibh); if (gfs2_is_stuffed(ip)) gfs2_buffer_clear_tail(dibh, sizeof(struct gfs2_dinode) + newsize); @@ -1488,7 +1494,9 @@ static int sweep_bh_for_rgrps(struct gfs2_inode *ip, struct gfs2_holder *rd_gh, struct buffer_head *bh, __be64 *start, __be64 *end, bool meta, u32 *btotal) { - struct gfs2_sbd *sdp = GFS2_SB(&ip->i_inode); + struct inode *inode = &ip->i_inode; + struct gfs2_glock *gl = gfs2_inode_glock(inode); + struct gfs2_sbd *sdp = GFS2_SB(inode); struct gfs2_rgrpd *rgd; struct gfs2_trans *tr; __be64 *p; @@ -1546,7 +1554,7 @@ more_rgrps: jblocks_rqsted = rgd->rd_length + RES_DINODE + RES_INDIRECT; - isize_blks = gfs2_get_inode_blocks(&ip->i_inode); + isize_blks = gfs2_get_inode_blocks(inode); if (isize_blks > atomic_read(&sdp->sd_log_thresh2)) jblocks_rqsted += atomic_read(&sdp->sd_log_thresh2); @@ -1587,7 +1595,7 @@ more_rgrps: goto out_unlock; } - gfs2_trans_add_meta(ip->i_gl, bh); + gfs2_trans_add_meta(gl, bh); buf_in_tr = true; *p = 0; if (bstart + blen == bn) { @@ -1597,7 +1605,7 @@ more_rgrps: if (bstart) { __gfs2_free_blocks(ip, rgd, bstart, (u32)blen, meta); (*btotal) += blen; - gfs2_add_inode_blocks(&ip->i_inode, -blen); + gfs2_add_inode_blocks(inode, -blen); } bstart = bn; blen = 1; @@ -1605,7 +1613,7 @@ more_rgrps: if (bstart) { __gfs2_free_blocks(ip, rgd, bstart, (u32)blen, meta); (*btotal) += blen; - gfs2_add_inode_blocks(&ip->i_inode, -blen); + gfs2_add_inode_blocks(inode, -blen); } out_unlock: if (!ret && blks_outside_rgrp) { /* If buffer still has non-zero blocks @@ -1620,8 +1628,8 @@ out_unlock: /* Every transaction boundary, we rewrite the dinode to keep its di_blocks current in case of failure. */ - inode_set_mtime_to_ts(&ip->i_inode, inode_set_ctime_current(&ip->i_inode)); - gfs2_trans_add_meta(ip->i_gl, dibh); + inode_set_mtime_to_ts(inode, inode_set_ctime_current(inode)); + gfs2_trans_add_meta(gl, dibh); gfs2_dinode_out(ip, dibh->b_data); brelse(dibh); up_write(&ip->i_rw_mutex); @@ -1691,7 +1699,7 @@ enum dealloc_states { }; static inline void -metapointer_range(struct metapath *mp, int height, +metapointer_range(struct metapath *mp, unsigned int height, __u16 *start_list, unsigned int start_aligned, __u16 *end_list, unsigned int end_aligned, __be64 **start, __be64 **end) @@ -1746,7 +1754,9 @@ static inline bool walk_done(struct gfs2_sbd *sdp, */ static int punch_hole(struct gfs2_inode *ip, u64 offset, u64 length) { - struct gfs2_sbd *sdp = GFS2_SB(&ip->i_inode); + struct inode *inode = &ip->i_inode; + struct gfs2_glock *gl = gfs2_inode_glock(inode); + struct gfs2_sbd *sdp = GFS2_SB(inode); u64 maxsize = sdp->sd_heightsize[ip->i_height]; struct metapath mp = {}; struct buffer_head *dibh, *bh; @@ -1760,7 +1770,7 @@ static int punch_hole(struct gfs2_inode *ip, u64 offset, u64 length) unsigned int strip_h = ip->i_height - 1; u32 btotal = 0; int ret, state; - int mp_h; /* metapath buffers are read in to this height */ + unsigned int mp_h; /* metapath buffers are read in to this height */ u64 prev_bnr = 0; __be64 *start, *end; @@ -1833,7 +1843,7 @@ static int punch_hole(struct gfs2_inode *ip, u64 offset, u64 length) for (mp_h = 0; mp_h < mp.mp_aheight - 1; mp_h++) { metapointer_range(&mp, mp_h, start_list, start_aligned, end_list, end_aligned, &start, &end); - gfs2_metapath_ra(ip->i_gl, start, end); + gfs2_metapath_ra(gl, start, end); } if (mp.mp_aheight == ip->i_height) @@ -1953,7 +1963,7 @@ static int punch_hole(struct gfs2_inode *ip, u64 offset, u64 length) start_list, start_aligned, end_list, end_aligned, &start, &end); - gfs2_metapath_ra(ip->i_gl, start, end); + gfs2_metapath_ra(gl, start, end); } } @@ -1985,10 +1995,9 @@ static int punch_hole(struct gfs2_inode *ip, u64 offset, u64 length) down_write(&ip->i_rw_mutex); } gfs2_statfs_change(sdp, 0, +btotal, 0); - gfs2_quota_change(ip, -(s64)btotal, ip->i_inode.i_uid, - ip->i_inode.i_gid); - inode_set_mtime_to_ts(&ip->i_inode, inode_set_ctime_current(&ip->i_inode)); - gfs2_trans_add_meta(ip->i_gl, dibh); + gfs2_quota_change(ip, -(s64)btotal, inode->i_uid, inode->i_gid); + inode_set_mtime_to_ts(inode, inode_set_ctime_current(inode)); + gfs2_trans_add_meta(gl, dibh); gfs2_dinode_out(ip, dibh->b_data); up_write(&ip->i_rw_mutex); gfs2_trans_end(sdp); @@ -2010,7 +2019,9 @@ out_metapath: static int trunc_end(struct gfs2_inode *ip) { - struct gfs2_sbd *sdp = GFS2_SB(&ip->i_inode); + struct inode *inode = &ip->i_inode; + struct gfs2_glock *gl = gfs2_inode_glock(inode); + struct gfs2_sbd *sdp = GFS2_SB(inode); struct buffer_head *dibh; int error; @@ -2024,16 +2035,16 @@ static int trunc_end(struct gfs2_inode *ip) if (error) goto out; - if (!i_size_read(&ip->i_inode)) { + if (!i_size_read(inode)) { ip->i_height = 0; ip->i_goal = ip->i_no_addr; gfs2_buffer_clear_tail(dibh, sizeof(struct gfs2_dinode)); gfs2_ordered_del_inode(ip); } - inode_set_mtime_to_ts(&ip->i_inode, inode_set_ctime_current(&ip->i_inode)); + inode_set_mtime_to_ts(inode, inode_set_ctime_current(inode)); ip->i_diskflags &= ~GFS2_DIF_TRUNC_IN_PROG; - gfs2_trans_add_meta(ip->i_gl, dibh); + gfs2_trans_add_meta(gl, dibh); gfs2_dinode_out(ip, dibh->b_data); brelse(dibh); @@ -2094,6 +2105,7 @@ static int do_shrink(struct inode *inode, u64 newsize) static int do_grow(struct inode *inode, u64 size) { + struct gfs2_glock *gl = gfs2_inode_glock(inode); struct gfs2_inode *ip = GFS2_I(inode); struct gfs2_sbd *sdp = GFS2_SB(inode); struct gfs2_alloc_parms ap = { .target = 1, }; @@ -2138,7 +2150,7 @@ static int do_grow(struct inode *inode, u64 size) truncate_setsize(inode, size); inode_set_mtime_to_ts(&ip->i_inode, inode_set_ctime_current(&ip->i_inode)); - gfs2_trans_add_meta(ip->i_gl, dibh); + gfs2_trans_add_meta(gl, dibh); gfs2_dinode_out(ip, dibh->b_data); brelse(dibh); @@ -2377,6 +2389,7 @@ int gfs2_write_alloc_required(struct gfs2_inode *ip, u64 offset, static int stuffed_zero_range(struct inode *inode, loff_t offset, loff_t length) { + struct gfs2_glock *gl = gfs2_inode_glock(inode); struct gfs2_inode *ip = GFS2_I(inode); struct buffer_head *dibh; int error; @@ -2389,7 +2402,7 @@ static int stuffed_zero_range(struct inode *inode, loff_t offset, loff_t length) error = gfs2_meta_inode_buffer(ip, &dibh); if (error) return error; - gfs2_trans_add_meta(ip->i_gl, dibh); + gfs2_trans_add_meta(gl, dibh); memset(dibh->b_data + sizeof(struct gfs2_dinode) + offset, 0, length); brelse(dibh); @@ -2416,7 +2429,7 @@ static int gfs2_journaled_truncate_range(struct inode *inode, loff_t offset, if (offs && chunk > PAGE_SIZE) chunk = offs + ((chunk - offs) & PAGE_MASK); - truncate_pagecache_range(inode, offset, chunk); + truncate_pagecache_range(inode, offset, offset + chunk - 1); offset += chunk; length -= chunk; diff --git a/fs/gfs2/dentry.c b/fs/gfs2/dentry.c index 95050e719233..7b344461658e 100644 --- a/fs/gfs2/dentry.c +++ b/fs/gfs2/dentry.c @@ -35,8 +35,8 @@ static int gfs2_drevalidate(struct inode *dir, const struct qstr *name, struct dentry *dentry, unsigned int flags) { + struct gfs2_glock *gl = gfs2_inode_glock(dir); struct gfs2_sbd *sdp = GFS2_SB(dir); - struct gfs2_inode *dip = GFS2_I(dir); struct inode *inode; struct gfs2_holder d_gh; struct gfs2_inode *ip = NULL; @@ -57,9 +57,9 @@ static int gfs2_drevalidate(struct inode *dir, const struct qstr *name, if (sdp->sd_lockstruct.ls_ops->lm_mount == NULL) return 1; - had_lock = (gfs2_glock_is_locked_by_me(dip->i_gl) != NULL); + had_lock = (gfs2_glock_is_locked_by_me(gl) != NULL); if (!had_lock) { - error = gfs2_glock_nq_init(dip->i_gl, LM_ST_SHARED, 0, &d_gh); + error = gfs2_glock_nq_init(gl, LM_ST_SHARED, 0, &d_gh); if (error) return 0; } diff --git a/fs/gfs2/dir.c b/fs/gfs2/dir.c index 0237b36b9eb1..6cfe335fd590 100644 --- a/fs/gfs2/dir.c +++ b/fs/gfs2/dir.c @@ -90,10 +90,11 @@ typedef int (*gfs2_dscan_t)(const struct gfs2_dirent *dent, int gfs2_dir_get_new_buffer(struct gfs2_inode *ip, u64 block, struct buffer_head **bhp) { + struct gfs2_glock *gl = gfs2_inode_glock(&ip->i_inode); struct buffer_head *bh; - bh = gfs2_meta_new(ip->i_gl, block); - gfs2_trans_add_meta(ip->i_gl, bh); + bh = gfs2_meta_new(gl, block); + gfs2_trans_add_meta(gl, bh); gfs2_metatype_set(bh, GFS2_METATYPE_JD, GFS2_FORMAT_JD); gfs2_buffer_clear_tail(bh, sizeof(struct gfs2_meta_header)); *bhp = bh; @@ -103,10 +104,11 @@ int gfs2_dir_get_new_buffer(struct gfs2_inode *ip, u64 block, static int gfs2_dir_get_existing_buffer(struct gfs2_inode *ip, u64 block, struct buffer_head **bhp) { + struct gfs2_glock *gl = gfs2_inode_glock(&ip->i_inode); struct buffer_head *bh; int error; - error = gfs2_meta_read(ip->i_gl, block, DIO_WAIT, 0, &bh); + error = gfs2_meta_read(gl, block, DIO_WAIT, 0, &bh); if (error) return error; if (gfs2_metatype_check(GFS2_SB(&ip->i_inode), bh, GFS2_METATYPE_JD)) { @@ -120,6 +122,7 @@ static int gfs2_dir_get_existing_buffer(struct gfs2_inode *ip, u64 block, static int gfs2_dir_write_stuffed(struct gfs2_inode *ip, const char *buf, unsigned int offset, unsigned int size) { + struct gfs2_glock *gl = gfs2_inode_glock(&ip->i_inode); struct buffer_head *dibh; int error; @@ -127,7 +130,7 @@ static int gfs2_dir_write_stuffed(struct gfs2_inode *ip, const char *buf, if (error) return error; - gfs2_trans_add_meta(ip->i_gl, dibh); + gfs2_trans_add_meta(gl, dibh); memcpy(dibh->b_data + offset + sizeof(struct gfs2_dinode), buf, size); if (ip->i_inode.i_size < offset + size) i_size_write(&ip->i_inode, offset + size); @@ -153,6 +156,7 @@ static int gfs2_dir_write_stuffed(struct gfs2_inode *ip, const char *buf, static int gfs2_dir_write_data(struct gfs2_inode *ip, const char *buf, u64 offset, unsigned int size) { + struct gfs2_glock *gl = gfs2_inode_glock(&ip->i_inode); struct gfs2_sbd *sdp = GFS2_SB(&ip->i_inode); struct buffer_head *dibh; u64 lblock, dblock; @@ -208,7 +212,7 @@ static int gfs2_dir_write_data(struct gfs2_inode *ip, const char *buf, if (error) goto fail; - gfs2_trans_add_meta(ip->i_gl, bh); + gfs2_trans_add_meta(gl, bh); memcpy(bh->b_data + o, buf, amount); brelse(bh); @@ -230,7 +234,7 @@ out: i_size_write(&ip->i_inode, offset + copied); inode_set_mtime_to_ts(&ip->i_inode, inode_set_ctime_current(&ip->i_inode)); - gfs2_trans_add_meta(ip->i_gl, dibh); + gfs2_trans_add_meta(gl, dibh); gfs2_dinode_out(ip, dibh->b_data); brelse(dibh); @@ -268,6 +272,7 @@ static int gfs2_dir_read_stuffed(struct gfs2_inode *ip, __be64 *buf, static int gfs2_dir_read_data(struct gfs2_inode *ip, __be64 *buf, unsigned int size) { + struct gfs2_glock *gl = gfs2_inode_glock(&ip->i_inode); struct gfs2_sbd *sdp = GFS2_SB(&ip->i_inode); u64 lblock, dblock; u32 extlen = 0; @@ -299,9 +304,9 @@ static int gfs2_dir_read_data(struct gfs2_inode *ip, __be64 *buf, if (error || !dblock) goto fail; BUG_ON(extlen < 1); - bh = gfs2_meta_ra(ip->i_gl, dblock, extlen); + bh = gfs2_meta_ra(gl, dblock, extlen); } else { - error = gfs2_meta_read(ip->i_gl, dblock, DIO_WAIT, 0, &bh); + error = gfs2_meta_read(gl, dblock, DIO_WAIT, 0, &bh); if (error) goto fail; } @@ -672,6 +677,7 @@ static int dirent_next(struct gfs2_inode *dip, struct buffer_head *bh, static void dirent_del(struct gfs2_inode *dip, struct buffer_head *bh, struct gfs2_dirent *prev, struct gfs2_dirent *cur) { + struct gfs2_glock *gl = gfs2_inode_glock(&dip->i_inode); u16 cur_rec_len, prev_rec_len; if (gfs2_dirent_sentinel(cur)) { @@ -679,7 +685,7 @@ static void dirent_del(struct gfs2_inode *dip, struct buffer_head *bh, return; } - gfs2_trans_add_meta(dip->i_gl, bh); + gfs2_trans_add_meta(gl, bh); /* If there is no prev entry, this is the first entry in the block. The de_rec_len is already as big as it needs to be. Just zero @@ -712,13 +718,13 @@ static struct gfs2_dirent *do_init_dirent(struct inode *inode, struct buffer_head *bh, unsigned offset) { - struct gfs2_inode *ip = GFS2_I(inode); + struct gfs2_glock *gl = gfs2_inode_glock(inode); struct gfs2_dirent *ndent; unsigned totlen; totlen = be16_to_cpu(dent->de_rec_len); BUG_ON(offset + name->len > totlen); - gfs2_trans_add_meta(ip->i_gl, bh); + gfs2_trans_add_meta(gl, bh); ndent = (struct gfs2_dirent *)((char *)dent + offset); dent->de_rec_len = cpu_to_be16(offset); gfs2_qstr2dirent(name, totlen - offset, ndent); @@ -759,10 +765,13 @@ static struct gfs2_dirent *gfs2_dirent_split_alloc(struct inode *inode, static int get_leaf(struct gfs2_inode *dip, u64 leaf_no, struct buffer_head **bhp) { + struct inode *inode = &dip->i_inode; + struct gfs2_glock *gl = gfs2_inode_glock(inode); + struct gfs2_sbd *sdp = GFS2_SB(inode); int error; - error = gfs2_meta_read(dip->i_gl, leaf_no, DIO_WAIT, 0, bhp); - if (!error && gfs2_metatype_check(GFS2_SB(&dip->i_inode), *bhp, GFS2_METATYPE_LF)) { + error = gfs2_meta_read(gl, leaf_no, DIO_WAIT, 0, bhp); + if (!error && gfs2_metatype_check(sdp, *bhp, GFS2_METATYPE_LF)) { /* pr_info("block num=%llu\n", leaf_no); */ error = -EIO; } @@ -863,6 +872,7 @@ got_dent: static struct gfs2_leaf *new_leaf(struct inode *inode, struct buffer_head **pbh, u16 depth) { + struct gfs2_glock *gl = gfs2_inode_glock(inode); struct gfs2_inode *ip = GFS2_I(inode); unsigned int n = 1; u64 bn; @@ -875,12 +885,12 @@ static struct gfs2_leaf *new_leaf(struct inode *inode, struct buffer_head **pbh, error = gfs2_alloc_blocks(ip, &bn, &n, 0); if (error) return NULL; - bh = gfs2_meta_new(ip->i_gl, bn); + bh = gfs2_meta_new(gl, bn); if (!bh) return NULL; gfs2_trans_remove_revoke(GFS2_SB(inode), bn, 1); - gfs2_trans_add_meta(ip->i_gl, bh); + gfs2_trans_add_meta(gl, bh); gfs2_metatype_set(bh, GFS2_METATYPE_LF, GFS2_FORMAT_LF); leaf = (struct gfs2_leaf *)bh->b_data; leaf->lf_depth = cpu_to_be16(depth); @@ -907,6 +917,7 @@ static struct gfs2_leaf *new_leaf(struct inode *inode, struct buffer_head **pbh, static int dir_make_exhash(struct inode *inode) { + struct gfs2_glock *gl = gfs2_inode_glock(inode); struct gfs2_inode *dip = GFS2_I(inode); struct gfs2_sbd *sdp = GFS2_SB(inode); struct gfs2_dirent *dent; @@ -968,7 +979,7 @@ static int dir_make_exhash(struct inode *inode) /* We're done with the new leaf block, now setup the new hash table. */ - gfs2_trans_add_meta(dip->i_gl, dibh); + gfs2_trans_add_meta(gl, dibh); gfs2_buffer_clear_tail(dibh, sizeof(struct gfs2_dinode)); lp = (__be64 *)(dibh->b_data + sizeof(struct gfs2_dinode)); @@ -998,6 +1009,7 @@ static int dir_make_exhash(struct inode *inode) static int dir_split_leaf(struct inode *inode, const struct qstr *name) { + struct gfs2_glock *gl = gfs2_inode_glock(inode); struct gfs2_inode *dip = GFS2_I(inode); struct buffer_head *nbh, *obh, *dibh; struct gfs2_leaf *nleaf, *oleaf; @@ -1025,7 +1037,7 @@ static int dir_split_leaf(struct inode *inode, const struct qstr *name) return 1; /* can't split */ } - gfs2_trans_add_meta(dip->i_gl, obh); + gfs2_trans_add_meta(gl, obh); nleaf = new_leaf(inode, &nbh, be16_to_cpu(oleaf->lf_depth) + 1); if (!nleaf) { @@ -1118,7 +1130,7 @@ static int dir_split_leaf(struct inode *inode, const struct qstr *name) error = gfs2_meta_inode_buffer(dip, &dibh); if (!gfs2_assert_withdraw(GFS2_SB(&dip->i_inode), !error)) { - gfs2_trans_add_meta(dip->i_gl, dibh); + gfs2_trans_add_meta(gl, dibh); gfs2_add_inode_blocks(&dip->i_inode, 1); gfs2_dinode_out(dip, dibh->b_data); brelse(dibh); @@ -1480,8 +1492,8 @@ out: static void gfs2_dir_readahead(struct inode *inode, unsigned hsize, u32 index, struct file_ra_state *f_ra) { + struct gfs2_glock *gl = gfs2_inode_glock(inode); struct gfs2_inode *ip = GFS2_I(inode); - struct gfs2_glock *gl = ip->i_gl; struct buffer_head *bh; u64 blocknr = 0, last; unsigned count; @@ -1722,6 +1734,7 @@ out: static int dir_new_leaf(struct inode *inode, const struct qstr *name) { struct buffer_head *bh, *obh; + struct gfs2_glock *gl = gfs2_inode_glock(inode); struct gfs2_inode *ip = GFS2_I(inode); struct gfs2_leaf *leaf, *oleaf; u32 dist = 1; @@ -1745,7 +1758,7 @@ static int dir_new_leaf(struct inode *inode, const struct qstr *name) return error; } while(1); - gfs2_trans_add_meta(ip->i_gl, obh); + gfs2_trans_add_meta(gl, obh); leaf = new_leaf(inode, &bh, be16_to_cpu(oleaf->lf_depth)); if (!leaf) { @@ -1760,7 +1773,7 @@ static int dir_new_leaf(struct inode *inode, const struct qstr *name) error = gfs2_meta_inode_buffer(ip, &bh); if (error) return error; - gfs2_trans_add_meta(ip->i_gl, bh); + gfs2_trans_add_meta(gl, bh); gfs2_add_inode_blocks(&ip->i_inode, 1); gfs2_dinode_out(ip, bh->b_data); brelse(bh); @@ -1935,6 +1948,7 @@ int gfs2_dir_del(struct gfs2_inode *dip, const struct dentry *dentry) int gfs2_dir_mvino(struct gfs2_inode *dip, const struct qstr *filename, const struct gfs2_inode *nip, unsigned int new_type) { + struct gfs2_glock *gl = gfs2_inode_glock(&dip->i_inode); struct buffer_head *bh; struct gfs2_dirent *dent; @@ -1946,7 +1960,7 @@ int gfs2_dir_mvino(struct gfs2_inode *dip, const struct qstr *filename, if (IS_ERR(dent)) return PTR_ERR(dent); - gfs2_trans_add_meta(dip->i_gl, bh); + gfs2_trans_add_meta(gl, bh); gfs2_inum_out(nip, dent); dent->de_type = cpu_to_be16(new_type); brelse(bh); @@ -1972,6 +1986,7 @@ static int leaf_dealloc(struct gfs2_inode *dip, u32 index, u32 len, u64 leaf_no, struct buffer_head *leaf_bh, int last_dealloc) { + struct gfs2_glock *gl = gfs2_inode_glock(&dip->i_inode); struct gfs2_sbd *sdp = GFS2_SB(&dip->i_inode); struct gfs2_leaf *tmp_leaf; struct gfs2_rgrp_list rlist; @@ -2066,7 +2081,7 @@ static int leaf_dealloc(struct gfs2_inode *dip, u32 index, u32 len, if (error) goto out_end_trans; - gfs2_trans_add_meta(dip->i_gl, dibh); + gfs2_trans_add_meta(gl, dibh); /* On the last dealloc, make this a regular file in case we crash. (We don't want to free these blocks a second time.) */ if (last_dealloc) diff --git a/fs/gfs2/export.c b/fs/gfs2/export.c index 3334c394ce9c..970bf2d72388 100644 --- a/fs/gfs2/export.c +++ b/fs/gfs2/export.c @@ -87,7 +87,8 @@ static int gfs2_get_name(struct dentry *parent, char *name, { struct inode *dir = d_inode(parent); struct inode *inode = d_inode(child); - struct gfs2_inode *dip, *ip; + struct gfs2_glock *gl; + struct gfs2_inode *ip; struct get_name_filldir gnfd = { .ctx.actor = get_name_filldir, .name = name @@ -102,14 +103,14 @@ static int gfs2_get_name(struct dentry *parent, char *name, if (!S_ISDIR(dir->i_mode) || !inode) return -EINVAL; - dip = GFS2_I(dir); + gl = gfs2_inode_glock(dir); ip = GFS2_I(inode); *name = 0; gnfd.inum.no_addr = ip->i_no_addr; gnfd.inum.no_formal_ino = ip->i_no_formal_ino; - error = gfs2_glock_nq_init(dip->i_gl, LM_ST_SHARED, 0, &gh); + error = gfs2_glock_nq_init(gl, LM_ST_SHARED, 0, &gh); if (error) return error; diff --git a/fs/gfs2/file.c b/fs/gfs2/file.c index b8c10de113ba..1efd0679badd 100644 --- a/fs/gfs2/file.c +++ b/fs/gfs2/file.c @@ -57,13 +57,13 @@ static loff_t gfs2_llseek(struct file *file, loff_t offset, int whence) { - struct gfs2_inode *ip = GFS2_I(file->f_mapping->host); + struct gfs2_glock *gl = gfs2_inode_glock(file->f_mapping->host); struct gfs2_holder i_gh; loff_t error; switch (whence) { case SEEK_END: - error = gfs2_glock_nq_init(ip->i_gl, LM_ST_SHARED, LM_FLAG_ANY, + error = gfs2_glock_nq_init(gl, LM_ST_SHARED, LM_FLAG_ANY, &i_gh); if (!error) { error = generic_file_llseek(file, offset, whence); @@ -105,11 +105,11 @@ static loff_t gfs2_llseek(struct file *file, loff_t offset, int whence) static int gfs2_readdir(struct file *file, struct dir_context *ctx) { struct inode *dir = file->f_mapping->host; - struct gfs2_inode *dip = GFS2_I(dir); + struct gfs2_glock *gl = gfs2_inode_glock(dir); struct gfs2_holder d_gh; int error; - error = gfs2_glock_nq_init(dip->i_gl, LM_ST_SHARED, 0, &d_gh); + error = gfs2_glock_nq_init(gl, LM_ST_SHARED, 0, &d_gh); if (error) return error; @@ -158,6 +158,7 @@ static inline u32 gfs2_gfsflags_to_fsflags(struct inode *inode, u32 gfsflags) int gfs2_fileattr_get(struct dentry *dentry, struct file_kattr *fa) { struct inode *inode = d_inode(dentry); + struct gfs2_glock *gl = gfs2_inode_glock(inode); struct gfs2_inode *ip = GFS2_I(inode); struct gfs2_holder gh; int error; @@ -166,7 +167,7 @@ int gfs2_fileattr_get(struct dentry *dentry, struct file_kattr *fa) if (d_is_special(dentry)) return -ENOTTY; - gfs2_holder_init(ip->i_gl, LM_ST_SHARED, 0, &gh); + gfs2_holder_init(gl, LM_ST_SHARED, 0, &gh); error = gfs2_glock_nq(&gh); if (error) goto out_uninit; @@ -218,6 +219,7 @@ void gfs2_set_inode_flags(struct inode *inode) */ static int do_gfs2_set_flags(struct inode *inode, u32 reqflags, u32 mask) { + struct gfs2_glock *gl = gfs2_inode_glock(inode); struct gfs2_inode *ip = GFS2_I(inode); struct gfs2_sbd *sdp = GFS2_SB(inode); struct buffer_head *bh; @@ -225,7 +227,7 @@ static int do_gfs2_set_flags(struct inode *inode, u32 reqflags, u32 mask) int error; u32 new_flags, flags; - error = gfs2_glock_nq_init(ip->i_gl, LM_ST_EXCLUSIVE, 0, &gh); + error = gfs2_glock_nq_init(gl, LM_ST_EXCLUSIVE, 0, &gh); if (error) return error; @@ -242,7 +244,7 @@ static int do_gfs2_set_flags(struct inode *inode, u32 reqflags, u32 mask) } if ((flags ^ new_flags) & GFS2_DIF_JDATA) { if (new_flags & GFS2_DIF_JDATA) - gfs2_log_flush(sdp, ip->i_gl, + gfs2_log_flush(sdp, gl, GFS2_LOG_HEAD_FLUSH_NORMAL | GFS2_LFC_SET_FLAGS); error = filemap_fdatawrite(inode->i_mapping); @@ -262,7 +264,7 @@ static int do_gfs2_set_flags(struct inode *inode, u32 reqflags, u32 mask) if (error) goto out_trans_end; inode_set_ctime_current(inode); - gfs2_trans_add_meta(ip->i_gl, bh); + gfs2_trans_add_meta(gl, bh); ip->i_diskflags = new_flags; gfs2_dinode_out(ip, bh->b_data); brelse(bh); @@ -275,7 +277,7 @@ out: return error; } -int gfs2_fileattr_set(struct mnt_idmap *idmap, +int gfs2_fileattr_set(const struct mnt_idmap *idmap, struct dentry *dentry, struct file_kattr *fa) { struct inode *inode = d_inode(dentry); @@ -417,6 +419,7 @@ static vm_fault_t gfs2_page_mkwrite(struct vm_fault *vmf) { struct folio *folio = page_folio(vmf->page); struct inode *inode = file_inode(vmf->vma->vm_file); + struct gfs2_glock *gl = gfs2_inode_glock(inode); struct gfs2_inode *ip = GFS2_I(inode); struct gfs2_sbd *sdp = GFS2_SB(inode); struct gfs2_alloc_parms ap = {}; @@ -430,7 +433,7 @@ static vm_fault_t gfs2_page_mkwrite(struct vm_fault *vmf) sb_start_pagefault(inode->i_sb); - gfs2_holder_init(ip->i_gl, LM_ST_EXCLUSIVE, 0, &gh); + gfs2_holder_init(gl, LM_ST_EXCLUSIVE, 0, &gh); err = gfs2_glock_nq(&gh); if (err) { ret = vmf_fs_error(err); @@ -455,7 +458,7 @@ static vm_fault_t gfs2_page_mkwrite(struct vm_fault *vmf) gfs2_size_hint(vmf->vma->vm_file, pos, length); - set_bit(GLF_DIRTY, &ip->i_gl->gl_flags); + set_bit(GLF_DIRTY, &gl->gl_flags); set_bit(GIF_SW_PAGED, &ip->i_flags); /* @@ -552,12 +555,12 @@ out_uninit: static vm_fault_t gfs2_fault(struct vm_fault *vmf) { struct inode *inode = file_inode(vmf->vma->vm_file); - struct gfs2_inode *ip = GFS2_I(inode); + struct gfs2_glock *gl = gfs2_inode_glock(inode); struct gfs2_holder gh; vm_fault_t ret; int err; - gfs2_holder_init(ip->i_gl, LM_ST_SHARED, 0, &gh); + gfs2_holder_init(gl, LM_ST_SHARED, 0, &gh); err = gfs2_glock_nq(&gh); if (err) { ret = vmf_fs_error(err); @@ -590,6 +593,7 @@ static const struct vm_operations_struct gfs2_vm_ops = { static int gfs2_mmap(struct file *file, struct vm_area_struct *vma) { + struct gfs2_glock *gl = gfs2_inode_glock(file->f_mapping->host); struct gfs2_inode *ip = GFS2_I(file->f_mapping->host); if (!(file->f_flags & O_NOATIME) && @@ -597,7 +601,7 @@ static int gfs2_mmap(struct file *file, struct vm_area_struct *vma) struct gfs2_holder i_gh; int error; - error = gfs2_glock_nq_init(ip->i_gl, LM_ST_SHARED, LM_FLAG_ANY, + error = gfs2_glock_nq_init(gl, LM_ST_SHARED, LM_FLAG_ANY, &i_gh); if (error) return error; @@ -674,13 +678,14 @@ fail: static int gfs2_open(struct inode *inode, struct file *file) { + struct gfs2_glock *gl = gfs2_inode_glock(inode); struct gfs2_inode *ip = GFS2_I(inode); struct gfs2_holder i_gh; int error; bool need_unlock = false; if (S_ISREG(ip->i_inode.i_mode)) { - error = gfs2_glock_nq_init(ip->i_gl, LM_ST_SHARED, LM_FLAG_ANY, + error = gfs2_glock_nq_init(gl, LM_ST_SHARED, LM_FLAG_ANY, &i_gh); if (error) return error; @@ -745,6 +750,7 @@ static int gfs2_fsync(struct file *file, loff_t start, loff_t end, struct address_space *mapping = file->f_mapping; struct inode *inode = mapping->host; int sync_state = inode_state_read_once(inode) & I_DIRTY; + struct gfs2_glock *gl = gfs2_inode_glock(inode); struct gfs2_inode *ip = GFS2_I(inode); int ret = 0, ret1 = 0; @@ -767,7 +773,7 @@ static int gfs2_fsync(struct file *file, loff_t start, loff_t end, ret = file_write_and_wait(file); if (ret) return ret; - gfs2_ail_flush(ip->i_gl, 1); + gfs2_ail_flush(gl, 1); } if (mapping->nrpages) @@ -812,7 +818,8 @@ static ssize_t gfs2_file_direct_read(struct kiocb *iocb, struct iov_iter *to, struct gfs2_holder *gh) { struct file *file = iocb->ki_filp; - struct gfs2_inode *ip = GFS2_I(file->f_mapping->host); + struct inode *inode = file->f_mapping->host; + struct gfs2_glock *gl = gfs2_inode_glock(inode); size_t prev_count = 0, window_size = 0; size_t read = 0; ssize_t ret; @@ -837,7 +844,7 @@ static ssize_t gfs2_file_direct_read(struct kiocb *iocb, struct iov_iter *to, if (!iov_iter_count(to)) return 0; /* skip atime */ - gfs2_holder_init(ip->i_gl, LM_ST_DEFERRED, 0, gh); + gfs2_holder_init(gl, LM_ST_DEFERRED, 0, gh); retry: ret = gfs2_glock_nq(gh); if (ret) @@ -876,7 +883,7 @@ static ssize_t gfs2_file_direct_write(struct kiocb *iocb, struct iov_iter *from, { struct file *file = iocb->ki_filp; struct inode *inode = file->f_mapping->host; - struct gfs2_inode *ip = GFS2_I(inode); + struct gfs2_glock *gl = gfs2_inode_glock(inode); size_t prev_count = 0, window_size = 0; size_t written = 0; bool enough_retries; @@ -900,13 +907,13 @@ static ssize_t gfs2_file_direct_write(struct kiocb *iocb, struct iov_iter *from, * unfortunately, have the option of only flushing a range like the * VFS does. */ - gfs2_holder_init(ip->i_gl, LM_ST_DEFERRED, 0, gh); + gfs2_holder_init(gl, LM_ST_DEFERRED, 0, gh); retry: ret = gfs2_glock_nq(gh); if (ret) goto out_uninit; /* Silently fall back to buffered I/O when writing beyond EOF */ - if (iocb->ki_pos + iov_iter_count(from) > i_size_read(&ip->i_inode)) + if (iocb->ki_pos + iov_iter_count(from) > i_size_read(inode)) goto out_unlock; from->nofault = true; @@ -948,7 +955,7 @@ out_uninit: static ssize_t gfs2_file_read_iter(struct kiocb *iocb, struct iov_iter *to) { - struct gfs2_inode *ip; + struct gfs2_glock *gl; struct gfs2_holder gh; size_t prev_count = 0, window_size = 0; size_t read = 0; @@ -979,8 +986,8 @@ static ssize_t gfs2_file_read_iter(struct kiocb *iocb, struct iov_iter *to) if (iocb->ki_flags & IOCB_NOWAIT) return ret; } - ip = GFS2_I(iocb->ki_filp->f_mapping->host); - gfs2_holder_init(ip->i_gl, LM_ST_SHARED, 0, &gh); + gl = gfs2_inode_glock(iocb->ki_filp->f_mapping->host); + gfs2_holder_init(gl, LM_ST_SHARED, 0, &gh); retry: ret = gfs2_glock_nq(&gh); if (ret) @@ -1013,7 +1020,7 @@ static ssize_t gfs2_file_buffered_write(struct kiocb *iocb, { struct file *file = iocb->ki_filp; struct inode *inode = file_inode(file); - struct gfs2_inode *ip = GFS2_I(inode); + struct gfs2_glock *gl = gfs2_inode_glock(inode); struct gfs2_sbd *sdp = GFS2_SB(inode); struct gfs2_holder *statfs_gh = NULL; size_t prev_count = 0, window_size = 0; @@ -1034,7 +1041,7 @@ static ssize_t gfs2_file_buffered_write(struct kiocb *iocb, return -ENOMEM; } - gfs2_holder_init(ip->i_gl, LM_ST_EXCLUSIVE, 0, gh); + gfs2_holder_init(gl, LM_ST_EXCLUSIVE, 0, gh); if (should_fault_in_pages(from, iocb, &prev_count, &window_size)) { retry: window_size -= fault_in_iov_iter_readable(from, window_size); @@ -1049,9 +1056,9 @@ retry: goto out_uninit; if (inode == sdp->sd_rindex) { - struct gfs2_inode *m_ip = GFS2_I(sdp->sd_statfs_inode); + struct gfs2_glock *m_gl = gfs2_inode_glock(sdp->sd_statfs_inode); - ret = gfs2_glock_nq_init(m_ip->i_gl, LM_ST_EXCLUSIVE, + ret = gfs2_glock_nq_init(m_gl, LM_ST_EXCLUSIVE, GL_NOCACHE, statfs_gh); if (ret) goto out_unlock; @@ -1105,14 +1112,15 @@ static ssize_t gfs2_file_write_iter(struct kiocb *iocb, struct iov_iter *from) { struct file *file = iocb->ki_filp; struct inode *inode = file_inode(file); - struct gfs2_inode *ip = GFS2_I(inode); struct gfs2_holder gh; ssize_t ret; gfs2_size_hint(file, iocb->ki_pos, iov_iter_count(from)); if (iocb->ki_flags & IOCB_APPEND) { - ret = gfs2_glock_nq_init(ip->i_gl, LM_ST_SHARED, 0, &gh); + struct gfs2_glock *gl = gfs2_inode_glock(inode); + + ret = gfs2_glock_nq_init(gl, LM_ST_SHARED, 0, &gh); if (ret) return ret; gfs2_glock_dq_uninit(&gh); @@ -1180,6 +1188,7 @@ out_unlock: static int fallocate_chunk(struct inode *inode, loff_t offset, loff_t len) { struct super_block *sb = inode->i_sb; + struct gfs2_glock *gl = gfs2_inode_glock(inode); struct gfs2_inode *ip = GFS2_I(inode); loff_t end = offset + len; struct buffer_head *dibh; @@ -1189,7 +1198,7 @@ static int fallocate_chunk(struct inode *inode, loff_t offset, loff_t len) if (unlikely(error)) return error; - gfs2_trans_add_meta(ip->i_gl, dibh); + gfs2_trans_add_meta(gl, dibh); if (gfs2_is_stuffed(ip)) { error = gfs2_unstuff_dinode(ip); @@ -1378,6 +1387,7 @@ static long gfs2_fallocate(struct file *file, int mode, loff_t offset, loff_t le { struct inode *inode = file_inode(file); struct gfs2_sbd *sdp = GFS2_SB(inode); + struct gfs2_glock *gl = gfs2_inode_glock(inode); struct gfs2_inode *ip = GFS2_I(inode); struct gfs2_holder gh; int ret; @@ -1390,7 +1400,7 @@ static long gfs2_fallocate(struct file *file, int mode, loff_t offset, loff_t le inode_lock(inode); - gfs2_holder_init(ip->i_gl, LM_ST_EXCLUSIVE, 0, &gh); + gfs2_holder_init(gl, LM_ST_EXCLUSIVE, 0, &gh); ret = gfs2_glock_nq(&gh); if (ret) goto out_uninit; diff --git a/fs/gfs2/glock.c b/fs/gfs2/glock.c index d59ea71a84db..d22a088c66cd 100644 --- a/fs/gfs2/glock.c +++ b/fs/gfs2/glock.c @@ -329,11 +329,6 @@ static void gfs2_holder_wake(struct gfs2_holder *gh) clear_bit(HIF_WAIT, &gh->gh_iflags); smp_mb__after_atomic(); wake_up_bit(&gh->gh_iflags, HIF_WAIT); - if (gh->gh_flags & GL_ASYNC) { - struct gfs2_sbd *sdp = glock_sbd(gh->gh_gl); - - wake_up(&sdp->sd_async_glock_wait); - } } /** @@ -512,11 +507,9 @@ static void state_change(struct gfs2_glock *gl, unsigned int new_state) static void gfs2_set_demote(int nr, struct gfs2_glock *gl) { - struct gfs2_sbd *sdp = glock_sbd(gl); - set_bit(nr, &gl->gl_flags); - smp_mb(); - wake_up(&sdp->sd_async_glock_wait); + smp_mb__after_atomic(); + wake_up_bit(&gl->gl_flags, GLF_DEMOTE); } static void gfs2_demote_wake(struct gfs2_glock *gl) @@ -891,7 +884,9 @@ static void gfs2_try_to_evict(struct gfs2_glock *gl) /* If the inode was evicted, gl->gl_object will now be NULL. */ ip = gfs2_grab_existing_inode(gl); if (ip) { - gfs2_glock_poke(ip->i_gl); + struct gfs2_glock *gl = gfs2_inode_glock(&ip->i_inode); + + gfs2_glock_poke(gl); iput(&ip->i_inode); } } @@ -1270,11 +1265,14 @@ static int glocks_pending(unsigned int num_gh, struct gfs2_holder *ghs) int gfs2_glock_async_wait(unsigned int num_gh, struct gfs2_holder *ghs, unsigned int retries) { - struct gfs2_sbd *sdp = glock_sbd(ghs[0].gh_gl); unsigned long start_time = jiffies; - int i, ret = 0; - long timeout; + struct wait_queue_head *waitq[4]; + struct wait_queue_entry wait[4]; + long ret, timeout; + int i; + BUILD_BUG_ON(ARRAY_SIZE(waitq) != ARRAY_SIZE(wait)); + BUG_ON(num_gh > ARRAY_SIZE(wait)); might_sleep(); timeout = GL_GLOCK_MIN_HOLD; @@ -1292,14 +1290,38 @@ int gfs2_glock_async_wait(unsigned int num_gh, struct gfs2_holder *ghs, timeout += (incr / 3) + get_random_long() % (incr / 3); } - if (!wait_event_interruptible_timeout(sdp->sd_async_glock_wait, - !glocks_pending(num_gh, ghs), timeout)) { - ret = -ESTALE; /* request timed out. */ - goto out; + ret = timeout; + for (i = 0; i < num_gh; i++) { + waitq[i] = bit_waitqueue(&ghs[i].gh_iflags, HIF_WAIT); + init_wait(wait + i); } - if (signal_pending(current)) - goto interrupted; + for (;;) { + for (i = 0; i < num_gh; i++) + prepare_to_wait(waitq[i], wait + i, TASK_INTERRUPTIBLE); + if (!glocks_pending(num_gh, ghs)) + break; + if (signal_pending(current)) { + ret = -EINTR; + break; + } + ret = schedule_timeout(ret); + if (!glocks_pending(num_gh, ghs)) + break; + if (!ret) { + ret = -ESTALE; /* request timed out. */ + break; + } + if (signal_pending(current)) { + ret = -EINTR; + break; + } + } + for (i = 0; i < num_gh; i++) + finish_wait(waitq[i], wait + i); + if (ret < 0) + goto out; + ret = 0; for (i = 0; i < num_gh; i++) { struct gfs2_holder *gh = &ghs[i]; int ret2; diff --git a/fs/gfs2/glops.c b/fs/gfs2/glops.c index 28f32424ee64..2597e1cf2315 100644 --- a/fs/gfs2/glops.c +++ b/fs/gfs2/glops.c @@ -393,6 +393,7 @@ static int gfs2_dinode_in(struct gfs2_inode *ip, const void *buf) umode_t mode = be32_to_cpu(str->di_mode); struct inode *inode = &ip->i_inode; bool is_new = inode_state_read_once(inode) & I_NEW; + u64 size; if (unlikely(ip->i_no_addr != be64_to_cpu(str->di_num.no_addr))) { gfs2_consist_inode(ip); @@ -418,7 +419,12 @@ static int gfs2_dinode_in(struct gfs2_inode *ip, const void *buf) i_uid_write(inode, be32_to_cpu(str->di_uid)); i_gid_write(inode, be32_to_cpu(str->di_gid)); set_nlink(inode, be32_to_cpu(str->di_nlink)); - i_size_write(inode, be64_to_cpu(str->di_size)); + size = be64_to_cpu(str->di_size); + if (unlikely(size > inode->i_sb->s_maxbytes)) { + gfs2_consist_inode(ip); + return -EIO; + } + i_size_write(inode, size); gfs2_set_inode_blocks(inode, be64_to_cpu(str->di_blocks)); atime.tv_sec = be64_to_cpu(str->di_atime); atime.tv_nsec = be32_to_cpu(str->di_atime_nsec); @@ -462,7 +468,7 @@ static int gfs2_dinode_in(struct gfs2_inode *ip, const void *buf) return -EIO; } - if (gfs2_is_stuffed(ip) && inode->i_size > gfs2_max_stuffed_size(ip)) { + if (gfs2_is_stuffed(ip) && size > gfs2_max_stuffed_size(ip)) { gfs2_consist_inode(ip); return -EIO; } @@ -602,8 +608,7 @@ static void freeze_go_callback(struct gfs2_glock *gl, bool remote) static int freeze_go_xmote_bh(struct gfs2_glock *gl) { struct gfs2_sbd *sdp = glock_sbd(gl); - struct gfs2_inode *ip = GFS2_I(sdp->sd_jdesc->jd_inode); - struct gfs2_glock *j_gl = ip->i_gl; + struct gfs2_glock *j_gl = gfs2_inode_glock(sdp->sd_jdesc->jd_inode); struct gfs2_log_header_host head; int error; diff --git a/fs/gfs2/incore.h b/fs/gfs2/incore.h index dadb4d3c9d3d..49cc232942a2 100644 --- a/fs/gfs2/incore.h +++ b/fs/gfs2/incore.h @@ -391,7 +391,7 @@ struct gfs2_inode { u64 i_generation; u64 i_eattr; unsigned long i_flags; /* GIF_... */ - struct gfs2_glock *i_gl; + struct gfs2_glock __rcu *i_gl; struct gfs2_holder i_iopen_gh; struct gfs2_qadata *i_qadata; /* quota allocation data */ struct gfs2_holder i_rgd_gh; @@ -716,7 +716,6 @@ struct gfs2_sbd { struct work_struct sd_freeze_work; struct work_struct sd_withdraw_work; wait_queue_head_t sd_kill_wait; - wait_queue_head_t sd_async_glock_wait; atomic_t sd_glock_disposal; struct completion sd_locking_init; struct completion sd_withdraw_helper; @@ -879,5 +878,8 @@ static inline unsigned gfs2_max_stuffed_size(const struct gfs2_inode *ip) return GFS2_SB(&ip->i_inode)->sd_sb.sb_bsize - sizeof(struct gfs2_dinode); } +static inline struct gfs2_glock *gfs2_inode_glock(struct inode *inode) +{ + return rcu_dereference_protected(GFS2_I(inode)->i_gl, 1); +} #endif /* __INCORE_DOT_H__ */ - diff --git a/fs/gfs2/inode.c b/fs/gfs2/inode.c index f361876c5583..2c2dc459c037 100644 --- a/fs/gfs2/inode.c +++ b/fs/gfs2/inode.c @@ -130,6 +130,7 @@ struct inode *gfs2_inode_lookup(struct super_block *sb, unsigned int type, { struct inode *inode; struct gfs2_inode *ip; + struct gfs2_glock *gl = NULL; struct gfs2_holder i_gh; int error; @@ -147,9 +148,10 @@ struct inode *gfs2_inode_lookup(struct super_block *sb, unsigned int type, gfs2_setup_inode(inode); error = gfs2_glock_get(sdp, no_addr, &gfs2_inode_glops, CREATE, - &ip->i_gl); + &gl); if (unlikely(error)) goto fail; + rcu_assign_pointer(ip->i_gl, gl); error = gfs2_glock_get(sdp, no_addr, &gfs2_iopen_glops, CREATE, &io_gl); @@ -178,14 +180,14 @@ struct inode *gfs2_inode_lookup(struct super_block *sb, unsigned int type, * block. We read the inode when instantiating it * after possibly checking the block type. */ - error = gfs2_glock_nq_init(ip->i_gl, LM_ST_EXCLUSIVE, + error = gfs2_glock_nq_init(gl, LM_ST_EXCLUSIVE, GL_SKIP, &i_gh); if (error) goto fail; error = -ESTALE; if (no_formal_ino && - gfs2_inode_already_deleted(ip->i_gl, no_formal_ino)) + gfs2_inode_already_deleted(gl, no_formal_ino)) goto fail; if (blktype != GFS2_BLKST_FREE) { @@ -196,20 +198,20 @@ struct inode *gfs2_inode_lookup(struct super_block *sb, unsigned int type, } } - set_bit(GLF_INSTANTIATE_NEEDED, &ip->i_gl->gl_flags); + set_bit(GLF_INSTANTIATE_NEEDED, &gl->gl_flags); /* Lowest possible timestamp; will be overwritten in gfs2_dinode_in. */ inode_set_atime(inode, 1LL << (8 * sizeof(inode_get_atime_sec(inode)) - 1), 0); - glock_set_object(ip->i_gl, ip); + glock_set_object(gl, ip); if (type == DT_UNKNOWN) { /* Inode glock must be locked already */ error = gfs2_instantiate(&i_gh); if (error) { - glock_clear_object(ip->i_gl, ip); + glock_clear_object(gl, ip); goto fail; } } else { @@ -240,9 +242,9 @@ fail: gfs2_glock_dq_uninit(&ip->i_iopen_gh); if (gfs2_holder_initialized(&i_gh)) gfs2_glock_dq_uninit(&i_gh); - if (ip->i_gl) { - gfs2_glock_put(ip->i_gl); - ip->i_gl = NULL; + if (gl) { + gfs2_glock_put(gl); + rcu_assign_pointer(ip->i_gl, NULL); } iget_failed(inode); return ERR_PTR(error); @@ -322,8 +324,8 @@ struct inode *gfs2_lookup_meta(struct inode *dip, const char *name) struct inode *gfs2_lookupi(struct inode *dir, const struct qstr *name, int is_root) { + struct gfs2_glock *gl = gfs2_inode_glock(dir); struct super_block *sb = dir->i_sb; - struct gfs2_inode *dip = GFS2_I(dir); struct gfs2_holder d_gh; int error = 0; struct inode *inode = NULL; @@ -339,8 +341,8 @@ struct inode *gfs2_lookupi(struct inode *dir, const struct qstr *name, return dir; } - if (gfs2_glock_is_locked_by_me(dip->i_gl) == NULL) { - error = gfs2_glock_nq_init(dip->i_gl, LM_ST_SHARED, 0, &d_gh); + if (gfs2_glock_is_locked_by_me(gl) == NULL) { + error = gfs2_glock_nq_init(gl, LM_ST_SHARED, 0, &d_gh); if (error) return ERR_PTR(error); } @@ -456,7 +458,7 @@ out: static void gfs2_final_release_pages(struct gfs2_inode *ip) { struct inode *inode = &ip->i_inode; - struct gfs2_glock *gl = ip->i_gl; + struct gfs2_glock *gl = gfs2_inode_glock(inode); /* This can only happen during incomplete inode creation. */ if (unlikely(!gl)) @@ -546,12 +548,13 @@ static void gfs2_init_dir(struct buffer_head *dibh, static void gfs2_init_xattr(struct gfs2_inode *ip) { + struct gfs2_glock *gl = gfs2_inode_glock(&ip->i_inode); struct gfs2_sbd *sdp = GFS2_SB(&ip->i_inode); struct buffer_head *bh; struct gfs2_ea_header *ea; - bh = gfs2_meta_new(ip->i_gl, ip->i_eattr); - gfs2_trans_add_meta(ip->i_gl, bh); + bh = gfs2_meta_new(gl, ip->i_eattr); + gfs2_trans_add_meta(gl, bh); gfs2_metatype_set(bh, GFS2_METATYPE_EA, GFS2_FORMAT_EA); gfs2_buffer_clear_tail(bh, sizeof(struct gfs2_meta_header)); @@ -574,11 +577,12 @@ static void gfs2_init_xattr(struct gfs2_inode *ip) static void init_dinode(struct gfs2_inode *dip, struct gfs2_inode *ip, const char *symname) { + struct gfs2_glock *gl = gfs2_inode_glock(&ip->i_inode); struct gfs2_dinode *di; struct buffer_head *dibh; - dibh = gfs2_meta_new(ip->i_gl, ip->i_no_addr); - gfs2_trans_add_meta(ip->i_gl, dibh); + dibh = gfs2_meta_new(gl, ip->i_no_addr); + gfs2_trans_add_meta(gl, dibh); di = (struct gfs2_dinode *)dibh->b_data; gfs2_dinode_out(ip, di); @@ -708,7 +712,7 @@ static int gfs2_create_inode(struct inode *dir, struct dentry *dentry, struct inode *inode = NULL; struct gfs2_inode *dip = GFS2_I(dir), *ip; struct gfs2_sbd *sdp = GFS2_SB(&dip->i_inode); - struct gfs2_glock *io_gl; + struct gfs2_glock *gl = NULL, *io_gl; int error, dealloc_error; u32 aflags = 0; unsigned blocks = 1; @@ -726,7 +730,8 @@ static int gfs2_create_inode(struct inode *dir, struct dentry *dentry, if (error) goto fail; - error = gfs2_glock_nq_init(dip->i_gl, LM_ST_EXCLUSIVE, 0, &d_gh); + error = gfs2_glock_nq_init(gfs2_inode_glock(dir), LM_ST_EXCLUSIVE, 0, + &d_gh); if (error) goto fail; gfs2_holder_mark_uninitialized(&gh); @@ -735,7 +740,7 @@ static int gfs2_create_inode(struct inode *dir, struct dentry *dentry, if (error) goto fail_gunlock; - inode = gfs2_dir_search(dir, &dentry->d_name, !S_ISREG(mode) || excl); + inode = gfs2_dir_search(dir, &dentry->d_name, excl); error = PTR_ERR(inode); if (!IS_ERR(inode)) { if (file && (file->f_flags & __O_REGULAR) && @@ -832,9 +837,10 @@ static int gfs2_create_inode(struct inode *dir, struct dentry *dentry, gfs2_set_inode_blocks(inode, blocks); - error = gfs2_glock_get(sdp, ip->i_no_addr, &gfs2_inode_glops, CREATE, &ip->i_gl); + error = gfs2_glock_get(sdp, ip->i_no_addr, &gfs2_inode_glops, CREATE, &gl); if (error) goto fail_dealloc_inode; + rcu_assign_pointer(ip->i_gl, gl); error = gfs2_glock_get(sdp, ip->i_no_addr, &gfs2_iopen_glops, CREATE, &io_gl); if (error) @@ -854,10 +860,10 @@ retry: if (error) goto fail_gunlock2; - error = gfs2_glock_nq_init(ip->i_gl, LM_ST_EXCLUSIVE, GL_SKIP, &gh); + error = gfs2_glock_nq_init(gl, LM_ST_EXCLUSIVE, GL_SKIP, &gh); if (error) goto fail_gunlock3; - clear_bit(GLF_INSTANTIATE_NEEDED, &ip->i_gl->gl_flags); + clear_bit(GLF_INSTANTIATE_NEEDED, &gl->gl_flags); error = gfs2_trans_begin(sdp, blocks, 0); if (error) @@ -870,7 +876,7 @@ retry: init_dinode(dip, ip, symname); gfs2_trans_end(sdp); - glock_set_object(ip->i_gl, ip); + glock_set_object(gl, ip); glock_set_object(io_gl, ip); gfs2_set_iop(inode); @@ -914,7 +920,7 @@ retry: return error; fail_gunlock4: - glock_clear_object(ip->i_gl, ip); + glock_clear_object(gl, ip); glock_clear_object(io_gl, ip); fail_gunlock3: gfs2_glock_dq_uninit(&ip->i_iopen_gh); @@ -932,9 +938,9 @@ fail_dealloc_inode: fs_warn(sdp, "%s: %d\n", __func__, dealloc_error); ip->i_no_addr = 0; fail_free_inode: - if (ip->i_gl) { - gfs2_glock_put(ip->i_gl); - ip->i_gl = NULL; + if (gl) { + gfs2_glock_put(gl); + rcu_assign_pointer(ip->i_gl, NULL); } gfs2_rs_deltree(&ip->i_res); gfs2_qa_put(ip); @@ -967,7 +973,7 @@ fail: * Returns: errno */ -static int gfs2_create(struct mnt_idmap *idmap, struct inode *dir, +static int gfs2_create(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *dentry, umode_t mode) { return gfs2_create_inode(dir, dentry, NULL, S_IFREG | mode, 0, NULL, 0, 1); @@ -996,7 +1002,7 @@ static struct dentry *__gfs2_lookup(struct inode *dir, struct dentry *dentry, if (inode == NULL || IS_ERR(inode)) return d_splice_alias(inode, dentry); - gl = GFS2_I(inode)->i_gl; + gl = gfs2_inode_glock(inode); error = gfs2_glock_nq_init(gl, LM_ST_SHARED, LM_FLAG_ANY, &gh); if (error) { iput(inode); @@ -1056,8 +1062,8 @@ static int gfs2_link(struct dentry *old_dentry, struct inode *dir, if (error) return error; - gfs2_holder_init(dip->i_gl, LM_ST_EXCLUSIVE, 0, &d_gh); - gfs2_holder_init(ip->i_gl, LM_ST_EXCLUSIVE, 0, &gh); + gfs2_holder_init(gfs2_inode_glock(dir), LM_ST_EXCLUSIVE, 0, &d_gh); + gfs2_holder_init(gfs2_inode_glock(inode), LM_ST_EXCLUSIVE, 0, &gh); error = gfs2_glock_nq(&d_gh); if (error) @@ -1130,7 +1136,7 @@ static int gfs2_link(struct dentry *old_dentry, struct inode *dir, if (error) goto out_brelse; - gfs2_trans_add_meta(ip->i_gl, dibh); + gfs2_trans_add_meta(gfs2_inode_glock(inode), dibh); inc_nlink(&ip->i_inode); inode_set_ctime_current(&ip->i_inode); ihold(inode); @@ -1256,8 +1262,8 @@ static int gfs2_unlink(struct inode *dir, struct dentry *dentry) error = -EROFS; - gfs2_holder_init(dip->i_gl, LM_ST_EXCLUSIVE, 0, &d_gh); - gfs2_holder_init(ip->i_gl, LM_ST_EXCLUSIVE, 0, &gh); + gfs2_holder_init(gfs2_inode_glock(dir), LM_ST_EXCLUSIVE, 0, &d_gh); + gfs2_holder_init(gfs2_inode_glock(inode), LM_ST_EXCLUSIVE, 0, &gh); rgd = gfs2_blk2rgrpd(sdp, ip->i_no_addr, 1); if (!rgd) @@ -1323,7 +1329,7 @@ out_inodes: * Returns: errno */ -static int gfs2_symlink(struct mnt_idmap *idmap, struct inode *dir, +static int gfs2_symlink(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *dentry, const char *symname) { unsigned int size; @@ -1332,7 +1338,7 @@ static int gfs2_symlink(struct mnt_idmap *idmap, struct inode *dir, if (size >= gfs2_max_stuffed_size(GFS2_I(dir))) return -ENAMETOOLONG; - return gfs2_create_inode(dir, dentry, NULL, S_IFLNK | S_IRWXUGO, 0, symname, size, 0); + return gfs2_create_inode(dir, dentry, NULL, S_IFLNK | S_IRWXUGO, 0, symname, size, 1); } /** @@ -1345,12 +1351,12 @@ static int gfs2_symlink(struct mnt_idmap *idmap, struct inode *dir, * Returns: the dentry, or ERR_PTR(errno) */ -static struct dentry *gfs2_mkdir(struct mnt_idmap *idmap, struct inode *dir, +static struct dentry *gfs2_mkdir(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *dentry, umode_t mode) { unsigned dsize = gfs2_max_stuffed_size(GFS2_I(dir)); - return ERR_PTR(gfs2_create_inode(dir, dentry, NULL, mode, 0, NULL, dsize, 0)); + return ERR_PTR(gfs2_create_inode(dir, dentry, NULL, mode, 0, NULL, dsize, 1)); } /** @@ -1363,10 +1369,10 @@ static struct dentry *gfs2_mkdir(struct mnt_idmap *idmap, struct inode *dir, * */ -static int gfs2_mknod(struct mnt_idmap *idmap, struct inode *dir, +static int gfs2_mknod(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *dentry, umode_t mode, dev_t dev) { - return gfs2_create_inode(dir, dentry, NULL, mode, dev, NULL, 0, 0); + return gfs2_create_inode(dir, dentry, NULL, mode, dev, NULL, 0, 1); } /** @@ -1531,18 +1537,19 @@ static int gfs2_rename(struct inode *odir, struct dentry *odentry, } num_gh = 1; - gfs2_holder_init(odip->i_gl, LM_ST_EXCLUSIVE, GL_ASYNC, ghs); + gfs2_holder_init(gfs2_inode_glock(odir), LM_ST_EXCLUSIVE, GL_ASYNC, ghs); if (odip != ndip) { - gfs2_holder_init(ndip->i_gl, LM_ST_EXCLUSIVE,GL_ASYNC, + gfs2_holder_init(gfs2_inode_glock(ndir), LM_ST_EXCLUSIVE,GL_ASYNC, ghs + num_gh); num_gh++; } - gfs2_holder_init(ip->i_gl, LM_ST_EXCLUSIVE, GL_ASYNC, ghs + num_gh); + gfs2_holder_init(gfs2_inode_glock(&ip->i_inode), LM_ST_EXCLUSIVE, + GL_ASYNC, ghs + num_gh); num_gh++; if (nip) { - gfs2_holder_init(nip->i_gl, LM_ST_EXCLUSIVE, GL_ASYNC, - ghs + num_gh); + gfs2_holder_init(gfs2_inode_glock(&nip->i_inode), + LM_ST_EXCLUSIVE, GL_ASYNC, ghs + num_gh); num_gh++; } @@ -1777,16 +1784,19 @@ static int gfs2_exchange(struct inode *odir, struct dentry *odentry, } num_gh = 1; - gfs2_holder_init(odip->i_gl, LM_ST_EXCLUSIVE, GL_ASYNC, ghs); + gfs2_holder_init(gfs2_inode_glock(odir), LM_ST_EXCLUSIVE, GL_ASYNC, + ghs); if (odip != ndip) { - gfs2_holder_init(ndip->i_gl, LM_ST_EXCLUSIVE, GL_ASYNC, - ghs + num_gh); + gfs2_holder_init(gfs2_inode_glock(ndir), LM_ST_EXCLUSIVE, + GL_ASYNC, ghs + num_gh); num_gh++; } - gfs2_holder_init(oip->i_gl, LM_ST_EXCLUSIVE, GL_ASYNC, ghs + num_gh); + gfs2_holder_init(gfs2_inode_glock(&oip->i_inode), LM_ST_EXCLUSIVE, + GL_ASYNC, ghs + num_gh); num_gh++; - gfs2_holder_init(nip->i_gl, LM_ST_EXCLUSIVE, GL_ASYNC, ghs + num_gh); + gfs2_holder_init(gfs2_inode_glock(&nip->i_inode), LM_ST_EXCLUSIVE, + GL_ASYNC, ghs + num_gh); num_gh++; again: @@ -1877,7 +1887,7 @@ out: return error; } -static int gfs2_rename2(struct mnt_idmap *idmap, struct inode *odir, +static int gfs2_rename2(const struct mnt_idmap *idmap, struct inode *odir, struct dentry *odentry, struct inode *ndir, struct dentry *ndentry, unsigned int flags) { @@ -1900,13 +1910,14 @@ static int gfs2_rename2(struct mnt_idmap *idmap, struct inode *odir, * * This can handle symlinks of any size. * - * Returns: 0 on success or error code + * Returns: the link target on success, an ERR_PTR() on failure */ static const char *gfs2_get_link(struct dentry *dentry, struct inode *inode, struct delayed_call *done) { + struct gfs2_glock *gl = gfs2_inode_glock(inode); struct gfs2_inode *ip = GFS2_I(inode); struct gfs2_holder i_gh; struct buffer_head *dibh; @@ -1917,7 +1928,7 @@ static const char *gfs2_get_link(struct dentry *dentry, if (!dentry) return ERR_PTR(-ECHILD); - gfs2_holder_init(ip->i_gl, LM_ST_SHARED, 0, &i_gh); + gfs2_holder_init(gl, LM_ST_SHARED, 0, &i_gh); error = gfs2_glock_nq(&i_gh); if (error) { gfs2_holder_uninit(&i_gh); @@ -1963,7 +1974,7 @@ out: * Returns: errno */ -int gfs2_permission(struct mnt_idmap *idmap, struct inode *inode, +int gfs2_permission(const struct mnt_idmap *idmap, struct inode *inode, int mask) { int may_not_block = mask & MAY_NOT_BLOCK; @@ -2095,10 +2106,11 @@ out: * Returns: errno */ -static int gfs2_setattr(struct mnt_idmap *idmap, +static int gfs2_setattr(const struct mnt_idmap *idmap, struct dentry *dentry, struct iattr *attr) { struct inode *inode = d_inode(dentry); + struct gfs2_glock *gl = gfs2_inode_glock(inode); struct gfs2_inode *ip = GFS2_I(inode); struct gfs2_holder i_gh; int error; @@ -2107,7 +2119,7 @@ static int gfs2_setattr(struct mnt_idmap *idmap, if (error) return error; - error = gfs2_glock_nq_init(ip->i_gl, LM_ST_EXCLUSIVE, 0, &i_gh); + error = gfs2_glock_nq_init(gl, LM_ST_EXCLUSIVE, 0, &i_gh); if (error) goto out; @@ -2156,19 +2168,20 @@ out: * Returns: errno */ -static int gfs2_getattr(struct mnt_idmap *idmap, +static int gfs2_getattr(const struct mnt_idmap *idmap, const struct path *path, struct kstat *stat, u32 request_mask, unsigned int flags) { struct inode *inode = d_inode(path->dentry); + struct gfs2_glock *gl = gfs2_inode_glock(inode); struct gfs2_inode *ip = GFS2_I(inode); struct gfs2_holder gh; u32 gfsflags; int error; gfs2_holder_mark_uninitialized(&gh); - if (gfs2_glock_is_locked_by_me(ip->i_gl) == NULL) { - error = gfs2_glock_nq_init(ip->i_gl, LM_ST_SHARED, LM_FLAG_ANY, &gh); + if (gfs2_glock_is_locked_by_me(gl) == NULL) { + error = gfs2_glock_nq_init(gl, LM_ST_SHARED, LM_FLAG_ANY, &gh); if (error) return error; } @@ -2204,14 +2217,14 @@ static bool fault_in_fiemap(struct fiemap_extent_info *fi) static int gfs2_fiemap(struct inode *inode, struct fiemap_extent_info *fieinfo, u64 start, u64 len) { - struct gfs2_inode *ip = GFS2_I(inode); + struct gfs2_glock *gl = gfs2_inode_glock(inode); struct gfs2_holder gh; int ret; inode_lock_shared(inode); retry: - ret = gfs2_glock_nq_init(ip->i_gl, LM_ST_SHARED, 0, &gh); + ret = gfs2_glock_nq_init(gl, LM_ST_SHARED, 0, &gh); if (ret) goto out; @@ -2234,12 +2247,12 @@ out: loff_t gfs2_seek_data(struct file *file, loff_t offset) { struct inode *inode = file->f_mapping->host; - struct gfs2_inode *ip = GFS2_I(inode); + struct gfs2_glock *gl = gfs2_inode_glock(inode); struct gfs2_holder gh; loff_t ret; inode_lock_shared(inode); - ret = gfs2_glock_nq_init(ip->i_gl, LM_ST_SHARED, 0, &gh); + ret = gfs2_glock_nq_init(gl, LM_ST_SHARED, 0, &gh); if (!ret) ret = iomap_seek_data(inode, offset, &gfs2_iomap_ops); gfs2_glock_dq_uninit(&gh); @@ -2253,12 +2266,12 @@ loff_t gfs2_seek_data(struct file *file, loff_t offset) loff_t gfs2_seek_hole(struct file *file, loff_t offset) { struct inode *inode = file->f_mapping->host; - struct gfs2_inode *ip = GFS2_I(inode); + struct gfs2_glock *gl = gfs2_inode_glock(inode); struct gfs2_holder gh; loff_t ret; inode_lock_shared(inode); - ret = gfs2_glock_nq_init(ip->i_gl, LM_ST_SHARED, 0, &gh); + ret = gfs2_glock_nq_init(gl, LM_ST_SHARED, 0, &gh); if (!ret) ret = iomap_seek_hole(inode, offset, &gfs2_iomap_ops); gfs2_glock_dq_uninit(&gh); @@ -2272,8 +2285,7 @@ loff_t gfs2_seek_hole(struct file *file, loff_t offset) static int gfs2_update_time(struct inode *inode, enum fs_update_time type, unsigned int flags) { - struct gfs2_inode *ip = GFS2_I(inode); - struct gfs2_glock *gl = ip->i_gl; + struct gfs2_glock *gl = gfs2_inode_glock(inode); struct gfs2_holder *gh; int error; diff --git a/fs/gfs2/inode.h b/fs/gfs2/inode.h index 2fcd96dd1361..99196d3114e4 100644 --- a/fs/gfs2/inode.h +++ b/fs/gfs2/inode.h @@ -97,7 +97,7 @@ int gfs2_dinode_dealloc(struct gfs2_inode *ip); struct inode *gfs2_lookupi(struct inode *dir, const struct qstr *name, int is_root); -int gfs2_permission(struct mnt_idmap *idmap, +int gfs2_permission(const struct mnt_idmap *idmap, struct inode *inode, int mask); struct inode *gfs2_lookup_meta(struct inode *dip, const char *name); void gfs2_dinode_out(const struct gfs2_inode *ip, void *buf); @@ -109,7 +109,7 @@ extern const struct file_operations gfs2_file_fops_nolock; extern const struct file_operations gfs2_dir_fops_nolock; int gfs2_fileattr_get(struct dentry *dentry, struct file_kattr *fa); -int gfs2_fileattr_set(struct mnt_idmap *idmap, +int gfs2_fileattr_set(const struct mnt_idmap *idmap, struct dentry *dentry, struct file_kattr *fa); void gfs2_set_inode_flags(struct inode *inode); diff --git a/fs/gfs2/lock_dlm.c b/fs/gfs2/lock_dlm.c index cef4bdee92a8..47eec0c13025 100644 --- a/fs/gfs2/lock_dlm.c +++ b/fs/gfs2/lock_dlm.c @@ -431,7 +431,7 @@ static void gdlm_cancel(struct gfs2_glock *gl) * 10. gfs2_control sets control_lock lvb = new gen + bits for failed jids * 12. gfs2_recover does journal recoveries for failed jids identified above * 14. gfs2_control clears control_lock lvb bits for recovered jids - * 15. gfs2_control checks if recover_block == recover_start (step 3 occured + * 15. gfs2_control checks if recover_block == recover_start (step 3 occurred * again) then do nothing, otherwise if recover_start > recover_block * then clear BLOCK_LOCKS. * @@ -809,7 +809,7 @@ static void gfs2_control_func(struct work_struct *work) /* * No more jid bits set in lvb, all recovery is done, unblock locks - * (unless a new recover_prep callback has occured blocking locks + * (unless a new recover_prep callback has occurred blocking locks * again while working above) */ diff --git a/fs/gfs2/log.c b/fs/gfs2/log.c index 78bba8cc10b8..b55daf1e6381 100644 --- a/fs/gfs2/log.c +++ b/fs/gfs2/log.c @@ -107,7 +107,7 @@ __acquires(&sdp->sd_ail_lock) gfs2_assert(sdp, bd->bd_tr == tr); if (!buffer_busy(bh)) { - if (buffer_uptodate(bh)) { + if (!buffer_write_io_error(bh)) { list_move(&bd->bd_ail_st_list, &tr->tr_ail2_list); continue; @@ -321,7 +321,7 @@ static int gfs2_ail1_empty_one(struct gfs2_sbd *sdp, struct gfs2_trans *tr, active_count++; continue; } - if (!buffer_uptodate(bh) && + if (buffer_write_io_error(bh) && !cmpxchg(&sdp->sd_log_error, 0, -EIO)) gfs2_io_error_bh(sdp, bh); /* @@ -1038,7 +1038,6 @@ void gfs2_remove_from_journal(struct buffer_head *bh, int meta) set_bit(TR_TOUCHED, &tr->tr_flags); } was_pinned = 1; - brelse(bh); } if (bd) { if (bd->bd_tr) { @@ -1056,6 +1055,8 @@ void gfs2_remove_from_journal(struct buffer_head *bh, int meta) } clear_buffer_dirty(bh); clear_buffer_uptodate(bh); + if (was_pinned) + brelse(bh); } /** diff --git a/fs/gfs2/lops.c b/fs/gfs2/lops.c index 6dabe73ad790..88c84895ec6f 100644 --- a/fs/gfs2/lops.c +++ b/fs/gfs2/lops.c @@ -48,7 +48,7 @@ void gfs2_pin(struct gfs2_sbd *sdp, struct buffer_head *bh) clear_buffer_dirty(bh); if (test_set_buffer_pinned(bh)) gfs2_assert_withdraw(sdp, 0); - if (!buffer_uptodate(bh)) + if (!buffer_uptodate(bh) || buffer_write_io_error(bh)) gfs2_io_error_bh(sdp, bh); bd = bh->b_private; /* If this buffer is in the AIL and it has already been written @@ -179,6 +179,8 @@ static void gfs2_end_log_write_bh(struct gfs2_sbd *sdp, struct folio *folio, do { if (error) mark_buffer_write_io_error(bh); + else + clear_buffer_write_io_error(bh); unlock_buffer(bh); next = bh->b_this_page; size -= bh->b_size; @@ -775,9 +777,8 @@ static int buf_lo_scan_elements(struct gfs2_jdesc *jd, u32 start, struct gfs2_log_descriptor *ld, __be64 *ptr, int pass) { - struct gfs2_inode *ip = GFS2_I(jd->jd_inode); + struct gfs2_glock *gl = gfs2_inode_glock(jd->jd_inode); struct gfs2_sbd *sdp = GFS2_SB(jd->jd_inode); - struct gfs2_glock *gl = ip->i_gl; unsigned int blks = be32_to_cpu(ld->ld_data1); struct buffer_head *bh_log, *bh_ip; u64 blkno; @@ -828,17 +829,17 @@ static int buf_lo_scan_elements(struct gfs2_jdesc *jd, u32 start, static void buf_lo_after_scan(struct gfs2_jdesc *jd, int error, int pass) { - struct gfs2_inode *ip = GFS2_I(jd->jd_inode); + struct gfs2_glock *gl = gfs2_inode_glock(jd->jd_inode); struct gfs2_sbd *sdp = GFS2_SB(jd->jd_inode); if (error) { - gfs2_inode_metasync(ip->i_gl); + gfs2_inode_metasync(gl); return; } if (pass != 1) return; - gfs2_inode_metasync(ip->i_gl); + gfs2_inode_metasync(gl); fs_info(sdp, "jid=%u: Replayed %u of %u blocks\n", jd->jd_jid, jd->jd_replayed_blocks, jd->jd_found_blocks); @@ -1000,8 +1001,7 @@ static int databuf_lo_scan_elements(struct gfs2_jdesc *jd, u32 start, struct gfs2_log_descriptor *ld, __be64 *ptr, int pass) { - struct gfs2_inode *ip = GFS2_I(jd->jd_inode); - struct gfs2_glock *gl = ip->i_gl; + struct gfs2_glock *gl = gfs2_inode_glock(jd->jd_inode); unsigned int blks = be32_to_cpu(ld->ld_data1); struct buffer_head *bh_log, *bh_ip; u64 blkno; @@ -1048,18 +1048,18 @@ static int databuf_lo_scan_elements(struct gfs2_jdesc *jd, u32 start, static void databuf_lo_after_scan(struct gfs2_jdesc *jd, int error, int pass) { - struct gfs2_inode *ip = GFS2_I(jd->jd_inode); + struct gfs2_glock *gl = gfs2_inode_glock(jd->jd_inode); struct gfs2_sbd *sdp = GFS2_SB(jd->jd_inode); if (error) { - gfs2_inode_metasync(ip->i_gl); + gfs2_inode_metasync(gl); return; } if (pass != 1) return; /* data sync? */ - gfs2_inode_metasync(ip->i_gl); + gfs2_inode_metasync(gl); fs_info(sdp, "jid=%u: Replayed %u of %u data blocks\n", jd->jd_jid, jd->jd_replayed_blocks, jd->jd_found_blocks); diff --git a/fs/gfs2/meta_io.c b/fs/gfs2/meta_io.c index a87cfbf0df38..63591a4be10a 100644 --- a/fs/gfs2/meta_io.c +++ b/fs/gfs2/meta_io.c @@ -402,18 +402,19 @@ static struct buffer_head *gfs2_getjdatabuf(struct gfs2_inode *ip, u64 blkno) void gfs2_journal_wipe(struct gfs2_inode *ip, u64 bstart, u32 blen) { + struct gfs2_glock *gl = gfs2_inode_glock(&ip->i_inode); struct gfs2_sbd *sdp = GFS2_SB(&ip->i_inode); struct buffer_head *bh; int ty; /* This can only happen during incomplete inode creation. */ - if (!ip->i_gl) + if (!gl) return; gfs2_ail1_wipe(sdp, bstart, blen); while (blen) { ty = REMOVE_META; - bh = gfs2_getbuf(ip->i_gl, bstart, NO_CREATE); + bh = gfs2_getbuf(gl, bstart, NO_CREATE); if (!bh && gfs2_is_jdata(ip)) { bh = gfs2_getjdatabuf(ip, bstart); ty = REMOVE_JDATA; @@ -447,8 +448,8 @@ void gfs2_journal_wipe(struct gfs2_inode *ip, u64 bstart, u32 blen) int gfs2_meta_buffer(struct gfs2_inode *ip, u32 mtype, u64 num, struct buffer_head **bhp) { + struct gfs2_glock *gl = gfs2_inode_glock(&ip->i_inode); struct gfs2_sbd *sdp = GFS2_SB(&ip->i_inode); - struct gfs2_glock *gl = ip->i_gl; struct buffer_head *bh; int ret = 0; int rahead = 0; diff --git a/fs/gfs2/ops_fstype.c b/fs/gfs2/ops_fstype.c index 718e0da7dfce..6ec47376aecf 100644 --- a/fs/gfs2/ops_fstype.c +++ b/fs/gfs2/ops_fstype.c @@ -90,7 +90,6 @@ static struct gfs2_sbd *init_sbd(struct super_block *sb) gfs2_tune_init(&sdp->sd_tune); init_waitqueue_head(&sdp->sd_kill_wait); - init_waitqueue_head(&sdp->sd_async_glock_wait); atomic_set(&sdp->sd_glock_disposal, 0); init_completion(&sdp->sd_locking_init); init_completion(&sdp->sd_withdraw_helper); @@ -534,7 +533,7 @@ static void gfs2_others_may_mount(struct gfs2_sbd *sdp) static int gfs2_jindex_hold(struct gfs2_sbd *sdp, struct gfs2_holder *ji_gh) { - struct gfs2_inode *dip = GFS2_I(sdp->sd_jindex); + struct gfs2_glock *gl = gfs2_inode_glock(sdp->sd_jindex); struct qstr name; char buf[20]; struct gfs2_jdesc *jd; @@ -545,7 +544,7 @@ static int gfs2_jindex_hold(struct gfs2_sbd *sdp, struct gfs2_holder *ji_gh) mutex_lock(&sdp->sd_jindex_mutex); for (;;) { - error = gfs2_glock_nq_init(dip->i_gl, LM_ST_SHARED, 0, ji_gh); + error = gfs2_glock_nq_init(gl, LM_ST_SHARED, 0, ji_gh); if (error) break; @@ -611,7 +610,6 @@ static int init_statfs(struct gfs2_sbd *sdp) struct inode *pn = NULL; char buf[30]; struct gfs2_jdesc *jd; - struct gfs2_inode *ip; sdp->sd_statfs_inode = gfs2_lookup_meta(master, "statfs"); if (IS_ERR(sdp->sd_statfs_inode)) { @@ -656,15 +654,15 @@ static int init_statfs(struct gfs2_sbd *sdp) iput(pn); pn = NULL; - ip = GFS2_I(sdp->sd_sc_inode); - error = gfs2_glock_nq_init(ip->i_gl, LM_ST_EXCLUSIVE, GL_NOPID, - &sdp->sd_sc_gh); + error = gfs2_glock_nq_init(gfs2_inode_glock(sdp->sd_sc_inode), + LM_ST_EXCLUSIVE, GL_NOPID, &sdp->sd_sc_gh); if (error) { fs_err(sdp, "can't lock local \"sc\" file: %d\n", error); goto free_local; } /* read in the local statfs buffer - other nodes don't change it. */ - error = gfs2_meta_inode_buffer(ip, &sdp->sd_sc_bh); + error = gfs2_meta_inode_buffer(GFS2_I(sdp->sd_sc_inode), + &sdp->sd_sc_bh); if (error) { fs_err(sdp, "Cannot read in local statfs: %d\n", error); goto unlock_sd_gh; @@ -697,7 +695,7 @@ static int init_journal(struct gfs2_sbd *sdp, int undo) { struct inode *master = d_inode(sdp->sd_master_dir); struct gfs2_holder ji_gh; - struct gfs2_inode *ip; + struct gfs2_glock *gl; int error = 0; gfs2_holder_mark_uninitialized(&ji_gh); @@ -751,8 +749,8 @@ static int init_journal(struct gfs2_sbd *sdp, int undo) goto fail_jindex; } - ip = GFS2_I(sdp->sd_jdesc->jd_inode); - error = gfs2_glock_nq_init(ip->i_gl, LM_ST_SHARED, + gl = gfs2_inode_glock(sdp->sd_jdesc->jd_inode); + error = gfs2_glock_nq_init(gl, LM_ST_SHARED, LM_FLAG_RECOVER | GL_EXACT | GL_NOCACHE | GL_NOPID, &sdp->sd_jinode_gh); @@ -893,7 +891,7 @@ static int init_per_node(struct gfs2_sbd *sdp, int undo) struct inode *pn = NULL; char buf[30]; int error = 0; - struct gfs2_inode *ip; + struct gfs2_glock *gl; struct inode *master = d_inode(sdp->sd_master_dir); if (sdp->sd_args.ar_spectator) @@ -920,8 +918,8 @@ static int init_per_node(struct gfs2_sbd *sdp, int undo) iput(pn); pn = NULL; - ip = GFS2_I(sdp->sd_qc_inode); - error = gfs2_glock_nq_init(ip->i_gl, LM_ST_EXCLUSIVE, GL_NOPID, + gl = gfs2_inode_glock(sdp->sd_qc_inode); + error = gfs2_glock_nq_init(gl, LM_ST_EXCLUSIVE, GL_NOPID, &sdp->sd_qc_gh); if (error) { fs_err(sdp, "can't lock local \"qc\" file: %d\n", error); diff --git a/fs/gfs2/quota.c b/fs/gfs2/quota.c index 001c8b39ca55..b2984fb7b35c 100644 --- a/fs/gfs2/quota.c +++ b/fs/gfs2/quota.c @@ -408,7 +408,7 @@ static int bh_get(struct gfs2_quota_data *qd) { struct gfs2_sbd *sdp = qd->qd_sbd; struct inode *inode = sdp->sd_qc_inode; - struct gfs2_inode *ip = GFS2_I(inode); + struct gfs2_glock *gl = gfs2_inode_glock(inode); unsigned int block, offset; struct buffer_head *bh = NULL; struct iomap iomap = { }; @@ -434,7 +434,7 @@ static int bh_get(struct gfs2_quota_data *qd) if (iomap.type != IOMAP_MAPPED) return error; - error = gfs2_meta_read(ip->i_gl, iomap.addr >> inode->i_blkbits, + error = gfs2_meta_read(gl, iomap.addr >> inode->i_blkbits, DIO_WAIT, 0, &bh); if (error) return error; @@ -562,6 +562,8 @@ int gfs2_qa_get(struct gfs2_inode *ip) if (sdp->sd_args.ar_quota == GFS2_QUOTA_OFF) return 0; + if (ip->i_diskflags & GFS2_DIF_SYSTEM) + return 0; spin_lock(&inode->i_lock); if (ip->i_qadata == NULL) { @@ -605,7 +607,7 @@ int gfs2_quota_hold(struct gfs2_inode *ip, kuid_t uid, kgid_t gid) return 0; error = gfs2_qa_get(ip); - if (error) + if (error || !ip->i_qadata) return error; qd = ip->i_qadata->qa_qd; @@ -687,12 +689,12 @@ static int sort_qd(const void *a, const void *b) static void do_qc(struct gfs2_quota_data *qd, s64 change) { struct gfs2_sbd *sdp = qd->qd_sbd; - struct gfs2_inode *ip = GFS2_I(sdp->sd_qc_inode); + struct gfs2_glock *gl = gfs2_inode_glock(sdp->sd_qc_inode); struct gfs2_quota_change *qc = qd->qd_bh_qc; bool needs_put = false; s64 x; - gfs2_trans_add_meta(ip->i_gl, qd->qd_bh); + gfs2_trans_add_meta(gl, qd->qd_bh); /* * The QDF_CHANGE flag indicates that the slot in the quota change file @@ -740,8 +742,8 @@ static void do_qc(struct gfs2_quota_data *qd, s64 change) static int gfs2_write_buf_to_page(struct gfs2_sbd *sdp, unsigned long index, unsigned off, void *buf, unsigned bytes) { - struct gfs2_inode *ip = GFS2_I(sdp->sd_quota_inode); - struct inode *inode = &ip->i_inode; + struct inode *inode = sdp->sd_quota_inode; + struct gfs2_glock *gl = gfs2_inode_glock(inode); struct address_space *mapping = inode->i_mapping; struct folio *folio; struct buffer_head *bh; @@ -780,7 +782,7 @@ static int gfs2_write_buf_to_page(struct gfs2_sbd *sdp, unsigned long index, set_buffer_uptodate(bh); if (bh_read(bh, REQ_META | REQ_PRIO) < 0) goto unlock_out; - gfs2_trans_add_data(ip->i_gl, bh); + gfs2_trans_add_data(gl, bh); /* If we need to write to the next block as well */ if (to_write > (bsize - boff)) { @@ -908,7 +910,9 @@ static int do_sync(unsigned int num_qd, struct gfs2_quota_data **qda, u64 sync_gen) { struct gfs2_sbd *sdp = (*qda)->qd_sbd; - struct gfs2_inode *ip = GFS2_I(sdp->sd_quota_inode); + struct inode *inode = sdp->sd_quota_inode; + struct gfs2_glock *gl = gfs2_inode_glock(inode); + struct gfs2_inode *ip = GFS2_I(inode); struct gfs2_alloc_parms ap = {}; unsigned int data_blocks, ind_blocks; struct gfs2_holder *ghs, i_gh; @@ -927,7 +931,7 @@ static int do_sync(unsigned int num_qd, struct gfs2_quota_data **qda, return -ENOMEM; sort(qda, num_qd, sizeof(struct gfs2_quota_data *), sort_qd, NULL); - inode_lock(&ip->i_inode); + inode_lock(inode); for (qx = 0; qx < num_qd; qx++) { error = gfs2_glock_nq_init(qda[qx]->qd_gl, LM_ST_EXCLUSIVE, GL_NOCACHE, &ghs[qx]); @@ -935,7 +939,7 @@ static int do_sync(unsigned int num_qd, struct gfs2_quota_data **qda, goto out_dq; } - error = gfs2_glock_nq_init(ip->i_gl, LM_ST_EXCLUSIVE, 0, &i_gh); + error = gfs2_glock_nq_init(gl, LM_ST_EXCLUSIVE, 0, &i_gh); if (error) goto out_dq; @@ -991,10 +995,9 @@ out_alloc: out_dq: while (qx--) gfs2_glock_dq_uninit(&ghs[qx]); - inode_unlock(&ip->i_inode); + inode_unlock(inode); kfree(ghs); - gfs2_log_flush(glock_sbd(ip->i_gl), ip->i_gl, - GFS2_LOG_HEAD_FLUSH_NORMAL | GFS2_LFC_DO_SYNC); + gfs2_log_flush(sdp, gl, GFS2_LOG_HEAD_FLUSH_NORMAL | GFS2_LFC_DO_SYNC); if (!error) { for (x = 0; x < num_qd; x++) { qd = qda[x]; @@ -1038,7 +1041,7 @@ static int do_glock(struct gfs2_quota_data *qd, int force_refresh, struct gfs2_holder *q_gh) { struct gfs2_sbd *sdp = qd->qd_sbd; - struct gfs2_inode *ip = GFS2_I(sdp->sd_quota_inode); + struct gfs2_glock *gl = gfs2_inode_glock(sdp->sd_quota_inode); struct gfs2_holder i_gh; int error; @@ -1062,7 +1065,7 @@ restart: if (error) return error; - error = gfs2_glock_nq_init(ip->i_gl, LM_ST_SHARED, 0, &i_gh); + error = gfs2_glock_nq_init(gl, LM_ST_SHARED, 0, &i_gh); if (error) goto fail; @@ -1098,6 +1101,9 @@ int gfs2_quota_lock(struct gfs2_inode *ip, kuid_t uid, kgid_t gid) error = gfs2_quota_hold(ip, uid, gid); if (error) return error; + /* Meta inodes never get quota data (see gfs2_qa_get()); nothing to lock. */ + if (!ip->i_qadata) + return 0; sort(ip->i_qadata->qa_qd, ip->i_qadata->qa_qd_num, sizeof(struct gfs2_quota_data *), sort_qd, NULL); @@ -1402,7 +1408,8 @@ int gfs2_quota_refresh(struct gfs2_sbd *sdp, struct kqid qid) int gfs2_quota_init(struct gfs2_sbd *sdp) { - struct gfs2_inode *ip = GFS2_I(sdp->sd_qc_inode); + struct inode *inode = sdp->sd_qc_inode; + struct gfs2_glock *gl = gfs2_inode_glock(inode); u64 size = i_size_read(sdp->sd_qc_inode); unsigned int blocks = size >> sdp->sd_sb.sb_bsize_shift; unsigned int x, slot = 0; @@ -1434,12 +1441,12 @@ int gfs2_quota_init(struct gfs2_sbd *sdp) if (!extlen) { extlen = 32; - error = gfs2_get_extent(&ip->i_inode, x, &dblock, &extlen); + error = gfs2_get_extent(inode, x, &dblock, &extlen); if (error) goto fail; } error = -EIO; - bh = gfs2_meta_ra(ip->i_gl, dblock, extlen); + bh = gfs2_meta_ra(gl, dblock, extlen); if (!bh) goto fail; if (gfs2_metatype_check(sdp, bh, GFS2_METATYPE_QC)) @@ -1473,7 +1480,7 @@ int gfs2_quota_init(struct gfs2_sbd *sdp) spin_lock_bucket(hash); old_qd = gfs2_qd_search_bucket_noref(hash, sdp, qc_id); if (old_qd) { - fs_err(sdp, "Corruption found in quota_change%u" + fs_err(sdp, "Corruption found in quota_change%u " "file: duplicate identifier in " "slot %u\n", sdp->sd_jdesc->jd_jid, slot); @@ -1714,7 +1721,9 @@ static int gfs2_set_dqblk(struct super_block *sb, struct kqid qid, struct qc_dqblk *fdq) { struct gfs2_sbd *sdp = sb->s_fs_info; - struct gfs2_inode *ip = GFS2_I(sdp->sd_quota_inode); + struct inode *inode = sdp->sd_quota_inode; + struct gfs2_glock *gl = gfs2_inode_glock(inode); + struct gfs2_inode *ip = GFS2_I(inode); struct gfs2_quota_data *qd; struct gfs2_holder q_gh, i_gh; unsigned int data_blocks, ind_blocks; @@ -1741,11 +1750,11 @@ static int gfs2_set_dqblk(struct super_block *sb, struct kqid qid, if (error) goto out_put; - inode_lock(&ip->i_inode); + inode_lock(inode); error = gfs2_glock_nq_init(qd->qd_gl, LM_ST_EXCLUSIVE, 0, &q_gh); if (error) goto out_unlockput; - error = gfs2_glock_nq_init(ip->i_gl, LM_ST_EXCLUSIVE, 0, &i_gh); + error = gfs2_glock_nq_init(gl, LM_ST_EXCLUSIVE, 0, &i_gh); if (error) goto out_q; @@ -1807,7 +1816,7 @@ out_q: gfs2_glock_dq_uninit(&q_gh); out_unlockput: gfs2_qa_put(ip); - inode_unlock(&ip->i_inode); + inode_unlock(inode); out_put: qd_put(qd); return error; diff --git a/fs/gfs2/recovery.c b/fs/gfs2/recovery.c index 616c46aa3434..84db5bb7dc5c 100644 --- a/fs/gfs2/recovery.c +++ b/fs/gfs2/recovery.c @@ -32,14 +32,15 @@ struct workqueue_struct *gfs2_recovery_wq; int gfs2_replay_read_block(struct gfs2_jdesc *jd, unsigned int blk, struct buffer_head **bh) { - struct gfs2_inode *ip = GFS2_I(jd->jd_inode); - struct gfs2_glock *gl = ip->i_gl; + struct inode *inode = jd->jd_inode; + struct gfs2_glock *gl = gfs2_inode_glock(inode); + struct gfs2_inode *ip = GFS2_I(inode); u64 dblock; u32 extlen; int error; extlen = 32; - error = gfs2_get_extent(&ip->i_inode, blk, &dblock, &extlen); + error = gfs2_get_extent(inode, blk, &dblock, &extlen); if (error) return error; if (!dblock) { @@ -155,7 +156,7 @@ int __get_log_header(struct gfs2_sbd *sdp, const struct gfs2_log_header *lh, * @blk: the block to look at * @head: the log header to return * - * Read the log header for a given segement in a given journal. Do a few + * Read the log header for a given segment in a given journal. Do a few * sanity checks on it. * * Returns: 0 on success, @@ -306,15 +307,13 @@ static int update_statfs_inode(struct gfs2_jdesc *jd, struct inode *inode) { struct gfs2_sbd *sdp = GFS2_SB(jd->jd_inode); - struct gfs2_inode *ip; struct buffer_head *bh; struct gfs2_statfs_change_host sc; int error = 0; BUG_ON(!inode); - ip = GFS2_I(inode); - error = gfs2_meta_inode_buffer(ip, &bh); + error = gfs2_meta_inode_buffer(GFS2_I(inode), &bh); if (error) goto out; @@ -345,7 +344,7 @@ static int update_statfs_inode(struct gfs2_jdesc *jd, mark_buffer_dirty(bh); brelse(bh); - gfs2_inode_metasync(ip->i_gl); + gfs2_inode_metasync(gfs2_inode_glock(inode)); out: return error; @@ -398,7 +397,7 @@ out: void gfs2_recover_func(struct work_struct *work) { struct gfs2_jdesc *jd = container_of(work, struct gfs2_jdesc, jd_work); - struct gfs2_inode *ip = GFS2_I(jd->jd_inode); + struct gfs2_glock *gl = gfs2_inode_glock(jd->jd_inode); struct gfs2_sbd *sdp = GFS2_SB(jd->jd_inode); struct gfs2_log_header_host head; struct gfs2_holder j_gh, ji_gh; @@ -440,7 +439,7 @@ void gfs2_recover_func(struct work_struct *work) goto fail; } - error = gfs2_glock_nq_init(ip->i_gl, LM_ST_SHARED, + error = gfs2_glock_nq_init(gl, LM_ST_SHARED, LM_FLAG_RECOVER | GL_NOCACHE, &ji_gh); if (error) diff --git a/fs/gfs2/rgrp.c b/fs/gfs2/rgrp.c index 5988a165a830..3ac751eea055 100644 --- a/fs/gfs2/rgrp.c +++ b/fs/gfs2/rgrp.c @@ -745,6 +745,7 @@ void gfs2_clear_rgrpd(struct gfs2_sbd *sdp) /** * compute_bitstructs - Compute the bitmap sizes + * @sb: The superblock * @rgd: The resource group descriptor * * Calculates bitmap descriptors, one for each block that contains bitmap data @@ -752,84 +753,74 @@ void gfs2_clear_rgrpd(struct gfs2_sbd *sdp) * Returns: errno */ -static int compute_bitstructs(struct gfs2_rgrpd *rgd) +static int compute_bitstructs(struct super_block *sb, struct gfs2_rgrpd *rgd) { struct gfs2_sbd *sdp = rgd->rd_sbd; struct gfs2_bitmap *bi; - u32 length = rgd->rd_length; /* # blocks in hdr & bitmap */ + u32 expected_length; u32 bytes_left, bytes; + u64 data_end; int x; - if (!length) - return -EINVAL; + /* + * The first resource group block has a gfs2_rgrp header; the remaining + * blocks have a gfs2_meta_header header. The rest of each block is + * filled with bitmap data. + */ + + if (rgd->rd_addr <= (GFS2_SB_ADDR >> sdp->sd_fsb2bb_shift)) { + gfs2_consist_rgrpd(rgd); + return -EIO; + } + if (check_add_overflow(rgd->rd_data0, rgd->rd_data, &data_end) || + rgd->rd_data == 0 || data_end > sb_bdev_nr_blocks(sb)) { + gfs2_consist_rgrpd(rgd); + return -EIO; + } + if (rgd->rd_bitbytes != DIV_ROUND_UP(rgd->rd_data, GFS2_NBBY)) { + gfs2_consist_rgrpd(rgd); + return -EIO; + } + expected_length = DIV_ROUND_UP(rgd->rd_bitbytes + + sizeof(struct gfs2_rgrp) - sizeof(struct gfs2_meta_header), + sdp->sd_sb.sb_bsize - sizeof(struct gfs2_meta_header)); + if (rgd->rd_length != expected_length) { + gfs2_consist_rgrpd(rgd); + return -EIO; + } + if (rgd->rd_data0 < rgd->rd_addr + rgd->rd_length) { + gfs2_consist_rgrpd(rgd); + return -EIO; + } - rgd->rd_bits = kzalloc_objs(struct gfs2_bitmap, length, GFP_NOFS); + rgd->rd_bits = kzalloc_objs(struct gfs2_bitmap, rgd->rd_length, GFP_NOFS); if (!rgd->rd_bits) return -ENOMEM; bytes_left = rgd->rd_bitbytes; - for (x = 0; x < length; x++) { + for (x = 0; x < rgd->rd_length; x++) { bi = rgd->rd_bits + x; bi->bi_flags = 0; - /* small rgrp; bitmap stored completely in header block */ - if (length == 1) { - bytes = bytes_left; - bi->bi_offset = sizeof(struct gfs2_rgrp); + if (x == 0) { + /* header block */ bi->bi_start = 0; - bi->bi_bytes = bytes; - bi->bi_blocks = bytes * GFS2_NBBY; - /* header block */ - } else if (x == 0) { - bytes = sdp->sd_sb.sb_bsize - sizeof(struct gfs2_rgrp); bi->bi_offset = sizeof(struct gfs2_rgrp); - bi->bi_start = 0; - bi->bi_bytes = bytes; - bi->bi_blocks = bytes * GFS2_NBBY; - /* last block */ - } else if (x + 1 == length) { - bytes = bytes_left; - bi->bi_offset = sizeof(struct gfs2_meta_header); - bi->bi_start = rgd->rd_bitbytes - bytes_left; - bi->bi_bytes = bytes; - bi->bi_blocks = bytes * GFS2_NBBY; - /* other blocks */ } else { - bytes = sdp->sd_sb.sb_bsize - - sizeof(struct gfs2_meta_header); + /* bitmap-only block */ + struct gfs2_bitmap *prev = bi - 1; + + bi->bi_start = prev->bi_start + prev->bi_bytes; bi->bi_offset = sizeof(struct gfs2_meta_header); - bi->bi_start = rgd->rd_bitbytes - bytes_left; - bi->bi_bytes = bytes; - bi->bi_blocks = bytes * GFS2_NBBY; } - + bytes = sdp->sd_sb.sb_bsize - bi->bi_offset; + if (bytes > bytes_left) + bytes = bytes_left; + bi->bi_bytes = bytes; + bi->bi_blocks = bytes * GFS2_NBBY; bytes_left -= bytes; } - - if (bytes_left) { - gfs2_consist_rgrpd(rgd); - return -EIO; - } - bi = rgd->rd_bits + (length - 1); - if ((bi->bi_start + bi->bi_bytes) * GFS2_NBBY != rgd->rd_data) { - gfs2_lm(sdp, - "ri_addr=%llu " - "ri_length=%u " - "ri_data0=%llu " - "ri_data=%u " - "ri_bitbytes=%u " - "start=%u len=%u offset=%u\n", - (unsigned long long)rgd->rd_addr, - rgd->rd_length, - (unsigned long long)rgd->rd_data0, - rgd->rd_data, - rgd->rd_bitbytes, - bi->bi_start, bi->bi_bytes, bi->bi_offset); - gfs2_consist_rgrpd(rgd); - return -EIO; - } - return 0; } @@ -864,6 +855,7 @@ static int rgd_insert(struct gfs2_rgrpd *rgd) { struct gfs2_sbd *sdp = rgd->rd_sbd; struct rb_node **newn = &sdp->sd_rindex_tree.rb_node, *parent = NULL; + struct rb_node *prevn; /* Figure out where to put new node */ while (*newn) { @@ -882,6 +874,19 @@ static int rgd_insert(struct gfs2_rgrpd *rgd) rb_link_node(&rgd->rd_node, parent, newn); rb_insert_color(&rgd->rd_node, &sdp->sd_rindex_tree); sdp->sd_rgrps++; + + prevn = rb_prev(&rgd->rd_node); + if (prevn) { + struct gfs2_rgrpd *prev = + rb_entry(prevn, struct gfs2_rgrpd, rd_node); + + if (prev->rd_data0 + prev->rd_data > rgd->rd_addr) { + fs_err(sdp, "overlapping resource groups.\n"); + rb_erase(&rgd->rd_node, &sdp->sd_rindex_tree); + return -ENOENT; + } + } + return 0; } @@ -928,7 +933,7 @@ static int read_rindex_entry(struct gfs2_inode *ip) if (error) goto fail; - error = compute_bitstructs(rgd); + error = compute_bitstructs(sdp->sd_vfs, rgd); if (error) goto fail_glock; @@ -944,7 +949,9 @@ static int read_rindex_entry(struct gfs2_inode *ip) return 0; } - error = 0; /* someone else read in the rgrp; free it and ignore it */ + /* If someone else read in the rgrp, free it and ignore it. */ + if (error == -EEXIST) + error = 0; fail_glock: gfs2_glock_put(rgd->rd_gl); @@ -1033,8 +1040,8 @@ static int gfs2_ri_update(struct gfs2_inode *ip) int gfs2_rindex_update(struct gfs2_sbd *sdp) { + struct gfs2_glock *gl = gfs2_inode_glock(sdp->sd_rindex); struct gfs2_inode *ip = GFS2_I(sdp->sd_rindex); - struct gfs2_glock *gl = ip->i_gl; struct gfs2_holder ri_gh; int error = 0; int unlock_required = 0; @@ -2453,7 +2460,7 @@ int gfs2_alloc_blocks(struct gfs2_inode *ip, u64 *bn, unsigned int *nblocks, if (error == 0) { struct gfs2_dinode *di = (struct gfs2_dinode *)dibh->b_data; - gfs2_trans_add_meta(ip->i_gl, dibh); + gfs2_trans_add_meta(gfs2_inode_glock(&ip->i_inode), dibh); di->di_goal_meta = di->di_goal_data = cpu_to_be64(ip->i_goal); brelse(dibh); diff --git a/fs/gfs2/super.c b/fs/gfs2/super.c index 06302c29340f..04bb4cf787d4 100644 --- a/fs/gfs2/super.c +++ b/fs/gfs2/super.c @@ -132,8 +132,7 @@ int gfs2_jdesc_check(struct gfs2_jdesc *jd) int gfs2_make_fs_rw(struct gfs2_sbd *sdp) { - struct gfs2_inode *ip = GFS2_I(sdp->sd_jdesc->jd_inode); - struct gfs2_glock *j_gl = ip->i_gl; + struct gfs2_glock *j_gl = gfs2_inode_glock(sdp->sd_jdesc->jd_inode); int error; j_gl->gl_ops->go_inval(j_gl, DIO_METADATA); @@ -176,6 +175,7 @@ void gfs2_statfs_change_out(const struct gfs2_statfs_change_host *sc, void *buf) int gfs2_statfs_init(struct gfs2_sbd *sdp) { + struct gfs2_glock *gl = gfs2_inode_glock(sdp->sd_statfs_inode); struct gfs2_inode *m_ip = GFS2_I(sdp->sd_statfs_inode); struct gfs2_statfs_change_host *m_sc = &sdp->sd_statfs_master; struct gfs2_statfs_change_host *l_sc = &sdp->sd_statfs_local; @@ -183,7 +183,7 @@ int gfs2_statfs_init(struct gfs2_sbd *sdp) struct gfs2_holder gh; int error; - error = gfs2_glock_nq_init(m_ip->i_gl, LM_ST_EXCLUSIVE, GL_NOCACHE, + error = gfs2_glock_nq_init(gl, LM_ST_EXCLUSIVE, GL_NOCACHE, &gh); if (error) return error; @@ -216,13 +216,13 @@ out: void gfs2_statfs_change(struct gfs2_sbd *sdp, s64 total, s64 free, s64 dinodes) { - struct gfs2_inode *l_ip = GFS2_I(sdp->sd_sc_inode); + struct gfs2_glock *gl = gfs2_inode_glock(sdp->sd_sc_inode); struct gfs2_statfs_change_host *l_sc = &sdp->sd_statfs_local; struct gfs2_statfs_change_host *m_sc = &sdp->sd_statfs_master; s64 x, y; int need_sync = 0; - gfs2_trans_add_meta(l_ip->i_gl, sdp->sd_sc_bh); + gfs2_trans_add_meta(gl, sdp->sd_sc_bh); spin_lock(&sdp->sd_statfs_spin); l_sc->sc_total += total; @@ -244,13 +244,13 @@ void gfs2_statfs_change(struct gfs2_sbd *sdp, s64 total, s64 free, void update_statfs(struct gfs2_sbd *sdp, struct buffer_head *m_bh) { - struct gfs2_inode *m_ip = GFS2_I(sdp->sd_statfs_inode); - struct gfs2_inode *l_ip = GFS2_I(sdp->sd_sc_inode); + struct gfs2_glock *m_gl = gfs2_inode_glock(sdp->sd_statfs_inode); + struct gfs2_glock *l_gl = gfs2_inode_glock(sdp->sd_sc_inode); struct gfs2_statfs_change_host *m_sc = &sdp->sd_statfs_master; struct gfs2_statfs_change_host *l_sc = &sdp->sd_statfs_local; - gfs2_trans_add_meta(l_ip->i_gl, sdp->sd_sc_bh); - gfs2_trans_add_meta(m_ip->i_gl, m_bh); + gfs2_trans_add_meta(l_gl, sdp->sd_sc_bh); + gfs2_trans_add_meta(m_gl, m_bh); spin_lock(&sdp->sd_statfs_spin); m_sc->sc_total += l_sc->sc_total; @@ -273,8 +273,8 @@ int gfs2_statfs_sync(struct super_block *sb, int type) struct buffer_head *m_bh; int error; - error = gfs2_glock_nq_init(m_ip->i_gl, LM_ST_EXCLUSIVE, GL_NOCACHE, - &gh); + error = gfs2_glock_nq_init(gfs2_inode_glock(&m_ip->i_inode), + LM_ST_EXCLUSIVE, GL_NOCACHE, &gh); if (error) goto out; @@ -323,7 +323,6 @@ struct lfcc { static int gfs2_lock_fs_check_clean(struct gfs2_sbd *sdp) { - struct gfs2_inode *ip; struct gfs2_jdesc *jd; struct lfcc *lfcc; LIST_HEAD(list); @@ -336,13 +335,14 @@ static int gfs2_lock_fs_check_clean(struct gfs2_sbd *sdp) */ list_for_each_entry(jd, &sdp->sd_jindex_list, jd_list) { + struct gfs2_glock *gl = gfs2_inode_glock(jd->jd_inode); + lfcc = kmalloc_obj(struct lfcc); if (!lfcc) { error = -ENOMEM; goto out; } - ip = GFS2_I(jd->jd_inode); - error = gfs2_glock_nq_init(ip->i_gl, LM_ST_SHARED, 0, &lfcc->gh); + error = gfs2_glock_nq_init(gl, LM_ST_SHARED, 0, &lfcc->gh); if (error) { kfree(lfcc); goto out; @@ -438,15 +438,16 @@ void gfs2_dinode_out(const struct gfs2_inode *ip, void *buf) static int gfs2_write_inode(struct inode *inode, struct writeback_control *wbc) { + struct gfs2_glock *gl = gfs2_inode_glock(inode); struct gfs2_inode *ip = GFS2_I(inode); struct gfs2_sbd *sdp = GFS2_SB(inode); - struct address_space *metamapping = gfs2_glock2aspace(ip->i_gl); + struct address_space *metamapping = gfs2_glock2aspace(gl); struct backing_dev_info *bdi = inode_to_bdi(metamapping->host); int ret = 0; bool flush_all = (wbc->sync_mode == WB_SYNC_ALL || gfs2_is_jdata(ip)); if (flush_all) - gfs2_log_flush(GFS2_SB(inode), ip->i_gl, + gfs2_log_flush(GFS2_SB(inode), gl, GFS2_LOG_HEAD_FLUSH_NORMAL | GFS2_LFC_WRITE_INODE); if (bdi_wb_dirty_exceeded(bdi)) @@ -481,6 +482,7 @@ static int gfs2_write_inode(struct inode *inode, struct writeback_control *wbc) static void gfs2_dirty_inode(struct inode *inode, int flags) { + struct gfs2_glock *gl = gfs2_inode_glock(inode); struct gfs2_inode *ip = GFS2_I(inode); struct gfs2_sbd *sdp = GFS2_SB(inode); struct buffer_head *bh; @@ -490,20 +492,20 @@ static void gfs2_dirty_inode(struct inode *inode, int flags) int ret; /* This can only happen during incomplete inode creation. */ - if (unlikely(!ip->i_gl)) + if (unlikely(!gl)) return; if (gfs2_withdrawn(sdp)) return; - if (!gfs2_glock_is_locked_by_me(ip->i_gl)) { - ret = gfs2_glock_nq_init(ip->i_gl, LM_ST_EXCLUSIVE, 0, &gh); + if (!gfs2_glock_is_locked_by_me(gl)) { + ret = gfs2_glock_nq_init(gl, LM_ST_EXCLUSIVE, 0, &gh); if (ret) { fs_err(sdp, "dirty_inode: glock %d\n", ret); - gfs2_dump_glock(NULL, ip->i_gl, true); + gfs2_dump_glock(NULL, gl, true); return; } need_unlock = 1; - } else if (WARN_ON_ONCE(ip->i_gl->gl_state != LM_ST_EXCLUSIVE)) + } else if (WARN_ON_ONCE(gl->gl_state != LM_ST_EXCLUSIVE)) return; if (current->journal_info == NULL) { @@ -517,7 +519,7 @@ static void gfs2_dirty_inode(struct inode *inode, int flags) ret = gfs2_meta_inode_buffer(ip, &bh); if (ret == 0) { - gfs2_trans_add_meta(ip->i_gl, bh); + gfs2_trans_add_meta(gl, bh); gfs2_dinode_out(ip, bh->b_data); brelse(bh); } @@ -1176,9 +1178,12 @@ static void gfs2_glock_put_eventually(struct gfs2_glock *gl) static enum evict_behavior gfs2_upgrade_iopen_glock(struct inode *inode) { + struct gfs2_glock *gl = gfs2_inode_glock(inode); struct gfs2_inode *ip = GFS2_I(inode); - struct gfs2_sbd *sdp = GFS2_SB(inode); struct gfs2_holder *gh = &ip->i_iopen_gh; + struct wait_queue_head *holder_waitq, *glock_waitq; + struct wait_queue_entry holder_wait, glock_wait; + long ret = 5 * HZ; int error; gh->gh_flags |= GL_NOCACHE; @@ -1209,13 +1214,31 @@ static enum evict_behavior gfs2_upgrade_iopen_glock(struct inode *inode) if (error) return EVICT_SHOULD_SKIP_DELETE; - wait_event_interruptible_timeout(sdp->sd_async_glock_wait, - !test_bit(HIF_WAIT, &gh->gh_iflags) || - glock_needs_demote(ip->i_gl), - 5 * HZ); + holder_waitq = bit_waitqueue(&gh->gh_iflags, HIF_WAIT); + glock_waitq = bit_waitqueue(&gl->gl_flags, GLF_DEMOTE); + init_wait(&holder_wait); + init_wait(&glock_wait); + for (;;) { + prepare_to_wait(holder_waitq, &holder_wait, TASK_INTERRUPTIBLE); + prepare_to_wait(glock_waitq, &glock_wait, TASK_INTERRUPTIBLE); + if (gfs2_glock_poll(gh) || glock_needs_demote(gl)) + break; + if (signal_pending(current)) + break; + ret = schedule_timeout(ret); + if (gfs2_glock_poll(gh) || glock_needs_demote(gl)) + break; + if (!ret) + break; + if (signal_pending(current)) + break; + } + finish_wait(holder_waitq, &holder_wait); + finish_wait(glock_waitq, &glock_wait); + if (!test_bit(HIF_HOLDER, &gh->gh_iflags)) { gfs2_glock_dq(gh); - if (glock_needs_demote(ip->i_gl)) + if (glock_needs_demote(gl)) return EVICT_SHOULD_SKIP_DELETE; return EVICT_SHOULD_DEFER_DELETE; } @@ -1238,6 +1261,7 @@ static enum evict_behavior gfs2_upgrade_iopen_glock(struct inode *inode) static enum evict_behavior evict_should_delete(struct inode *inode, struct gfs2_holder *gh) { + struct gfs2_glock *gl = gfs2_inode_glock(inode); struct gfs2_inode *ip = GFS2_I(inode); struct super_block *sb = inode->i_sb; struct gfs2_sbd *sdp = sb->s_fs_info; @@ -1255,11 +1279,11 @@ static enum evict_behavior evict_should_delete(struct inode *inode, return EVICT_SHOULD_DEFER_DELETE; /* Must not read inode block until block type has been verified */ - ret = gfs2_glock_nq_init(ip->i_gl, LM_ST_EXCLUSIVE, GL_SKIP, gh); + ret = gfs2_glock_nq_init(gl, LM_ST_EXCLUSIVE, GL_SKIP, gh); if (unlikely(ret)) return EVICT_SHOULD_SKIP_DELETE; - if (gfs2_inode_already_deleted(ip->i_gl, ip->i_no_formal_ino)) + if (gfs2_inode_already_deleted(gl, ip->i_no_formal_ino)) return EVICT_SHOULD_SKIP_DELETE; ret = gfs2_check_blk_type(sdp, ip->i_no_addr, GFS2_BLKST_UNLINKED); if (ret) @@ -1288,8 +1312,8 @@ static enum evict_behavior evict_should_delete(struct inode *inode, */ static int evict_unlinked_inode(struct inode *inode, struct gfs2_holder *gh) { + struct gfs2_glock *gl = gfs2_inode_glock(inode); struct gfs2_inode *ip = GFS2_I(inode); - struct gfs2_glock *gl = ip->i_gl; int ret; /* The inode glock must be held exclusively and be instantiated. */ @@ -1390,8 +1414,7 @@ static int evict_linked_inode(struct inode *inode, struct gfs2_holder *gh) { struct super_block *sb = inode->i_sb; struct gfs2_sbd *sdp = sb->s_fs_info; - struct gfs2_inode *ip = GFS2_I(inode); - struct gfs2_glock *gl = ip->i_gl; + struct gfs2_glock *gl = gfs2_inode_glock(inode); struct address_space *metamapping = gfs2_glock2aspace(gl); int ret; @@ -1446,13 +1469,14 @@ static void gfs2_evict_inode(struct inode *inode) { struct super_block *sb = inode->i_sb; struct gfs2_sbd *sdp = sb->s_fs_info; + struct gfs2_glock *gl = gfs2_inode_glock(inode); struct gfs2_inode *ip = GFS2_I(inode); struct gfs2_holder gh; enum evict_behavior behavior; int ret; gfs2_holder_mark_uninitialized(&gh); - if (sb_rdonly(sb) || !ip->i_no_addr || !ip->i_gl) + if (sb_rdonly(sb) || !ip->i_no_addr || !gl) goto out; /* @@ -1505,10 +1529,10 @@ out: gfs2_glock_dq_uninit(&ip->i_iopen_gh); gfs2_glock_put_eventually(gl); } - if (ip->i_gl) { - glock_clear_object(ip->i_gl, ip); + if (gl) { + glock_clear_object(gl, ip); wait_on_bit_io(&ip->i_flags, GIF_GLOP_PENDING, TASK_UNINTERRUPTIBLE); - gfs2_glock_put_eventually(ip->i_gl); + gfs2_glock_put_eventually(gl); rcu_assign_pointer(ip->i_gl, NULL); } } diff --git a/fs/gfs2/trace_gfs2.h b/fs/gfs2/trace_gfs2.h index bc40320ef239..0c60d755ba76 100644 --- a/fs/gfs2/trace_gfs2.h +++ b/fs/gfs2/trace_gfs2.h @@ -457,7 +457,7 @@ TRACE_EVENT(gfs2_bmap, ), TP_fast_assign( - __entry->dev = glock_sbd(ip->i_gl)->sd_vfs->s_dev; + __entry->dev = ip->i_inode.i_sb->s_dev; __entry->lblock = lblock; __entry->pblock = buffer_mapped(bh) ? bh->b_blocknr : 0; __entry->inum = ip->i_no_addr; @@ -493,7 +493,7 @@ TRACE_EVENT(gfs2_iomap_start, ), TP_fast_assign( - __entry->dev = glock_sbd(ip->i_gl)->sd_vfs->s_dev; + __entry->dev = ip->i_inode.i_sb->s_dev; __entry->inum = ip->i_no_addr; __entry->pos = pos; __entry->length = length; @@ -525,7 +525,7 @@ TRACE_EVENT(gfs2_iomap_end, ), TP_fast_assign( - __entry->dev = glock_sbd(ip->i_gl)->sd_vfs->s_dev; + __entry->dev = ip->i_inode.i_sb->s_dev; __entry->inum = ip->i_no_addr; __entry->offset = iomap->offset; __entry->length = iomap->length; diff --git a/fs/gfs2/util.c b/fs/gfs2/util.c index 83b8bb6446e5..61b0668ecabc 100644 --- a/fs/gfs2/util.c +++ b/fs/gfs2/util.c @@ -55,10 +55,9 @@ int check_journal_clean(struct gfs2_sbd *sdp, struct gfs2_jdesc *jd, int error; struct gfs2_holder j_gh; struct gfs2_log_header_host head; - struct gfs2_inode *ip; + struct gfs2_glock *gl = gfs2_inode_glock(jd->jd_inode); - ip = GFS2_I(jd->jd_inode); - error = gfs2_glock_nq_init(ip->i_gl, LM_ST_SHARED, LM_FLAG_RECOVER | + error = gfs2_glock_nq_init(gl, LM_ST_SHARED, LM_FLAG_RECOVER | GL_EXACT | GL_NOCACHE, &j_gh); if (error) { if (verbose) @@ -333,6 +332,7 @@ void gfs2_consist_i(struct gfs2_sbd *sdp, const char *function, void gfs2_consist_inode_i(struct gfs2_inode *ip, const char *function, char *file, unsigned int line) { + struct gfs2_glock *gl = gfs2_inode_glock(&ip->i_inode); struct gfs2_sbd *sdp = GFS2_SB(&ip->i_inode); gfs2_lm(sdp, @@ -342,7 +342,7 @@ void gfs2_consist_inode_i(struct gfs2_inode *ip, (unsigned long long)ip->i_no_formal_ino, (unsigned long long)ip->i_no_addr, function, file, line); - gfs2_dump_glock(NULL, ip->i_gl, 1); + gfs2_dump_glock(NULL, gl, 1); gfs2_withdraw(sdp); } diff --git a/fs/gfs2/xattr.c b/fs/gfs2/xattr.c index b9f48d6f10a9..c26f180419d9 100644 --- a/fs/gfs2/xattr.c +++ b/fs/gfs2/xattr.c @@ -128,11 +128,12 @@ static int ea_foreach_i(struct gfs2_inode *ip, struct buffer_head *bh, static int ea_foreach(struct gfs2_inode *ip, ea_call_t ea_call, void *data) { + struct gfs2_glock *gl = gfs2_inode_glock(&ip->i_inode); struct buffer_head *bh, *eabh; __be64 *eablk, *end; int error; - error = gfs2_meta_read(ip->i_gl, ip->i_eattr, DIO_WAIT, 0, &bh); + error = gfs2_meta_read(gl, ip->i_eattr, DIO_WAIT, 0, &bh); if (error) return error; @@ -156,7 +157,7 @@ static int ea_foreach(struct gfs2_inode *ip, ea_call_t ea_call, void *data) break; bn = be64_to_cpu(*eablk); - error = gfs2_meta_read(ip->i_gl, bn, DIO_WAIT, 0, &eabh); + error = gfs2_meta_read(gl, bn, DIO_WAIT, 0, &eabh); if (error) break; error = ea_foreach_i(ip, eabh, ea_call, data); @@ -279,7 +280,7 @@ static int ea_dealloc_unstuffed(struct gfs2_inode *ip, struct buffer_head *bh, if (error) goto out_gunlock; - gfs2_trans_add_meta(ip->i_gl, bh); + gfs2_trans_add_meta(gfs2_inode_glock(&ip->i_inode), bh); dataptrs = GFS2_EA2DATAPTRS(ea); for (x = 0; x < ea->ea_num_ptrs; x++, dataptrs++) { @@ -426,7 +427,8 @@ ssize_t gfs2_listxattr(struct dentry *dentry, char *buffer, size_t size) er.er_data_len = size; } - error = gfs2_glock_nq_init(ip->i_gl, LM_ST_SHARED, LM_FLAG_ANY, &i_gh); + error = gfs2_glock_nq_init(gfs2_inode_glock(&ip->i_inode), LM_ST_SHARED, + LM_FLAG_ANY, &i_gh); if (error) return error; @@ -457,6 +459,7 @@ ssize_t gfs2_listxattr(struct dentry *dentry, char *buffer, size_t size) static int gfs2_iter_unstuffed(struct gfs2_inode *ip, struct gfs2_ea_header *ea, const char *din, char *dout) { + struct gfs2_glock *gl = gfs2_inode_glock(&ip->i_inode); struct gfs2_sbd *sdp = GFS2_SB(&ip->i_inode); struct buffer_head **bh; unsigned int amount = GFS2_EA_DATA_LEN(ea); @@ -472,7 +475,7 @@ static int gfs2_iter_unstuffed(struct gfs2_inode *ip, struct gfs2_ea_header *ea, return -ENOMEM; for (x = 0; x < nptrs; x++) { - error = gfs2_meta_read(ip->i_gl, be64_to_cpu(*dataptrs), 0, 0, + error = gfs2_meta_read(gl, be64_to_cpu(*dataptrs), 0, 0, bh + x); if (error) { while (x--) @@ -505,7 +508,7 @@ static int gfs2_iter_unstuffed(struct gfs2_inode *ip, struct gfs2_ea_header *ea, } if (din) { - gfs2_trans_add_meta(ip->i_gl, bh[x]); + gfs2_trans_add_meta(gl, bh[x]); memcpy(pos, din, cp_size); din += sdp->sd_jbsize; } @@ -608,14 +611,14 @@ static int gfs2_xattr_get(const struct xattr_handler *handler, struct dentry *unused, struct inode *inode, const char *name, void *buffer, size_t size) { - struct gfs2_inode *ip = GFS2_I(inode); + struct gfs2_glock *gl = gfs2_inode_glock(inode); struct gfs2_holder gh; int ret; /* During lookup, SELinux calls this function with the glock locked. */ - if (!gfs2_glock_is_locked_by_me(ip->i_gl)) { - ret = gfs2_glock_nq_init(ip->i_gl, LM_ST_SHARED, LM_FLAG_ANY, &gh); + if (!gfs2_glock_is_locked_by_me(gl)) { + ret = gfs2_glock_nq_init(gl, LM_ST_SHARED, LM_FLAG_ANY, &gh); if (ret) return ret; } else { @@ -637,6 +640,7 @@ static int gfs2_xattr_get(const struct xattr_handler *handler, static int ea_alloc_blk(struct gfs2_inode *ip, struct buffer_head **bhp) { + struct gfs2_glock *gl = gfs2_inode_glock(&ip->i_inode); struct gfs2_sbd *sdp = GFS2_SB(&ip->i_inode); struct gfs2_ea_header *ea; unsigned int n = 1; @@ -647,8 +651,8 @@ static int ea_alloc_blk(struct gfs2_inode *ip, struct buffer_head **bhp) if (error) return error; gfs2_trans_remove_revoke(sdp, block, 1); - *bhp = gfs2_meta_new(ip->i_gl, block); - gfs2_trans_add_meta(ip->i_gl, *bhp); + *bhp = gfs2_meta_new(gl, block); + gfs2_trans_add_meta(gl, *bhp); gfs2_metatype_set(*bhp, GFS2_METATYPE_EA, GFS2_FORMAT_EA); gfs2_buffer_clear_tail(*bhp, sizeof(struct gfs2_meta_header)); @@ -678,6 +682,7 @@ static int ea_alloc_blk(struct gfs2_inode *ip, struct buffer_head **bhp) static int ea_write(struct gfs2_inode *ip, struct gfs2_ea_header *ea, struct gfs2_ea_request *er) { + struct gfs2_glock *gl = gfs2_inode_glock(&ip->i_inode); struct gfs2_sbd *sdp = GFS2_SB(&ip->i_inode); int error; @@ -709,8 +714,8 @@ static int ea_write(struct gfs2_inode *ip, struct gfs2_ea_header *ea, if (error) return error; gfs2_trans_remove_revoke(sdp, block, 1); - bh = gfs2_meta_new(ip->i_gl, block); - gfs2_trans_add_meta(ip->i_gl, bh); + bh = gfs2_meta_new(gl, block); + gfs2_trans_add_meta(gl, bh); gfs2_metatype_set(bh, GFS2_METATYPE_ED, GFS2_FORMAT_ED); gfs2_add_inode_blocks(&ip->i_inode, 1); @@ -841,11 +846,12 @@ static struct gfs2_ea_header *ea_split_ea(struct gfs2_ea_header *ea) static void ea_set_remove_stuffed(struct gfs2_inode *ip, struct gfs2_ea_location *el) { + struct gfs2_glock *gl = gfs2_inode_glock(&ip->i_inode); struct gfs2_ea_header *ea = el->el_ea; struct gfs2_ea_header *prev = el->el_prev; u32 len; - gfs2_trans_add_meta(ip->i_gl, el->el_bh); + gfs2_trans_add_meta(gl, el->el_bh); if (!prev || !GFS2_EA_IS_STUFFED(ea)) { ea->ea_type = GFS2_EATYPE_UNUSED; @@ -875,6 +881,7 @@ struct ea_set { static int ea_set_simple_noalloc(struct gfs2_inode *ip, struct buffer_head *bh, struct gfs2_ea_header *ea, struct ea_set *es) { + struct gfs2_glock *gl = gfs2_inode_glock(&ip->i_inode); struct gfs2_ea_request *er = es->es_er; int error; @@ -882,7 +889,7 @@ static int ea_set_simple_noalloc(struct gfs2_inode *ip, struct buffer_head *bh, if (error) return error; - gfs2_trans_add_meta(ip->i_gl, bh); + gfs2_trans_add_meta(gl, bh); if (es->ea_split) ea = ea_split_ea(ea); @@ -902,11 +909,12 @@ static int ea_set_simple_noalloc(struct gfs2_inode *ip, struct buffer_head *bh, static int ea_set_simple_alloc(struct gfs2_inode *ip, struct gfs2_ea_request *er, void *private) { + struct gfs2_glock *gl = gfs2_inode_glock(&ip->i_inode); struct ea_set *es = private; struct gfs2_ea_header *ea = es->es_ea; int error; - gfs2_trans_add_meta(ip->i_gl, es->es_bh); + gfs2_trans_add_meta(gl, es->es_bh); if (es->ea_split) ea = ea_split_ea(ea); @@ -971,6 +979,7 @@ static int ea_set_simple(struct gfs2_inode *ip, struct buffer_head *bh, static int ea_set_block(struct gfs2_inode *ip, struct gfs2_ea_request *er, void *private) { + struct gfs2_glock *gl = gfs2_inode_glock(&ip->i_inode); struct gfs2_sbd *sdp = GFS2_SB(&ip->i_inode); struct buffer_head *indbh, *newbh; __be64 *eablk; @@ -980,7 +989,7 @@ static int ea_set_block(struct gfs2_inode *ip, struct gfs2_ea_request *er, if (ip->i_diskflags & GFS2_DIF_EA_INDIRECT) { __be64 *end; - error = gfs2_meta_read(ip->i_gl, ip->i_eattr, DIO_WAIT, 0, + error = gfs2_meta_read(gl, ip->i_eattr, DIO_WAIT, 0, &indbh); if (error) return error; @@ -1002,7 +1011,7 @@ static int ea_set_block(struct gfs2_inode *ip, struct gfs2_ea_request *er, goto out; } - gfs2_trans_add_meta(ip->i_gl, indbh); + gfs2_trans_add_meta(gl, indbh); } else { u64 blk; unsigned int n = 1; @@ -1010,8 +1019,8 @@ static int ea_set_block(struct gfs2_inode *ip, struct gfs2_ea_request *er, if (error) return error; gfs2_trans_remove_revoke(sdp, blk, 1); - indbh = gfs2_meta_new(ip->i_gl, blk); - gfs2_trans_add_meta(ip->i_gl, indbh); + indbh = gfs2_meta_new(gl, blk); + gfs2_trans_add_meta(gl, indbh); gfs2_metatype_set(indbh, GFS2_METATYPE_IN, GFS2_FORMAT_IN); gfs2_buffer_clear_tail(indbh, mh_size); @@ -1088,6 +1097,7 @@ static int ea_set_remove_unstuffed(struct gfs2_inode *ip, static int ea_remove_stuffed(struct gfs2_inode *ip, struct gfs2_ea_location *el) { + struct gfs2_glock *gl = gfs2_inode_glock(&ip->i_inode); struct gfs2_ea_header *ea = el->el_ea; struct gfs2_ea_header *prev = el->el_prev; int error; @@ -1096,7 +1106,7 @@ static int ea_remove_stuffed(struct gfs2_inode *ip, struct gfs2_ea_location *el) if (error) return error; - gfs2_trans_add_meta(ip->i_gl, el->el_bh); + gfs2_trans_add_meta(gl, el->el_bh); if (prev) { u32 len; @@ -1229,11 +1239,12 @@ int __gfs2_xattr_set(struct inode *inode, const char *name, } static int gfs2_xattr_set(const struct xattr_handler *handler, - struct mnt_idmap *idmap, + const struct mnt_idmap *idmap, struct dentry *unused, struct inode *inode, const char *name, const void *value, size_t size, int flags) { + struct gfs2_glock *gl = gfs2_inode_glock(inode); struct gfs2_inode *ip = GFS2_I(inode); struct gfs2_holder gh; int ret; @@ -1244,12 +1255,12 @@ static int gfs2_xattr_set(const struct xattr_handler *handler, /* May be called from gfs_setattr with the glock locked. */ - if (!gfs2_glock_is_locked_by_me(ip->i_gl)) { - ret = gfs2_glock_nq_init(ip->i_gl, LM_ST_EXCLUSIVE, 0, &gh); + if (!gfs2_glock_is_locked_by_me(gl)) { + ret = gfs2_glock_nq_init(gl, LM_ST_EXCLUSIVE, 0, &gh); if (ret) goto out; } else { - if (WARN_ON_ONCE(ip->i_gl->gl_state != LM_ST_EXCLUSIVE)) { + if (WARN_ON_ONCE(gl->gl_state != LM_ST_EXCLUSIVE)) { ret = -EIO; goto out; } @@ -1265,6 +1276,7 @@ out: static int ea_dealloc_indirect(struct gfs2_inode *ip) { + struct gfs2_glock *gl = gfs2_inode_glock(&ip->i_inode); struct gfs2_sbd *sdp = GFS2_SB(&ip->i_inode); struct gfs2_rgrp_list rlist; struct gfs2_rgrpd *rgd; @@ -1283,7 +1295,7 @@ static int ea_dealloc_indirect(struct gfs2_inode *ip) memset(&rlist, 0, sizeof(struct gfs2_rgrp_list)); - error = gfs2_meta_read(ip->i_gl, ip->i_eattr, DIO_WAIT, 0, &indbh); + error = gfs2_meta_read(gl, ip->i_eattr, DIO_WAIT, 0, &indbh); if (error) return error; @@ -1333,7 +1345,7 @@ static int ea_dealloc_indirect(struct gfs2_inode *ip) if (error) goto out_gunlock; - gfs2_trans_add_meta(ip->i_gl, indbh); + gfs2_trans_add_meta(gl, indbh); eablk = (__be64 *)(indbh->b_data + sizeof(struct gfs2_meta_header)); bstart = 0; @@ -1367,7 +1379,7 @@ static int ea_dealloc_indirect(struct gfs2_inode *ip) error = gfs2_meta_inode_buffer(ip, &dibh); if (!error) { - gfs2_trans_add_meta(ip->i_gl, dibh); + gfs2_trans_add_meta(gl, dibh); gfs2_dinode_out(ip, dibh->b_data); brelse(dibh); } @@ -1385,6 +1397,7 @@ out: static int ea_dealloc_block(struct gfs2_inode *ip, bool initialized) { + struct gfs2_glock *gl = gfs2_inode_glock(&ip->i_inode); struct gfs2_sbd *sdp = GFS2_SB(&ip->i_inode); struct gfs2_rgrpd *rgd; struct buffer_head *dibh; @@ -1419,7 +1432,7 @@ static int ea_dealloc_block(struct gfs2_inode *ip, bool initialized) if (initialized) { error = gfs2_meta_inode_buffer(ip, &dibh); if (!error) { - gfs2_trans_add_meta(ip->i_gl, dibh); + gfs2_trans_add_meta(gl, dibh); gfs2_dinode_out(ip, dibh->b_data); brelse(dibh); } diff --git a/fs/hfs/attr.c b/fs/hfs/attr.c index f8395cdd1adf..6d737a085461 100644 --- a/fs/hfs/attr.c +++ b/fs/hfs/attr.c @@ -121,7 +121,7 @@ static int hfs_xattr_get(const struct xattr_handler *handler, } static int hfs_xattr_set(const struct xattr_handler *handler, - struct mnt_idmap *idmap, + const struct mnt_idmap *idmap, struct dentry *unused, struct inode *inode, const char *name, const void *value, size_t size, int flags) diff --git a/fs/hfs/dir.c b/fs/hfs/dir.c index e1f1fb351464..f6b97da19788 100644 --- a/fs/hfs/dir.c +++ b/fs/hfs/dir.c @@ -183,7 +183,7 @@ static int hfs_dir_release(struct inode *inode, struct file *file) * a directory and return a corresponding inode, given the inode for * the directory and the name (and its length) of the new file. */ -static int hfs_create(struct mnt_idmap *idmap, struct inode *dir, +static int hfs_create(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *dentry, umode_t mode) { struct inode *inode; @@ -213,7 +213,7 @@ static int hfs_create(struct mnt_idmap *idmap, struct inode *dir, * in a directory, given the inode for the parent directory and the * name (and its length) of the new directory. */ -static struct dentry *hfs_mkdir(struct mnt_idmap *idmap, struct inode *dir, +static struct dentry *hfs_mkdir(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *dentry, umode_t mode) { struct inode *inode; @@ -281,7 +281,7 @@ static int hfs_remove(struct inode *dir, struct dentry *dentry) * new file/directory. * XXX: how do you handle must_be dir? */ -static int hfs_rename(struct mnt_idmap *idmap, struct inode *old_dir, +static int hfs_rename(const struct mnt_idmap *idmap, struct inode *old_dir, struct dentry *old_dentry, struct inode *new_dir, struct dentry *new_dentry, unsigned int flags) { diff --git a/fs/hfs/hfs_fs.h b/fs/hfs/hfs_fs.h index e250f87a5e33..fdfa5d303d5e 100644 --- a/fs/hfs/hfs_fs.h +++ b/fs/hfs/hfs_fs.h @@ -212,7 +212,7 @@ extern struct inode *hfs_new_inode(struct inode *dir, const struct qstr *name, extern void hfs_inode_write_fork(struct inode *inode, struct hfs_extent *ext, __be32 *log_size, __be32 *phys_size); extern int hfs_write_inode(struct inode *inode, struct writeback_control *wbc); -extern int hfs_inode_setattr(struct mnt_idmap *idmap, struct dentry *dentry, +extern int hfs_inode_setattr(const struct mnt_idmap *idmap, struct dentry *dentry, struct iattr *attr); extern void hfs_inode_read_fork(struct inode *inode, struct hfs_extent *ext, __be32 __log_size, __be32 phys_size, diff --git a/fs/hfs/inode.c b/fs/hfs/inode.c index 2aef3c36a150..c81314f668ac 100644 --- a/fs/hfs/inode.c +++ b/fs/hfs/inode.c @@ -643,7 +643,7 @@ static int hfs_file_release(struct inode *inode, struct file *file) return 0; } -int hfs_inode_setattr(struct mnt_idmap *idmap, struct dentry *dentry, +int hfs_inode_setattr(const struct mnt_idmap *idmap, struct dentry *dentry, struct iattr *attr) { struct inode *inode = d_inode(dentry); diff --git a/fs/hfsplus/dir.c b/fs/hfsplus/dir.c index 51fcba2e6d40..b3a1193a491f 100644 --- a/fs/hfsplus/dir.c +++ b/fs/hfsplus/dir.c @@ -460,7 +460,7 @@ out: return res; } -static int hfsplus_symlink(struct mnt_idmap *idmap, struct inode *dir, +static int hfsplus_symlink(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *dentry, const char *symname) { struct hfsplus_sb_info *sbi = HFSPLUS_SB(dir->i_sb); @@ -511,7 +511,7 @@ out: return res; } -static int hfsplus_mknod(struct mnt_idmap *idmap, struct inode *dir, +static int hfsplus_mknod(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *dentry, umode_t mode, dev_t rdev) { struct hfsplus_sb_info *sbi = HFSPLUS_SB(dir->i_sb); @@ -561,19 +561,19 @@ out: return res; } -static int hfsplus_create(struct mnt_idmap *idmap, struct inode *dir, +static int hfsplus_create(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *dentry, umode_t mode) { return hfsplus_mknod(&nop_mnt_idmap, dir, dentry, mode, 0); } -static struct dentry *hfsplus_mkdir(struct mnt_idmap *idmap, struct inode *dir, +static struct dentry *hfsplus_mkdir(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *dentry, umode_t mode) { return ERR_PTR(hfsplus_mknod(&nop_mnt_idmap, dir, dentry, mode, 0)); } -static int hfsplus_rename(struct mnt_idmap *idmap, +static int hfsplus_rename(const struct mnt_idmap *idmap, struct inode *old_dir, struct dentry *old_dentry, struct inode *new_dir, struct dentry *new_dentry, unsigned int flags) diff --git a/fs/hfsplus/hfsplus_fs.h b/fs/hfsplus/hfsplus_fs.h index 1e5b58e6a13f..d55cb16e3897 100644 --- a/fs/hfsplus/hfsplus_fs.h +++ b/fs/hfsplus/hfsplus_fs.h @@ -459,13 +459,13 @@ void hfsplus_inode_write_fork(struct inode *inode, struct hfsplus_fork_raw *fork); int hfsplus_cat_read_inode(struct inode *inode, struct hfs_find_data *fd); int hfsplus_cat_write_inode(struct inode *inode); -int hfsplus_getattr(struct mnt_idmap *idmap, const struct path *path, +int hfsplus_getattr(const struct mnt_idmap *idmap, const struct path *path, struct kstat *stat, u32 request_mask, unsigned int query_flags); int hfsplus_file_fsync(struct file *file, loff_t start, loff_t end, int datasync); int hfsplus_fileattr_get(struct dentry *dentry, struct file_kattr *fa); -int hfsplus_fileattr_set(struct mnt_idmap *idmap, +int hfsplus_fileattr_set(const struct mnt_idmap *idmap, struct dentry *dentry, struct file_kattr *fa); /* ioctl.c */ diff --git a/fs/hfsplus/inode.c b/fs/hfsplus/inode.c index 2ce6de574fa6..aed0866499b2 100644 --- a/fs/hfsplus/inode.c +++ b/fs/hfsplus/inode.c @@ -305,7 +305,7 @@ static int hfsplus_file_release(struct inode *inode, struct file *file) return 0; } -static int hfsplus_setattr(struct mnt_idmap *idmap, +static int hfsplus_setattr(const struct mnt_idmap *idmap, struct dentry *dentry, struct iattr *attr) { struct inode *inode = d_inode(dentry); @@ -335,7 +335,7 @@ static int hfsplus_setattr(struct mnt_idmap *idmap, return 0; } -int hfsplus_getattr(struct mnt_idmap *idmap, const struct path *path, +int hfsplus_getattr(const struct mnt_idmap *idmap, const struct path *path, struct kstat *stat, u32 request_mask, unsigned int query_flags) { @@ -797,7 +797,7 @@ int hfsplus_fileattr_get(struct dentry *dentry, struct file_kattr *fa) return 0; } -int hfsplus_fileattr_set(struct mnt_idmap *idmap, +int hfsplus_fileattr_set(const struct mnt_idmap *idmap, struct dentry *dentry, struct file_kattr *fa) { struct inode *inode = d_inode(dentry); diff --git a/fs/hfsplus/xattr.c b/fs/hfsplus/xattr.c index 21a1c196c71f..71364e093fa7 100644 --- a/fs/hfsplus/xattr.c +++ b/fs/hfsplus/xattr.c @@ -1008,7 +1008,7 @@ static int hfsplus_osx_getxattr(const struct xattr_handler *handler, } static int hfsplus_osx_setxattr(const struct xattr_handler *handler, - struct mnt_idmap *idmap, + const struct mnt_idmap *idmap, struct dentry *unused, struct inode *inode, const char *name, const void *buffer, size_t size, int flags) diff --git a/fs/hfsplus/xattr_security.c b/fs/hfsplus/xattr_security.c index 90f68ec119cd..1969919c12cb 100644 --- a/fs/hfsplus/xattr_security.c +++ b/fs/hfsplus/xattr_security.c @@ -23,7 +23,7 @@ static int hfsplus_security_getxattr(const struct xattr_handler *handler, } static int hfsplus_security_setxattr(const struct xattr_handler *handler, - struct mnt_idmap *idmap, + const struct mnt_idmap *idmap, struct dentry *unused, struct inode *inode, const char *name, const void *buffer, size_t size, int flags) diff --git a/fs/hfsplus/xattr_trusted.c b/fs/hfsplus/xattr_trusted.c index fdbaebc1c49a..c140a95ab3f0 100644 --- a/fs/hfsplus/xattr_trusted.c +++ b/fs/hfsplus/xattr_trusted.c @@ -22,7 +22,7 @@ static int hfsplus_trusted_getxattr(const struct xattr_handler *handler, } static int hfsplus_trusted_setxattr(const struct xattr_handler *handler, - struct mnt_idmap *idmap, + const struct mnt_idmap *idmap, struct dentry *unused, struct inode *inode, const char *name, const void *buffer, size_t size, int flags) diff --git a/fs/hfsplus/xattr_user.c b/fs/hfsplus/xattr_user.c index 6464b6c3d58d..7e5da15f9937 100644 --- a/fs/hfsplus/xattr_user.c +++ b/fs/hfsplus/xattr_user.c @@ -22,7 +22,7 @@ static int hfsplus_user_getxattr(const struct xattr_handler *handler, } static int hfsplus_user_setxattr(const struct xattr_handler *handler, - struct mnt_idmap *idmap, + const struct mnt_idmap *idmap, struct dentry *unused, struct inode *inode, const char *name, const void *buffer, size_t size, int flags) diff --git a/fs/hostfs/hostfs_kern.c b/fs/hostfs/hostfs_kern.c index 7add056d47d8..613146e76dec 100644 --- a/fs/hostfs/hostfs_kern.c +++ b/fs/hostfs/hostfs_kern.c @@ -592,7 +592,7 @@ static struct inode *hostfs_iget(struct super_block *sb, char *name) return inode; } -static int hostfs_create(struct mnt_idmap *idmap, struct inode *dir, +static int hostfs_create(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *dentry, umode_t mode) { struct inode *inode; @@ -673,7 +673,7 @@ static int hostfs_unlink(struct inode *ino, struct dentry *dentry) return err; } -static int hostfs_symlink(struct mnt_idmap *idmap, struct inode *ino, +static int hostfs_symlink(const struct mnt_idmap *idmap, struct inode *ino, struct dentry *dentry, const char *to) { char *file; @@ -686,7 +686,7 @@ static int hostfs_symlink(struct mnt_idmap *idmap, struct inode *ino, return err; } -static struct dentry *hostfs_mkdir(struct mnt_idmap *idmap, struct inode *ino, +static struct dentry *hostfs_mkdir(const struct mnt_idmap *idmap, struct inode *ino, struct dentry *dentry, umode_t mode) { struct inode *inode; @@ -719,7 +719,7 @@ static int hostfs_rmdir(struct inode *ino, struct dentry *dentry) return err; } -static int hostfs_mknod(struct mnt_idmap *idmap, struct inode *dir, +static int hostfs_mknod(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *dentry, umode_t mode, dev_t dev) { struct inode *inode; @@ -745,7 +745,7 @@ static int hostfs_mknod(struct mnt_idmap *idmap, struct inode *dir, return 0; } -static int hostfs_rename2(struct mnt_idmap *idmap, +static int hostfs_rename2(const struct mnt_idmap *idmap, struct inode *old_dir, struct dentry *old_dentry, struct inode *new_dir, struct dentry *new_dentry, unsigned int flags) @@ -774,7 +774,7 @@ static int hostfs_rename2(struct mnt_idmap *idmap, return err; } -static int hostfs_permission(struct mnt_idmap *idmap, +static int hostfs_permission(const struct mnt_idmap *idmap, struct inode *ino, int desired) { char *name; @@ -801,7 +801,7 @@ static int hostfs_permission(struct mnt_idmap *idmap, return err; } -static int hostfs_setattr(struct mnt_idmap *idmap, +static int hostfs_setattr(const struct mnt_idmap *idmap, struct dentry *dentry, struct iattr *attr) { struct inode *inode = d_inode(dentry); diff --git a/fs/hpfs/hpfs_fn.h b/fs/hpfs/hpfs_fn.h index 237c1c23e855..a398dd8bdf30 100644 --- a/fs/hpfs/hpfs_fn.h +++ b/fs/hpfs/hpfs_fn.h @@ -280,7 +280,7 @@ void hpfs_init_inode(struct inode *); void hpfs_read_inode(struct inode *); void hpfs_write_inode(struct inode *); void hpfs_write_inode_nolock(struct inode *); -int hpfs_setattr(struct mnt_idmap *, struct dentry *, struct iattr *); +int hpfs_setattr(const struct mnt_idmap *, struct dentry *, struct iattr *); void hpfs_write_if_changed(struct inode *); void hpfs_evict_inode(struct inode *); diff --git a/fs/hpfs/inode.c b/fs/hpfs/inode.c index 1b4fcf760aad..396773d0b669 100644 --- a/fs/hpfs/inode.c +++ b/fs/hpfs/inode.c @@ -257,7 +257,7 @@ void hpfs_write_inode_nolock(struct inode *i) brelse(bh); } -int hpfs_setattr(struct mnt_idmap *idmap, struct dentry *dentry, +int hpfs_setattr(const struct mnt_idmap *idmap, struct dentry *dentry, struct iattr *attr) { struct inode *inode = d_inode(dentry); diff --git a/fs/hpfs/namei.c b/fs/hpfs/namei.c index 9446f4038874..ac9b5e3e83fa 100644 --- a/fs/hpfs/namei.c +++ b/fs/hpfs/namei.c @@ -19,7 +19,7 @@ static void hpfs_update_directory_times(struct inode *dir) hpfs_write_inode_nolock(dir); } -static struct dentry *hpfs_mkdir(struct mnt_idmap *idmap, struct inode *dir, +static struct dentry *hpfs_mkdir(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *dentry, umode_t mode) { const unsigned char *name = dentry->d_name.name; @@ -128,7 +128,7 @@ bail: return ERR_PTR(err); } -static int hpfs_create(struct mnt_idmap *idmap, struct inode *dir, +static int hpfs_create(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *dentry, umode_t mode) { const unsigned char *name = dentry->d_name.name; @@ -215,7 +215,7 @@ bail: return err; } -static int hpfs_mknod(struct mnt_idmap *idmap, struct inode *dir, +static int hpfs_mknod(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *dentry, umode_t mode, dev_t rdev) { const unsigned char *name = dentry->d_name.name; @@ -289,7 +289,7 @@ bail: return err; } -static int hpfs_symlink(struct mnt_idmap *idmap, struct inode *dir, +static int hpfs_symlink(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *dentry, const char *symlink) { const unsigned char *name = dentry->d_name.name; @@ -500,7 +500,7 @@ const struct address_space_operations hpfs_symlink_aops = { .read_folio = hpfs_symlink_read_folio }; -static int hpfs_rename(struct mnt_idmap *idmap, struct inode *old_dir, +static int hpfs_rename(const struct mnt_idmap *idmap, struct inode *old_dir, struct dentry *old_dentry, struct inode *new_dir, struct dentry *new_dentry, unsigned int flags) { diff --git a/fs/hugetlbfs/inode.c b/fs/hugetlbfs/inode.c index ab1e4dec3f77..01c4d6f44ebe 100644 --- a/fs/hugetlbfs/inode.c +++ b/fs/hugetlbfs/inode.c @@ -827,7 +827,7 @@ out_nolock: return error; } -static int hugetlbfs_setattr(struct mnt_idmap *idmap, +static int hugetlbfs_setattr(const struct mnt_idmap *idmap, struct dentry *dentry, struct iattr *attr) { struct inode *inode = d_inode(dentry); @@ -891,7 +891,7 @@ static struct inode *hugetlbfs_get_root(struct super_block *sb, static struct lock_class_key hugetlbfs_i_mmap_rwsem_key; static struct inode *hugetlbfs_get_inode(struct super_block *sb, - struct mnt_idmap *idmap, + const struct mnt_idmap *idmap, struct inode *dir, umode_t mode, dev_t dev) { @@ -954,7 +954,7 @@ static struct inode *hugetlbfs_get_inode(struct super_block *sb, /* * File creation. Allocate an inode, and we're done.. */ -static int hugetlbfs_mknod(struct mnt_idmap *idmap, struct inode *dir, +static int hugetlbfs_mknod(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *dentry, umode_t mode, dev_t dev) { struct inode *inode; @@ -967,7 +967,7 @@ static int hugetlbfs_mknod(struct mnt_idmap *idmap, struct inode *dir, return 0; } -static struct dentry *hugetlbfs_mkdir(struct mnt_idmap *idmap, struct inode *dir, +static struct dentry *hugetlbfs_mkdir(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *dentry, umode_t mode) { int retval = hugetlbfs_mknod(idmap, dir, dentry, @@ -977,14 +977,14 @@ static struct dentry *hugetlbfs_mkdir(struct mnt_idmap *idmap, struct inode *dir return ERR_PTR(retval); } -static int hugetlbfs_create(struct mnt_idmap *idmap, +static int hugetlbfs_create(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *dentry, umode_t mode) { return hugetlbfs_mknod(idmap, dir, dentry, mode | S_IFREG, 0); } -static int hugetlbfs_tmpfile(struct mnt_idmap *idmap, +static int hugetlbfs_tmpfile(const struct mnt_idmap *idmap, struct inode *dir, struct file *file, umode_t mode) { @@ -998,7 +998,7 @@ static int hugetlbfs_tmpfile(struct mnt_idmap *idmap, return finish_open_simple(file, 0); } -static int hugetlbfs_symlink(struct mnt_idmap *idmap, +static int hugetlbfs_symlink(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *dentry, const char *symname) { diff --git a/fs/inode.c b/fs/inode.c index a9d37be390a1..1cb6293b237a 100644 --- a/fs/inode.c +++ b/fs/inode.c @@ -1770,7 +1770,7 @@ EXPORT_SYMBOL(ilookup); * function must never block --- find_inode() can block in * __wait_on_freeing_inode() --- or when the caller can not increment * the reference count because the resulting iput() might cause an - * inode eviction. The tradeoff is that the @match funtion must be + * inode eviction. The tradeoff is that the @match function must be * very carefully implemented. */ struct inode *find_inode_nowait(struct super_block *sb, @@ -2336,7 +2336,7 @@ EXPORT_SYMBOL(touch_atime); * response to write or truncate. Return 0 if nothing has to be changed. * Negative value on error (change should be denied). */ -int dentry_needs_remove_privs(struct mnt_idmap *idmap, +int dentry_needs_remove_privs(const struct mnt_idmap *idmap, struct dentry *dentry) { struct inode *inode = d_inode(dentry); @@ -2355,7 +2355,7 @@ int dentry_needs_remove_privs(struct mnt_idmap *idmap, return mask; } -static int __remove_privs(struct mnt_idmap *idmap, +static int __remove_privs(const struct mnt_idmap *idmap, struct dentry *dentry, int kill) { struct iattr newattrs; @@ -2715,7 +2715,7 @@ EXPORT_SYMBOL(init_special_inode); * and initializing i_uid and i_gid. On non-idmapped mounts or if permission * checking is to be performed on the raw inode simply pass @nop_mnt_idmap. */ -void inode_init_owner(struct mnt_idmap *idmap, struct inode *inode, +void inode_init_owner(const struct mnt_idmap *idmap, struct inode *inode, const struct inode *dir, umode_t mode) { inode_fsuid_set(inode, idmap); @@ -2745,7 +2745,7 @@ EXPORT_SYMBOL(inode_init_owner); * On non-idmapped mounts or if permission checking is to be performed on the * raw inode simply pass @nop_mnt_idmap. */ -bool inode_owner_or_capable(struct mnt_idmap *idmap, +bool inode_owner_or_capable(const struct mnt_idmap *idmap, const struct inode *inode) { vfsuid_t vfsuid; @@ -3032,7 +3032,7 @@ EXPORT_SYMBOL(inode_set_ctime_deleg); * * Return: true if the caller is sufficiently privileged, false if not. */ -bool in_group_or_capable(struct mnt_idmap *idmap, +bool in_group_or_capable(const struct mnt_idmap *idmap, const struct inode *inode, vfsgid_t vfsgid) { if (vfsgid_in_group_p(vfsgid)) @@ -3057,7 +3057,7 @@ EXPORT_SYMBOL(in_group_or_capable); * * Return: the new mode to use for the file */ -umode_t mode_strip_sgid(struct mnt_idmap *idmap, +umode_t mode_strip_sgid(const struct mnt_idmap *idmap, const struct inode *dir, umode_t mode) { if ((mode & (S_ISGID | S_IXGRP)) != (S_ISGID | S_IXGRP)) diff --git a/fs/internal.h b/fs/internal.h index c658c8a5ebd5..e833c7e6e14f 100644 --- a/fs/internal.h +++ b/fs/internal.h @@ -55,7 +55,7 @@ extern int filename_lookup(int dfd, struct filename *name, unsigned flags, struct path *path, const struct path *root); int filename_rmdir(int dfd, struct filename *name); int filename_unlinkat(int dfd, struct filename *name); -int may_linkat(struct mnt_idmap *idmap, const struct path *link); +int may_linkat(const struct mnt_idmap *idmap, const struct path *link); int filename_renameat2(int olddfd, struct filename *oldname, int newdfd, struct filename *newname, unsigned int flags); int filename_mkdirat(int dfd, struct filename *name, umode_t mode); @@ -63,7 +63,7 @@ int filename_mknodat(int dfd, struct filename *name, umode_t mode, unsigned int int filename_symlinkat(struct filename *from, int newdfd, struct filename *to); int filename_linkat(int olddfd, struct filename *old, int newdfd, struct filename *new, int flags); -int vfs_tmpfile(struct mnt_idmap *idmap, +int vfs_tmpfile(const struct mnt_idmap *idmap, const struct path *parentpath, struct file *file, umode_t mode); struct dentry *d_hash_and_lookup(struct dentry *, struct qstr *); @@ -198,6 +198,7 @@ extern struct file *do_file_open_root(const struct path *, extern struct open_how build_open_how(int flags, umode_t mode); extern int build_open_flags(const struct open_how *how, struct open_flags *op); struct file *file_close_fd_locked(struct files_struct *files, unsigned fd); +int filp_close_sync(struct file *filp, fl_owner_t id); int do_ftruncate(struct file *file, loff_t length, unsigned int flags); int chmod_common(const struct path *path, umode_t mode); @@ -205,13 +206,14 @@ int do_fchownat(int dfd, const char __user *filename, uid_t user, gid_t group, int flag); int chown_common(const struct path *path, uid_t user, gid_t group); extern int vfs_open(const struct path *, struct file *); +int vfs_open_consume(struct path *, struct file *); /* * inode.c */ extern long prune_icache_sb(struct super_block *sb, struct shrink_control *sc); -int dentry_needs_remove_privs(struct mnt_idmap *, struct dentry *dentry); -bool in_group_or_capable(struct mnt_idmap *idmap, +int dentry_needs_remove_privs(const struct mnt_idmap *, struct dentry *dentry); +bool in_group_or_capable(const struct mnt_idmap *idmap, const struct inode *inode, vfsgid_t vfsgid); /* @@ -299,21 +301,21 @@ int filename_setxattr(int dfd, struct filename *filename, int setxattr_copy(const char __user *name, struct kernel_xattr_ctx *ctx); int import_xattr_name(struct xattr_name *kname, const char __user *name); -int may_write_xattr(struct mnt_idmap *idmap, struct inode *inode); +int may_write_xattr(const struct mnt_idmap *idmap, struct inode *inode); #ifdef CONFIG_FS_POSIX_ACL -int do_set_acl(struct mnt_idmap *idmap, struct dentry *dentry, +int do_set_acl(const struct mnt_idmap *idmap, struct dentry *dentry, const char *acl_name, const void *kvalue, size_t size); -ssize_t do_get_acl(struct mnt_idmap *idmap, struct dentry *dentry, +ssize_t do_get_acl(const struct mnt_idmap *idmap, struct dentry *dentry, const char *acl_name, void *kvalue, size_t size); #else -static inline int do_set_acl(struct mnt_idmap *idmap, +static inline int do_set_acl(const struct mnt_idmap *idmap, struct dentry *dentry, const char *acl_name, const void *kvalue, size_t size) { return -EOPNOTSUPP; } -static inline ssize_t do_get_acl(struct mnt_idmap *idmap, +static inline ssize_t do_get_acl(const struct mnt_idmap *idmap, struct dentry *dentry, const char *acl_name, void *kvalue, size_t size) { @@ -327,8 +329,8 @@ ssize_t __kernel_write_iter(struct file *file, struct iov_iter *from, loff_t *po * fs/attr.c */ struct mnt_idmap *alloc_mnt_idmap(struct user_namespace *mnt_userns); -struct mnt_idmap *mnt_idmap_get(struct mnt_idmap *idmap); -void mnt_idmap_put(struct mnt_idmap *idmap); +const struct mnt_idmap *mnt_idmap_get(const struct mnt_idmap *idmap); +void mnt_idmap_put(const struct mnt_idmap *idmap); struct stashed_operations { struct dentry *(*stash_dentry)(struct dentry **stashed, struct dentry *dentry); @@ -354,12 +356,12 @@ static inline bool path_mounted(const struct path *path) } void file_f_owner_release(struct file *file); bool file_seek_cur_needs_f_lock(struct file *file); -int statmount_mnt_idmap(struct mnt_idmap *idmap, struct seq_file *seq, bool uid_map); +int statmount_mnt_idmap(const struct mnt_idmap *idmap, struct seq_file *seq, bool uid_map); struct dentry *find_next_child(struct dentry *parent, struct dentry *prev); -int anon_inode_getattr(struct mnt_idmap *idmap, const struct path *path, +int anon_inode_getattr(const struct mnt_idmap *idmap, const struct path *path, struct kstat *stat, u32 request_mask, unsigned int query_flags); -int anon_inode_setattr(struct mnt_idmap *idmap, struct dentry *dentry, +int anon_inode_setattr(const struct mnt_idmap *idmap, struct dentry *dentry, struct iattr *attr); void pidfs_get_root(struct path *path); void nsfs_get_root(struct path *path); diff --git a/fs/iomap/bio.c b/fs/iomap/bio.c index 48100c614431..d46c2f8ea18c 100644 --- a/fs/iomap/bio.c +++ b/fs/iomap/bio.c @@ -169,6 +169,7 @@ int iomap_bio_read_folio_range_sync(const struct iomap_iter *iter, { const struct iomap *srcmap = iomap_iter_srcmap(iter); sector_t sector = iomap_sector(srcmap, pos); + struct bvec_iter saved_iter; struct bio_vec bvec; struct bio bio; int error; @@ -178,10 +179,11 @@ int iomap_bio_read_folio_range_sync(const struct iomap_iter *iter, bio_add_folio_nofail(&bio, folio, len, offset_in_folio(folio, pos)); if (srcmap->flags & IOMAP_F_INTEGRITY) fs_bio_integrity_alloc(&bio); + saved_iter = bio.bi_iter; error = submit_bio_wait(&bio); if (bio_integrity(&bio)) { if (!error) - error = fs_bio_integrity_verify(&bio, sector, len); + error = fs_bio_integrity_verify(&bio, &saved_iter); fs_bio_integrity_free(&bio); } bio_uninit(&bio); diff --git a/fs/iomap/buffered-io.c b/fs/iomap/buffered-io.c index 6306ca747f3b..1e0de37b2411 100644 --- a/fs/iomap/buffered-io.c +++ b/fs/iomap/buffered-io.c @@ -143,8 +143,8 @@ static unsigned ifs_next_clean_block(struct folio *folio, blks + start_blk) - blks; } -static unsigned ifs_find_dirty_range(struct folio *folio, - struct iomap_folio_state *ifs, u64 *range_start, u64 range_end) +static unsigned ifs_find_dirty_range(struct folio *folio, u64 *range_start, + u64 range_end) { struct inode *inode = folio->mapping->host; unsigned start_blk = @@ -176,7 +176,7 @@ static unsigned iomap_find_dirty_range(struct folio *folio, u64 *range_start, return 0; if (ifs) - return ifs_find_dirty_range(folio, ifs, range_start, range_end); + return ifs_find_dirty_range(folio, range_start, range_end); return range_end - *range_start; } @@ -1708,7 +1708,7 @@ static int iomap_zero_iter(struct iomap_iter *iter, bool *did_zero, * @iomap_flags: Flags to set on the associated iomap to track the batch. * * Returns the folio count directly. Also returns the associated control flag if - * the the batch lookup is performed and the expected offset of a subsequent + * the batch lookup is performed and the expected offset of a subsequent * lookup via out params. The caller is responsible to set the flag on the * associated iomap. */ diff --git a/fs/iomap/direct-io.c b/fs/iomap/direct-io.c index 8b4039d16ce8..a431ceda9ebd 100644 --- a/fs/iomap/direct-io.c +++ b/fs/iomap/direct-io.c @@ -76,10 +76,19 @@ static void iomap_dio_submit_bio(const struct iomap_iter *iter, if (dio->dops && dio->dops->submit_io) { dio->dops->submit_io(iter, bio, pos); - } else { - WARN_ON_ONCE(iter->iomap.flags & IOMAP_F_ANON_WRITE); - blk_crypto_submit_bio(bio); + return; + } + + WARN_ON_ONCE(iter->iomap.flags & IOMAP_F_ANON_WRITE); + + if (iter->iomap.flags & IOMAP_F_INTEGRITY) { + if (dio->flags & IOMAP_DIO_WRITE) + fs_bio_integrity_generate(bio); + else + fs_bio_integrity_alloc(bio); } + + blk_crypto_submit_bio(bio); } static inline enum fserror_type iomap_dio_err_type(const struct iomap_dio *dio) @@ -246,8 +255,7 @@ static void __iomap_dio_bio_end_io(struct bio *bio, bool inline_completion) fs_bio_integrity_free(bio); if (dio->flags & IOMAP_DIO_BOUNCE) { - bio_iov_iter_unbounce(bio, !!dio->error, - dio->flags & IOMAP_DIO_USER_BACKED); + bio_free_folios(bio); bio_put(bio); } else if (dio->flags & IOMAP_DIO_USER_BACKED) { bio_check_pages_dirty(bio); @@ -336,6 +344,7 @@ static ssize_t iomap_dio_bio_iter_one(struct iomap_iter *iter, struct iomap_dio *dio, loff_t pos, unsigned int alignment, blk_opf_t op) { + unsigned int maxsize = iomap_max_bio_size(&iter->iomap); unsigned int nr_vecs; struct bio *bio; ssize_t ret; @@ -353,14 +362,12 @@ static ssize_t iomap_dio_bio_iter_one(struct iomap_iter *iter, bio->bi_private = dio; bio->bi_end_io = iomap_dio_bio_end_io; - if (dio->flags & IOMAP_DIO_BOUNCE) - ret = bio_iov_iter_bounce(bio, dio->submit.iter, - iomap_max_bio_size(&iter->iomap), alignment); + ret = bio_iov_iter_bounce_write(bio, dio->submit.iter, maxsize, + alignment); else - ret = bio_iov_iter_get_pages(bio, dio->submit.iter, - bdev_dma_alignment(bio->bi_bdev), - alignment - 1); + ret = bio_iov_iter_get_pages(bio, dio->submit.iter, maxsize, + bdev_dma_alignment(bio->bi_bdev), alignment - 1); if (unlikely(ret)) goto out_put_bio; ret = bio->bi_iter.bi_size; @@ -374,13 +381,6 @@ static ssize_t iomap_dio_bio_iter_one(struct iomap_iter *iter, goto out_bio_release_pages; } - if (iter->iomap.flags & IOMAP_F_INTEGRITY) { - if (dio->flags & IOMAP_DIO_WRITE) - fs_bio_integrity_generate(bio); - else - fs_bio_integrity_alloc(bio); - } - if (dio->flags & IOMAP_DIO_WRITE) task_io_account_write(ret); else if ((dio->flags & IOMAP_DIO_USER_BACKED) && @@ -397,7 +397,7 @@ static ssize_t iomap_dio_bio_iter_one(struct iomap_iter *iter, out_bio_release_pages: if (dio->flags & IOMAP_DIO_BOUNCE) - bio_iov_iter_unbounce(bio, true, false); + bio_free_folios(bio); else bio_release_pages(bio, false); out_put_bio: @@ -505,7 +505,7 @@ static int iomap_dio_bio_iter(struct iomap_iter *iter, struct iomap_dio *dio) * We can only do inline completion for pure overwrites that * don't require additional I/O at completion time. * - * This rules out writes that need zeroing or metdata updates to + * This rules out writes that need zeroing or metadata updates to * convert unwritten or shared extents. * * Writes that extend i_size are also not supported, but this is @@ -1034,9 +1034,9 @@ ssize_t __iomap_dio_read_simple(struct kiocb *iocb, struct iov_iter *iter, bio->bi_iter.bi_sector = iomap_sector(&iomi->iomap, iomi->pos); bio->bi_ioprio = iocb->ki_ioprio; - ret = bio_iov_iter_get_pages(bio, iter, - bdev_dma_alignment(bio->bi_bdev), - alignment - 1); + ret = bio_iov_iter_get_pages(bio, iter, BIO_MAX_SIZE, + bdev_dma_alignment(bio->bi_bdev), + alignment - 1); if (unlikely(ret)) goto out_bio_put; diff --git a/fs/iomap/ioend.c b/fs/iomap/ioend.c index 7bbbb417f915..bbebecc31670 100644 --- a/fs/iomap/ioend.c +++ b/fs/iomap/ioend.c @@ -25,6 +25,7 @@ struct iomap_ioend *iomap_init_ioend(struct inode *inode, ioend->io_parent = NULL; INIT_LIST_HEAD(&ioend->io_list); ioend->io_flags = ioend_flags; + ioend->io_bvec_offset = bio->bi_iter.bi_offset; ioend->io_inode = inode; ioend->io_offset = file_offset; ioend->io_size = bio->bi_iter.bi_size; @@ -149,7 +150,7 @@ int iomap_ioend_writeback_submit(struct iomap_writepage_ctx *wpc, int error) return error; } - if (wpc->iomap.flags & IOMAP_F_INTEGRITY) + if (ioend->io_flags & IOMAP_IOEND_INTEGRITY) fs_bio_integrity_generate(&ioend->io_bio); submit_bio(&ioend->io_bio); return 0; @@ -215,7 +216,7 @@ ssize_t iomap_add_to_ioend(struct iomap_writepage_ctx *wpc, struct folio *folio, { struct iomap_ioend *ioend = wpc->wb_ctx; size_t poff = offset_in_folio(folio, pos); - unsigned int ioend_flags = 0; + unsigned int ioend_flags = iomap_ioend_flags(&wpc->iomap); unsigned int map_len = min_t(u64, dirty_len, wpc->iomap.offset + wpc->iomap.length - pos); int error; @@ -225,20 +226,16 @@ ssize_t iomap_add_to_ioend(struct iomap_writepage_ctx *wpc, struct folio *folio, WARN_ON_ONCE(!folio->private && map_len < dirty_len); switch (wpc->iomap.type) { + case IOMAP_HOLE: + return map_len; case IOMAP_UNWRITTEN: - ioend_flags |= IOMAP_IOEND_UNWRITTEN; - break; case IOMAP_MAPPED: break; - case IOMAP_HOLE: - return map_len; default: WARN_ON_ONCE(1); return -EIO; } - if (wpc->iomap.flags & IOMAP_F_SHARED) - ioend_flags |= IOMAP_IOEND_SHARED; if (pos == wpc->iomap.offset && (wpc->iomap.flags & IOMAP_F_BOUNDARY)) ioend_flags |= IOMAP_IOEND_BOUNDARY; @@ -312,6 +309,16 @@ new_ioend: } EXPORT_SYMBOL_GPL(iomap_add_to_ioend); +#ifdef CONFIG_BLK_DEV_INTEGRITY +int iomap_ioend_integrity_verify(struct iomap_ioend *ioend) +{ + struct bvec_iter data_iter = BVEC_ITER_IOEND(ioend); + + return fs_bio_integrity_verify(&ioend->io_bio, &data_iter); +} +EXPORT_SYMBOL_GPL(iomap_ioend_integrity_verify); +#endif /* CONFIG_BLK_DEV_INTEGRITY */ + static u32 iomap_finish_ioend(struct iomap_ioend *ioend, int error) { if (ioend->io_parent) { @@ -327,13 +334,6 @@ static u32 iomap_finish_ioend(struct iomap_ioend *ioend, int error) if (!atomic_dec_and_test(&ioend->io_remaining)) return 0; - if (!ioend->io_error && - bio_integrity(&ioend->io_bio) && - bio_op(&ioend->io_bio) == REQ_OP_READ) { - ioend->io_error = fs_bio_integrity_verify(&ioend->io_bio, - ioend->io_sector, ioend->io_size); - } - if (ioend->io_flags & IOMAP_IOEND_DIRECT) return iomap_finish_ioend_direct(ioend); if (bio_op(&ioend->io_bio) == REQ_OP_READ) @@ -512,6 +512,96 @@ struct iomap_ioend *iomap_split_ioend(struct iomap_ioend *ioend, } EXPORT_SYMBOL_GPL(iomap_split_ioend); +void iomap_bounce_read(struct iomap_ioend *orig_ioend, unsigned int minsize, + void (*submit_ioend)(struct iomap_ioend *ioend)) +{ + struct inode *inode = orig_ioend->io_inode; + struct bio *orig_bio = &orig_ioend->io_bio; + loff_t file_offset = orig_ioend->io_offset; + sector_t sector = orig_ioend->io_sector; + size_t total_len = round_up(orig_ioend->io_size, minsize); + + WARN_ON_ONCE(!(orig_ioend->io_flags & IOMAP_IOEND_DIRECT)); + + /* We can't poll a bio that is not passed on to hardware */ + orig_bio->bi_opf &= ~REQ_POLLED; + + do { + struct iomap_ioend *ioend; + struct bio *bio; + int error; + + bio = bio_alloc_bioset(orig_bio->bi_bdev, + min(total_len / minsize, BIO_MAX_VECS), + orig_bio->bi_opf, GFP_KERNEL, + &iomap_ioend_split_bioset); + error = bio_alloc_bounce_folios(bio, total_len, minsize); + if (error) { + bio_put(bio); + orig_bio->bi_status = errno_to_blk_status(error); + break; + } + bio->bi_ioprio = orig_bio->bi_ioprio; + bio->bi_write_hint = orig_bio->bi_write_hint; + bio->bi_write_stream = orig_bio->bi_write_stream; + bio->bi_iter.bi_sector = sector; + + ioend = iomap_init_ioend(inode, bio, file_offset, + orig_ioend->io_flags); + + total_len -= bio->bi_iter.bi_size; + file_offset += bio->bi_iter.bi_size; + sector += (bio->bi_iter.bi_size >> SECTOR_SHIFT); + + bio->bi_private = orig_bio; + bio_inc_remaining(orig_bio); + submit_ioend(ioend); + } while (total_len > 0); + + bio_endio(&orig_ioend->io_bio); +} +EXPORT_SYMBOL_GPL(iomap_bounce_read); + +static void iomap_ioend_unbounce(struct iomap_ioend *orig_ioend, + struct iomap_ioend *ioend) +{ + struct bio *orig_bio = &orig_ioend->io_bio; + struct iov_iter to; + struct bio_vec *bv; + int i; + + iov_iter_bvec(&to, ITER_DEST, orig_bio->bi_io_vec, orig_bio->bi_vcnt, + orig_ioend->io_size); + to.iov_offset = orig_ioend->io_bvec_offset; + + if (ioend->io_offset != orig_ioend->io_offset) { + WARN_ON_ONCE(ioend->io_offset < orig_ioend->io_offset); + iov_iter_advance(&to, ioend->io_offset - orig_ioend->io_offset); + } + + /* copying to pinned pages should always work */ + bio_for_each_bvec_all(bv, &ioend->io_bio, i) + WARN_ON_ONCE(copy_to_iter(bvec_virt(bv), bv->bv_len, &to) != + bv->bv_len); +} + +void iomap_bounce_read_end_io(struct iomap_ioend *ioend, struct bio *orig_bio, + int error) +{ + if (error) + orig_bio->bi_status = errno_to_blk_status(error); + else + iomap_ioend_unbounce(iomap_ioend_from_bio(orig_bio), ioend); + + bio_free_folios(&ioend->io_bio); + if (bio_integrity(&ioend->io_bio)) + fs_bio_integrity_free(&ioend->io_bio); + bio_put(&ioend->io_bio); + + bio_endio(orig_bio); +} +EXPORT_SYMBOL_GPL(iomap_bounce_read_end_io); + static int __init iomap_ioend_init(void) { const unsigned int nr_mempool_entries = 4 * (PAGE_SIZE / SECTOR_SIZE); diff --git a/fs/isofs/compress.c b/fs/isofs/compress.c index f9869d62b850..2d23abaeb874 100644 --- a/fs/isofs/compress.c +++ b/fs/isofs/compress.c @@ -39,7 +39,7 @@ static DEFINE_MUTEX(zisofs_zlib_lock); */ static loff_t zisofs_uncompress_block(struct inode *inode, loff_t block_start, loff_t block_end, int pcount, - struct page **pages, unsigned poffset, + struct folio **folios, unsigned int poffset, int *errp) { unsigned int zisofs_block_shift = ISOFS_I(inode)->i_format_parm[1]; @@ -66,11 +66,12 @@ static loff_t zisofs_uncompress_block(struct inode *inode, loff_t block_start, if (block_size == 0) { for ( i = 0 ; i < pcount ; i++ ) { unsigned int off = i ? 0 : poffset; + struct folio *folio = folios[i]; - if (!pages[i]) + if (!folio) continue; - memzero_page(pages[i], off, PAGE_SIZE - off); - SetPageUptodate(pages[i]); + folio_zero_range(folio, off, folio_size(folio) - off); + folio_mark_uptodate(folio); } return (((loff_t)pcount) << PAGE_SHIFT) - poffset; } @@ -119,11 +120,12 @@ static loff_t zisofs_uncompress_block(struct inode *inode, loff_t block_start, while (curpage < pcount && curbh < haveblocks && zerr != Z_STREAM_END) { + struct folio *folio = folios[curpage]; + if (!stream.avail_out) { - if (pages[curpage]) { - stream.next_out = kmap_local_page(pages[curpage]) - + poffset; - stream.avail_out = PAGE_SIZE - poffset; + if (folio) { + stream.next_out = kmap_local_folio(folio, poffset); + stream.avail_out = folio_size(folio) - poffset; poffset = 0; } else { stream.next_out = (void *)&zisofs_sink_page; @@ -173,9 +175,9 @@ static loff_t zisofs_uncompress_block(struct inode *inode, loff_t block_start, if (!stream.avail_out) { /* This page completed */ - if (pages[curpage]) { - flush_dcache_page(pages[curpage]); - SetPageUptodate(pages[curpage]); + if (folio) { + flush_dcache_folio(folio); + folio_mark_uptodate(folio); } if (stream.next_out != (unsigned char *)zisofs_sink_page) { kunmap_local(stream.next_out); @@ -206,7 +208,7 @@ b_eio: * fills in other pages if we have data for them. */ static int zisofs_fill_pages(struct inode *inode, int full_page, int pcount, - struct page **pages) + struct folio **folios) { loff_t start_off, end_off; loff_t block_start, block_end; @@ -221,14 +223,14 @@ static int zisofs_fill_pages(struct inode *inode, int full_page, int pcount, int err; loff_t ret; - BUG_ON(!pages[full_page]); + BUG_ON(!folios[full_page]); /* * We want to read at least 'full_page' page. Because we have to * uncompress the whole compression block anyway, fill the surrounding * pages with the data we have anyway... */ - start_off = page_offset(pages[full_page]); + start_off = folio_pos(folios[full_page]); end_off = min_t(loff_t, start_off + PAGE_SIZE, inode->i_size); cstart_block = start_off >> zisofs_block_shift; @@ -267,9 +269,9 @@ static int zisofs_fill_pages(struct inode *inode, int full_page, int pcount, } err = 0; ret = zisofs_uncompress_block(inode, block_start, block_end, - pcount, pages, poffset, &err); + pcount, folios, poffset, &err); poffset += ret; - pages += poffset >> PAGE_SHIFT; + folios += poffset >> PAGE_SHIFT; pcount -= poffset >> PAGE_SHIFT; full_page -= poffset >> PAGE_SHIFT; poffset &= ~PAGE_MASK; @@ -289,9 +291,11 @@ static int zisofs_fill_pages(struct inode *inode, int full_page, int pcount, cstart_block++; } - if (poffset && *pages) { - memzero_page(*pages, poffset, PAGE_SIZE - poffset); - SetPageUptodate(*pages); + if (poffset && *folios) { + struct folio *folio = *folios; + + folio_zero_range(folio, poffset, folio_size(folio) - poffset); + folio_mark_uptodate(folio); } brelse(bh); return 0; @@ -312,7 +316,7 @@ static int zisofs_read_folio(struct file *file, struct folio *folio) unsigned int zisofs_pages_per_cblock = PAGE_SHIFT <= zisofs_block_shift ? (1 << (zisofs_block_shift - PAGE_SHIFT)) : 0; - struct page **pages; + struct folio **folios; pgoff_t index = folio->index, end_index; end_index = (inode->i_size + PAGE_SIZE - 1) >> PAGE_SHIFT; @@ -336,33 +340,38 @@ static int zisofs_read_folio(struct file *file, struct folio *folio) full_page = 0; pcount = 1; } - pages = kzalloc_objs(*pages, - max_t(unsigned int, zisofs_pages_per_cblock, 1)); - if (!pages) { + folios = kzalloc_objs(*folios, + max_t(unsigned int, zisofs_pages_per_cblock, 1)); + if (!folios) { folio_unlock(folio); return -ENOMEM; } - pages[full_page] = &folio->page; + folios[full_page] = folio; for (i = 0; i < pcount; i++, index++) { - if (i != full_page) - pages[i] = grab_cache_page_nowait(mapping, index); + if (i == full_page) + continue; + folios[i] = __filemap_get_folio(mapping, index, + FGP_LOCK | FGP_CREAT | FGP_NOWAIT, + mapping_gfp_mask(mapping)); + if (IS_ERR(folios[i])) + folios[i] = NULL; } - err = zisofs_fill_pages(inode, full_page, pcount, pages); + err = zisofs_fill_pages(inode, full_page, pcount, folios); - /* Release any residual pages, do not SetPageUptodate */ + /* Release any residual folios, do not mark them uptodate */ for (i = 0; i < pcount; i++) { - if (pages[i]) { - flush_dcache_page(pages[i]); - unlock_page(pages[i]); + if (folios[i]) { + flush_dcache_folio(folios[i]); + folio_unlock(folios[i]); if (i != full_page) - put_page(pages[i]); + folio_put(folios[i]); } - } + } /* At this point, err contains 0 or -EIO depending on the "critical" page */ - kfree(pages); + kfree(folios); return err; } diff --git a/fs/isofs/inode.c b/fs/isofs/inode.c index 337836a0a170..184350d2e6ad 100644 --- a/fs/isofs/inode.c +++ b/fs/isofs/inode.c @@ -821,6 +821,8 @@ root_found: if (!sb_set_blocksize(s, orig_zonesize)) goto out_freesbi; + sbi->s_session_start = (sector_t)vol_desc_start << + (ISOFS_BLOCK_BITS - s->s_blocksize_bits); sbi->s_nls_iocharset = NULL; #ifdef CONFIG_JOLIET @@ -1174,7 +1176,6 @@ static int isofs_read_level3_size(struct inode *inode) unsigned long block, offset, block_saved, offset_saved; int i = 0; int more_entries = 0; - struct iso_directory_record *tmpde = NULL; struct iso_inode_info *ei = ISOFS_I(inode); inode->i_size = 0; @@ -1198,9 +1199,12 @@ static int isofs_read_level3_size(struct inode *inode) goto out_noread; } de = (struct iso_directory_record *) (bh->b_data + offset); - de_len = *(unsigned char *) de; - if (de_len == 0) { + /* + * If we are at the end of a block (or at its zero-padded + * tail), move on to the next block. + */ + if (offset >= bufsize || de->length[0] == 0) { brelse(bh); bh = NULL; ++block; @@ -1208,32 +1212,18 @@ static int isofs_read_level3_size(struct inode *inode) continue; } + if (!isofs_dir_record_valid(de, offset, bufsize)) { + printk(KERN_NOTICE "iso9660: Corrupted directory entry in block %lu of inode %llu\n", + block, inode->i_ino); + brelse(bh); + return -EIO; + } + + de_len = de->length[0]; block_saved = block; offset_saved = offset; offset += de_len; - /* Make sure we have a full directory entry */ - if (offset >= bufsize) { - int slop = bufsize - offset + de_len; - if (!tmpde) { - tmpde = kmalloc(256, GFP_KERNEL); - if (!tmpde) - goto out_nomem; - } - memcpy(tmpde, de, slop); - offset &= bufsize - 1; - block++; - brelse(bh); - bh = NULL; - if (offset) { - bh = sb_bread(inode->i_sb, block); - if (!bh) - goto out_noread; - memcpy((void *)tmpde+slop, bh->b_data, offset); - } - de = tmpde; - } - inode->i_size += isonum_733(de->size); if (i == 1) { ei->i_next_section_block = block_saved; @@ -1247,17 +1237,11 @@ static int isofs_read_level3_size(struct inode *inode) goto out_toomany; } while (more_entries); out: - kfree(tmpde); brelse(bh); return 0; -out_nomem: - brelse(bh); - return -ENOMEM; - out_noread: printk(KERN_INFO "ISOFS: unable to read i-node block %lu\n", block); - kfree(tmpde); return -EIO; out_toomany: diff --git a/fs/isofs/isofs.h b/fs/isofs/isofs.h index dacb9cdae4fd..79ca0256843a 100644 --- a/fs/isofs/isofs.h +++ b/fs/isofs/isofs.h @@ -35,6 +35,8 @@ struct isofs_sb_info { unsigned long s_firstdatazone; unsigned long s_log_zone_size; unsigned long s_max_size; + /* Session start in filesystem block units. */ + sector_t s_session_start; int s_rock_offset; /* offset of SUSP fields within SU area */ s32 s_sbsector; diff --git a/fs/isofs/rock.c b/fs/isofs/rock.c index 2628f31bd3a5..84e0d764c210 100644 --- a/fs/isofs/rock.c +++ b/fs/isofs/rock.c @@ -9,6 +9,7 @@ #include <linux/slab.h> #include <linux/pagemap.h> +#include <linux/blkdev.h> #include "isofs.h" #include "rock.h" @@ -84,6 +85,10 @@ static void init_rock_state(struct rock_state *rs, struct inode *inode) */ static int rock_continue(struct rock_state *rs) { + struct super_block *sb = rs->inode->i_sb; + struct isofs_sb_info *sbi = ISOFS_SB(sb); + sector_t session_end = sbi->s_session_start + + ((sector_t)sbi->s_nzones << (ISOFS_BLOCK_BITS - sb->s_blocksize_bits)); int ret = 1; int blocksize = 1 << rs->inode->i_blkbits; const int min_de_size = offsetof(struct rock_ridge, u); @@ -101,11 +106,13 @@ static int rock_continue(struct rock_state *rs) goto out; } - if ((unsigned)rs->cont_extent >= ISOFS_SB(rs->inode->i_sb)->s_nzones) { + if (rs->cont_extent && + (rs->cont_extent < sbi->s_session_start || + rs->cont_extent >= session_end)) { printk(KERN_NOTICE "rock: corrupted directory entry. " "extent=%u out of volume (nzones=%lu)\n", (unsigned)rs->cont_extent, - ISOFS_SB(rs->inode->i_sb)->s_nzones); + sbi->s_nzones); ret = -EIO; goto out; } diff --git a/fs/jbd2/commit.c b/fs/jbd2/commit.c index 3029cb6f6d64..ebf6ba58ff4d 100644 --- a/fs/jbd2/commit.c +++ b/fs/jbd2/commit.c @@ -32,14 +32,14 @@ static void journal_end_buffer_io_sync(struct bio *bio) { struct buffer_head *bh; - bool uptodate = bio_endio_bh(bio, &bh); + bool success = bio_endio_bh(bio, &bh); struct buffer_head *orig_bh = bh->b_private; BUFFER_TRACE(bh, ""); - if (uptodate) - set_buffer_uptodate(bh); + if (success) + clear_buffer_write_io_error(bh); else - clear_buffer_uptodate(bh); + mark_buffer_write_io_error(bh); if (orig_bh) { clear_and_wake_up_bit(BH_Shadow, &orig_bh->b_state); } @@ -169,7 +169,7 @@ static int journal_wait_on_commit_record(journal_t *journal, clear_buffer_dirty(bh); wait_on_buffer(bh); - if (unlikely(!buffer_uptodate(bh))) + if (unlikely(buffer_write_io_error(bh))) ret = -EIO; put_bh(bh); /* One for getblk() */ @@ -330,9 +330,9 @@ static __u32 jbd2_checksum_data(__u32 crc32_sum, struct buffer_head *bh) char *addr; __u32 checksum; - addr = kmap_local_folio(bh->b_folio, bh_offset(bh)); + addr = kmap_local_bh(bh); checksum = crc32_be(crc32_sum, addr, bh->b_size); - kunmap_local(addr); + kunmap_local_bh(bh, addr); return checksum; } @@ -357,10 +357,10 @@ static void jbd2_block_tag_csum_set(journal_t *j, journal_block_tag_t *tag, return; seq = cpu_to_be32(sequence); - addr = kmap_local_folio(bh->b_folio, bh_offset(bh)); + addr = kmap_local_bh(bh); csum32 = jbd2_chksum(j->j_csum_seed, (__u8 *)&seq, sizeof(seq)); csum32 = jbd2_chksum(csum32, addr, bh->b_size); - kunmap_local(addr); + kunmap_local_bh(bh, addr); if (jbd2_has_feature_csum3(j)) tag3->t_checksum = cpu_to_be32(csum32); @@ -834,7 +834,7 @@ start_journal_io: wait_on_buffer(bh); cond_resched(); - if (unlikely(!buffer_uptodate(bh))) + if (unlikely(buffer_write_io_error(bh))) err = -EIO; jbd2_unfile_log_bh(bh); stats.run.rs_blocks_logged++; @@ -877,7 +877,7 @@ start_journal_io: wait_on_buffer(bh); cond_resched(); - if (unlikely(!buffer_uptodate(bh))) + if (unlikely(buffer_write_io_error(bh))) err = -EIO; BUFFER_TRACE(bh, "ph5: control buffer writeout done: unfile"); diff --git a/fs/jbd2/journal.c b/fs/jbd2/journal.c index 00f5a98f3d4f..cda1ff8851dc 100644 --- a/fs/jbd2/journal.c +++ b/fs/jbd2/journal.c @@ -328,8 +328,6 @@ int jbd2_journal_write_metadata_buffer(transaction_t *transaction, { int do_escape = 0; struct buffer_head *new_bh; - struct folio *new_folio; - unsigned int new_offset; struct buffer_head *bh_in = jh2bh(jh_in); journal_t *journal = transaction->t_journal; @@ -349,24 +347,31 @@ int jbd2_journal_write_metadata_buffer(transaction_t *transaction, /* keep subsequent assertions sane */ atomic_set(&new_bh->b_count, 1); + /* + * b_frozen_data is slab memory, not page cache, so when we use it the + * shadow buffer gets no folio at all: b_folio stays NULL from the + * allocation and b_data points straight at the copy. Pointing it at + * the slab folio instead would hand its overloaded ->mapping to + * anything that goes looking for an address_space. + */ + spin_lock(&jh_in->b_state_lock); /* * If a new transaction has already done a buffer copy-out, then * we use that version of the data for the commit. */ if (jh_in->b_frozen_data) { - new_folio = virt_to_folio(jh_in->b_frozen_data); - new_offset = offset_in_folio(new_folio, jh_in->b_frozen_data); do_escape = jbd2_data_needs_escaping(jh_in->b_frozen_data); if (do_escape) jbd2_data_do_escape(jh_in->b_frozen_data); + new_bh->b_data = jh_in->b_frozen_data; } else { + struct folio *folio = bh_in->b_folio; + unsigned int offset = offset_in_folio(folio, bh_in->b_data); char *tmp; char *mapped_data; - new_folio = bh_in->b_folio; - new_offset = offset_in_folio(new_folio, bh_in->b_data); - mapped_data = kmap_local_folio(new_folio, new_offset); + mapped_data = kmap_local_folio(folio, offset); /* * Fire data frozen trigger if data already wasn't frozen. Do * this before checking for escaping, as the trigger may modify @@ -380,8 +385,10 @@ int jbd2_journal_write_metadata_buffer(transaction_t *transaction, /* * Do we need to do a data copy? */ - if (!do_escape) + if (!do_escape) { + folio_set_bh(new_bh, folio, offset); goto escape_done; + } spin_unlock(&jh_in->b_state_lock); tmp = kmalloc(bh_in->b_size, GFP_NOFS | __GFP_NOFAIL); @@ -392,7 +399,7 @@ int jbd2_journal_write_metadata_buffer(transaction_t *transaction, } jh_in->b_frozen_data = tmp; - memcpy_from_folio(tmp, new_folio, new_offset, bh_in->b_size); + memcpy_from_folio(tmp, folio, offset, bh_in->b_size); /* * This isn't strictly necessary, as we're using frozen * data for the escaping, but it keeps consistency with @@ -401,13 +408,11 @@ int jbd2_journal_write_metadata_buffer(transaction_t *transaction, jh_in->b_frozen_triggers = jh_in->b_triggers; copy_done: - new_folio = virt_to_folio(jh_in->b_frozen_data); - new_offset = offset_in_folio(new_folio, jh_in->b_frozen_data); jbd2_data_do_escape(jh_in->b_frozen_data); + new_bh->b_data = jh_in->b_frozen_data; } escape_done: - folio_set_bh(new_bh, new_folio, new_offset); new_bh->b_size = bh_in->b_size; new_bh->b_bdev = journal->j_dev; new_bh->b_blocknr = blocknr; @@ -882,7 +887,7 @@ int jbd2_fc_wait_bufs(journal_t *journal, int num_blks) * Update j_fc_off so jbd2_fc_release_bufs can release remain * buffer head. */ - if (unlikely(!buffer_uptodate(bh))) { + if (unlikely(buffer_write_io_error(bh))) { journal->j_fc_off = i + 1; return -EIO; } diff --git a/fs/jbd2/transaction.c b/fs/jbd2/transaction.c index 5cc7d097b2ac..85d84d909f78 100644 --- a/fs/jbd2/transaction.c +++ b/fs/jbd2/transaction.c @@ -920,7 +920,7 @@ static void jbd2_freeze_jh_data(struct journal_head *jh) char *source; struct buffer_head *bh = jh2bh(jh); - J_EXPECT_JH(jh, buffer_uptodate(bh), "Possible IO failure.\n"); + J_EXPECT_JH(jh, buffer_uptodate(bh), "Buffer not uptodate!\n"); source = kmap_local_folio(bh->b_folio, bh_offset(bh)); /* Fire data frozen trigger just before we copy the data */ jbd2_buffer_frozen_trigger(jh, source, jh->b_triggers); diff --git a/fs/jffs2/acl.c b/fs/jffs2/acl.c index f0f8a4f57add..7548f44bf327 100644 --- a/fs/jffs2/acl.c +++ b/fs/jffs2/acl.c @@ -228,7 +228,7 @@ static int __jffs2_set_acl(struct inode *inode, int xprefix, struct posix_acl *a return rc; } -int jffs2_set_acl(struct mnt_idmap *idmap, struct dentry *dentry, +int jffs2_set_acl(const struct mnt_idmap *idmap, struct dentry *dentry, struct posix_acl *acl, int type) { int rc, xprefix; diff --git a/fs/jffs2/acl.h b/fs/jffs2/acl.h index e976b8cb82cf..bc5df521633f 100644 --- a/fs/jffs2/acl.h +++ b/fs/jffs2/acl.h @@ -28,7 +28,7 @@ struct jffs2_acl_header { #ifdef CONFIG_JFFS2_FS_POSIX_ACL struct posix_acl *jffs2_get_acl(struct inode *inode, int type, bool rcu); -int jffs2_set_acl(struct mnt_idmap *idmap, struct dentry *dentry, +int jffs2_set_acl(const struct mnt_idmap *idmap, struct dentry *dentry, struct posix_acl *acl, int type); extern int jffs2_init_acl_pre(struct inode *, struct inode *, umode_t *); extern int jffs2_init_acl_post(struct inode *); diff --git a/fs/jffs2/dir.c b/fs/jffs2/dir.c index 656c920864c5..23813f191281 100644 --- a/fs/jffs2/dir.c +++ b/fs/jffs2/dir.c @@ -25,20 +25,20 @@ static int jffs2_readdir (struct file *, struct dir_context *); -static int jffs2_create (struct mnt_idmap *, struct inode *, +static int jffs2_create (const struct mnt_idmap *, struct inode *, struct dentry *, umode_t); static struct dentry *jffs2_lookup (struct inode *,struct dentry *, unsigned int); static int jffs2_link (struct dentry *,struct inode *,struct dentry *); static int jffs2_unlink (struct inode *,struct dentry *); -static int jffs2_symlink (struct mnt_idmap *, struct inode *, +static int jffs2_symlink (const struct mnt_idmap *, struct inode *, struct dentry *, const char *); -static struct dentry *jffs2_mkdir (struct mnt_idmap *, struct inode *,struct dentry *, +static struct dentry *jffs2_mkdir (const struct mnt_idmap *, struct inode *,struct dentry *, umode_t); static int jffs2_rmdir (struct inode *,struct dentry *); -static int jffs2_mknod (struct mnt_idmap *, struct inode *,struct dentry *, +static int jffs2_mknod (const struct mnt_idmap *, struct inode *,struct dentry *, umode_t,dev_t); -static int jffs2_rename (struct mnt_idmap *, struct inode *, +static int jffs2_rename (const struct mnt_idmap *, struct inode *, struct dentry *, struct inode *, struct dentry *, unsigned int); @@ -162,7 +162,7 @@ static int jffs2_readdir(struct file *file, struct dir_context *ctx) /***********************************************************************/ -static int jffs2_create(struct mnt_idmap *idmap, struct inode *dir_i, +static int jffs2_create(const struct mnt_idmap *idmap, struct inode *dir_i, struct dentry *dentry, umode_t mode) { struct jffs2_raw_inode *ri; @@ -284,7 +284,7 @@ static int jffs2_link (struct dentry *old_dentry, struct inode *dir_i, struct de /***********************************************************************/ -static int jffs2_symlink (struct mnt_idmap *idmap, struct inode *dir_i, +static int jffs2_symlink (const struct mnt_idmap *idmap, struct inode *dir_i, struct dentry *dentry, const char *target) { struct jffs2_inode_info *f, *dir_f; @@ -448,7 +448,7 @@ static int jffs2_symlink (struct mnt_idmap *idmap, struct inode *dir_i, } -static struct dentry *jffs2_mkdir (struct mnt_idmap *idmap, struct inode *dir_i, +static struct dentry *jffs2_mkdir (const struct mnt_idmap *idmap, struct inode *dir_i, struct dentry *dentry, umode_t mode) { struct jffs2_inode_info *f, *dir_f; @@ -620,7 +620,7 @@ static int jffs2_rmdir (struct inode *dir_i, struct dentry *dentry) return ret; } -static int jffs2_mknod (struct mnt_idmap *idmap, struct inode *dir_i, +static int jffs2_mknod (const struct mnt_idmap *idmap, struct inode *dir_i, struct dentry *dentry, umode_t mode, dev_t rdev) { struct jffs2_inode_info *f, *dir_f; @@ -769,7 +769,7 @@ static int jffs2_mknod (struct mnt_idmap *idmap, struct inode *dir_i, return ret; } -static int jffs2_rename (struct mnt_idmap *idmap, +static int jffs2_rename (const struct mnt_idmap *idmap, struct inode *old_dir_i, struct dentry *old_dentry, struct inode *new_dir_i, struct dentry *new_dentry, unsigned int flags) diff --git a/fs/jffs2/fs.c b/fs/jffs2/fs.c index 6ada8369a762..05cf860307c7 100644 --- a/fs/jffs2/fs.c +++ b/fs/jffs2/fs.c @@ -190,7 +190,7 @@ int jffs2_do_setattr (struct inode *inode, struct iattr *iattr) return 0; } -int jffs2_setattr(struct mnt_idmap *idmap, struct dentry *dentry, +int jffs2_setattr(const struct mnt_idmap *idmap, struct dentry *dentry, struct iattr *iattr) { struct inode *inode = d_inode(dentry); diff --git a/fs/jffs2/os-linux.h b/fs/jffs2/os-linux.h index 86ab014a349c..bff2134d771d 100644 --- a/fs/jffs2/os-linux.h +++ b/fs/jffs2/os-linux.h @@ -164,7 +164,7 @@ long jffs2_ioctl(struct file *, unsigned int, unsigned long); extern const struct inode_operations jffs2_symlink_inode_operations; /* fs.c */ -int jffs2_setattr (struct mnt_idmap *, struct dentry *, struct iattr *); +int jffs2_setattr (const struct mnt_idmap *, struct dentry *, struct iattr *); int jffs2_do_setattr (struct inode *, struct iattr *); struct inode *jffs2_iget(struct super_block *, unsigned long); void jffs2_evict_inode (struct inode *); diff --git a/fs/jffs2/security.c b/fs/jffs2/security.c index 437f3a2c1b54..67330aeb8ae8 100644 --- a/fs/jffs2/security.c +++ b/fs/jffs2/security.c @@ -57,7 +57,7 @@ static int jffs2_security_getxattr(const struct xattr_handler *handler, } static int jffs2_security_setxattr(const struct xattr_handler *handler, - struct mnt_idmap *idmap, + const struct mnt_idmap *idmap, struct dentry *unused, struct inode *inode, const char *name, const void *buffer, size_t size, int flags) diff --git a/fs/jffs2/xattr_trusted.c b/fs/jffs2/xattr_trusted.c index b7c5da2d89bd..85133ad8b449 100644 --- a/fs/jffs2/xattr_trusted.c +++ b/fs/jffs2/xattr_trusted.c @@ -25,7 +25,7 @@ static int jffs2_trusted_getxattr(const struct xattr_handler *handler, } static int jffs2_trusted_setxattr(const struct xattr_handler *handler, - struct mnt_idmap *idmap, + const struct mnt_idmap *idmap, struct dentry *unused, struct inode *inode, const char *name, const void *buffer, size_t size, int flags) diff --git a/fs/jffs2/xattr_user.c b/fs/jffs2/xattr_user.c index f64edce4927b..dcfd3caf1d8b 100644 --- a/fs/jffs2/xattr_user.c +++ b/fs/jffs2/xattr_user.c @@ -25,7 +25,7 @@ static int jffs2_user_getxattr(const struct xattr_handler *handler, } static int jffs2_user_setxattr(const struct xattr_handler *handler, - struct mnt_idmap *idmap, + const struct mnt_idmap *idmap, struct dentry *unused, struct inode *inode, const char *name, const void *buffer, size_t size, int flags) diff --git a/fs/jfs/acl.c b/fs/jfs/acl.c index 16b71a23ff1e..6e0a7feb6c80 100644 --- a/fs/jfs/acl.c +++ b/fs/jfs/acl.c @@ -89,7 +89,7 @@ static int __jfs_set_acl(tid_t tid, struct inode *inode, int type, return rc; } -int jfs_set_acl(struct mnt_idmap *idmap, struct dentry *dentry, +int jfs_set_acl(const struct mnt_idmap *idmap, struct dentry *dentry, struct posix_acl *acl, int type) { int rc; diff --git a/fs/jfs/file.c b/fs/jfs/file.c index 246568cb9a6e..2f5bb79c0591 100644 --- a/fs/jfs/file.c +++ b/fs/jfs/file.c @@ -89,7 +89,7 @@ static int jfs_release(struct inode *inode, struct file *file) return 0; } -int jfs_setattr(struct mnt_idmap *idmap, struct dentry *dentry, +int jfs_setattr(const struct mnt_idmap *idmap, struct dentry *dentry, struct iattr *iattr) { struct inode *inode = d_inode(dentry); diff --git a/fs/jfs/ioctl.c b/fs/jfs/ioctl.c index 563f148be8af..27d39cddaca6 100644 --- a/fs/jfs/ioctl.c +++ b/fs/jfs/ioctl.c @@ -70,7 +70,7 @@ int jfs_fileattr_get(struct dentry *dentry, struct file_kattr *fa) return 0; } -int jfs_fileattr_set(struct mnt_idmap *idmap, +int jfs_fileattr_set(const struct mnt_idmap *idmap, struct dentry *dentry, struct file_kattr *fa) { struct inode *inode = d_inode(dentry); diff --git a/fs/jfs/jfs_acl.h b/fs/jfs/jfs_acl.h index f892e54d0fcd..bda26b333519 100644 --- a/fs/jfs/jfs_acl.h +++ b/fs/jfs/jfs_acl.h @@ -8,7 +8,7 @@ #ifdef CONFIG_JFS_POSIX_ACL struct posix_acl *jfs_get_acl(struct inode *inode, int type, bool rcu); -int jfs_set_acl(struct mnt_idmap *idmap, struct dentry *dentry, +int jfs_set_acl(const struct mnt_idmap *idmap, struct dentry *dentry, struct posix_acl *acl, int type); int jfs_init_acl(tid_t, struct inode *, struct inode *); diff --git a/fs/jfs/jfs_inode.h b/fs/jfs/jfs_inode.h index 2c6c81c8cb9f..5a118b07fbff 100644 --- a/fs/jfs/jfs_inode.h +++ b/fs/jfs/jfs_inode.h @@ -10,7 +10,7 @@ struct fid; extern struct inode *ialloc(struct inode *, umode_t); extern int jfs_fsync(struct file *, loff_t, loff_t, int); extern int jfs_fileattr_get(struct dentry *dentry, struct file_kattr *fa); -extern int jfs_fileattr_set(struct mnt_idmap *idmap, +extern int jfs_fileattr_set(const struct mnt_idmap *idmap, struct dentry *dentry, struct file_kattr *fa); extern long jfs_ioctl(struct file *, unsigned int, unsigned long); extern struct inode *jfs_iget(struct super_block *, unsigned long); @@ -28,7 +28,7 @@ extern struct dentry *jfs_fh_to_parent(struct super_block *sb, struct fid *fid, int fh_len, int fh_type); extern void jfs_set_inode_flags(struct inode *); extern int jfs_get_block(struct inode *, sector_t, struct buffer_head *, int); -extern int jfs_setattr(struct mnt_idmap *, struct dentry *, struct iattr *); +extern int jfs_setattr(const struct mnt_idmap *, struct dentry *, struct iattr *); extern const struct address_space_operations jfs_aops; extern const struct inode_operations jfs_dir_inode_operations; diff --git a/fs/jfs/namei.c b/fs/jfs/namei.c index 8a36c218f0f7..8ab2e952ce16 100644 --- a/fs/jfs/namei.c +++ b/fs/jfs/namei.c @@ -60,7 +60,7 @@ static inline void free_ea_wmap(struct inode *inode) * RETURN: Errors from subroutines * */ -static int jfs_create(struct mnt_idmap *idmap, struct inode *dip, +static int jfs_create(const struct mnt_idmap *idmap, struct inode *dip, struct dentry *dentry, umode_t mode) { int rc = 0; @@ -193,7 +193,7 @@ static int jfs_create(struct mnt_idmap *idmap, struct inode *dip, * note: * EACCES: user needs search+write permission on the parent directory */ -static struct dentry *jfs_mkdir(struct mnt_idmap *idmap, struct inode *dip, +static struct dentry *jfs_mkdir(const struct mnt_idmap *idmap, struct inode *dip, struct dentry *dentry, umode_t mode) { int rc = 0; @@ -876,7 +876,7 @@ static int jfs_link(struct dentry *old_dentry, * an intermediate result whose length exceeds PATH_MAX [XPG4.2] */ -static int jfs_symlink(struct mnt_idmap *idmap, struct inode *dip, +static int jfs_symlink(const struct mnt_idmap *idmap, struct inode *dip, struct dentry *dentry, const char *name) { int rc; @@ -1066,7 +1066,7 @@ static int jfs_symlink(struct mnt_idmap *idmap, struct inode *dip, * * FUNCTION: rename a file or directory */ -static int jfs_rename(struct mnt_idmap *idmap, struct inode *old_dir, +static int jfs_rename(const struct mnt_idmap *idmap, struct inode *old_dir, struct dentry *old_dentry, struct inode *new_dir, struct dentry *new_dentry, unsigned int flags) { @@ -1355,7 +1355,7 @@ static int jfs_rename(struct mnt_idmap *idmap, struct inode *old_dir, * * FUNCTION: Create a special file (device) */ -static int jfs_mknod(struct mnt_idmap *idmap, struct inode *dir, +static int jfs_mknod(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *dentry, umode_t mode, dev_t rdev) { struct jfs_inode_info *jfs_ip; diff --git a/fs/jfs/xattr.c b/fs/jfs/xattr.c index 11d7f74d207b..dcc4a69d44fe 100644 --- a/fs/jfs/xattr.c +++ b/fs/jfs/xattr.c @@ -956,7 +956,7 @@ static int jfs_xattr_get(const struct xattr_handler *handler, } static int jfs_xattr_set(const struct xattr_handler *handler, - struct mnt_idmap *idmap, + const struct mnt_idmap *idmap, struct dentry *unused, struct inode *inode, const char *name, const void *value, size_t size, int flags) @@ -975,7 +975,7 @@ static int jfs_xattr_get_os2(const struct xattr_handler *handler, } static int jfs_xattr_set_os2(const struct xattr_handler *handler, - struct mnt_idmap *idmap, + const struct mnt_idmap *idmap, struct dentry *unused, struct inode *inode, const char *name, const void *value, size_t size, int flags) diff --git a/fs/kernfs/dir.c b/fs/kernfs/dir.c index 82bbaeb326aa..324d61a00545 100644 --- a/fs/kernfs/dir.c +++ b/fs/kernfs/dir.c @@ -30,6 +30,8 @@ static char kernfs_pr_cont_buf[PATH_MAX]; /* protected by pr_cont_lock */ #define rb_to_kn(X) rb_entry((X), struct kernfs_node, rb) +static void kernfs_activate_one(struct kernfs_node *kn); + static bool __kernfs_active(struct kernfs_node *kn) { return atomic_read(&kn->active) >= 0; @@ -736,13 +738,19 @@ struct kernfs_node *kernfs_new_node(struct kernfs_node *parent, { struct kernfs_node *kn; - if (parent->mode & S_ISGID) { + /* + * The mode and the gid below are read unlocked on purpose: they feed + * a node that does not exist yet, so nothing orders a racing chmod or + * chown against this creation. + */ + if (READ_ONCE(parent->mode) & S_ISGID) { /* this code block imitates inode_init_owner() for * kernfs */ + struct kernfs_iattrs *attrs = READ_ONCE(parent->iattr); - if (parent->iattr) - gid = parent->iattr->ia_gid; + if (attrs) + gid = READ_ONCE(attrs->ia_gid); if (flags & KERNFS_DIR) mode |= S_ISGID; @@ -855,7 +863,6 @@ int kernfs_add_one(struct kernfs_node *kn) } up_write(&root->kernfs_iattr_rwsem); - up_write(&root->kernfs_rwsem); /* * Activate the new node unless CREATE_DEACTIVATED is requested. @@ -863,9 +870,15 @@ int kernfs_add_one(struct kernfs_node *kn) * activating the node with kernfs_activate(). A node which hasn't * been activated is not visible to userland and its removal won't * trigger deactivation. + * + * @kn has no children yet, so kernfs_activate() would walk only @kn. + * Do it here rather than dropping the write lock and taking it again + * for every new node. */ - if (!(kernfs_root(kn)->flags & KERNFS_ROOT_CREATE_DEACTIVATED)) - kernfs_activate(kn); + if (!(root->flags & KERNFS_ROOT_CREATE_DEACTIVATED)) + kernfs_activate_one(kn); + + up_write(&root->kernfs_rwsem); return 0; out_unlock: @@ -1171,23 +1184,18 @@ struct kernfs_node *kernfs_create_empty_dir(struct kernfs_node *parent, static int kernfs_dop_revalidate(struct inode *dir, const struct qstr *name, struct dentry *dentry, unsigned int flags) { - struct kernfs_node *kn, *parent; - struct kernfs_root *root; + struct kernfs_node *parent = dir->i_private; + struct kernfs_node *kn; + const char *kn_name; if (flags & LOOKUP_RCU) return -ECHILD; /* Negative hashed dentry? */ if (d_really_is_negative(dentry)) { - /* If the kernfs parent node has changed discard and - * proceed to ->lookup. - * - * There's nothing special needed here when getting the - * dentry parent, even if a concurrent rename is in - * progress. That's because the dentry is negative so - * it can only be the target of the rename and it will - * be doing a d_move() not a replace. Consequently the - * dentry d_parent won't change over the d_move(). + /* + * If the kernfs parent node has changed discard and proceed to + * ->lookup. * * Also kernfs negative dentries transitioning from * negative to positive during revalidate won't happen @@ -1195,50 +1203,41 @@ static int kernfs_dop_revalidate(struct inode *dir, const struct qstr *name, * changes and the lookup re-done so that a new positive * dentry can be properly created. */ - root = kernfs_root_from_sb(dentry->d_sb); - down_read(&root->kernfs_rwsem); - parent = kernfs_dentry_node(dentry->d_parent); - if (parent) { - if (kernfs_dir_changed(parent, dentry)) { - up_read(&root->kernfs_rwsem); - return 0; - } - } - up_read(&root->kernfs_rwsem); - - /* The kernfs parent node hasn't changed, leave the - * dentry negative and return success. - */ - return 1; + return !kernfs_dir_changed(parent, dentry); } kn = kernfs_dentry_node(dentry); - root = kernfs_root(kn); - down_read(&root->kernfs_rwsem); + + guard(rcu)(); /* The kernfs node has been deactivated */ - if (!kernfs_active(kn)) - goto out_bad; + if (!__kernfs_active(kn)) + return 0; - parent = kernfs_parent(kn); /* The kernfs node has been moved? */ - if (kernfs_dentry_node(dentry->d_parent) != parent) - goto out_bad; + if (kernfs_parent(kn) != parent) + return 0; /* The kernfs node has been renamed */ - if (strcmp(dentry->d_name.name, kernfs_rcu_name(kn)) != 0) - goto out_bad; + kn_name = kernfs_rcu_name(kn); + if (name->len != strlen(kn_name) || + memcmp(name->name, kn_name, name->len)) + return 0; - /* The kernfs node has been moved to a different namespace */ - if (parent && kernfs_ns_enabled(parent) && - kernfs_ns_id(kernfs_info(dentry->d_sb)->ns) != kernfs_ns_id(kn->ns)) - goto out_bad; + /* + * The kernfs node has been moved to a different namespace. + * + * KERNFS_NS is set by kernfs_enable_ns() while @parent still has no + * children, so it cannot change while a child of @parent is being + * revalidated. The other bits in that word, KERNFS_ACTIVATED and + * KERNFS_REMOVING, are updated under kernfs_rwsem and are not read + * here, so racing with them is intentional and harmless. + */ + if (data_race(kernfs_ns_enabled(parent)) && + kernfs_info(dir->i_sb)->ns != READ_ONCE(kn->ns)) + return 0; - up_read(&root->kernfs_rwsem); return 1; -out_bad: - up_read(&root->kernfs_rwsem); - return 0; } const struct dentry_operations kernfs_dops = { @@ -1288,7 +1287,7 @@ static struct dentry *kernfs_iop_lookup(struct inode *dir, return d_splice_alias(inode, dentry); } -static struct dentry *kernfs_iop_mkdir(struct mnt_idmap *idmap, +static struct dentry *kernfs_iop_mkdir(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *dentry, umode_t mode) { @@ -1326,7 +1325,7 @@ static int kernfs_iop_rmdir(struct inode *dir, struct dentry *dentry) return ret; } -static int kernfs_iop_rename(struct mnt_idmap *idmap, +static int kernfs_iop_rename(const struct mnt_idmap *idmap, struct inode *old_dir, struct dentry *old_dentry, struct inode *new_dir, struct dentry *new_dentry, unsigned int flags) @@ -1820,14 +1819,20 @@ int kernfs_rename_ns(struct kernfs_node *kn, struct kernfs_node *new_parent, const char *new_name, const struct ns_common *new_ns) { struct kernfs_node *old_parent; + const char *dup_name = NULL; + const char *put_name = NULL; struct kernfs_root *root; const char *old_name; + bool reparent; int error; /* can't move or rename root */ if (!rcu_access_pointer(kn->__parent)) return -EINVAL; + if (new_name) + dup_name = kstrdup_const(new_name, GFP_KERNEL); + root = kernfs_root(kn); down_write(&root->kernfs_rwsem); @@ -1859,9 +1864,10 @@ int kernfs_rename_ns(struct kernfs_node *kn, struct kernfs_node *new_parent, /* rename kernfs_node */ if (strcmp(old_name, new_name) != 0) { error = -ENOMEM; - new_name = kstrdup_const(new_name, GFP_KERNEL); - if (!new_name) + if (!dup_name) goto out; + new_name = dup_name; + dup_name = NULL; } else { new_name = NULL; } @@ -1871,35 +1877,39 @@ int kernfs_rename_ns(struct kernfs_node *kn, struct kernfs_node *new_parent, */ kernfs_unlink_sibling(kn); - /* rename_lock protects ->parent accessors */ - if (old_parent != new_parent) { + reparent = old_parent != new_parent; + if (reparent) kernfs_get(new_parent); - write_lock_irq(&root->kernfs_rename_lock); + /* + * kernfs_rename_lock protects ->__parent, ->ns and ->name, so take it + * even when the parent does not change. + */ + write_lock_irq(&root->kernfs_rename_lock); + + if (reparent) rcu_assign_pointer(kn->__parent, new_parent); + WRITE_ONCE(kn->ns, new_ns); + if (new_name) + rcu_assign_pointer(kn->name, new_name); - kn->ns = new_ns; - if (new_name) - rcu_assign_pointer(kn->name, new_name); + write_unlock_irq(&root->kernfs_rename_lock); - write_unlock_irq(&root->kernfs_rename_lock); + if (reparent) kernfs_put(old_parent); - } else { - /* name assignment is RCU protected, parent is the same */ - kn->ns = new_ns; - if (new_name) - rcu_assign_pointer(kn->name, new_name); - } kn->hash = kernfs_name_hash(new_name ?: old_name, kn->ns); kernfs_link_sibling(kn); if (new_name && !is_kernel_rodata((unsigned long)old_name)) - kfree_rcu_mightsleep(old_name); + put_name = old_name; error = 0; out: up_write(&root->kernfs_rwsem); + kfree_const(dup_name); + if (put_name) + kfree_rcu_mightsleep(put_name); return error; } @@ -1909,33 +1919,49 @@ static int kernfs_dir_fop_release(struct inode *inode, struct file *filp) return 0; } +/* + * Find where a listing left off. @resumed says whether @pos is still that + * entry; if not, the search falls back to @hash, keyed by @name if given. + */ static struct kernfs_node *kernfs_dir_pos(const struct ns_common *ns, - struct kernfs_node *parent, loff_t hash, struct kernfs_node *pos) + struct kernfs_node *parent, loff_t hash, struct kernfs_node *pos, + const char *name, bool *resumed) { + if (resumed) + *resumed = false; if (pos) { + /* + * A rename keeps the hash if the new name hashes the same, so + * check @name too. Otherwise the caller would step over the + * entry now sitting where @pos used to be. + */ int valid = kernfs_active(pos) && rcu_access_pointer(pos->__parent) == parent && - hash == pos->hash; + hash == pos->hash && + (!name || !strcmp(name, kernfs_rcu_name(pos))); kernfs_put(pos); if (!valid) pos = NULL; + else if (resumed) + *resumed = true; } if (!pos && (hash > 1) && (hash < INT_MAX)) { struct rb_node *node = parent->dir.children.rb_node; - u64 ns_id = kernfs_ns_id(ns); + + /* + * Keep a node only on the way left, so the search ends on the + * first entry after the key. An empty @name sorts before all + * entries sharing the hash, so it lands on the first of them. + */ while (node) { - pos = rb_to_kn(node); + struct kernfs_node *kn = rb_to_kn(node); - if (hash < pos->hash) + if (kernfs_name_compare(hash, name ?: "", ns, kn) < 0) { + pos = kn; node = node->rb_left; - else if (hash > pos->hash) + } else { node = node->rb_right; - else if (ns_id < kernfs_ns_id(pos->ns)) - node = node->rb_left; - else if (ns_id > kernfs_ns_id(pos->ns)) - node = node->rb_right; - else - break; + } } } /* Skip over entries which are dying/dead or in the wrong namespace */ @@ -1951,10 +1977,14 @@ static struct kernfs_node *kernfs_dir_pos(const struct ns_common *ns, } static struct kernfs_node *kernfs_dir_next_pos(const struct ns_common *ns, - struct kernfs_node *parent, ino_t ino, struct kernfs_node *pos) + struct kernfs_node *parent, loff_t hash, struct kernfs_node *pos, + const char *name) { - pos = kernfs_dir_pos(ns, parent, ino, pos); - if (pos) { + bool resumed; + + pos = kernfs_dir_pos(ns, parent, hash, pos, name, &resumed); + /* Step over @pos only if it survived; @name finds the spot if not. */ + if (pos && resumed) { do { struct rb_node *node = rb_next(&pos->rb); if (!node) @@ -1972,34 +2002,55 @@ static int kernfs_fop_readdir(struct file *file, struct dir_context *ctx) struct dentry *dentry = file->f_path.dentry; struct kernfs_node *parent = kernfs_dentry_node(dentry); struct kernfs_node *pos = file->private_data; + char *name __free(kfree) = NULL; struct kernfs_root *root; const struct ns_common *ns = NULL; if (!dir_emit_dots(file, ctx)) return 0; + /* + * One buffer for the call, holding the name of the entry the listing + * is on. PATH_MAX: kernfs bounds no single name. + */ + name = kmalloc(PATH_MAX, GFP_KERNEL); + if (!name) + return -ENOMEM; + root = kernfs_root(parent); down_read(&root->kernfs_rwsem); if (kernfs_ns_enabled(parent)) ns = kernfs_info(dentry->d_sb)->ns; - for (pos = kernfs_dir_pos(ns, parent, ctx->pos, pos); + for (pos = kernfs_dir_pos(ns, parent, ctx->pos, pos, NULL, NULL); pos; - pos = kernfs_dir_next_pos(ns, parent, ctx->pos, pos)) { - const char *name = kernfs_rcu_name(pos); + pos = kernfs_dir_next_pos(ns, parent, ctx->pos, pos, name)) { unsigned int type = fs_umode_to_dtype(pos->mode); - int len = strlen(name); ino_t ino = kernfs_ino(pos); + int len; + + /* + * The copy is also the resume key, so a truncated name would + * resume here again. getname() caps a path, so only an + * in-kernel caller can get here; end the listing instead. + */ + len = strscpy(name, kernfs_rcu_name(pos), PATH_MAX); + if (WARN_ON_ONCE(len < 0)) + break; ctx->pos = pos->hash; file->private_data = pos; kernfs_get(pos); - if (!dir_emit(ctx, name, len, ino, type)) { - up_read(&root->kernfs_rwsem); + /* + * dir_emit() can fault, so run it unlocked. @pos is pinned + * above and kernfs_dir_pos() rechecks it on the way back. + */ + up_read(&root->kernfs_rwsem); + if (!dir_emit(ctx, name, len, ino, type)) return 0; - } + down_read(&root->kernfs_rwsem); } up_read(&root->kernfs_rwsem); file->private_data = NULL; diff --git a/fs/kernfs/file.c b/fs/kernfs/file.c index 8e0e90c93372..cca9f83fc9b5 100644 --- a/fs/kernfs/file.c +++ b/fs/kernfs/file.c @@ -525,18 +525,31 @@ out_unlock: static int kernfs_get_open_node(struct kernfs_node *kn, struct kernfs_open_file *of) { - struct kernfs_open_node *on; + struct kernfs_open_node *on, *new_on = NULL; struct mutex *mutex; + /* + * Peek without the mutex: if nothing has this open, we will need a + * node and can allocate before taking a mutex shared by every node + * hashing to it. + */ + if (!rcu_access_pointer(kn->attr.open)) + new_on = kzalloc_obj(*new_on); + mutex = kernfs_open_file_mutex_lock(kn); on = kernfs_deref_open_node_locked(kn); if (!on) { /* not there, initialize a new one */ - on = kzalloc_obj(*on); + on = new_on; + new_on = NULL; if (!on) { - mutex_unlock(mutex); - return -ENOMEM; + /* the peek raced; rare, so allocate here */ + on = kzalloc_obj(*on); + if (!on) { + mutex_unlock(mutex); + return -ENOMEM; + } } atomic_set(&on->event, 1); init_waitqueue_head(&on->poll); @@ -549,6 +562,7 @@ static int kernfs_get_open_node(struct kernfs_node *kn, on->nr_to_release++; mutex_unlock(mutex); + kfree(new_on); return 0; } @@ -904,9 +918,12 @@ static loff_t kernfs_fop_llseek(struct file *file, loff_t offset, int whence) static void kernfs_notify_workfn(struct work_struct *work) { - struct kernfs_node *kn; + char name_buf[NAME_MAX + 1]; struct kernfs_super_info *info; + struct kernfs_node *kn; struct kernfs_root *root; + struct qstr name; + bool have_name; repeat: /* pop one off the notify_list */ spin_lock_irq(&kernfs_notify_lock); @@ -922,14 +939,20 @@ repeat: root = kernfs_root(kn); /* kick fsnotify */ + /* + * Sample the name once so kernfs_rwsem need not be held across the + * loop. A name that does not fit is reported without one; fsnotify() + * takes the name as optional, so a watcher loses the name and not the + * event. + */ + have_name = kernfs_name(kn, name_buf, sizeof(name_buf)) >= 0; + name = QSTR(name_buf); + down_read(&root->kernfs_supers_rwsem); - down_read(&root->kernfs_rwsem); - list_for_each_entry(info, &kernfs_root(kn)->supers, node) { + list_for_each_entry(info, &root->supers, node) { struct kernfs_node *parent; struct inode *p_inode = NULL; - const char *kn_name; struct inode *inode; - struct qstr name; /* * We want fsnotify_modify() on @kn but as the @@ -941,15 +964,14 @@ repeat: if (!inode) continue; - kn_name = kernfs_rcu_name(kn); - name = QSTR(kn_name); parent = kernfs_get_parent(kn); if (parent) { p_inode = ilookup(info->sb, kernfs_ino(parent)); if (p_inode) { fsnotify(FS_MODIFY | FS_EVENT_ON_CHILD, inode, FSNOTIFY_EVENT_INODE, - p_inode, &name, inode, 0); + p_inode, have_name ? &name : NULL, + inode, 0); iput(p_inode); } @@ -962,7 +984,6 @@ repeat: iput(inode); } - up_read(&root->kernfs_rwsem); up_read(&root->kernfs_supers_rwsem); kernfs_put(kn); goto repeat; diff --git a/fs/kernfs/inode.c b/fs/kernfs/inode.c index abb286bc3474..6630b29d7c07 100644 --- a/fs/kernfs/inode.c +++ b/fs/kernfs/inode.c @@ -107,7 +107,7 @@ int kernfs_setattr(struct kernfs_node *kn, const struct iattr *iattr) return ret; } -int kernfs_iop_setattr(struct mnt_idmap *idmap, struct dentry *dentry, +int kernfs_iop_setattr(const struct mnt_idmap *idmap, struct dentry *dentry, struct iattr *iattr) { struct inode *inode = d_inode(dentry); @@ -179,7 +179,7 @@ static void kernfs_refresh_inode(struct kernfs_node *kn, struct inode *inode) set_nlink(inode, kn->dir.subdirs + 2); } -int kernfs_iop_getattr(struct mnt_idmap *idmap, +int kernfs_iop_getattr(const struct mnt_idmap *idmap, const struct path *path, struct kstat *stat, u32 request_mask, unsigned int query_flags) { @@ -270,7 +270,7 @@ void kernfs_evict_inode(struct inode *inode) kernfs_put(kn); } -int kernfs_iop_permission(struct mnt_idmap *idmap, +int kernfs_iop_permission(const struct mnt_idmap *idmap, struct inode *inode, int mask) { struct kernfs_node *kn; @@ -342,7 +342,7 @@ static int kernfs_vfs_xattr_get(const struct xattr_handler *handler, } static int kernfs_vfs_xattr_set(const struct xattr_handler *handler, - struct mnt_idmap *idmap, + const struct mnt_idmap *idmap, struct dentry *unused, struct inode *inode, const char *suffix, const void *value, size_t size, int flags) @@ -354,7 +354,7 @@ static int kernfs_vfs_xattr_set(const struct xattr_handler *handler, } static int kernfs_vfs_user_xattr_set(const struct xattr_handler *handler, - struct mnt_idmap *idmap, + const struct mnt_idmap *idmap, struct dentry *unused, struct inode *inode, const char *suffix, const void *value, size_t size, int flags) diff --git a/fs/kernfs/kernfs-internal.h b/fs/kernfs/kernfs-internal.h index aa784b540b36..f1e93b09e27d 100644 --- a/fs/kernfs/kernfs-internal.h +++ b/fs/kernfs/kernfs-internal.h @@ -117,7 +117,14 @@ static inline bool kernfs_rename_is_locked(const struct kernfs_node *kn) static inline const char *kernfs_rcu_name(const struct kernfs_node *kn) { - return rcu_dereference_check(kn->name, kernfs_root_is_locked(kn)); + /* + * Like kernfs_node::__parent below, the name is only replaced under + * both kernfs_root::kernfs_rwsem and kernfs_root::kernfs_rename_lock, + * so either one keeps it, and the string it points at, stable. + */ + return rcu_dereference_check(kn->name, + kernfs_root_is_locked(kn) || + kernfs_rename_is_locked(kn)); } static inline struct kernfs_node *kernfs_parent(const struct kernfs_node *kn) @@ -147,20 +154,19 @@ static inline struct kernfs_node *kernfs_dentry_node(struct dentry *dentry) static inline void kernfs_set_rev(struct kernfs_node *parent, struct dentry *dentry) { - dentry->d_time = parent->dir.rev; + WRITE_ONCE(dentry->d_time, READ_ONCE(parent->dir.rev)); } static inline void kernfs_inc_rev(struct kernfs_node *parent) { - parent->dir.rev++; + lockdep_assert_held_write(&parent->dir.root->kernfs_rwsem); + WRITE_ONCE(parent->dir.rev, parent->dir.rev + 1); } static inline bool kernfs_dir_changed(struct kernfs_node *parent, struct dentry *dentry) { - if (parent->dir.rev != dentry->d_time) - return true; - return false; + return READ_ONCE(parent->dir.rev) != READ_ONCE(dentry->d_time); } extern const struct super_operations kernfs_sops; @@ -171,11 +177,11 @@ extern struct kmem_cache *kernfs_node_cache, *kernfs_iattrs_cache; */ extern const struct xattr_handler * const kernfs_xattr_handlers[]; void kernfs_evict_inode(struct inode *inode); -int kernfs_iop_permission(struct mnt_idmap *idmap, +int kernfs_iop_permission(const struct mnt_idmap *idmap, struct inode *inode, int mask); -int kernfs_iop_setattr(struct mnt_idmap *idmap, struct dentry *dentry, +int kernfs_iop_setattr(const struct mnt_idmap *idmap, struct dentry *dentry, struct iattr *iattr); -int kernfs_iop_getattr(struct mnt_idmap *idmap, +int kernfs_iop_getattr(const struct mnt_idmap *idmap, const struct path *path, struct kstat *stat, u32 request_mask, unsigned int query_flags); ssize_t kernfs_iop_listxattr(struct dentry *dentry, char *buf, size_t size); diff --git a/fs/kernfs/mount.c b/fs/kernfs/mount.c index a57399021c8b..a0be784bfb06 100644 --- a/fs/kernfs/mount.c +++ b/fs/kernfs/mount.c @@ -124,22 +124,32 @@ static struct dentry *__kernfs_fh_to_dentry(struct super_block *sb, return NULL; } - kn = kernfs_find_and_get_node_by_id(info->root, id); - if (!kn) - return ERR_PTR(-ESTALE); + /* + * Hold kernfs_rwsem across the lookup as well as kernfs_get_inode(). + * __kernfs_remove() deactivates the subtree and clears i_nlink on its + * inodes under the write lock, so under the read lock either + * kernfs_find_and_get_node_by_id() refuses the node, or the inode is + * in the inode hash before the ilookup() pass goes looking for it. + */ + scoped_guard(rwsem_read, &info->root->kernfs_rwsem) { + kn = kernfs_find_and_get_node_by_id(info->root, id); + if (!kn) + return ERR_PTR(-ESTALE); - if (get_parent) { - struct kernfs_node *parent; + if (get_parent) { + struct kernfs_node *parent; - parent = kernfs_get_parent(kn); + parent = kernfs_get_parent(kn); + kernfs_put(kn); + kn = parent; + if (!kn) + return ERR_PTR(-ESTALE); + } + + inode = kernfs_get_inode(sb, kn); kernfs_put(kn); - kn = parent; - if (!kn) - return ERR_PTR(-ESTALE); } - inode = kernfs_get_inode(sb, kn); - kernfs_put(kn); return d_obtain_alias(inode); } diff --git a/fs/kernfs/symlink.c b/fs/kernfs/symlink.c index 90e2b3221b83..3e53105d3abf 100644 --- a/fs/kernfs/symlink.c +++ b/fs/kernfs/symlink.c @@ -31,9 +31,20 @@ struct kernfs_node *kernfs_create_link(struct kernfs_node *parent, kuid_t uid = GLOBAL_ROOT_UID; kgid_t gid = GLOBAL_ROOT_GID; - if (target->iattr) { - uid = target->iattr->ia_uid; - gid = target->iattr->ia_gid; + /* + * A symlink takes its owner from its target, so both fields have to + * come from the same moment: read them under kernfs_iattr_rwsem, or + * a chown of the target racing this could leave the link with the + * old uid and the new gid. The section ends before kernfs_add_one() + * takes kernfs_rwsem. + */ + scoped_guard(rwsem_read, &kernfs_root(target)->kernfs_iattr_rwsem) { + struct kernfs_iattrs *attrs = READ_ONCE(target->iattr); + + if (attrs) { + uid = attrs->ia_uid; + gid = attrs->ia_gid; + } } kn = kernfs_new_node(parent, name, S_IFLNK|0777, uid, gid, KERNFS_LINK); diff --git a/fs/libfs.c b/fs/libfs.c index 27d7dc16fcb0..8e2cc627bc7f 100644 --- a/fs/libfs.c +++ b/fs/libfs.c @@ -29,7 +29,7 @@ #include "internal.h" -int simple_getattr(struct mnt_idmap *idmap, const struct path *path, +int simple_getattr(const struct mnt_idmap *idmap, const struct path *path, struct kstat *stat, u32 request_mask, unsigned int query_flags) { @@ -867,7 +867,7 @@ int simple_rename_exchange(struct inode *old_dir, struct dentry *old_dentry, } EXPORT_SYMBOL_GPL(simple_rename_exchange); -int simple_rename(struct mnt_idmap *idmap, struct inode *old_dir, +int simple_rename(const struct mnt_idmap *idmap, struct inode *old_dir, struct dentry *old_dentry, struct inode *new_dir, struct dentry *new_dentry, unsigned int flags) { @@ -913,7 +913,7 @@ EXPORT_SYMBOL(simple_rename); * on simple regular filesystems. Anything that needs to change on-disk * or wire state on size changes needs its own setattr method. */ -int simple_setattr(struct mnt_idmap *idmap, struct dentry *dentry, +int simple_setattr(const struct mnt_idmap *idmap, struct dentry *dentry, struct iattr *iattr) { struct inode *inode = d_inode(dentry); @@ -1727,7 +1727,7 @@ static struct dentry *empty_dir_lookup(struct inode *dir, struct dentry *dentry, return ERR_PTR(-ENOENT); } -static int empty_dir_setattr(struct mnt_idmap *idmap, +static int empty_dir_setattr(const struct mnt_idmap *idmap, struct dentry *dentry, struct iattr *attr) { return -EPERM; diff --git a/fs/lockd/svc.c b/fs/lockd/svc.c index ee90e743064a..f0e1a58c9106 100644 --- a/fs/lockd/svc.c +++ b/fs/lockd/svc.c @@ -36,7 +36,6 @@ #include <net/ip.h> #include <net/addrconf.h> #include <net/ipv6.h> -#include <linux/nfs.h> #include "lockd.h" #include "netns.h" diff --git a/fs/lockd/svclock.c b/fs/lockd/svclock.c index e628b5d35507..495eacb3264f 100644 --- a/fs/lockd/svclock.c +++ b/fs/lockd/svclock.c @@ -295,14 +295,17 @@ restart: list_for_each_entry_safe(block, next, &file->f_blocks, b_flist) { if (!match(block->b_host, host)) continue; - /* Do not destroy blocks that are not on - * the global retry list - why? */ + /* + * nlmsvc_retry_blocked() holds f_mutex while the block + * is off nlm_blocked, so a block off the list here has + * been retired. + */ if (list_empty(&block->b_list)) continue; kref_get(&block->b_count); spin_unlock(&nlm_blocked_lock); - mutex_unlock(&file->f_mutex); nlmsvc_unlink_block(block); + mutex_unlock(&file->f_mutex); nlmsvc_release_block(block); goto restart; } @@ -1012,6 +1015,8 @@ nlmsvc_retry_blocked(struct svc_rqst *rqstp) { unsigned long timeout = MAX_SCHEDULE_TIMEOUT; struct nlm_block *block; + struct nlm_file *file; + bool due; spin_lock(&nlm_blocked_lock); while (!list_empty(&nlm_blocked) && !svc_thread_should_stop(rqstp)) { @@ -1023,8 +1028,28 @@ nlmsvc_retry_blocked(struct svc_rqst *rqstp) timeout = block->b_when - jiffies; break; } + kref_get(&block->b_count); spin_unlock(&nlm_blocked_lock); + /* + * Hold f_mutex so nlmsvc_traverse_blocks() cannot scan + * the file while the retry has the block off nlm_blocked. + */ + file = block->b_file; + mutex_lock(&file->f_mutex); + spin_lock(&nlm_blocked_lock); + due = !list_empty(&block->b_list) && + block->b_when != NLM_NEVER && + !time_after(block->b_when, jiffies); + spin_unlock(&nlm_blocked_lock); + + if (!due) { + mutex_unlock(&file->f_mutex); + nlmsvc_release_block(block); + spin_lock(&nlm_blocked_lock); + continue; + } + dprintk("nlmsvc_retry_blocked(%p, when=%ld)\n", block, block->b_when); if (block->b_flags & B_QUEUED) { @@ -1033,6 +1058,8 @@ nlmsvc_retry_blocked(struct svc_rqst *rqstp) retry_deferred_block(block); } else nlmsvc_grant_blocked(block); + mutex_unlock(&file->f_mutex); + nlmsvc_release_block(block); spin_lock(&nlm_blocked_lock); } spin_unlock(&nlm_blocked_lock); diff --git a/fs/lockd/trace.h b/fs/lockd/trace.h index a11d04e8c835..1f79955ea0f5 100644 --- a/fs/lockd/trace.h +++ b/fs/lockd/trace.h @@ -7,7 +7,6 @@ #include <linux/tracepoint.h> #include <linux/crc32.h> -#include <linux/nfs.h> #include "lockd.h" diff --git a/fs/lockd/xdr.h b/fs/lockd/xdr.h index a1126cca98c6..56b9796aa39d 100644 --- a/fs/lockd/xdr.h +++ b/fs/lockd/xdr.h @@ -10,7 +10,7 @@ #include <linux/fs.h> #include <linux/filelock.h> -#include <linux/nfs.h> +#include <linux/nfs_fh.h> #include <linux/sunrpc/xdr.h> #define SM_MAXSTRLEN 1024 diff --git a/fs/minix/file.c b/fs/minix/file.c index 02aabbdb5dea..c0929fc38fbc 100644 --- a/fs/minix/file.c +++ b/fs/minix/file.c @@ -23,7 +23,7 @@ const struct file_operations minix_file_operations = { .splice_read = filemap_splice_read, }; -static int minix_setattr(struct mnt_idmap *idmap, +static int minix_setattr(const struct mnt_idmap *idmap, struct dentry *dentry, struct iattr *attr) { struct inode *inode = d_inode(dentry); diff --git a/fs/minix/inode.c b/fs/minix/inode.c index daf83e4ff25c..670179173645 100644 --- a/fs/minix/inode.c +++ b/fs/minix/inode.c @@ -724,7 +724,7 @@ out: return err; } -int minix_getattr(struct mnt_idmap *idmap, const struct path *path, +int minix_getattr(const struct mnt_idmap *idmap, const struct path *path, struct kstat *stat, u32 request_mask, unsigned int flags) { struct super_block *sb = path->dentry->d_sb; diff --git a/fs/minix/minix.h b/fs/minix/minix.h index 78722ce22e1e..db92cf9e0e1b 100644 --- a/fs/minix/minix.h +++ b/fs/minix/minix.h @@ -55,7 +55,7 @@ unsigned long minix_count_free_inodes(struct super_block *sb); int minix_new_block(struct inode *inode); void minix_free_block(struct inode *inode, unsigned long block); unsigned long minix_count_free_blocks(struct super_block *sb); -int minix_getattr(struct mnt_idmap *, const struct path *, +int minix_getattr(const struct mnt_idmap *, const struct path *, struct kstat *, u32, unsigned int); int minix_prepare_chunk(struct folio *folio, loff_t pos, unsigned len); struct mapping_metadata_bhs *minix_get_metadata_bhs(struct inode *inode); diff --git a/fs/minix/namei.c b/fs/minix/namei.c index 5525ba367ed7..f450b11b9860 100644 --- a/fs/minix/namei.c +++ b/fs/minix/namei.c @@ -33,7 +33,7 @@ static struct dentry *minix_lookup(struct inode * dir, struct dentry *dentry, un return d_splice_alias(inode, dentry); } -static int minix_mknod(struct mnt_idmap *idmap, struct inode *dir, +static int minix_mknod(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *dentry, umode_t mode, dev_t rdev) { struct inode *inode; @@ -50,7 +50,7 @@ static int minix_mknod(struct mnt_idmap *idmap, struct inode *dir, return add_nondir(dentry, inode); } -static int minix_tmpfile(struct mnt_idmap *idmap, struct inode *dir, +static int minix_tmpfile(const struct mnt_idmap *idmap, struct inode *dir, struct file *file, umode_t mode) { struct inode *inode = minix_new_inode(dir, mode); @@ -63,13 +63,13 @@ static int minix_tmpfile(struct mnt_idmap *idmap, struct inode *dir, return finish_open_simple(file, 0); } -static int minix_create(struct mnt_idmap *idmap, struct inode *dir, +static int minix_create(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *dentry, umode_t mode) { return minix_mknod(&nop_mnt_idmap, dir, dentry, mode, 0); } -static int minix_symlink(struct mnt_idmap *idmap, struct inode *dir, +static int minix_symlink(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *dentry, const char *symname) { int i = strlen(symname)+1; @@ -104,7 +104,7 @@ static int minix_link(struct dentry * old_dentry, struct inode * dir, return add_nondir(dentry, inode); } -static struct dentry *minix_mkdir(struct mnt_idmap *idmap, struct inode *dir, +static struct dentry *minix_mkdir(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *dentry, umode_t mode) { struct inode * inode; @@ -187,7 +187,7 @@ out: return err; } -static int minix_rename(struct mnt_idmap *idmap, +static int minix_rename(const struct mnt_idmap *idmap, struct inode *old_dir, struct dentry *old_dentry, struct inode *new_dir, struct dentry *new_dentry, unsigned int flags) diff --git a/fs/mnt_idmapping.c b/fs/mnt_idmapping.c index cb61fbdb52e9..bed57094cef0 100644 --- a/fs/mnt_idmapping.c +++ b/fs/mnt_idmapping.c @@ -28,7 +28,7 @@ struct mnt_idmap { * mapping. This means that {g,u}id 0 is mapped to {g,u}id 0, {g,u}id 1 is * mapped to {g,u}id 1, [...], {g,u}id 1000 to {g,u}id 1000, [...]. */ -struct mnt_idmap nop_mnt_idmap = { +const struct mnt_idmap nop_mnt_idmap = { .count = REFCOUNT_INIT(1), }; EXPORT_SYMBOL_GPL(nop_mnt_idmap); @@ -37,7 +37,7 @@ EXPORT_SYMBOL_GPL(nop_mnt_idmap); * Carries the invalid idmapping of a full 0-4294967295 {g,u}id range. * This means that all {g,u}ids are mapped to INVALID_VFS{G,U}ID. */ -struct mnt_idmap invalid_mnt_idmap = { +const struct mnt_idmap invalid_mnt_idmap = { .count = REFCOUNT_INIT(1), }; EXPORT_SYMBOL_GPL(invalid_mnt_idmap); @@ -77,7 +77,7 @@ static inline bool initial_idmapping(const struct user_namespace *ns) * returned. */ -vfsuid_t make_vfsuid(struct mnt_idmap *idmap, +vfsuid_t make_vfsuid(const struct mnt_idmap *idmap, struct user_namespace *fs_userns, kuid_t kuid) { @@ -117,7 +117,7 @@ EXPORT_SYMBOL_GPL(make_vfsuid); * If @kgid has no mapping in either @idmap or @fs_userns INVALID_GID is * returned. */ -vfsgid_t make_vfsgid(struct mnt_idmap *idmap, +vfsgid_t make_vfsgid(const struct mnt_idmap *idmap, struct user_namespace *fs_userns, kgid_t kgid) { gid_t gid; @@ -147,7 +147,7 @@ EXPORT_SYMBOL_GPL(make_vfsgid); * * Return: @vfsuid mapped into the filesystem idmapping */ -kuid_t from_vfsuid(struct mnt_idmap *idmap, +kuid_t from_vfsuid(const struct mnt_idmap *idmap, struct user_namespace *fs_userns, vfsuid_t vfsuid) { uid_t uid; @@ -176,7 +176,7 @@ EXPORT_SYMBOL_GPL(from_vfsuid); * * Return: @vfsgid mapped into the filesystem idmapping */ -kgid_t from_vfsgid(struct mnt_idmap *idmap, +kgid_t from_vfsgid(const struct mnt_idmap *idmap, struct user_namespace *fs_userns, vfsgid_t vfsgid) { gid_t gid; @@ -312,10 +312,12 @@ struct mnt_idmap *alloc_mnt_idmap(struct user_namespace *mnt_userns) * * Return: @idmap with reference count bumped if @not_mnt_idmap isn't passed. */ -struct mnt_idmap *mnt_idmap_get(struct mnt_idmap *idmap) +const struct mnt_idmap *mnt_idmap_get(const struct mnt_idmap *idmap) { + struct mnt_idmap *nonconst_idmap = (struct mnt_idmap *)idmap; + if (idmap != &nop_mnt_idmap && idmap != &invalid_mnt_idmap) - refcount_inc(&idmap->count); + refcount_inc(&nonconst_idmap->count); return idmap; } @@ -328,17 +330,20 @@ EXPORT_SYMBOL_GPL(mnt_idmap_get); * If this is a non-initial idmapping, put the reference count when a mount is * released and free it if we're the last user. */ -void mnt_idmap_put(struct mnt_idmap *idmap) +void mnt_idmap_put(const struct mnt_idmap *idmap) { + struct mnt_idmap *nonconst_idmap = (struct mnt_idmap *)idmap; + if (idmap != &nop_mnt_idmap && idmap != &invalid_mnt_idmap && - refcount_dec_and_test(&idmap->count)) - free_mnt_idmap(idmap); + refcount_dec_and_test(&nonconst_idmap->count)) + free_mnt_idmap(nonconst_idmap); } EXPORT_SYMBOL_GPL(mnt_idmap_put); -int statmount_mnt_idmap(struct mnt_idmap *idmap, struct seq_file *seq, bool uid_map) +int statmount_mnt_idmap(const struct mnt_idmap *idmap, struct seq_file *seq, bool uid_map) { - struct uid_gid_map *map, *map_up; + const struct uid_gid_map *map; + struct uid_gid_map *map_up; u32 idx, nr_mappings; if (!is_valid_mnt_idmap(idmap)) @@ -358,7 +363,7 @@ int statmount_mnt_idmap(struct mnt_idmap *idmap, struct seq_file *seq, bool uid_ for (idx = 0, nr_mappings = 0; idx < map->nr_extents; idx++) { uid_t lower; - struct uid_gid_extent *extent; + const struct uid_gid_extent *extent; if (map->nr_extents <= UID_GID_MAP_MAX_BASE_EXTENTS) extent = &map->extent[idx]; diff --git a/fs/mount.h b/fs/mount.h index 94fcc306d21e..85f136786bbc 100644 --- a/fs/mount.h +++ b/fs/mount.h @@ -33,7 +33,8 @@ struct mnt_namespace { } __randomize_layout; struct mnt_pcp { - int mnt_count; + unsigned int mnt_gets; + unsigned int mnt_puts; int mnt_writers; }; diff --git a/fs/namei.c b/fs/namei.c index 20a6534ea3ef..59c8a669081a 100644 --- a/fs/namei.c +++ b/fs/namei.c @@ -371,7 +371,7 @@ struct filename *complete_getname(struct delayed_filename *v) * On non-idmapped mounts or if permission checking is to be performed on the * raw inode simply pass @nop_mnt_idmap. */ -static int check_acl(struct mnt_idmap *idmap, +static int check_acl(const struct mnt_idmap *idmap, struct inode *inode, int mask) { #ifdef CONFIG_FS_POSIX_ACL @@ -435,7 +435,7 @@ static inline bool no_acl_inode(struct inode *inode) * On non-idmapped mounts or if permission checking is to be performed on the * raw inode simply pass @nop_mnt_idmap. */ -static int acl_permission_check(struct mnt_idmap *idmap, +static int acl_permission_check(const struct mnt_idmap *idmap, struct inode *inode, int mask) { unsigned int mode = inode->i_mode; @@ -518,7 +518,7 @@ static int acl_permission_check(struct mnt_idmap *idmap, * On non-idmapped mounts or if permission checking is to be performed on the * raw inode simply pass @nop_mnt_idmap. */ -int generic_permission(struct mnt_idmap *idmap, struct inode *inode, +int generic_permission(const struct mnt_idmap *idmap, struct inode *inode, int mask) { int ret; @@ -575,7 +575,7 @@ EXPORT_SYMBOL(generic_permission); * flag in inode->i_opflags, that says "this has not special * permission function, use the fast case". */ -static inline int do_inode_permission(struct mnt_idmap *idmap, +static inline int do_inode_permission(const struct mnt_idmap *idmap, struct inode *inode, int mask) { if (unlikely(!(inode->i_opflags & IOP_FASTPERM))) { @@ -625,7 +625,7 @@ static int sb_permission(struct super_block *sb, struct inode *inode, int mask) * * When checking for MAY_APPEND, MAY_WRITE must also be set in @mask. */ -int inode_permission(struct mnt_idmap *idmap, +int inode_permission(const struct mnt_idmap *idmap, struct inode *inode, int mask) { int retval; @@ -680,7 +680,7 @@ EXPORT_SYMBOL(inode_permission); * on IOP_FASTPERM can still get the optimization if they set IOP_FASTPERM_MAY_EXEC * on their directory inodes. */ -static __always_inline int lookup_inode_permission_may_exec(struct mnt_idmap *idmap, +static __always_inline int lookup_inode_permission_may_exec(const struct mnt_idmap *idmap, struct inode *inode, int mask) { /* Lookup already checked this to return -ENOTDIR */ @@ -1273,7 +1273,7 @@ fs_initcall(init_fs_namei_sysctls); */ static inline int may_follow_link(struct nameidata *nd, const struct inode *inode) { - struct mnt_idmap *idmap; + const struct mnt_idmap *idmap; vfsuid_t vfsuid; if (!sysctl_protected_symlinks) @@ -1314,7 +1314,7 @@ static inline int may_follow_link(struct nameidata *nd, const struct inode *inod * * Otherwise returns true. */ -static bool safe_hardlink_source(struct mnt_idmap *idmap, +static bool safe_hardlink_source(const struct mnt_idmap *idmap, struct inode *inode) { umode_t mode = inode->i_mode; @@ -1357,7 +1357,7 @@ static bool safe_hardlink_source(struct mnt_idmap *idmap, * * Returns 0 if successful, -ve on error. */ -int may_linkat(struct mnt_idmap *idmap, const struct path *link) +int may_linkat(const struct mnt_idmap *idmap, const struct path *link) { struct inode *inode = link->dentry->d_inode; @@ -1407,7 +1407,7 @@ int may_linkat(struct mnt_idmap *idmap, const struct path *link) * * Returns 0 if the open is allowed, -ve on error. */ -static int may_create_in_sticky(struct mnt_idmap *idmap, struct nameidata *nd, +static int may_create_in_sticky(const struct mnt_idmap *idmap, struct nameidata *nd, struct inode *const inode) { umode_t dir_mode = nd->dir_mode; @@ -1933,7 +1933,7 @@ static noinline struct dentry *lookup_slow(const struct qstr *name, struct inode *inode = dir->d_inode; struct dentry *res; inode_lock_shared(inode); - res = __lookup_slow(name, dir, flags); + res = __lookup_slow(name, dir, flags | LOOKUP_SHARED); inode_unlock_shared(inode); return res; } @@ -1947,12 +1947,12 @@ static struct dentry *lookup_slow_killable(const struct qstr *name, if (inode_lock_shared_killable(inode)) return ERR_PTR(-EINTR); - res = __lookup_slow(name, dir, flags); + res = __lookup_slow(name, dir, flags | LOOKUP_SHARED); inode_unlock_shared(inode); return res; } -static inline int may_lookup(struct mnt_idmap *idmap, +static inline int may_lookup(const struct mnt_idmap *idmap, struct nameidata *restrict nd) { int err, mask; @@ -2596,7 +2596,7 @@ static int link_path_walk(const char *name, struct nameidata *nd) /* At this point we know we have a real path component. */ for(;;) { - struct mnt_idmap *idmap; + const struct mnt_idmap *idmap; const char *link; unsigned long lastword; @@ -2946,8 +2946,8 @@ struct dentry *start_dirop(struct dentry *parent, struct qstr *name, * end_dirop - signal completion of a dirop * @de: the dentry which was returned by start_dirop or similar. * - * If the de is an error, nothing happens. Otherwise any lock taken to - * protect the dentry is dropped and the dentry itself is release (dput()). + * If the @de is an error, nothing happens. Otherwise any lock taken to + * protect the dentry is dropped and the dentry itself is released (dput()). */ void end_dirop(struct dentry *de) { @@ -3111,7 +3111,7 @@ int lookup_noperm_common(struct qstr *qname, struct dentry *base) return 0; } -static int lookup_one_common(struct mnt_idmap *idmap, +static int lookup_one_common(const struct mnt_idmap *idmap, struct qstr *qname, struct dentry *base) { int err; @@ -3190,7 +3190,7 @@ EXPORT_SYMBOL(lookup_noperm); * * The caller must hold base->i_rwsem. */ -struct dentry *lookup_one(struct mnt_idmap *idmap, struct qstr *name, +struct dentry *lookup_one(const struct mnt_idmap *idmap, struct qstr *name, struct dentry *base) { struct dentry *dentry; @@ -3210,7 +3210,7 @@ EXPORT_SYMBOL(lookup_one); /** * lookup_one_unlocked - lookup single pathname component * @idmap: idmap of the mount the lookup is performed from - * @name: qstr olding pathname component to lookup + * @name: qstr holding pathname component to lookup * @base: base directory to lookup from * * This can be used for in-kernel filesystem clients such as file servers. @@ -3223,7 +3223,7 @@ EXPORT_SYMBOL(lookup_one); * - ERR_PTR(-ENOENT) if parent has been removed, or * - ERR_PTR(-EACCES) if parent directory is not searchable. */ -struct dentry *lookup_one_unlocked(struct mnt_idmap *idmap, struct qstr *name, +struct dentry *lookup_one_unlocked(const struct mnt_idmap *idmap, struct qstr *name, struct dentry *base) { int err; @@ -3243,7 +3243,7 @@ EXPORT_SYMBOL(lookup_one_unlocked); /** * lookup_one_positive_killable - lookup single pathname component * @idmap: idmap of the mount the lookup is performed from - * @name: qstr olding pathname component to lookup + * @name: qstr holding pathname component to lookup * @base: base directory to lookup from * * This helper will yield ERR_PTR(-ENOENT) on negatives. The helper returns @@ -3259,11 +3259,11 @@ EXPORT_SYMBOL(lookup_one_unlocked); * the i_rwsem itself if necessary. If a fatal signal is pending or * delivered, it will return %-EINTR if the lock is needed. * - * Returns: A dentry, possibly negative, or + * Returns: A positive dentry, or * - same errors as lookup_one_unlocked() or * - ERR_PTR(-EINTR) if a fatal signal is pending. */ -struct dentry *lookup_one_positive_killable(struct mnt_idmap *idmap, +struct dentry *lookup_one_positive_killable(const struct mnt_idmap *idmap, struct qstr *name, struct dentry *base) { @@ -3306,7 +3306,7 @@ EXPORT_SYMBOL(lookup_one_positive_killable); * - ERR_PTR(-ENOENT) if the name could not be found, or * - same errors as lookup_one_unlocked(). */ -struct dentry *lookup_one_positive_unlocked(struct mnt_idmap *idmap, +struct dentry *lookup_one_positive_unlocked(const struct mnt_idmap *idmap, struct qstr *name, struct dentry *base) { @@ -3381,7 +3381,7 @@ struct dentry *lookup_noperm_positive_unlocked(struct qstr *name, EXPORT_SYMBOL(lookup_noperm_positive_unlocked); /** - * start_creating - prepare to create a given name with permission checking + * start_creating - prepare to access or create a given name with permission checking * @idmap: idmap of the mount * @parent: directory in which to prepare to create the name * @name: the name to be created @@ -3396,7 +3396,7 @@ EXPORT_SYMBOL(lookup_noperm_positive_unlocked); * * Returns: a negative or positive dentry, or an error. */ -struct dentry *start_creating(struct mnt_idmap *idmap, struct dentry *parent, +struct dentry *start_creating(const struct mnt_idmap *idmap, struct dentry *parent, struct qstr *name) { int err = lookup_one_common(idmap, name, parent); @@ -3413,8 +3413,8 @@ EXPORT_SYMBOL(start_creating); * @parent: directory in which to find the name * @name: the name to be removed * - * Locks are taken and a lookup in performed prior to removing - * an object from a directory. Permission checking (MAY_EXEC) is performed + * Locks are taken and a lookup is performed prior to removing an object + * from a directory. Permission checking (MAY_EXEC) is performed * against @idmap. * * If the name doesn't exist, an error is returned. @@ -3423,7 +3423,7 @@ EXPORT_SYMBOL(start_creating); * * Returns: a positive dentry, or an error. */ -struct dentry *start_removing(struct mnt_idmap *idmap, struct dentry *parent, +struct dentry *start_removing(const struct mnt_idmap *idmap, struct dentry *parent, struct qstr *name) { int err = lookup_one_common(idmap, name, parent); @@ -3440,7 +3440,7 @@ EXPORT_SYMBOL(start_removing); * @parent: directory in which to prepare to create the name * @name: the name to be created * - * Locks are taken and a lookup in performed prior to creating + * Locks are taken and a lookup is performed prior to creating * an object in a directory. Permission checking (MAY_EXEC) is performed * against @idmap. * @@ -3451,7 +3451,7 @@ EXPORT_SYMBOL(start_removing); * * Returns: a negative or positive dentry, or an error. */ -struct dentry *start_creating_killable(struct mnt_idmap *idmap, +struct dentry *start_creating_killable(const struct mnt_idmap *idmap, struct dentry *parent, struct qstr *name) { @@ -3469,7 +3469,7 @@ EXPORT_SYMBOL(start_creating_killable); * @parent: directory in which to find the name * @name: the name to be removed * - * Locks are taken and a lookup in performed prior to removing + * Locks are taken and a lookup is performed prior to removing * an object from a directory. Permission checking (MAY_EXEC) is performed * against @idmap. * @@ -3482,7 +3482,7 @@ EXPORT_SYMBOL(start_creating_killable); * * Returns: a positive dentry, or an error. */ -struct dentry *start_removing_killable(struct mnt_idmap *idmap, +struct dentry *start_removing_killable(const struct mnt_idmap *idmap, struct dentry *parent, struct qstr *name) { @@ -3499,7 +3499,7 @@ EXPORT_SYMBOL(start_removing_killable); * @parent: directory in which to prepare to create the name * @name: the name to be created * - * Locks are taken and a lookup in performed prior to creating + * Locks are taken and a lookup is performed prior to creating * an object in a directory. * * If the name already exists, a positive dentry is returned. @@ -3522,7 +3522,7 @@ EXPORT_SYMBOL(start_creating_noperm); * @parent: directory in which to find the name * @name: the name to be removed * - * Locks are taken and a lookup in performed prior to removing + * Locks are taken and a lookup is performed prior to removing * an object from a directory. * * If the name doesn't exist, an error is returned. @@ -3543,11 +3543,11 @@ struct dentry *start_removing_noperm(struct dentry *parent, EXPORT_SYMBOL(start_removing_noperm); /** - * start_creating_dentry - prepare to create a given dentry - * @parent: directory from which dentry should be removed - * @child: the dentry to be removed + * start_creating_dentry - prepare to access or create a given dentry + * @parent: directory of dentry + * @child: the dentry to be prepared * - * A lock is taken to protect the dentry again other dirops and + * A lock is taken to protect the dentry against other dirops and * the validity of the dentry is checked: correct parent and still hashed. * * If the dentry is valid and negative a reference is taken and @@ -3580,7 +3580,7 @@ EXPORT_SYMBOL(start_creating_dentry); * @parent: directory from which dentry should be removed * @child: the dentry to be removed * - * A lock is taken to protect the dentry again other dirops and + * A lock is taken to protect the dentry against other dirops and * the validity of the dentry is checked: correct parent and still hashed. * * If the dentry is valid and positive, a reference is taken and @@ -3642,7 +3642,7 @@ int user_path_at(int dfd, const char __user *name, unsigned flags, } EXPORT_SYMBOL(user_path_at); -int __check_sticky(struct mnt_idmap *idmap, struct inode *dir, +int __check_sticky(const struct mnt_idmap *idmap, struct inode *dir, struct inode *inode) { kuid_t fsuid = current_fsuid(); @@ -3675,7 +3675,7 @@ EXPORT_SYMBOL(__check_sticky); * 11. We don't allow removal of NFS sillyrenamed files; it's handled by * nfs_async_unlink(). */ -int may_delete_dentry(struct mnt_idmap *idmap, struct inode *dir, +int may_delete_dentry(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *victim, bool isdir) { struct inode *inode = d_backing_inode(victim); @@ -3728,7 +3728,7 @@ EXPORT_SYMBOL(may_delete_dentry); * 4. We should have write and exec permissions on dir * 5. We can't do it if dir is immutable (done in permission()) */ -int may_create_dentry(struct mnt_idmap *idmap, +int may_create_dentry(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *child) { audit_inode_child(dir, child, AUDIT_TYPE_CHILD_CREATE); @@ -4142,7 +4142,7 @@ EXPORT_SYMBOL(end_renaming); * * Returns: mode to be passed to the filesystem */ -static inline umode_t vfs_prepare_mode(struct mnt_idmap *idmap, +static inline umode_t vfs_prepare_mode(const struct mnt_idmap *idmap, const struct inode *dir, umode_t mode, umode_t mask_perms, umode_t type) { @@ -4174,7 +4174,7 @@ static inline umode_t vfs_prepare_mode(struct mnt_idmap *idmap, * On non-idmapped mounts or if permission checking is to be performed on the * raw inode simply pass @nop_mnt_idmap. */ -int vfs_create(struct mnt_idmap *idmap, struct dentry *dentry, umode_t mode, +int vfs_create(const struct mnt_idmap *idmap, struct dentry *dentry, umode_t mode, struct delegated_inode *di) { struct inode *dir = d_inode(dentry->d_parent); @@ -4228,7 +4228,7 @@ bool may_open_dev(const struct path *path) !(path->mnt->mnt_sb->s_iflags & SB_I_NODEV); } -static int may_open(struct mnt_idmap *idmap, const struct path *path, +static int may_open(const struct mnt_idmap *idmap, const struct path *path, int acc_mode, int flag) { struct dentry *dentry = path->dentry; @@ -4287,7 +4287,7 @@ static int may_open(struct mnt_idmap *idmap, const struct path *path, return 0; } -static int handle_truncate(struct mnt_idmap *idmap, struct file *filp) +static int handle_truncate(const struct mnt_idmap *idmap, struct file *filp) { const struct path *path = &filp->f_path; struct inode *inode = path->dentry->d_inode; @@ -4312,7 +4312,7 @@ static inline int open_to_namei_flags(int flag) return flag; } -static int may_o_create(struct mnt_idmap *idmap, +static int may_o_create(const struct mnt_idmap *idmap, const struct path *dir, struct dentry *dentry, umode_t mode) { @@ -4432,7 +4432,7 @@ static struct dentry *lookup_open(struct nameidata *nd, struct file *file, const struct open_flags *op) { struct delegated_inode delegated_inode = { }; - struct mnt_idmap *idmap; + const struct mnt_idmap *idmap; struct dentry *dir = nd->path.dentry; struct inode *dir_inode = dir->d_inode; int open_flag; @@ -4440,12 +4440,14 @@ static struct dentry *lookup_open(struct nameidata *nd, struct file *file, int error, create_error; umode_t mode; bool got_write; + unsigned int shared_flag; retry: open_flag = op->open_flag; got_write = false; mode = op->mode; create_error = 0; + shared_flag = (open_flag & O_CREAT) ? 0 : LOOKUP_SHARED; if (open_flag & (O_CREAT | O_TRUNC | O_WRONLY | O_RDWR)) { got_write = !mnt_want_write(nd->path.mnt); @@ -4454,10 +4456,10 @@ retry: * a different error; we'll be dropping this one anyway. */ } - if (open_flag & O_CREAT) - inode_lock(dir_inode); - else + if (shared_flag) inode_lock_shared(dir_inode); + else + inode_lock(dir_inode); if (unlikely(IS_DEADDIR(dir_inode))) { dentry = ERR_PTR(-ENOENT); @@ -4526,7 +4528,7 @@ retry: if (d_in_lookup(dentry)) { struct dentry *res = dir_inode->i_op->lookup(dir_inode, dentry, - nd->flags); + nd->flags | shared_flag); d_lookup_done(dentry); if (unlikely(res)) { if (IS_ERR(res)) { @@ -4574,10 +4576,10 @@ out: if (file->f_mode & FMODE_OPENED) fsnotify_open(file); } - if ((open_flag & O_CREAT) || create_error) - inode_unlock(dir_inode); - else + if (shared_flag) inode_unlock_shared(dir_inode); + else + inode_unlock(dir_inode); if (got_write) mnt_drop_write(nd->path.mnt); @@ -4789,7 +4791,8 @@ finish_lookup: static int do_open(struct nameidata *nd, struct file *file, const struct open_flags *op) { - struct mnt_idmap *idmap; + struct vfsmount *mnt; + const struct mnt_idmap *idmap; int open_flag = op->open_flag; bool do_truncate; int acc_mode; @@ -4830,11 +4833,17 @@ static int do_open(struct nameidata *nd, error = mnt_want_write(nd->path.mnt); if (error) return error; + /* + * A dedicated reference is needed because after the call to + * vfs_open_consume() we no longer own the reference in nd->path.mnt + * while we need to undo write acess below. + */ + mnt = mntget(nd->path.mnt); do_truncate = true; } error = may_open(idmap, &nd->path, acc_mode, open_flag); if (!error && !(file->f_mode & FMODE_OPENED)) - error = vfs_open(&nd->path, file); + error = vfs_open_consume(&nd->path, file); if (!error) error = security_file_post_open(file, op->acc_mode); if (!error && do_truncate) @@ -4843,8 +4852,10 @@ static int do_open(struct nameidata *nd, WARN_ON(1); error = -EINVAL; } - if (do_truncate) - mnt_drop_write(nd->path.mnt); + if (do_truncate) { + mnt_drop_write(mnt); + mntput(mnt); + } return error; } @@ -4863,7 +4874,7 @@ static int do_open(struct nameidata *nd, * On non-idmapped mounts or if permission checking is to be performed on the * raw inode simply pass @nop_mnt_idmap. */ -int vfs_tmpfile(struct mnt_idmap *idmap, +int vfs_tmpfile(const struct mnt_idmap *idmap, const struct path *parentpath, struct file *file, umode_t mode) { @@ -4921,7 +4932,7 @@ int vfs_tmpfile(struct mnt_idmap *idmap, * hence this is only for kernel internal use, and must not be installed into * file tables or such. */ -struct file *kernel_tmpfile_open(struct mnt_idmap *idmap, +struct file *kernel_tmpfile_open(const struct mnt_idmap *idmap, const struct path *parentpath, umode_t mode, int open_flag, const struct cred *cred) @@ -5169,7 +5180,7 @@ struct file *dentry_create(struct path *path, int flags, umode_t mode, struct dentry *orig_dentry = dentry; struct dentry *dir = dentry->d_parent; struct inode *dir_inode = d_inode(dir); - struct mnt_idmap *idmap; + const struct mnt_idmap *idmap; int error, create_error; file = alloc_empty_file(flags, cred); @@ -5211,6 +5222,8 @@ struct file *dentry_create(struct path *path, int flags, umode_t mode, error = vfs_create(mnt_idmap(path->mnt), path->dentry, mode, NULL); if (!error) error = vfs_open(path, file); + if (!error) + file->f_mode |= FMODE_CREATED; } if (unlikely(error)) return ERR_PTR(error); @@ -5236,7 +5249,7 @@ EXPORT_SYMBOL(dentry_create); * On non-idmapped mounts or if permission checking is to be performed on the * raw inode simply pass @nop_mnt_idmap. */ -int vfs_mknod(struct mnt_idmap *idmap, struct inode *dir, +int vfs_mknod(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *dentry, umode_t mode, dev_t dev, struct delegated_inode *delegated_inode) { @@ -5294,7 +5307,7 @@ int filename_mknodat(int dfd, struct filename *name, umode_t mode, unsigned int dev) { struct delegated_inode di = { }; - struct mnt_idmap *idmap; + const struct mnt_idmap *idmap; struct dentry *dentry; struct path path; int error; @@ -5378,7 +5391,7 @@ SYSCALL_DEFINE3(mknod, const char __user *, filename, umode_t, mode, unsigned, d * * In case of an error the dentry is dput() and an ERR_PTR() is returned. */ -struct dentry *vfs_mkdir(struct mnt_idmap *idmap, struct inode *dir, +struct dentry *vfs_mkdir(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *dentry, umode_t mode, struct delegated_inode *delegated_inode) { @@ -5485,7 +5498,7 @@ SYSCALL_DEFINE2(mkdir, const char __user *, pathname, umode_t, mode) * On non-idmapped mounts or if permission checking is to be performed on the * raw inode simply pass @nop_mnt_idmap. */ -int vfs_rmdir(struct mnt_idmap *idmap, struct inode *dir, +int vfs_rmdir(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *dentry, struct delegated_inode *delegated_inode) { int error = may_delete_dentry(idmap, dir, dentry, true); @@ -5620,7 +5633,7 @@ SYSCALL_DEFINE1(rmdir, const char __user *, pathname) * On non-idmapped mounts or if permission checking is to be performed on the * raw inode simply pass @nop_mnt_idmap. */ -int vfs_unlink(struct mnt_idmap *idmap, struct inode *dir, +int vfs_unlink(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *dentry, struct delegated_inode *delegated_inode) { struct inode *target = dentry->d_inode; @@ -5770,7 +5783,7 @@ SYSCALL_DEFINE1(unlink, const char __user *, pathname) * On non-idmapped mounts or if permission checking is to be performed on the * raw inode simply pass @nop_mnt_idmap. */ -int vfs_symlink(struct mnt_idmap *idmap, struct inode *dir, +int vfs_symlink(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *dentry, const char *oldname, struct delegated_inode *delegated_inode) { @@ -5872,7 +5885,7 @@ SYSCALL_DEFINE2(symlink, const char __user *, oldname, const char __user *, newn * On non-idmapped mounts or if permission checking is to be performed on the * raw inode simply pass @nop_mnt_idmap. */ -int vfs_link(struct dentry *old_dentry, struct mnt_idmap *idmap, +int vfs_link(struct dentry *old_dentry, const struct mnt_idmap *idmap, struct inode *dir, struct dentry *new_dentry, struct delegated_inode *delegated_inode) { @@ -5949,7 +5962,7 @@ EXPORT_SYMBOL(vfs_link); int filename_linkat(int olddfd, struct filename *old, int newdfd, struct filename *new, int flags) { - struct mnt_idmap *idmap; + const struct mnt_idmap *idmap; struct dentry *new_dentry; struct path old_path, new_path; struct delegated_inode delegated_inode = { }; diff --git a/fs/namespace.c b/fs/namespace.c index 580877e46b1a..973efee4b968 100644 --- a/fs/namespace.c +++ b/fs/namespace.c @@ -109,7 +109,7 @@ struct mount_kattr { unsigned int lookup_flags; enum mount_kattr_flags_t kflags; struct user_namespace *mnt_userns; - struct mnt_idmap *mnt_idmap; + const struct mnt_idmap *mnt_idmap; }; /* /sys/fs */ @@ -249,16 +249,24 @@ void mnt_release_group_id(struct mount *mnt) mnt->mnt_group_id = 0; } -/* - * vfsmount lock must be held for read - */ -static inline void mnt_add_count(struct mount *mnt, int n) +static inline void mnt_inc_count(struct mount *mnt) { #ifdef CONFIG_SMP - this_cpu_add(mnt->mnt_pcp->mnt_count, n); + this_cpu_inc(mnt->mnt_pcp->mnt_gets); #else preempt_disable(); - mnt->mnt_count += n; + mnt->mnt_count++; + preempt_enable(); +#endif +} + +static inline void mnt_dec_count(struct mount *mnt) +{ +#ifdef CONFIG_SMP + this_cpu_inc(mnt->mnt_pcp->mnt_puts); +#else + preempt_disable(); + mnt->mnt_count--; preempt_enable(); #endif } @@ -269,14 +277,17 @@ static inline void mnt_add_count(struct mount *mnt, int n) int mnt_get_count(struct mount *mnt) { #ifdef CONFIG_SMP - int count = 0; + unsigned int gets = 0, puts = 0; int cpu; - for_each_possible_cpu(cpu) { - count += per_cpu_ptr(mnt->mnt_pcp, cpu)->mnt_count; - } + /* puts first, so a put counted here has its get counted below */ + for_each_possible_cpu(cpu) + puts += per_cpu_ptr(mnt->mnt_pcp, cpu)->mnt_puts; + smp_mb(); /* pairs with the smp_wmb() in mntput_no_expire() */ + for_each_possible_cpu(cpu) + gets += per_cpu_ptr(mnt->mnt_pcp, cpu)->mnt_gets; - return count; + return gets - puts; #else return mnt->mnt_count; #endif @@ -305,7 +316,7 @@ static struct mount *alloc_vfsmnt(const char *name) if (!mnt->mnt_pcp) goto out_free_devname; - this_cpu_add(mnt->mnt_pcp->mnt_count, 1); + this_cpu_inc(mnt->mnt_pcp->mnt_gets); #else mnt->mnt_count = 1; mnt->mnt_writers = 0; @@ -746,13 +757,13 @@ int __legitimize_mnt(struct vfsmount *bastard, unsigned seq) if (bastard == NULL) return 0; mnt = real_mount(bastard); - mnt_add_count(mnt, 1); - smp_mb(); // see mntput_no_expire() and do_umount() + mnt_inc_count(mnt); + smp_mb(); /* see mntput_no_expire_slowpath() and do_umount() */ if (likely(!read_seqretry(&mount_lock, seq))) return 0; lock_mount_hash(); if (unlikely(bastard->mnt_flags & (MNT_SYNC_UMOUNT | MNT_DOOMED))) { - mnt_add_count(mnt, -1); + mnt_dec_count(mnt); unlock_mount_hash(); return 1; } @@ -1254,6 +1265,7 @@ static struct mount *clone_mnt(struct mount *old, struct dentry *root, mnt->mnt.mnt_flags = READ_ONCE(old->mnt.mnt_flags) & ~MNT_INTERNAL_FLAGS; + mnt->mnt_t_flags = old->mnt_t_flags & T_UNBINDABLE; if (flag & (CL_SLAVE | CL_PRIVATE)) mnt->mnt_group_id = 0; /* not a peer of original */ @@ -1347,7 +1359,7 @@ static void noinline mntput_no_expire_slowpath(struct mount *mnt) * mount_lock, we'll see their refcount increment here. */ smp_mb(); - mnt_add_count(mnt, -1); + mnt_dec_count(mnt); count = mnt_get_count(mnt); if (count != 0) { WARN_ON(count < 0); @@ -1404,7 +1416,8 @@ static void mntput_no_expire(struct mount *mnt) * non-NULL under rcu_read_lock(), the reference * we are dropping is not the final one. */ - mnt_add_count(mnt, -1); + smp_wmb(); /* pairs with the smp_mb() in mnt_get_count() */ + mnt_dec_count(mnt); rcu_read_unlock(); return; } @@ -1426,7 +1439,7 @@ EXPORT_SYMBOL(mntput); struct vfsmount *mntget(struct vfsmount *mnt) { if (mnt) - mnt_add_count(real_mount(mnt), 1); + mnt_inc_count(real_mount(mnt)); return mnt; } EXPORT_SYMBOL(mntget); @@ -3467,7 +3480,7 @@ static int do_set_group(const struct path *from_path, const struct path *to_path return -EINVAL; /* Setting sharing groups is only allowed on private mounts */ - if (IS_MNT_SHARED(to) || IS_MNT_SLAVE(to)) + if (IS_MNT_SHARED(to) || IS_MNT_SLAVE(to) || IS_MNT_UNBINDABLE(to)) return -EINVAL; /* From should not be private */ @@ -4115,7 +4128,7 @@ int path_mount(const char *dev_name, const struct path *path, if (flags & SB_MANDLOCK) warn_mandlock(); - /* Default to relatime unless overriden */ + /* Default to relatime unless overridden */ if (!(flags & MS_NOATIME)) mnt_flags |= MNT_RELATIME; @@ -4247,8 +4260,6 @@ struct mnt_namespace *copy_mnt_ns(u64 flags, struct mnt_namespace *ns, struct mount *new; int copy_flags; - BUG_ON(!ns); - if (likely(!(flags & CLONE_NEWNS))) { get_mnt_ns(ns); return ns; @@ -4544,16 +4555,16 @@ SYSCALL_DEFINE3(fsmount, int, fs_fd, unsigned int, flags, FD_PREPARE(fdf, (flags & FSMOUNT_CLOEXEC) ? O_CLOEXEC : 0, dentry_open(&new_path, O_PATH, fc->cred)); - if (fdf.err) { + if (fdf->fd < 0) { dissolve_on_fput(new_path.mnt); - return fdf.err; + return fdf->fd; } /* * Attach to an apparent O_PATH fd with a note that we * need to unmount it, not just simply put it. */ - fd_prepare_file(fdf)->f_mode |= FMODE_NEED_UNMOUNT; + fdf->file->f_mode |= FMODE_NEED_UNMOUNT; return fd_publish(fdf); } @@ -4898,7 +4909,7 @@ static int mount_setattr_prepare(struct mount_kattr *kattr, struct mount *mnt) static void do_idmap_mount(const struct mount_kattr *kattr, struct mount *mnt) { - struct mnt_idmap *old_idmap; + const struct mnt_idmap *old_idmap; if (!kattr->mnt_idmap) return; @@ -4941,7 +4952,7 @@ static int do_mount_setattr(const struct path *path, struct mount_kattr *kattr) return -EINVAL; if (kattr->mnt_userns) { - struct mnt_idmap *mnt_idmap; + const struct mnt_idmap *mnt_idmap; mnt_idmap = alloc_mnt_idmap(kattr->mnt_userns); if (IS_ERR(mnt_idmap)) @@ -5198,12 +5209,12 @@ SYSCALL_DEFINE5(open_tree_attr, int, dfd, const char __user *, filename, return -EINVAL; FD_PREPARE(fdf, flags, vfs_open_tree(dfd, filename, flags)); - if (fdf.err) - return fdf.err; + if (fdf->fd < 0) + return fdf->fd; if (uattr) { struct mount_kattr kattr = {}; - struct file *file = fd_prepare_file(fdf); + struct file *file = fdf->file; int ret; if (flags & OPEN_TREE_CLONE) @@ -5246,7 +5257,7 @@ struct kstatmount { struct statmount __user *buf; size_t bufsize; struct vfsmount *mnt; - struct mnt_idmap *idmap; + const struct mnt_idmap *idmap; u64 mask; struct path root; struct seq_file seq; diff --git a/fs/netfs/Kconfig b/fs/netfs/Kconfig index 7701c037c328..d0e7b0971fa3 100644 --- a/fs/netfs/Kconfig +++ b/fs/netfs/Kconfig @@ -22,6 +22,9 @@ config NETFS_STATS between CPUs. On the other hand, the stats are very useful for debugging purposes. Saying 'Y' here is recommended. +config NETFS_PGPRIV2 + bool + config NETFS_DEBUG bool "Enable dynamic debugging netfslib and FS-Cache" depends on NETFS_SUPPORT diff --git a/fs/netfs/Makefile b/fs/netfs/Makefile index b43188d64bd8..54834cde7e56 100644 --- a/fs/netfs/Makefile +++ b/fs/netfs/Makefile @@ -11,7 +11,6 @@ netfs-y := \ misc.o \ objects.o \ read_collect.o \ - read_pgpriv2.o \ read_retry.o \ read_single.o \ rolling_buffer.o \ @@ -19,6 +18,7 @@ netfs-y := \ write_issue.o \ write_retry.o +netfs-$(CONFIG_NETFS_PGPRIV2) += read_pgpriv2.o netfs-$(CONFIG_NETFS_STATS) += stats.o netfs-$(CONFIG_FSCACHE) += \ diff --git a/fs/netfs/buffered_read.c b/fs/netfs/buffered_read.c index 105194de6e13..e30bde80276a 100644 --- a/fs/netfs/buffered_read.c +++ b/fs/netfs/buffered_read.c @@ -10,9 +10,9 @@ #include "internal.h" static void netfs_cache_expand_readahead(struct netfs_io_request *rreq, - unsigned long long *_start, - unsigned long long *_len, - unsigned long long i_size) + uoff_t *_start, + uoff_t *_len, + uoff_t i_size) { struct netfs_cache_resources *cres = &rreq->cache_resources; @@ -137,21 +137,6 @@ static ssize_t netfs_prepare_read_iterator(struct netfs_io_subrequest *subreq) return subreq->len; } -static enum netfs_io_source netfs_cache_prepare_read(struct netfs_io_request *rreq, - struct netfs_io_subrequest *subreq, - loff_t i_size) -{ - struct netfs_cache_resources *cres = &rreq->cache_resources; - enum netfs_io_source source; - - if (!cres->ops) - return NETFS_DOWNLOAD_FROM_SERVER; - source = cres->ops->prepare_read(subreq, i_size); - trace_netfs_sreq(subreq, netfs_sreq_trace_prepare); - return source; - -} - /* * Issue a read against the cache. * - Eats the caller's ref on subreq. @@ -166,6 +151,19 @@ static void netfs_read_cache_to_pagecache(struct netfs_io_request *rreq, netfs_cache_read_terminated, subreq); } +int netfs_read_query_cache(struct netfs_io_request *rreq, struct fscache_occupancy *occ) +{ + struct netfs_cache_resources *cres = &rreq->cache_resources; + + occ->granularity = PAGE_SIZE; + if (occ->query_from >= occ->query_to) + return 0; + if (!cres->ops) + return 0; + occ->query_from = round_up(occ->query_from, occ->granularity); + return cres->ops->query_occupancy(cres, occ); +} + void netfs_queue_read(struct netfs_io_request *rreq, struct netfs_io_subrequest *subreq) { @@ -242,7 +240,7 @@ static void netfs_mark_copy_to_cache(struct netfs_io_request *rreq, if (overlap > 0 && copy) { folio = folioq_folio(*fq, *slot); - if (unlikely(test_bit(NETFS_RREQ_USE_PGPRIV2, &rreq->flags))) { + if (netfs_using_pgpriv2(rreq)) { if (!folio_test_private_2(folio)) folio_start_private_2(folio); } else { @@ -268,18 +266,113 @@ static void netfs_mark_copy_to_cache(struct netfs_io_request *rreq, */ static void netfs_read_to_pagecache(struct netfs_io_request *rreq) { + struct fscache_occupancy _occ = { + .query_from = rreq->start, + .query_to = rreq->start + rreq->len, + .cached_from[0] = 0, + .cached_to[0] = 0, + .cached_from[1] = ULLONG_MAX, + .cached_to[1] = ULLONG_MAX, + }; + struct fscache_occupancy *occ = &_occ; struct folio_queue *fq = rreq->buffer.tail; - unsigned long long start = rreq->start; unsigned int offset = 0; ssize_t size = rreq->len; + uoff_t start = rreq->start; int ret = 0, slot = 0; do { + int (*prepare_read)(struct netfs_io_subrequest *subreq) = NULL; struct netfs_io_subrequest *subreq; - enum netfs_io_source source = NETFS_SOURCE_UNKNOWN; + enum netfs_io_source source; ssize_t slice; + uoff_t hole_to, cache_to; + size_t len = size; + bool copy = false; + + /* If we don't have any, find out the next couple of data + * extents from the cache, containing of following the + * specified start offset. Holes have to be fetched from the + * server; data regions from the cache. + */ + hole_to = occ->cached_from[0]; + cache_to = occ->cached_to[0]; + if (start >= cache_to) { + /* Extent exhausted; shuffle down. */ + int i; + + for (i = 0; i < ARRAY_SIZE(occ->cached_from) - 1; i++) { + occ->cached_from[i] = occ->cached_from[i + 1]; + occ->cached_to[i] = occ->cached_to[i + 1]; + occ->cached_type[i] = occ->cached_type[i + 1]; + } + occ->cached_from[i] = ULLONG_MAX; + occ->cached_to[i] = ULLONG_MAX; + + if (occ->cached_from[0] != ULLONG_MAX) + continue; + + /* Get new extents */ + ret = netfs_read_query_cache(rreq, occ); + if (ret < 0) + break; + continue; + } - subreq = netfs_alloc_subrequest(rreq); + uoff_t zero_point = netfs_read_zero_point(rreq->inode); + uoff_t zlimit = umin(zero_point, rreq->i_size); + + _debug("rsub %llx %llx-%llx", start, hole_to, cache_to); + + if (start >= hole_to && start < cache_to) { + /* Overlap with a cached region, where the cache may + * record a block of zeroes. + */ + _debug("cached s=%llx c=%llx l=%zx", start, cache_to, size); + len = umin(cache_to - start, size); + len = round_up(len, occ->granularity); + if (occ->cached_type[0] == FSCACHE_EXTENT_ZERO) { + source = NETFS_FILL_WITH_ZEROES; + netfs_stat(&netfs_n_rh_zero); + } else { + source = NETFS_READ_FROM_CACHE; + prepare_read = rreq->cache_resources.ops->prepare_read; + } + } else if (start >= zlimit && size > 0) { + /* If this range lies beyond the zero-point, that part + * can just be cleared locally. + */ + _debug("zero %llx-%llx", start, start + size); + len = size; + source = NETFS_FILL_WITH_ZEROES; + if (rreq->cache_resources.ops) + copy = true; + netfs_stat(&netfs_n_rh_zero); + } else { + /* Read a cache hole from the server. If any part of + * this range lies beyond the zero-point or the EOF, + * that part can just be cleared locally. + */ + uoff_t limit = min3(zlimit, start + size, hole_to); + + _debug("limit %llx %llx", rreq->i_size, zero_point); + _debug("download %llx-%llx", start, start + size); + len = umin(limit - start, ULONG_MAX); + source = NETFS_DOWNLOAD_FROM_SERVER; + prepare_read = rreq->netfs_ops->prepare_read; + if (rreq->cache_resources.ops) + copy = true; + netfs_stat(&netfs_n_rh_download); + } + + if (len == 0) { + pr_err("ZERO-LEN READ: R=%08x l=%zx/%zx s=%llx z=%llx i=%llx", + rreq->debug_id, len, size, + start, zero_point, rreq->i_size); + break; + } + + subreq = netfs_alloc_subrequest(rreq, source); if (!subreq) { ret = -ENOMEM; break; @@ -287,66 +380,23 @@ static void netfs_read_to_pagecache(struct netfs_io_request *rreq) subreq->start = start; subreq->len = size; + if (copy) + __set_bit(NETFS_SREQ_COPY_TO_CACHE, &subreq->flags); netfs_queue_read(rreq, subreq); - source = netfs_cache_prepare_read(rreq, subreq, rreq->i_size); - subreq->source = source; - if (source == NETFS_DOWNLOAD_FROM_SERVER) { - unsigned long long zero_point = netfs_read_zero_point(rreq->inode); - unsigned long long zp = umin(zero_point, rreq->i_size); - size_t len = subreq->len; - - if (unlikely(rreq->origin == NETFS_READ_SINGLE)) - zp = rreq->i_size; - if (subreq->start >= zp) { - subreq->source = source = NETFS_FILL_WITH_ZEROES; - goto fill_with_zeroes; - } + rreq->io_streams[0].sreq_max_len = MAX_RW_COUNT; + rreq->io_streams[0].sreq_max_segs = INT_MAX; - if (len > zp - subreq->start) - len = zp - subreq->start; - if (len == 0) { - pr_err("ZERO-LEN READ: R=%08x[%x] l=%zx/%zx s=%llx z=%llx i=%llx", - rreq->debug_id, subreq->debug_index, - subreq->len, size, - subreq->start, zero_point, rreq->i_size); + if (prepare_read) { + ret = prepare_read(subreq); + if (ret < 0) { netfs_cancel_read(subreq, ret); break; } - subreq->len = len; - - netfs_stat(&netfs_n_rh_download); - if (rreq->netfs_ops->prepare_read) { - ret = rreq->netfs_ops->prepare_read(subreq); - if (ret < 0) { - netfs_cancel_read(subreq, ret); - break; - } - trace_netfs_sreq(subreq, netfs_sreq_trace_prepare); - } - goto issue; - } - - fill_with_zeroes: - if (source == NETFS_FILL_WITH_ZEROES) { - subreq->source = NETFS_FILL_WITH_ZEROES; - trace_netfs_sreq(subreq, netfs_sreq_trace_submit); - netfs_stat(&netfs_n_rh_zero); - goto issue; - } - - if (source == NETFS_READ_FROM_CACHE) { - trace_netfs_sreq(subreq, netfs_sreq_trace_submit); - goto issue; + trace_netfs_sreq(subreq, netfs_sreq_trace_prepare); } - pr_err("Unexpected read source %u\n", source); - WARN_ON_ONCE(1); - netfs_cancel_read(subreq, ret); - break; - - issue: slice = netfs_prepare_read_iterator(subreq); if (slice < 0) { ret = slice; @@ -355,18 +405,16 @@ static void netfs_read_to_pagecache(struct netfs_io_request *rreq) } start += slice; size -= slice; - if (size <= 0) { - smp_wmb(); /* Write lists before ALL_QUEUED. */ - set_bit(NETFS_RREQ_ALL_QUEUED, &rreq->flags); - } + if (size <= 0) + netfs_all_subreqs_queued(rreq); if (fq) { /* See if the cache indicated this should be cached. */ - bool copy = test_bit(NETFS_SREQ_COPY_TO_CACHE, &subreq->flags); - + copy = test_bit(NETFS_SREQ_COPY_TO_CACHE, &subreq->flags); netfs_mark_copy_to_cache(rreq, &fq, &slot, &offset, slice, copy); } + trace_netfs_sreq(subreq, netfs_sreq_trace_submit); netfs_issue_read(rreq, subreq); netfs_maybe_bulk_drop_ra_refs(rreq); @@ -378,8 +426,7 @@ static void netfs_read_to_pagecache(struct netfs_io_request *rreq) } while (size > 0); if (unlikely(size > 0)) { - smp_wmb(); /* Write lists before ALL_QUEUED. */ - set_bit(NETFS_RREQ_ALL_QUEUED, &rreq->flags); + netfs_all_subreqs_queued(rreq); netfs_wake_collector(rreq); } @@ -646,11 +693,11 @@ EXPORT_SYMBOL(netfs_read_folio); * If any of these criteria are met, then zero out the unwritten parts * of the folio and return true. Otherwise, return false. */ -static bool netfs_skip_folio_read(struct folio *folio, loff_t pos, size_t len, +static bool netfs_skip_folio_read(struct folio *folio, uoff_t pos, size_t len, bool always_fill) { struct inode *inode = folio_inode(folio); - loff_t i_size = i_size_read(inode); + uoff_t i_size = i_size_read(inode); size_t offset = offset_in_folio(folio, pos); size_t plen = folio_size(folio); @@ -715,7 +762,7 @@ zero_out: */ int netfs_write_begin(struct netfs_inode *ctx, struct file *file, struct address_space *mapping, - loff_t pos, unsigned int len, struct folio **_folio, + uoff_t pos, unsigned int len, struct folio **_folio, void **_fsdata) { struct netfs_io_request *rreq; @@ -811,7 +858,7 @@ int netfs_prefetch_for_write(struct file *file, struct folio *folio, struct netfs_io_request *rreq; struct address_space *mapping = folio->mapping; struct netfs_inode *ctx = netfs_inode(mapping->host); - unsigned long long start = folio_pos(folio); + uoff_t start = folio_pos(folio); size_t flen = folio_size(folio); int ret; diff --git a/fs/netfs/buffered_write.c b/fs/netfs/buffered_write.c index 2cdb68e6b16f..49b47252f675 100644 --- a/fs/netfs/buffered_write.c +++ b/fs/netfs/buffered_write.c @@ -17,7 +17,7 @@ * as possible to hold as much of the remaining length as possible in one go. */ static struct folio *netfs_grab_folio_for_write(struct address_space *mapping, - loff_t pos, size_t part) + uoff_t pos, size_t part) { pgoff_t index = pos / PAGE_SIZE; fgf_t fgp_flags = FGP_WRITEBEGIN; @@ -35,9 +35,9 @@ static struct folio *netfs_grab_folio_for_write(struct address_space *mapping, * the values actually are. */ void netfs_update_i_size(struct netfs_inode *ctx, struct inode *inode, - loff_t pos, size_t copied) + uoff_t pos, size_t copied) { - loff_t i_size, end = pos + copied; + uoff_t i_size, end = pos + copied; blkcnt_t add; size_t gap; @@ -54,9 +54,6 @@ void netfs_update_i_size(struct netfs_inode *ctx, struct inode *inode, i_size = i_size_read(inode); if (end > i_size) { i_size_write(inode, end); -#if IS_ENABLED(CONFIG_FSCACHE) - fscache_update_cookie(ctx->cache, NULL, &end); -#endif gap = SECTOR_SIZE - (i_size & (SECTOR_SIZE - 1)); if (copied > gap) { @@ -91,50 +88,20 @@ ssize_t netfs_perform_write(struct kiocb *iocb, struct iov_iter *iter, struct inode *inode = file_inode(file); struct address_space *mapping = inode->i_mapping; struct netfs_inode *ctx = netfs_inode(inode); - struct writeback_control wbc = { - .sync_mode = WB_SYNC_NONE, - .for_sync = true, - .nr_to_write = LONG_MAX, - .range_start = iocb->ki_pos, - .range_end = iocb->ki_pos + iter->count, - }; - struct netfs_io_request *wreq = NULL; - struct folio *folio = NULL, *writethrough = NULL; + struct folio *folio = NULL; unsigned int bdp_flags = (iocb->ki_flags & IOCB_NOWAIT) ? BDP_ASYNC : 0; - ssize_t written = 0, ret, ret2; - loff_t pos = iocb->ki_pos; + ssize_t written = 0, ret; + uoff_t pos = iocb->ki_pos; size_t max_chunk = mapping_max_folio_size(mapping); bool maybe_trouble = false; - if (unlikely(iocb->ki_flags & (IOCB_DSYNC | IOCB_SYNC)) - ) { - wbc_attach_fdatawrite_inode(&wbc, mapping->host); - - ret = filemap_write_and_wait_range(mapping, pos, pos + iter->count); - if (ret < 0) { - wbc_detach_inode(&wbc); - goto out; - } - - wreq = netfs_begin_writethrough(iocb, iter->count); - if (IS_ERR(wreq)) { - wbc_detach_inode(&wbc); - ret = PTR_ERR(wreq); - wreq = NULL; - goto out; - } - if (!is_sync_kiocb(iocb)) - wreq->iocb = iocb; - netfs_stat(&netfs_n_wh_writethrough); - } else { - netfs_stat(&netfs_n_wh_buffered_write); - } + netfs_stat(&netfs_n_wh_buffered_write); do { enum netfs_folio_trace trace; struct netfs_folio *finfo; struct netfs_group *group; - unsigned long long fpos; + uoff_t fpos; size_t flen; size_t offset; /* Offset into pagecache folio */ size_t part; /* Bytes to write to folio */ @@ -390,15 +357,8 @@ ssize_t netfs_perform_write(struct kiocb *iocb, struct iov_iter *iter, pos += copied; written += copied; - if (likely(!wreq)) { - folio_mark_dirty(folio); - folio_unlock(folio); - } else { - netfs_advance_writethrough(wreq, &wbc, folio, copied, - offset + copied == flen, - &writethrough); - /* Folio unlocked */ - } + folio_mark_dirty(folio); + folio_unlock(folio); retry: folio_put(folio); folio = NULL; @@ -420,15 +380,6 @@ out: ctx->ops->post_modify(inode); } - if (unlikely(wreq)) { - ret2 = netfs_end_writethrough(wreq, &wbc, writethrough); - wbc_detach_inode(&wbc); - if (ret2 == -EIOCBQUEUED) - return ret2; - if (ret == 0 && ret2 < 0) - ret = ret2; - } - iocb->ki_pos += written; _leave(" = %zd [%zd]", written, ret); return written ? written : ret; diff --git a/fs/netfs/direct_read.c b/fs/netfs/direct_read.c index 6a8fb0d55e04..8c15f3079723 100644 --- a/fs/netfs/direct_read.c +++ b/fs/netfs/direct_read.c @@ -47,15 +47,15 @@ static void netfs_prepare_dio_read_iterator(struct netfs_io_subrequest *subreq) */ static void netfs_dispatch_unbuffered_reads(struct netfs_io_request *rreq) { - unsigned long long start = rreq->start; ssize_t size = rreq->len; + uoff_t start = rreq->start; int ret; do { struct netfs_io_subrequest *subreq; ssize_t slice; - subreq = netfs_alloc_subrequest(rreq); + subreq = netfs_alloc_subrequest(rreq, NETFS_DOWNLOAD_FROM_SERVER); if (!subreq) { /* Stash the error in the request if there's not * already an error set. @@ -64,7 +64,6 @@ static void netfs_dispatch_unbuffered_reads(struct netfs_io_request *rreq) break; } - subreq->source = NETFS_DOWNLOAD_FROM_SERVER; subreq->start = start; subreq->len = size; @@ -84,10 +83,8 @@ static void netfs_dispatch_unbuffered_reads(struct netfs_io_request *rreq) size -= slice; start += slice; rreq->submitted += slice; - if (size <= 0) { - smp_wmb(); /* Write lists before ALL_QUEUED. */ - set_bit(NETFS_RREQ_ALL_QUEUED, &rreq->flags); - } + if (size <= 0) + netfs_all_subreqs_queued(rreq); rreq->netfs_ops->issue_read(subreq); @@ -99,8 +96,7 @@ static void netfs_dispatch_unbuffered_reads(struct netfs_io_request *rreq) } while (size > 0); if (unlikely(size > 0)) { - smp_wmb(); /* Write lists before ALL_QUEUED. */ - set_bit(NETFS_RREQ_ALL_QUEUED, &rreq->flags); + netfs_all_subreqs_queued(rreq); netfs_wake_collector(rreq); } } diff --git a/fs/netfs/direct_write.c b/fs/netfs/direct_write.c index 2361277416c7..32200c10d2a4 100644 --- a/fs/netfs/direct_write.c +++ b/fs/netfs/direct_write.c @@ -225,9 +225,9 @@ ssize_t netfs_unbuffered_write_iter_locked(struct kiocb *iocb, struct iov_iter * struct netfs_group *netfs_group) { struct netfs_io_request *wreq; - unsigned long long start = iocb->ki_pos; - unsigned long long end = start + iov_iter_count(iter); ssize_t ret, n; + uoff_t start = iocb->ki_pos; + uoff_t end = start + iov_iter_count(iter); size_t len = iov_iter_count(iter); bool async = !is_sync_kiocb(iocb); @@ -336,8 +336,8 @@ ssize_t netfs_unbuffered_write_iter(struct kiocb *iocb, struct iov_iter *from) struct inode *inode = mapping->host; struct netfs_inode *ictx = netfs_inode(inode); ssize_t ret; - loff_t pos = iocb->ki_pos; - unsigned long long end = pos + iov_iter_count(from) - 1; + uoff_t pos = iocb->ki_pos; + uoff_t end = pos + iov_iter_count(from) - 1; _enter("%llx,%zx,%llx", pos, iov_iter_count(from), i_size_read(inode)); diff --git a/fs/netfs/fscache_cookie.c b/fs/netfs/fscache_cookie.c index 3d56fc73435f..5a226f9cbdea 100644 --- a/fs/netfs/fscache_cookie.c +++ b/fs/netfs/fscache_cookie.c @@ -327,7 +327,7 @@ static struct fscache_cookie *fscache_alloc_cookie( u8 advice, const void *index_key, size_t index_key_len, const void *aux_data, size_t aux_data_len, - loff_t object_size) + uoff_t object_size) { struct fscache_cookie *cookie; @@ -452,7 +452,7 @@ struct fscache_cookie *__fscache_acquire_cookie( u8 advice, const void *index_key, size_t index_key_len, const void *aux_data, size_t aux_data_len, - loff_t object_size) + uoff_t object_size) { struct fscache_cookie *cookie; @@ -663,7 +663,7 @@ static void fscache_unuse_cookie_locked(struct fscache_cookie *cookie) * Stop using the cookie for I/O. */ void __fscache_unuse_cookie(struct fscache_cookie *cookie, - const void *aux_data, const loff_t *object_size) + const void *aux_data, const uoff_t *object_size) { unsigned int debug_id = cookie->debug_id; unsigned int r = refcount_read(&cookie->ref); @@ -1049,7 +1049,7 @@ static void fscache_perform_invalidation(struct fscache_cookie *cookie) * Invalidate an object. */ void __fscache_invalidate(struct fscache_cookie *cookie, - const void *aux_data, loff_t new_size, + const void *aux_data, uoff_t new_size, unsigned int flags) { bool is_caching; diff --git a/fs/netfs/fscache_internal.h b/fs/netfs/fscache_internal.h deleted file mode 100644 index a09b948fcef2..000000000000 --- a/fs/netfs/fscache_internal.h +++ /dev/null @@ -1,14 +0,0 @@ -/* SPDX-License-Identifier: GPL-2.0-or-later */ -/* Internal definitions for FS-Cache - * - * Copyright (C) 2021 Red Hat, Inc. All Rights Reserved. - * Written by David Howells (dhowells@redhat.com) - */ - -#include "internal.h" - -#ifdef pr_fmt -#undef pr_fmt -#endif - -#define pr_fmt(fmt) "FS-Cache: " fmt diff --git a/fs/netfs/fscache_io.c b/fs/netfs/fscache_io.c index 37f05b4d3469..056a2bae5d99 100644 --- a/fs/netfs/fscache_io.c +++ b/fs/netfs/fscache_io.c @@ -79,7 +79,7 @@ static int fscache_begin_operation(struct netfs_cache_resources *cres, cres->ops = NULL; cres->cache_priv = cookie; cres->cache_priv2 = NULL; - cres->debug_id = cookie->debug_id; + cres->cookie_id = cookie->debug_id; cres->inval_counter = cookie->inval_counter; if (!fscache_begin_cookie_access(cookie, why)) { @@ -162,7 +162,7 @@ EXPORT_SYMBOL(__fscache_begin_write_operation); struct fscache_write_request { struct netfs_cache_resources cache_resources; struct address_space *mapping; - loff_t start; + uoff_t start; size_t len; bool set_bits; bool using_pgpriv2; @@ -171,7 +171,7 @@ struct fscache_write_request { }; void __fscache_clear_page_bits(struct address_space *mapping, - loff_t start, size_t len) + uoff_t start, size_t len) { pgoff_t first = start / PAGE_SIZE; pgoff_t last = (start + len - 1) / PAGE_SIZE; @@ -208,7 +208,7 @@ static void fscache_wreq_done(void *priv, ssize_t transferred_or_error) void __fscache_write_to_cache(struct fscache_cookie *cookie, struct address_space *mapping, - loff_t start, size_t len, loff_t i_size, + uoff_t start, size_t len, uoff_t i_size, netfs_io_terminated_t term_func, void *term_func_priv, bool using_pgpriv2, bool cond) @@ -267,7 +267,7 @@ EXPORT_SYMBOL(__fscache_write_to_cache); /* * Change the size of a backing object. */ -void __fscache_resize_cookie(struct fscache_cookie *cookie, loff_t new_size) +void __fscache_resize_cookie(struct fscache_cookie *cookie, uoff_t new_size) { struct netfs_cache_resources cres; diff --git a/fs/netfs/internal.h b/fs/netfs/internal.h index c79c8e69d60c..b8591abc90a9 100644 --- a/fs/netfs/internal.h +++ b/fs/netfs/internal.h @@ -23,6 +23,8 @@ /* * buffered_read.c */ +int netfs_read_query_cache(struct netfs_io_request *rreq, + struct fscache_occupancy *occ); void netfs_queue_read(struct netfs_io_request *rreq, struct netfs_io_subrequest *subreq); void netfs_cache_read_terminated(void *priv, ssize_t transferred_or_error); @@ -33,7 +35,7 @@ int netfs_prefetch_for_write(struct file *file, struct folio *folio, * buffered_write.c */ void netfs_update_i_size(struct netfs_inode *ctx, struct inode *inode, - loff_t pos, size_t copied); + uoff_t pos, size_t copied); /* * main.c @@ -86,13 +88,14 @@ void netfs_wait_for_put_ra_refs(struct netfs_io_request *rreq); */ struct netfs_io_request *netfs_alloc_request(struct address_space *mapping, struct file *file, - loff_t start, size_t len, + uoff_t start, size_t len, enum netfs_io_origin origin); void netfs_get_request(struct netfs_io_request *rreq, enum netfs_rreq_ref_trace what); void netfs_clear_subrequests(struct netfs_io_request *rreq); void netfs_put_request(struct netfs_io_request *rreq, enum netfs_rreq_ref_trace what); void netfs_put_failed_request(struct netfs_io_request *rreq); -struct netfs_io_subrequest *netfs_alloc_subrequest(struct netfs_io_request *rreq); +struct netfs_io_subrequest *netfs_alloc_subrequest(struct netfs_io_request *rreq, + enum netfs_io_source source); static inline void netfs_see_request(struct netfs_io_request *rreq, enum netfs_rreq_ref_trace what) @@ -120,9 +123,37 @@ void netfs_cache_read_terminated(void *priv, ssize_t transferred_or_error); /* * read_pgpriv2.c */ +#ifdef CONFIG_NETFS_PGPRIV2 +int netfs_prepare_pgpriv2_write_buffer(struct netfs_io_subrequest *subreq, + unsigned int max_segs); void netfs_pgpriv2_copy_to_cache(struct netfs_io_request *rreq, struct folio *folio); void netfs_pgpriv2_end_copy_to_cache(struct netfs_io_request *rreq); bool netfs_pgpriv2_unlock_copied_folios(struct netfs_io_request *wreq); +static inline bool netfs_using_pgpriv2(const struct netfs_io_request *rreq) +{ + return unlikely(test_bit(NETFS_RREQ_USE_PGPRIV2, &rreq->flags)); +} +#else +static inline int netfs_prepare_pgpriv2_write_buffer(struct netfs_io_subrequest *subreq, + unsigned int max_segs) +{ + return -EIO; +} +static inline void netfs_pgpriv2_copy_to_cache(struct netfs_io_request *rreq, struct folio *folio) +{ +} +static inline void netfs_pgpriv2_end_copy_to_cache(struct netfs_io_request *rreq) +{ +} +static inline bool netfs_pgpriv2_unlock_copied_folios(struct netfs_io_request *wreq) +{ + return true; +} +static inline bool netfs_using_pgpriv2(const struct netfs_io_request *rreq) +{ + return false; +} +#endif /* * read_retry.c @@ -157,7 +188,6 @@ extern atomic_t netfs_n_rh_write_zskip; extern atomic_t netfs_n_rh_retry_read_req; extern atomic_t netfs_n_rh_retry_read_subreq; extern atomic_t netfs_n_wh_buffered_write; -extern atomic_t netfs_n_wh_writethrough; extern atomic_t netfs_n_wh_dio_write; extern atomic_t netfs_n_wh_writepages; extern atomic_t netfs_n_wh_copy_to_cache; @@ -203,11 +233,11 @@ void netfs_write_collection_worker(struct work_struct *work); */ struct netfs_io_request *netfs_create_write_req(struct address_space *mapping, struct file *file, - loff_t start, + uoff_t start, enum netfs_io_origin origin); void netfs_prepare_write(struct netfs_io_request *wreq, struct netfs_io_stream *stream, - loff_t start); + uoff_t start); void netfs_reissue_write(struct netfs_io_stream *stream, struct netfs_io_subrequest *subreq, struct iov_iter *source); @@ -215,13 +245,7 @@ void netfs_issue_write(struct netfs_io_request *wreq, struct netfs_io_stream *stream); size_t netfs_advance_write(struct netfs_io_request *wreq, struct netfs_io_stream *stream, - loff_t start, size_t len, bool to_eof); -struct netfs_io_request *netfs_begin_writethrough(struct kiocb *iocb, size_t len); -int netfs_advance_writethrough(struct netfs_io_request *wreq, struct writeback_control *wbc, - struct folio *folio, size_t copied, bool to_page_end, - struct folio **writethrough_cache); -ssize_t netfs_end_writethrough(struct netfs_io_request *wreq, struct writeback_control *wbc, - struct folio *writethrough_cache); + uoff_t start, size_t len, bool to_eof); /* * write_retry.c @@ -321,6 +345,26 @@ static inline bool netfs_check_subreq_in_progress(const struct netfs_io_subreque } /* + * Indicate that we've generated and queued all the subrequests we're going to. + */ +static inline void netfs_all_subreqs_queued(struct netfs_io_request *rreq) +{ + smp_wmb(); /* Write lists before ALL_QUEUED. */ + set_bit(NETFS_RREQ_ALL_QUEUED, &rreq->flags); + smp_mb__after_atomic(); + trace_netfs_rreq(rreq, netfs_rreq_trace_all_queued); +} + +/* + * Query if all subrequests are queued. + */ +static inline bool netfs_are_all_subreqs_queued(const struct netfs_io_request *rreq) +{ + /* Read lists after ALL_QUEUED. */ + return test_bit_acquire(NETFS_RREQ_ALL_QUEUED, &rreq->flags); +} + +/* * fscache-cache.c */ #ifdef CONFIG_PROC_FS diff --git a/fs/netfs/iterator.c b/fs/netfs/iterator.c index b375567e0520..eb1efb17f53a 100644 --- a/fs/netfs/iterator.c +++ b/fs/netfs/iterator.c @@ -209,7 +209,7 @@ static size_t netfs_limit_xarray(const struct iov_iter *iter, size_t start_offse { struct folio *folio; unsigned int nsegs = 0; - loff_t pos = iter->xarray_start + iter->iov_offset; + uoff_t pos = iter->xarray_start + iter->iov_offset; pgoff_t index = pos / PAGE_SIZE; size_t span = 0, n = iter->count; diff --git a/fs/netfs/main.c b/fs/netfs/main.c index 927badf3989d..609e22e8f76a 100644 --- a/fs/netfs/main.c +++ b/fs/netfs/main.c @@ -44,7 +44,6 @@ static const char *netfs_origins[nr__netfs_io_origin] = { [NETFS_DIO_READ] = "DR", [NETFS_WRITEBACK] = "WB", [NETFS_WRITEBACK_SINGLE] = "W1", - [NETFS_WRITETHROUGH] = "WT", [NETFS_UNBUFFERED_WRITE] = "UW", [NETFS_DIO_WRITE] = "DW", [NETFS_PGPRIV2_COPY_TO_CACHE] = "2C", diff --git a/fs/netfs/misc.c b/fs/netfs/misc.c index f5c1c463f4ff..a3cd76d584b8 100644 --- a/fs/netfs/misc.c +++ b/fs/netfs/misc.c @@ -193,7 +193,7 @@ void netfs_clear_inode_writeback(struct inode *inode, const void *aux) struct fscache_cookie *cookie = netfs_i_cookie(netfs_inode(inode)); if (inode_state_read_once(inode) & I_PINNING_NETFS_WB) { - loff_t i_size = i_size_read(inode); + uoff_t i_size = i_size_read(inode); fscache_unuse_cookie(cookie, aux, &i_size); } } @@ -218,8 +218,8 @@ void netfs_invalidate_folio(struct folio *folio, size_t offset, size_t length) _enter("{%lx},%zx,%zx", folio->index, offset, length); if (offset == 0 && length == flen) { - unsigned long long i_size, remote_i_size, zero_point; - unsigned long long fpos = folio_pos(folio), end; + uoff_t i_size, remote_i_size, zero_point; + uoff_t fpos = folio_pos(folio), end; netfs_read_sizes(inode, &i_size, &remote_i_size, &zero_point); end = umin(fpos + flen, i_size); @@ -305,7 +305,7 @@ bool netfs_release_folio(struct folio *folio, gfp_t gfp) { struct inode *inode = folio_inode(folio); struct netfs_inode *ctx = netfs_inode(inode); - unsigned long long i_size, remote_i_size, zero_point, end; + uoff_t i_size, remote_i_size, zero_point, end; if (folio_test_dirty(folio)) return false; @@ -424,7 +424,7 @@ static int netfs_collect_in_app(struct netfs_io_request *rreq, need_collect = true; break; } - if (subreq || !test_bit(NETFS_RREQ_ALL_QUEUED, &rreq->flags)) + if (subreq || !netfs_are_all_subreqs_queued(rreq)) done = false; } diff --git a/fs/netfs/objects.c b/fs/netfs/objects.c index ad549daa9c79..4b8d20559b0e 100644 --- a/fs/netfs/objects.c +++ b/fs/netfs/objects.c @@ -16,7 +16,7 @@ static void netfs_free_request(struct work_struct *work); */ struct netfs_io_request *netfs_alloc_request(struct address_space *mapping, struct file *file, - loff_t start, size_t len, + uoff_t start, size_t len, enum netfs_io_origin origin) { static atomic_t debug_ids; @@ -44,6 +44,7 @@ struct netfs_io_request *netfs_alloc_request(struct address_space *mapping, rreq->gfp = gfp; rreq->start = start; rreq->collected_to = start; + rreq->cache_coll_to = start; rreq->cleaned_to = start; rreq->len = len; rreq->progress_at = 0; @@ -207,7 +208,8 @@ void netfs_put_failed_request(struct netfs_io_request *rreq) /* * Allocate and partially initialise an I/O request structure. */ -struct netfs_io_subrequest *netfs_alloc_subrequest(struct netfs_io_request *rreq) +struct netfs_io_subrequest *netfs_alloc_subrequest(struct netfs_io_request *rreq, + enum netfs_io_source source) { struct netfs_io_subrequest *subreq; mempool_t *mempool = rreq->netfs_ops->subrequest_pool ?: &netfs_subrequest_pool; @@ -224,6 +226,7 @@ struct netfs_io_subrequest *netfs_alloc_subrequest(struct netfs_io_request *rreq INIT_WORK(&subreq->work, NULL); INIT_LIST_HEAD(&subreq->rreq_link); refcount_set(&subreq->ref, 2); + subreq->source = source; subreq->rreq = rreq; subreq->debug_index = atomic_inc_return(&rreq->subreq_counter); netfs_get_request(rreq, netfs_rreq_trace_get_subreq); diff --git a/fs/netfs/read_collect.c b/fs/netfs/read_collect.c index a94197ef0181..2625efd48a9b 100644 --- a/fs/netfs/read_collect.c +++ b/fs/netfs/read_collect.c @@ -38,7 +38,7 @@ static void netfs_clear_unread(struct netfs_io_subrequest *subreq) */ void netfs_cancel_copy_to_cache(struct netfs_io_request *rreq, struct folio *folio) { - if (!test_bit(NETFS_RREQ_USE_PGPRIV2, &rreq->flags)) { + if (!netfs_using_pgpriv2(rreq)) { if (folio_get_private(folio) == NETFS_FOLIO_COPY_TO_CACHE) { folio_detach_private(folio); trace_netfs_folio(folio, netfs_folio_trace_cancel_copy); @@ -81,7 +81,7 @@ static void netfs_unlock_read_folio(struct netfs_io_request *rreq, if (unlikely(test_bit(NETFS_RREQ_CANCEL_CACHING, &rreq->flags))) netfs_cancel_copy_to_cache(rreq, folio); - if (!test_bit(NETFS_RREQ_USE_PGPRIV2, &rreq->flags)) { + if (!netfs_using_pgpriv2(rreq)) { if (netfs_folio_group(folio) == NETFS_FOLIO_COPY_TO_CACHE) { trace_netfs_folio(folio, netfs_folio_trace_sched_copy); folio_mark_dirty(folio); @@ -153,8 +153,8 @@ static void netfs_read_unlock_folios(struct netfs_io_request *rreq, unsigned int *notes) { struct folio_queue *folioq = rreq->buffer.tail; - unsigned long long collected_to = rreq->collected_to; unsigned int slot = rreq->buffer.first_tail_slot; + uoff_t collected_to = rreq->collected_to; if (rreq->cleaned_to >= rreq->collected_to) return; @@ -179,7 +179,7 @@ static void netfs_read_unlock_folios(struct netfs_io_request *rreq, for (;;) { struct folio *folio; - unsigned long long fpos, fend; + uoff_t fpos, fend; size_t fsize; folio = folioq_folio(folioq, slot); @@ -192,7 +192,7 @@ static void netfs_read_unlock_folios(struct netfs_io_request *rreq, fpos = folio_pos(folio); fend = fpos + fsize; - trace_netfs_collect_folio(rreq, folio, fend, collected_to); + trace_netfs_collect_folio(rreq, folio); /* Unlock any folio we've transferred all of. */ if (collected_to < fend) @@ -467,10 +467,8 @@ bool netfs_read_collection(struct netfs_io_request *rreq) /* We're done when the app thread has finished posting subreqs and the * queue is empty. */ - if (!test_bit(NETFS_RREQ_ALL_QUEUED, &rreq->flags)) + if (!netfs_are_all_subreqs_queued(rreq)) return false; - smp_rmb(); /* Read ALL_QUEUED before subreq lists. */ - if (!list_empty(&stream->subrequests)) return false; diff --git a/fs/netfs/read_pgpriv2.c b/fs/netfs/read_pgpriv2.c index a4b7bb88cbdb..5280b606fda4 100644 --- a/fs/netfs/read_pgpriv2.c +++ b/fs/netfs/read_pgpriv2.c @@ -20,7 +20,7 @@ static void netfs_pgpriv2_copy_folio(struct netfs_io_request *creq, struct folio { struct netfs_io_stream *cache = &creq->io_streams[1]; size_t fsize = folio_size(folio), flen = fsize; - loff_t fpos = folio_pos(folio), i_size; + uoff_t fpos = folio_pos(folio), i_size; bool to_eof = false; _enter(""); @@ -158,8 +158,7 @@ void netfs_pgpriv2_end_copy_to_cache(struct netfs_io_request *rreq) return; netfs_issue_write(creq, &creq->io_streams[1]); - smp_wmb(); /* Write lists before ALL_QUEUED. */ - set_bit(NETFS_RREQ_ALL_QUEUED, &creq->flags); + netfs_all_subreqs_queued(creq); trace_netfs_rreq(rreq, netfs_rreq_trace_end_copy_to_cache); if (list_empty_careful(&creq->io_streams[1].subrequests)) netfs_wake_collector(creq); @@ -175,8 +174,8 @@ void netfs_pgpriv2_end_copy_to_cache(struct netfs_io_request *rreq) bool netfs_pgpriv2_unlock_copied_folios(struct netfs_io_request *creq) { struct folio_queue *folioq = creq->buffer.tail; - unsigned long long collected_to = creq->collected_to; unsigned int slot = creq->buffer.first_tail_slot; + uoff_t collected_to = creq->collected_to; bool made_progress = false; if (slot >= folioq_nr_slots(folioq)) { @@ -186,7 +185,7 @@ bool netfs_pgpriv2_unlock_copied_folios(struct netfs_io_request *creq) for (;;) { struct folio *folio; - unsigned long long fpos, fend; + uoff_t fpos, fend; size_t fsize, flen; folio = folioq_folio(folioq, slot); @@ -199,9 +198,9 @@ bool netfs_pgpriv2_unlock_copied_folios(struct netfs_io_request *creq) fsize = folio_size(folio); flen = fsize; - fend = min_t(unsigned long long, fpos + flen, creq->i_size); + fend = min_t(uoff_t, fpos + flen, creq->i_size); - trace_netfs_collect_folio(creq, folio, fend, collected_to); + trace_netfs_collect_folio(creq, folio); /* Unlock any folio we've transferred all of. */ if (collected_to < fend) diff --git a/fs/netfs/read_retry.c b/fs/netfs/read_retry.c index 4f6a36c6e214..5bd8dee5a834 100644 --- a/fs/netfs/read_retry.c +++ b/fs/netfs/read_retry.c @@ -75,7 +75,7 @@ static void netfs_retry_read_subrequests(struct netfs_io_request *rreq) do { struct netfs_io_subrequest *from, *to, *tmp; struct iov_iter source; - unsigned long long start, len; + uoff_t start, len; size_t part; bool boundary = false, subreq_superfluous = false; @@ -195,12 +195,11 @@ static void netfs_retry_read_subrequests(struct netfs_io_request *rreq) * and insert them after. */ do { - subreq = netfs_alloc_subrequest(rreq); + subreq = netfs_alloc_subrequest(rreq, NETFS_DOWNLOAD_FROM_SERVER); if (!subreq) { subreq = to; goto abandon_after; } - subreq->source = NETFS_DOWNLOAD_FROM_SERVER; subreq->start = start; subreq->len = len; subreq->stream_nr = stream->stream_nr; @@ -272,6 +271,7 @@ void netfs_retry_reads(struct netfs_io_request *rreq) struct netfs_io_stream *stream = &rreq->io_streams[0]; netfs_stat(&netfs_n_rh_retry_read_req); + trace_netfs_rreq(rreq, netfs_rreq_trace_retry_begin); /* Wait for all outstanding I/O to quiesce before performing retries as * we may need to renegotiate the I/O sizes. @@ -282,6 +282,7 @@ void netfs_retry_reads(struct netfs_io_request *rreq) trace_netfs_rreq(rreq, netfs_rreq_trace_resubmit); netfs_retry_read_subrequests(rreq); + trace_netfs_rreq(rreq, netfs_rreq_trace_retry_end); } /* diff --git a/fs/netfs/read_single.c b/fs/netfs/read_single.c index de67ac41548d..b248e34bd0c8 100644 --- a/fs/netfs/read_single.c +++ b/fs/netfs/read_single.c @@ -58,20 +58,6 @@ static int netfs_single_begin_cache_read(struct netfs_io_request *rreq, struct n return fscache_begin_read_operation(&rreq->cache_resources, netfs_i_cookie(ctx)); } -static void netfs_single_cache_prepare_read(struct netfs_io_request *rreq, - struct netfs_io_subrequest *subreq) -{ - struct netfs_cache_resources *cres = &rreq->cache_resources; - - if (!cres->ops) { - subreq->source = NETFS_DOWNLOAD_FROM_SERVER; - return; - } - subreq->source = cres->ops->prepare_read(subreq, rreq->i_size); - trace_netfs_sreq(subreq, netfs_sreq_trace_prepare); - -} - static void netfs_single_read_cache(struct netfs_io_request *rreq, struct netfs_io_subrequest *subreq) { @@ -89,21 +75,36 @@ static void netfs_single_read_cache(struct netfs_io_request *rreq, */ static int netfs_single_dispatch_read(struct netfs_io_request *rreq) { + struct fscache_occupancy occ = { + .query_from = 0, + .query_to = rreq->len, + .cached_from[0] = ULLONG_MAX, + .cached_to[0] = ULLONG_MAX, + .cached_from[1] = ULLONG_MAX, + .cached_to[1] = ULLONG_MAX, + }; struct netfs_io_subrequest *subreq; + enum netfs_io_source source = NETFS_DOWNLOAD_FROM_SERVER; int ret = 0; - subreq = netfs_alloc_subrequest(rreq); + /* Try to use the cache if the cache content matches the size of the + * remote file. + */ + netfs_read_query_cache(rreq, &occ); + if (occ.cached_from[0] == 0 && + occ.cached_to[0] >= rreq->len) + source = NETFS_READ_FROM_CACHE; + + subreq = netfs_alloc_subrequest(rreq, source); if (!subreq) return -ENOMEM; - subreq->source = NETFS_SOURCE_UNKNOWN; subreq->start = 0; subreq->len = rreq->len; subreq->io_iter = rreq->buffer.iter; netfs_queue_read(rreq, subreq); - netfs_single_cache_prepare_read(rreq, subreq); switch (subreq->source) { case NETFS_DOWNLOAD_FROM_SERVER: netfs_stat(&netfs_n_rh_download); @@ -113,14 +114,18 @@ static int netfs_single_dispatch_read(struct netfs_io_request *rreq) goto cancel; } - smp_wmb(); /* Write lists before ALL_QUEUED. */ - set_bit(NETFS_RREQ_ALL_QUEUED, &rreq->flags); + netfs_all_subreqs_queued(rreq); rreq->netfs_ops->issue_read(subreq); rreq->submitted += subreq->len; break; case NETFS_READ_FROM_CACHE: - smp_wmb(); /* Write lists before ALL_QUEUED. */ - set_bit(NETFS_RREQ_ALL_QUEUED, &rreq->flags); + if (rreq->cache_resources.ops->prepare_read) { + ret = rreq->cache_resources.ops->prepare_read(subreq); + if (ret < 0) + goto cancel; + } + + netfs_all_subreqs_queued(rreq); trace_netfs_sreq(subreq, netfs_sreq_trace_submit); netfs_single_read_cache(rreq, subreq); rreq->submitted += subreq->len; @@ -136,8 +141,7 @@ static int netfs_single_dispatch_read(struct netfs_io_request *rreq) return ret; cancel: netfs_cancel_read(subreq, ret); - smp_wmb(); /* Write lists before ALL_QUEUED. */ - set_bit(NETFS_RREQ_ALL_QUEUED, &rreq->flags); + netfs_all_subreqs_queued(rreq); netfs_wake_collector(rreq); return ret; } diff --git a/fs/netfs/stats.c b/fs/netfs/stats.c index ab6b916addc4..9a607c4e62dd 100644 --- a/fs/netfs/stats.c +++ b/fs/netfs/stats.c @@ -32,7 +32,6 @@ atomic_t netfs_n_rh_write_zskip; atomic_t netfs_n_rh_retry_read_req; atomic_t netfs_n_rh_retry_read_subreq; atomic_t netfs_n_wh_buffered_write; -atomic_t netfs_n_wh_writethrough; atomic_t netfs_n_wh_dio_write; atomic_t netfs_n_wh_writepages; atomic_t netfs_n_wh_copy_to_cache; @@ -58,9 +57,8 @@ int netfs_stats_show(struct seq_file *m, void *v) atomic_read(&netfs_n_rh_read_single), atomic_read(&netfs_n_rh_write_begin), atomic_read(&netfs_n_rh_write_zskip)); - seq_printf(m, "Writes : BW=%u WT=%u DW=%u WP=%u 2C=%u\n", + seq_printf(m, "Writes : BW=%u DW=%u WP=%u 2C=%u\n", atomic_read(&netfs_n_wh_buffered_write), - atomic_read(&netfs_n_wh_writethrough), atomic_read(&netfs_n_wh_dio_write), atomic_read(&netfs_n_wh_writepages), atomic_read(&netfs_n_wh_copy_to_cache)); diff --git a/fs/netfs/write_collect.c b/fs/netfs/write_collect.c index 210eb8f3958d..6e8ea534230d 100644 --- a/fs/netfs/write_collect.c +++ b/fs/netfs/write_collect.c @@ -56,7 +56,7 @@ static void netfs_dump_request(const struct netfs_io_request *rreq) */ int netfs_folio_written_back(struct folio *folio) { - enum netfs_folio_trace why = netfs_folio_trace_clear; + enum netfs_folio_trace why = netfs_folio_trace_endwb; struct inode *inode = folio_inode(folio); struct netfs_inode *ictx = netfs_inode(inode); struct netfs_folio *finfo; @@ -67,7 +67,7 @@ int netfs_folio_written_back(struct folio *folio) /* Streaming writes cannot be redirtied whilst under writeback, * so discard the streaming record. */ - unsigned long long fend; + uoff_t fend; fend = folio_pos(folio) + finfo->dirty_offset + finfo->dirty_len; spin_lock(&ictx->inode.i_lock); @@ -79,13 +79,13 @@ int netfs_folio_written_back(struct folio *folio) group = finfo->netfs_group; gcount++; kfree(finfo); - why = netfs_folio_trace_clear_s; + why = netfs_folio_trace_endwb_s; goto end_wb; } if ((group = netfs_folio_group(folio))) { if (group == NETFS_FOLIO_COPY_TO_CACHE) { - why = netfs_folio_trace_clear_cc; + why = netfs_folio_trace_endwb_cc; folio_detach_private(folio); goto end_wb; } @@ -98,7 +98,7 @@ int netfs_folio_written_back(struct folio *folio) if (!folio_test_dirty(folio)) { folio_detach_private(folio); gcount++; - why = netfs_folio_trace_clear_g; + why = netfs_folio_trace_endwb_g; } } @@ -115,8 +115,8 @@ static void netfs_writeback_unlock_folios(struct netfs_io_request *wreq, unsigned int *notes) { struct folio_queue *folioq = wreq->buffer.tail; - unsigned long long collected_to = wreq->collected_to; unsigned int slot = wreq->buffer.first_tail_slot; + uoff_t collected_to = wreq->collected_to; if (WARN_ON_ONCE(!folioq)) { pr_err("[!] Writeback unlock found empty rolling buffer!\n"); @@ -140,7 +140,7 @@ static void netfs_writeback_unlock_folios(struct netfs_io_request *wreq, for (;;) { struct folio *folio; struct netfs_folio *finfo; - unsigned long long fpos, fend; + uoff_t fpos, fend; size_t fsize, flen; folio = folioq_folio(folioq, slot); @@ -154,9 +154,9 @@ static void netfs_writeback_unlock_folios(struct netfs_io_request *wreq, finfo = netfs_folio_info(folio); flen = finfo ? finfo->dirty_offset + finfo->dirty_len : fsize; - fend = min_t(unsigned long long, fpos + flen, wreq->i_size); + fend = min_t(uoff_t, fpos + flen, wreq->i_size); - trace_netfs_collect_folio(wreq, folio, fend, collected_to); + trace_netfs_collect_folio(wreq, folio); /* Unlock any folio we've transferred all of. */ if (collected_to < fend) @@ -189,6 +189,26 @@ done: } /* + * Collect cache results. + */ +static void netfs_cache_collect(struct netfs_io_request *wreq, + struct netfs_io_stream *stream, + enum netfs_cache_collect block_type) +{ + struct netfs_cache_resources *cres = &wreq->cache_resources; + + if (stream->source != NETFS_WRITE_TO_CACHE || + wreq->cache_coll_to >= stream->collected_to) + return; + + if (cres->ops && cres->ops->collect_write) + cres->ops->collect_write(wreq, wreq->cache_coll_to, + stream->collected_to - wreq->cache_coll_to, + block_type); + wreq->cache_coll_to = stream->collected_to; +} + +/* * Collect and assess the results of various write subrequests. We may need to * retry some of the results - or even do an RMW cycle for content crypto. * @@ -201,8 +221,8 @@ static void netfs_collect_write_results(struct netfs_io_request *wreq) { struct netfs_io_subrequest *front, *remove; struct netfs_io_stream *stream; - unsigned long long collected_to, issued_to; unsigned int notes; + uoff_t collected_to, issued_to; int s; _enter("%llx-%llx", wreq->start, wreq->start + wreq->len); @@ -214,7 +234,6 @@ reassess_streams: smp_rmb(); collected_to = ULLONG_MAX; if (wreq->origin == NETFS_WRITEBACK || - wreq->origin == NETFS_WRITETHROUGH || wreq->origin == NETFS_PGPRIV2_COPY_TO_CACHE) notes = NEED_UNLOCK; else @@ -236,13 +255,19 @@ reassess_streams: /* Read first subreq pointer before IN_PROGRESS flag. */ while (front) { + enum netfs_cache_collect cache_collect; + trace_netfs_collect_sreq(wreq, front); //_debug("sreq [%x] %llx %zx/%zx", // front->debug_index, front->start, front->transferred, front->len); if (stream->collected_to < front->start) { trace_netfs_collect_gap(wreq, stream, issued_to, 'F'); + if (stream->cache_collect != NETFS_CACHE_COLLECT_WRITE_GAP) + netfs_cache_collect(wreq, stream, stream->cache_collect); stream->collected_to = front->start; + netfs_cache_collect(wreq, stream, NETFS_CACHE_COLLECT_WRITE_GAP); + stream->cache_collect = NETFS_CACHE_COLLECT_WRITE_GAP; } /* Stall if the front is still undergoing I/O. */ @@ -250,7 +275,6 @@ reassess_streams: notes |= HIT_PENDING; break; } - smp_rmb(); /* Read counters after I-P flag. */ if (stream->failed) { stream->collected_to = front->start + front->len; @@ -263,15 +287,44 @@ reassess_streams: stream->transferred_valid = true; notes |= MADE_PROGRESS; } - if (test_bit(NETFS_SREQ_FAILED, &front->flags)) { - stream->failed = true; - stream->error = front->error; - if (stream->source == NETFS_UPLOAD_TO_SERVER) - mapping_set_error(wreq->mapping, front->error); - notes |= NEED_REASSESS | SAW_FAILURE; + + /* Handle failed or cancelled subreqs. Failure of + * cache writes are handled differently to upload + * failures. Cache writes aren't fatal, provided we're + * not doing disconnected operation, and so we can kind + * of treat them as if they had succeeded - except that + * we need to log any holes they cause. + */ + switch (stream->source) { + case NETFS_UPLOAD_TO_SERVER: + if (test_bit(NETFS_SREQ_FAILED, &front->flags)) { + if (!stream->failed) { + stream->failed = true; + stream->error = front->error; + mapping_set_error(wreq->mapping, front->error); + break; + } + notes |= NEED_REASSESS | SAW_FAILURE; + } + break; + + case NETFS_WRITE_TO_CACHE: + cache_collect = test_bit(NETFS_SREQ_CANCELLED, &front->flags) ? + NETFS_CACHE_COLLECT_WRITE_CANCEL : + NETFS_CACHE_COLLECT_WRITE_DATA; + if (cache_collect != stream->cache_collect && + stream->cache_collect != NETFS_CACHE_COLLECT_WRITE_GAP) { + trace_netfs_rreq(wreq, netfs_rreq_trace_cache_fail_collect); + netfs_cache_collect(wreq, stream, stream->cache_collect); + } + stream->cache_collect = cache_collect; + break; + + default: + WARN_ON(1); break; } - if (front->transferred < front->len) { + if (test_bit(NETFS_SREQ_NEED_RETRY, &front->flags)) { stream->need_retry = true; notes |= NEED_RETRY | MADE_PROGRESS; break; @@ -360,6 +413,7 @@ need_retry: */ bool netfs_write_collection(struct netfs_io_request *wreq) { + struct netfs_io_stream *cstream = &wreq->io_streams[1]; struct netfs_inode *ictx = netfs_inode(wreq->inode); size_t transferred; bool transferred_valid = false; @@ -372,9 +426,8 @@ bool netfs_write_collection(struct netfs_io_request *wreq) /* We're done when the app thread has finished posting subreqs and all * the queues in all the streams are empty. */ - if (!test_bit(NETFS_RREQ_ALL_QUEUED, &wreq->flags)) + if (!netfs_are_all_subreqs_queued(wreq)) return false; - smp_rmb(); /* Read ALL_QUEUED before lists. */ transferred = LONG_MAX; for (s = 0; s < NR_IO_STREAMS; s++) { @@ -395,13 +448,19 @@ bool netfs_write_collection(struct netfs_io_request *wreq) wreq->transferred = transferred; trace_netfs_rreq(wreq, netfs_rreq_trace_write_done); - if (wreq->io_streams[1].active && - wreq->io_streams[1].failed && - ictx->ops->invalidate_cache) { - /* Cache write failure doesn't prevent writeback completion - * unless we're in disconnected mode. - */ - ictx->ops->invalidate_cache(wreq); + if (cstream->active) { + if (test_bit(NETFS_RREQ_CACHE_ERROR, &wreq->flags)) { + if (ictx->ops->invalidate_cache) { + /* Cache write failure doesn't prevent + * writeback completion unless we're in + * disconnected mode. + */ + trace_netfs_rreq(wreq, netfs_rreq_trace_inval_cache); + ictx->ops->invalidate_cache(wreq); + } + } else if (!cstream->failed) { + netfs_cache_collect(wreq, cstream, cstream->cache_collect); + } } _debug("finished"); @@ -411,7 +470,6 @@ bool netfs_write_collection(struct netfs_io_request *wreq) switch (wreq->origin) { case NETFS_WRITEBACK: case NETFS_WRITEBACK_SINGLE: - case NETFS_WRITETHROUGH: netfs_wb_end(ictx); break; default: @@ -486,24 +544,51 @@ void netfs_write_subrequest_terminated(void *_op, ssize_t transferred_or_error) if (IS_ERR_VALUE(transferred_or_error)) { subreq->error = transferred_or_error; - /* if need retry is set, error should not matter */ - if (!test_bit(NETFS_SREQ_NEED_RETRY, &subreq->flags)) { - set_bit(NETFS_SREQ_FAILED, &subreq->flags); - trace_netfs_failure(wreq, subreq, transferred_or_error, netfs_fail_write); - } switch (subreq->source) { case NETFS_WRITE_TO_CACHE: + /* We don't mark a cache-write subreq as failed. + * Instead we tell the issuer to produce dummy subreqs + * instead and make a note if we need to invalidate the + * cache at the end. We also don't pause the loop that + * grabs pages and launches upload subreqs. + * + * Note that we need to distinguish between -ENOBUFS + * (no space available in the cache) and other errors. + * In the former case, we can keep the data we have, + * though we might have to change the way the on-disk + * data is tracked. + */ netfs_stat(&netfs_n_wh_write_failed); + if (test_bit(NETFS_SREQ_NEED_RETRY, &subreq->flags)) + break; + + trace_netfs_failure(wreq, subreq, transferred_or_error, netfs_fail_write); + __set_bit(NETFS_SREQ_CANCELLED, &subreq->flags); + set_bit(NETFS_RREQ_CACHE_STOP, &wreq->flags); + if (transferred_or_error == -ENOBUFS) + trace_netfs_rreq(wreq, netfs_rreq_trace_cache_no_space); + else if (!test_and_set_bit(NETFS_RREQ_CACHE_ERROR, &wreq->flags)) + trace_netfs_rreq(wreq, netfs_rreq_trace_cache_failed); + subreq->transferred = subreq->len; break; + case NETFS_UPLOAD_TO_SERVER: + /* If need_retry is set, error should not matter */ + if (!test_bit(NETFS_SREQ_NEED_RETRY, &subreq->flags)) { + set_bit(NETFS_SREQ_FAILED, &subreq->flags); + trace_netfs_failure(wreq, subreq, transferred_or_error, + netfs_fail_upload); + } + + set_bit(NETFS_RREQ_PAUSE, &wreq->flags); + trace_netfs_rreq(wreq, netfs_rreq_trace_set_pause); netfs_stat(&netfs_n_wh_upload_failed); break; + default: break; } - trace_netfs_rreq(wreq, netfs_rreq_trace_set_pause); - set_bit(NETFS_RREQ_PAUSE, &wreq->flags); } else { if (WARN(transferred_or_error > subreq->len - subreq->transferred, "Subreq excess write: R=%x[%x] %zd > %zu - %zu", diff --git a/fs/netfs/write_issue.c b/fs/netfs/write_issue.c index 851f6f93ad45..3989b4ec0c4b 100644 --- a/fs/netfs/write_issue.c +++ b/fs/netfs/write_issue.c @@ -89,14 +89,13 @@ static void netfs_kill_dirty_pages(struct address_space *mapping, */ struct netfs_io_request *netfs_create_write_req(struct address_space *mapping, struct file *file, - loff_t start, + uoff_t start, enum netfs_io_origin origin) { struct netfs_io_request *wreq; struct netfs_inode *ictx; bool is_cacheable = (origin == NETFS_WRITEBACK || origin == NETFS_WRITEBACK_SINGLE || - origin == NETFS_WRITETHROUGH || origin == NETFS_PGPRIV2_COPY_TO_CACHE); wreq = netfs_alloc_request(mapping, file, start, 0, origin); @@ -112,6 +111,8 @@ struct netfs_io_request *netfs_create_write_req(struct address_space *mapping, goto nomem; wreq->cleaned_to = wreq->start; + if (wreq->cache_resources.dio_size > 1) + wreq->cache_coll_to = round_down(wreq->start, wreq->cache_resources.dio_size); wreq->io_streams[0].stream_nr = 0; wreq->io_streams[0].source = NETFS_UPLOAD_TO_SERVER; @@ -156,7 +157,7 @@ EXPORT_SYMBOL(netfs_prepare_write_failed); */ void netfs_prepare_write(struct netfs_io_request *wreq, struct netfs_io_stream *stream, - loff_t start) + uoff_t start) { struct netfs_io_subrequest *subreq; struct iov_iter *wreq_iter = &wreq->buffer.iter; @@ -169,10 +170,9 @@ void netfs_prepare_write(struct netfs_io_request *wreq, wreq_iter->folioq_slot >= folioq_nr_slots(wreq_iter->folioq)) rolling_buffer_make_space(&wreq->buffer, wreq->gfp); - subreq = netfs_alloc_subrequest(wreq); + subreq = netfs_alloc_subrequest(wreq, stream->source); if (!subreq) return; - subreq->source = stream->source; subreq->start = start; subreq->stream_nr = stream->stream_nr; subreq->io_iter = *wreq_iter; @@ -233,6 +233,21 @@ static void netfs_do_issue_write(struct netfs_io_stream *stream, _enter("R=%x[%x],%zx", wreq->debug_id, subreq->debug_index, subreq->len); + if (stream->source == NETFS_WRITE_TO_CACHE && + unlikely(test_bit(NETFS_RREQ_CACHE_STOP, &wreq->flags))) { + size_t dio_size = wreq->cache_resources.dio_size; + size_t len, disp; + + disp = subreq->start & (dio_size - 1); + len = round_up(subreq->len + disp, dio_size); + + subreq->start -= disp; + subreq->len = len; + + __set_bit(NETFS_SREQ_CANCELLED, &subreq->flags); + return netfs_write_subrequest_terminated(subreq, subreq->len); + } + if (test_bit(NETFS_SREQ_FAILED, &subreq->flags)) return netfs_write_subrequest_terminated(subreq, subreq->error); @@ -266,6 +281,7 @@ void netfs_issue_write(struct netfs_io_request *wreq, if (!subreq) return; + stream->construct = NULL; subreq->io_iter.count = subreq->len; netfs_do_issue_write(stream, subreq); @@ -279,7 +295,7 @@ void netfs_issue_write(struct netfs_io_request *wreq, */ size_t netfs_advance_write(struct netfs_io_request *wreq, struct netfs_io_stream *stream, - loff_t start, size_t len, bool to_eof) + uoff_t start, size_t len, bool to_eof) { struct netfs_io_subrequest *subreq = stream->construct; size_t part; @@ -330,7 +346,7 @@ static int netfs_write_folio(struct netfs_io_request *wreq, struct netfs_folio *finfo; size_t iter_off = 0; size_t fsize = folio_size(folio), flen = fsize, foff = 0; - loff_t fpos = folio_pos(folio), i_size; + uoff_t fpos = folio_pos(folio), i_size; bool to_eof = false, streamw = false; bool debug = false; @@ -367,11 +383,7 @@ static int netfs_write_folio(struct netfs_io_request *wreq, streamw = true; } - if (wreq->origin == NETFS_WRITETHROUGH) { - to_eof = false; - if (flen > i_size - fpos) - flen = i_size - fpos; - } else if (flen > i_size - fpos) { + if (flen > i_size - fpos) { flen = i_size - fpos; if (!streamw) folio_zero_segment(folio, flen, fsize); @@ -525,8 +537,7 @@ static void netfs_end_issue_write(struct netfs_io_request *wreq) { bool needs_poke = true; - smp_wmb(); /* Write subreq lists before ALL_QUEUED. */ - set_bit(NETFS_RREQ_ALL_QUEUED, &wreq->flags); + netfs_all_subreqs_queued(wreq); for (int s = 0; s < NR_IO_STREAMS; s++) { struct netfs_io_stream *stream = &wreq->io_streams[s]; @@ -614,103 +625,6 @@ out: EXPORT_SYMBOL(netfs_writepages); /* - * Begin a write operation for writing through the pagecache. - */ -struct netfs_io_request *netfs_begin_writethrough(struct kiocb *iocb, size_t len) -{ - struct netfs_io_request *wreq = NULL; - struct netfs_inode *ictx = netfs_inode(file_inode(iocb->ki_filp)); - - netfs_wb_begin(ictx, false); - - wreq = netfs_create_write_req(iocb->ki_filp->f_mapping, iocb->ki_filp, - iocb->ki_pos, NETFS_WRITETHROUGH); - if (IS_ERR(wreq)) { - netfs_wb_end(ictx); - return wreq; - } - - wreq->io_streams[0].avail = true; - __set_bit(NETFS_RREQ_OFFLOAD_COLLECTION, &wreq->flags); - trace_netfs_write(wreq, netfs_write_trace_writethrough); - return wreq; -} - -/* - * Advance the state of the write operation used when writing through the - * pagecache. Data has been copied into the pagecache that we need to append - * to the request. If we've added more than wsize then we need to create a new - * subrequest. - */ -int netfs_advance_writethrough(struct netfs_io_request *wreq, struct writeback_control *wbc, - struct folio *folio, size_t copied, bool to_page_end, - struct folio **writethrough_cache) -{ - int ret; - - _enter("R=%x ic=%zu ws=%u cp=%zu tp=%u", - wreq->debug_id, wreq->buffer.iter.count, wreq->wsize, copied, to_page_end); - - /* The folio is locked. */ - - if (*writethrough_cache != folio) { - if (*writethrough_cache) { - /* Did the folio get moved? */ - folio_put(*writethrough_cache); - *writethrough_cache = NULL; - } - /* We can make multiple writes to the folio... */ - if (wreq->len == 0) - trace_netfs_folio(folio, netfs_folio_trace_wthru); - else - trace_netfs_folio(folio, netfs_folio_trace_wthru_plus); - *writethrough_cache = folio; - folio_get(folio); - } - - wreq->len += copied; - - if (!to_page_end) { - folio_mark_dirty(folio); - folio_unlock(folio); - return 0; - } - - ret = netfs_write_folio(wreq, wbc, folio); - folio_put(*writethrough_cache); - *writethrough_cache = NULL; - wreq->submitted = wreq->len; - return ret; -} - -/* - * End a write operation used when writing through the pagecache. - */ -ssize_t netfs_end_writethrough(struct netfs_io_request *wreq, struct writeback_control *wbc, - struct folio *writethrough_cache) -{ - ssize_t ret; - - _enter("R=%x", wreq->debug_id); - - if (writethrough_cache) { - folio_lock(writethrough_cache); - netfs_write_folio(wreq, wbc, writethrough_cache); - folio_put(writethrough_cache); - wreq->submitted = wreq->len; - } - - netfs_end_issue_write(wreq); - - if (wreq->iocb) - ret = -EIOCBQUEUED; - else - ret = netfs_wait_for_write(wreq); - netfs_put_request(wreq, netfs_rreq_trace_put_return); - return ret; -} - -/* * Write some of a pending folio data back to the server and/or the cache. */ static int netfs_write_folio_single(struct netfs_io_request *wreq, @@ -721,7 +635,7 @@ static int netfs_write_folio_single(struct netfs_io_request *wreq, struct netfs_io_stream *stream; size_t iter_off = 0; size_t fsize = folio_size(folio), flen; - loff_t fpos = folio_pos(folio); + uoff_t fpos = folio_pos(folio); ssize_t ret; bool to_eof = false; bool no_debug = false; @@ -891,8 +805,7 @@ int netfs_writeback_single(struct address_space *mapping, stop: for (int s = 0; s < NR_IO_STREAMS; s++) netfs_issue_write(wreq, &wreq->io_streams[s]); - smp_wmb(); /* Write lists before ALL_QUEUED. */ - set_bit(NETFS_RREQ_ALL_QUEUED, &wreq->flags); + netfs_all_subreqs_queued(wreq); netfs_wake_collector(wreq); diff --git a/fs/netfs/write_retry.c b/fs/netfs/write_retry.c index 058bc7a166a5..2f20577563e1 100644 --- a/fs/netfs/write_retry.c +++ b/fs/netfs/write_retry.c @@ -55,7 +55,7 @@ static void netfs_retry_write_stream(struct netfs_io_request *wreq, do { struct netfs_io_subrequest *subreq = NULL, *from, *to, *tmp; struct iov_iter source; - unsigned long long start, len; + uoff_t start, len; size_t part; bool boundary = false; @@ -149,8 +149,7 @@ static void netfs_retry_write_stream(struct netfs_io_request *wreq, * and insert them after. */ do { - subreq = netfs_alloc_subrequest(wreq); - subreq->source = to->source; + subreq = netfs_alloc_subrequest(wreq, stream->source); subreq->start = start; subreq->stream_nr = to->stream_nr; subreq->retry_count = 1; @@ -211,6 +210,7 @@ void netfs_retry_writes(struct netfs_io_request *wreq) int s; netfs_stat(&netfs_n_wh_retry_write_req); + trace_netfs_rreq(wreq, netfs_rreq_trace_retry_begin); /* Wait for all outstanding I/O to quiesce before performing retries as * we may need to renegotiate the I/O sizes. @@ -235,4 +235,6 @@ void netfs_retry_writes(struct netfs_io_request *wreq) netfs_retry_write_stream(wreq, stream); } } + + trace_netfs_rreq(wreq, netfs_rreq_trace_retry_end); } diff --git a/fs/nfs/Kconfig b/fs/nfs/Kconfig index 6bb30543eff0..fd430718d02b 100644 --- a/fs/nfs/Kconfig +++ b/fs/nfs/Kconfig @@ -123,8 +123,23 @@ config PNFS_FILE_LAYOUT config PNFS_BLOCK tristate - depends on NFS_V4 && BLK_DEV_DM + +config PNFS_BLOCK_LAYOUT + bool "NFS client support for pNFS block layouts" + depends on NFS_V4 && BLOCK && BLK_DEV_DM + select PNFS_BLOCK + help + Enable support for the pNFS block-volume layout type (RFC 5663). + + If unsure, say N. + +config PNFS_SCSI_LAYOUT + bool "NFS client support for pNFS SCSI layouts" + depends on NFS_V4 && BLOCK default NFS_V4 + select PNFS_BLOCK + help + Enable suport for the pNFS SCSI layout type (RFC 8154). config PNFS_FLEXFILE_LAYOUT tristate @@ -174,6 +189,7 @@ config NFS_FSCACHE bool "Provide NFS client caching support" depends on NFS_FS select NETFS_SUPPORT + select NETFS_PGPRIV2 select FSCACHE help Say Y here if you want NFS data to be cached locally on disc through diff --git a/fs/nfs/blocklayout/Makefile b/fs/nfs/blocklayout/Makefile index 7668a1bfb5fa..3403cb7fe201 100644 --- a/fs/nfs/blocklayout/Makefile +++ b/fs/nfs/blocklayout/Makefile @@ -4,4 +4,5 @@ # obj-$(CONFIG_PNFS_BLOCK) += blocklayoutdriver.o -blocklayoutdriver-y += blocklayout.o dev.o extent_tree.o rpc_pipefs.o +blocklayoutdriver-y += blocklayout.o dev.o extent_tree.o +blocklayoutdriver-$(CONFIG_PNFS_BLOCK_LAYOUT) += rpc_pipefs.o diff --git a/fs/nfs/blocklayout/blocklayout.c b/fs/nfs/blocklayout/blocklayout.c index d54a141a89b3..82873ff370ef 100644 --- a/fs/nfs/blocklayout/blocklayout.c +++ b/fs/nfs/blocklayout/blocklayout.c @@ -470,18 +470,6 @@ static struct pnfs_layout_hdr *__bl_alloc_layout_hdr(struct inode *inode, return &bl->bl_layout; } -static struct pnfs_layout_hdr *bl_alloc_layout_hdr(struct inode *inode, - gfp_t gfp_flags) -{ - return __bl_alloc_layout_hdr(inode, gfp_flags, false); -} - -static struct pnfs_layout_hdr *sl_alloc_layout_hdr(struct inode *inode, - gfp_t gfp_flags) -{ - return __bl_alloc_layout_hdr(inode, gfp_flags, true); -} - static void bl_free_lseg(struct pnfs_layout_segment *lseg) { dprintk("%s enter\n", __func__); @@ -859,7 +847,8 @@ bl_pg_init_read(struct nfs_pageio_descriptor *pgio, struct nfs_page *req) if (pgio->pg_lseg && test_bit(NFS_LSEG_UNAVAILABLE, &pgio->pg_lseg->pls_flags)) { - pnfs_error_mark_layout_for_return(pgio->pg_inode, pgio->pg_lseg); + pnfs_error_mark_layout_for_return(pgio->pg_inode, pgio->pg_lseg, + NULL); pnfs_set_lo_fail(pgio->pg_lseg); nfs_pageio_reset_read_mds(pgio); } @@ -921,7 +910,8 @@ bl_pg_init_write(struct nfs_pageio_descriptor *pgio, struct nfs_page *req) if (pgio->pg_lseg && test_bit(NFS_LSEG_UNAVAILABLE, &pgio->pg_lseg->pls_flags)) { - pnfs_error_mark_layout_for_return(pgio->pg_inode, pgio->pg_lseg); + pnfs_error_mark_layout_for_return(pgio->pg_inode, pgio->pg_lseg, + NULL); pnfs_set_lo_fail(pgio->pg_lseg); nfs_pageio_reset_write_mds(pgio); } @@ -954,6 +944,13 @@ static const struct nfs_pageio_ops bl_pg_write_ops = { .pg_cleanup = pnfs_generic_pg_cleanup, }; +#ifdef CONFIG_PNFS_BLOCK_LAYOUT +static struct pnfs_layout_hdr *bl_alloc_layout_hdr(struct inode *inode, + gfp_t gfp_flags) +{ + return __bl_alloc_layout_hdr(inode, gfp_flags, false); +} + static struct pnfs_layoutdriver_type blocklayout_type = { .id = LAYOUT_BLOCK_VOLUME, .name = "LAYOUT_BLOCK_VOLUME", @@ -978,6 +975,48 @@ static struct pnfs_layoutdriver_type blocklayout_type = { .sync = pnfs_generic_sync, }; +static int __init pnfs_register_blocklayout(void) +{ + int ret; + + ret = bl_init_pipefs(); + if (ret) + return ret; + + ret = pnfs_register_layoutdriver(&blocklayout_type); + if (ret) { + bl_cleanup_pipefs(); + return ret; + } + + return 0; +} + +static void __exit pnfs_unregister_blocklayout(void) +{ + pnfs_unregister_layoutdriver(&blocklayout_type); + bl_cleanup_pipefs(); +} + +MODULE_ALIAS("nfs-layouttype4-3"); +#else +static int __init pnfs_register_blocklayout(void) +{ + return 0; +} + +static void __exit pnfs_unregister_blocklayout(void) +{ +} +#endif /* CONFIG_PNFS_BLOCK_LAYOUT */ + +#ifdef CONFIG_PNFS_SCSI_LAYOUT +static struct pnfs_layout_hdr *sl_alloc_layout_hdr(struct inode *inode, + gfp_t gfp_flags) +{ + return __bl_alloc_layout_hdr(inode, gfp_flags, true); +} + static struct pnfs_layoutdriver_type scsilayout_type = { .id = LAYOUT_SCSI, .name = "LAYOUT_SCSI", @@ -1002,6 +1041,28 @@ static struct pnfs_layoutdriver_type scsilayout_type = { .sync = pnfs_generic_sync, }; +static int __init pnfs_register_scsilayout(void) +{ + return pnfs_register_layoutdriver(&scsilayout_type); +} + +static void __exit pnfs_unregister_scsilayout(void) +{ + pnfs_unregister_layoutdriver(&scsilayout_type); +} + +MODULE_ALIAS("nfs-layouttype4-5"); +#else +static int __init pnfs_register_scsilayout(void) +{ + return 0; +} + +static void __exit pnfs_unregister_scsilayout(void) +{ +} +#endif /* CONFIG_PNFS_SCSI_LAYOUT */ + static int __init nfs4blocklayout_init(void) { @@ -1009,25 +1070,17 @@ static int __init nfs4blocklayout_init(void) dprintk("%s: NFSv4 Block Layout Driver Registering...\n", __func__); - ret = bl_init_pipefs(); + ret = pnfs_register_blocklayout(); if (ret) - goto out; + return ret; - ret = pnfs_register_layoutdriver(&blocklayout_type); - if (ret) - goto out_cleanup_pipe; + ret = pnfs_register_scsilayout(); + if (ret) { + pnfs_unregister_blocklayout(); + return ret; + } - ret = pnfs_register_layoutdriver(&scsilayout_type); - if (ret) - goto out_unregister_block; return 0; - -out_unregister_block: - pnfs_unregister_layoutdriver(&blocklayout_type); -out_cleanup_pipe: - bl_cleanup_pipefs(); -out: - return ret; } static void __exit nfs4blocklayout_exit(void) @@ -1035,13 +1088,9 @@ static void __exit nfs4blocklayout_exit(void) dprintk("%s: NFSv4 Block Layout Driver Unregistering...\n", __func__); - pnfs_unregister_layoutdriver(&scsilayout_type); - pnfs_unregister_layoutdriver(&blocklayout_type); - bl_cleanup_pipefs(); + pnfs_unregister_scsilayout(); + pnfs_unregister_blocklayout(); } -MODULE_ALIAS("nfs-layouttype4-3"); -MODULE_ALIAS("nfs-layouttype4-5"); - module_init(nfs4blocklayout_init); module_exit(nfs4blocklayout_exit); diff --git a/fs/nfs/blocklayout/dev.c b/fs/nfs/blocklayout/dev.c index c926b7e43827..00462e6affb1 100644 --- a/fs/nfs/blocklayout/dev.c +++ b/fs/nfs/blocklayout/dev.c @@ -15,6 +15,7 @@ #define NFSDBG_FACILITY NFSDBG_PNFS_LD +#ifdef CONFIG_PNFS_SCSI_LAYOUT static void bl_unregister_scsi(struct pnfs_block_dev *dev) { struct block_device *bdev = file_bdev(dev->bdev_file); @@ -45,6 +46,16 @@ static bool bl_register_scsi(struct pnfs_block_dev *dev) trace_bl_pr_key_reg(bdev, dev->pr_key); return true; } +#else +static void bl_unregister_scsi(struct pnfs_block_dev *dev) +{ +} + +static bool bl_register_scsi(struct pnfs_block_dev *dev) +{ + return false; +} +#endif /* CONFIG_PNFS_SCSI_LAYOUT */ static void bl_unregister_dev(struct pnfs_block_dev *dev) { @@ -292,7 +303,7 @@ static int bl_parse_deviceid(struct nfs_server *server, struct pnfs_block_dev *d, struct pnfs_block_volume *volumes, int idx, gfp_t gfp_mask); - +#ifdef CONFIG_PNFS_BLOCK_LAYOUT static int bl_parse_simple(struct nfs_server *server, struct pnfs_block_dev *d, struct pnfs_block_volume *volumes, int idx, gfp_t gfp_mask) @@ -320,7 +331,17 @@ bl_parse_simple(struct nfs_server *server, struct pnfs_block_dev *d, file_bdev(bdev_file)->bd_disk->disk_name); return 0; } +#else +static int +bl_parse_simple(struct nfs_server *server, struct pnfs_block_dev *d, + struct pnfs_block_volume *volumes, int idx, gfp_t gfp_mask) +{ + dprintk("unsupported volume type: %d\n", PNFS_BLOCK_VOLUME_SIMPLE); + return -EIO; +} +#endif /* CONFIG_PNFS_BLOCK_LAYOUT */ +#ifdef CONFIG_PNFS_SCSI_LAYOUT static bool bl_validate_designator(struct pnfs_block_volume *v) { @@ -449,6 +470,15 @@ out_blkdev_put: d->bdev_file = NULL; return error; } +#else +static int +bl_parse_scsi(struct nfs_server *server, struct pnfs_block_dev *d, + struct pnfs_block_volume *volumes, int idx, gfp_t gfp_mask) +{ + dprintk("unsupported volume type: %d\n", PNFS_BLOCK_VOLUME_SCSI); + return -EIO; +} +#endif /* CONFIG_PNFS_SCSI_LAYOUT */ static int bl_parse_slice(struct nfs_server *server, struct pnfs_block_dev *d, diff --git a/fs/nfs/callback_proc.c b/fs/nfs/callback_proc.c index 3fb10c8e4271..9d836cb72e78 100644 --- a/fs/nfs/callback_proc.c +++ b/fs/nfs/callback_proc.c @@ -292,7 +292,7 @@ static u32 initiate_file_draining(struct nfs_client *clp, switch (pnfs_mark_matching_lsegs_return(lo, &free_me_list, &args->cbl_range, be32_to_cpu(args->cbl_stateid.seqid), - args->cbl_layoutchanged)) { + args->cbl_layoutchanged, NULL)) { case 0: case -EBUSY: /* There are layout segments that need to be returned */ @@ -391,7 +391,26 @@ __be32 nfs4_callback_devicenotify(void *argp, void *resp, if (!ld) continue; } - nfs4_delete_deviceid(ld, cps->clp, &dev->cbd_dev_id); + /* + * Bump the epoch before touching the cache so a + * GETDEVICEINFO already in flight can detect that it + * predates the notification. A referenced DELETE may be + * racing revocation, so defer it to the state manager -- + * this thread cannot issue fore-channel RPCs. + */ + nfs4_deviceid_bump_change_epoch(cps->clp); + if (dev->cbd_notify_type == NOTIFY_DEVICEID4_CHANGE) { + nfs4_delete_deviceid(ld, cps->clp, &dev->cbd_dev_id); + pnfs_layout_reresolve_deviceid_byclid(cps->clp, ld, + &dev->cbd_dev_id, + dev->cbd_immediate); + } else if (pnfs_layout_deviceid_referenced_byclid(cps->clp, + ld, &dev->cbd_dev_id)) { + pnfs_deviceid_delete_mark(cps->clp, ld, + &dev->cbd_dev_id); + } else { + nfs4_delete_deviceid(ld, cps->clp, &dev->cbd_dev_id); + } } pnfs_put_layoutdriver(ld); out: diff --git a/fs/nfs/callback_xdr.c b/fs/nfs/callback_xdr.c index 88e1af0fda01..5a1f05469268 100644 --- a/fs/nfs/callback_xdr.c +++ b/fs/nfs/callback_xdr.c @@ -271,6 +271,13 @@ __be32 decode_devicenotify_args(struct svc_rqst *rqstp, if (n == 0) goto out; + /* sanity check the count against the remaining stream */ + if (n > xdr_stream_remaining(xdr) / + ((4 * sizeof(uint32_t)) + NFS4_DEVICEID4_SIZE)) { + status = htonl(NFS4ERR_BADXDR); + goto out; + } + args->devs = kmalloc_objs(*args->devs, n); if (!args->devs) { status = htonl(NFS4ERR_DELAY); @@ -312,7 +319,7 @@ __be32 decode_devicenotify_args(struct svc_rqst *rqstp, memcpy(dev->cbd_dev_id.data, p, NFS4_DEVICEID4_SIZE); p += XDR_QUADLEN(NFS4_DEVICEID4_SIZE); - if (dev->cbd_layout_type == NOTIFY_DEVICEID4_CHANGE) { + if (dev->cbd_notify_type == NOTIFY_DEVICEID4_CHANGE) { p = xdr_inline_decode(xdr, sizeof(uint32_t)); if (unlikely(p == NULL)) { status = htonl(NFS4ERR_BADXDR); diff --git a/fs/nfs/client.c b/fs/nfs/client.c index 60386330aeec..dbb5131375f8 100644 --- a/fs/nfs/client.c +++ b/fs/nfs/client.c @@ -1295,7 +1295,8 @@ void nfs_clients_init(struct net *net) INIT_LIST_HEAD(&nn->nfs_volume_list); #if IS_ENABLED(CONFIG_NFS_V4) idr_init(&nn->cb_ident_idr); - INIT_LIST_HEAD(&nn->nfs4_data_server_cache); + for (int i = 0; i < NFS4_DS_CACHE_HASH_SIZE; i++) + INIT_HLIST_HEAD(&nn->nfs4_data_server_cache[i]); spin_lock_init(&nn->nfs4_data_server_lock); #endif /* CONFIG_NFS_V4 */ spin_lock_init(&nn->nfs_client_lock); @@ -1315,7 +1316,8 @@ void nfs_clients_exit(struct net *net) WARN_ON_ONCE(!list_empty(&nn->nfs_client_list)); WARN_ON_ONCE(!list_empty(&nn->nfs_volume_list)); #if IS_ENABLED(CONFIG_NFS_V4) - WARN_ON_ONCE(!list_empty(&nn->nfs4_data_server_cache)); + for (int i = 0; i < NFS4_DS_CACHE_HASH_SIZE; i++) + WARN_ON_ONCE(!hlist_empty(&nn->nfs4_data_server_cache[i])); #endif /* CONFIG_NFS_V4 */ } diff --git a/fs/nfs/dir.c b/fs/nfs/dir.c index 49394123bd09..354f986e60c4 100644 --- a/fs/nfs/dir.c +++ b/fs/nfs/dir.c @@ -2437,7 +2437,7 @@ out_err: return error; } -int nfs_create(struct mnt_idmap *idmap, struct inode *dir, +int nfs_create(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *dentry, umode_t mode) { return nfs_do_create(dir, dentry, mode, O_EXCL); @@ -2448,7 +2448,7 @@ EXPORT_SYMBOL_GPL(nfs_create); * See comments for nfs_proc_create regarding failed operations. */ int -nfs_mknod(struct mnt_idmap *idmap, struct inode *dir, +nfs_mknod(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *dentry, umode_t mode, dev_t rdev) { struct iattr attr; @@ -2475,7 +2475,7 @@ EXPORT_SYMBOL_GPL(nfs_mknod); /* * See comments for nfs_proc_create regarding failed operations. */ -struct dentry *nfs_mkdir(struct mnt_idmap *idmap, struct inode *dir, +struct dentry *nfs_mkdir(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *dentry, umode_t mode) { struct iattr attr; @@ -2641,7 +2641,7 @@ EXPORT_SYMBOL_GPL(nfs_unlink); * now have a new file handle and can instantiate an in-core NFS inode * and move the raw page into its mapping. */ -int nfs_symlink(struct mnt_idmap *idmap, struct inode *dir, +int nfs_symlink(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *dentry, const char *symname) { struct folio *folio; @@ -2771,7 +2771,7 @@ static bool nfs_rename_is_unsafe_cross_dir(struct dentry *old_dentry, * If these conditions are met, we can drop the dentries before doing * the rename. */ -int nfs_rename(struct mnt_idmap *idmap, struct inode *old_dir, +int nfs_rename(const struct mnt_idmap *idmap, struct inode *old_dir, struct dentry *old_dentry, struct inode *new_dir, struct dentry *new_dentry, unsigned int flags) { @@ -3393,7 +3393,7 @@ static int nfs_execute_ok(struct inode *inode, int mask) return ret; } -int nfs_permission(struct mnt_idmap *idmap, +int nfs_permission(const struct mnt_idmap *idmap, struct inode *inode, int mask) { diff --git a/fs/nfs/direct.c b/fs/nfs/direct.c index ccafdc1ce64d..a3c6e8f4ea06 100644 --- a/fs/nfs/direct.c +++ b/fs/nfs/direct.c @@ -145,11 +145,55 @@ static void nfs_direct_file_adjust_size_locked(struct inode *inode, } } -static void nfs_direct_release_pages(struct page **pages, unsigned int npages) +static void nfs_direct_release_pages(struct page **pages, unsigned int npages, + bool pinned) { - unsigned int i; - for (i = 0; i < npages; i++) - put_page(pages[i]); + if (pinned) + unpin_user_pages(pages, npages); +} + +static ssize_t nfs_direct_extract_pages(struct nfs_direct_req *dreq, + struct iov_iter *iter, + size_t size, loff_t *pos, + struct list_head *list) +{ + bool pinned = iov_iter_extract_will_pin(iter); + struct page **pagevec = NULL; + ssize_t result, bytes = 0; + int err = 0; + unsigned int npages, i; + size_t pgbase; + + result = iov_iter_extract_pages(iter, &pagevec, size, ~0U, 0, &pgbase); + if (result <= 0) + return result; + + npages = (result + pgbase + PAGE_SIZE - 1) >> PAGE_SHIFT; + for (i = 0; i < npages; i++) { + struct nfs_page *req; + unsigned int req_len = min_t(size_t, result - bytes, PAGE_SIZE - pgbase); + + req = nfs_page_create_from_page(dreq->ctx, pagevec[i], + pinned, pgbase, *pos, + req_len); + if (IS_ERR(req)) { + err = PTR_ERR(req); + break; + } + + list_add_tail(&req->wb_list, list); + pgbase = 0; + bytes += req_len; + *pos += req_len; + } + + if (i < npages) { + iov_iter_revert(iter, result - bytes); + nfs_direct_release_pages(pagevec + i, npages - i, pinned); + } + + kvfree(pagevec); + return bytes ? bytes : err; } void nfs_init_cinfo_from_dreq(struct nfs_commit_info *cinfo, @@ -288,13 +332,7 @@ out_put: static void nfs_read_sync_pgio_error(struct list_head *head, int error) { - struct nfs_page *req; - - while (!list_empty(head)) { - req = nfs_list_entry(head->next); - nfs_list_remove_request(req); - nfs_release_request(req); - } + nfs_release_request_list(head); } static void nfs_direct_pgio_init(struct nfs_pgio_header *hdr) @@ -326,6 +364,7 @@ static ssize_t nfs_direct_read_schedule_iovec(struct nfs_direct_req *dreq, ssize_t result = -EINVAL; size_t requested_bytes = 0; size_t rsize = max_t(size_t, NFS_SERVER(inode)->rsize, PAGE_SIZE); + LIST_HEAD(nfs_page_list); nfs_pageio_init_read(&desc, dreq->inode, false, &nfs_direct_read_completion_ops); @@ -334,40 +373,23 @@ static ssize_t nfs_direct_read_schedule_iovec(struct nfs_direct_req *dreq, inode_dio_begin(inode); while (iov_iter_count(iter)) { - struct page **pagevec; - size_t bytes; - size_t pgbase; - unsigned npages, i; - - result = iov_iter_get_pages_alloc2(iter, &pagevec, - rsize, &pgbase); + result = nfs_direct_extract_pages(dreq, iter, rsize, &pos, &nfs_page_list); if (result < 0) break; - - bytes = result; - npages = (result + pgbase + PAGE_SIZE - 1) / PAGE_SIZE; - for (i = 0; i < npages; i++) { - struct nfs_page *req; - unsigned int req_len = min_t(size_t, bytes, PAGE_SIZE - pgbase); - /* XXX do we need to do the eof zeroing found in async_filler? */ - req = nfs_page_create_from_page(dreq->ctx, pagevec[i], - pgbase, pos, req_len); - if (IS_ERR(req)) { - result = PTR_ERR(req); - break; - } + + while (!list_empty(&nfs_page_list)) { + struct nfs_page *req = nfs_list_entry(nfs_page_list.next); + size_t req_len = req->wb_bytes; + + nfs_list_remove_request(req); if (!nfs_pageio_add_request(&desc, req)) { result = desc.pg_error; nfs_release_request(req); + nfs_release_request_list(&nfs_page_list); break; } - pgbase = 0; - bytes -= req_len; requested_bytes += req_len; - pos += req_len; } - nfs_direct_release_pages(pagevec, npages); - kvfree(pagevec); if (result < 0) break; } @@ -858,6 +880,7 @@ static ssize_t nfs_direct_write_schedule_iovec(struct nfs_direct_req *dreq, ssize_t result = 0; size_t requested_bytes = 0; size_t wsize = max_t(size_t, NFS_SERVER(inode)->wsize, PAGE_SIZE); + LIST_HEAD(nfs_page_list); bool defer = false; trace_nfs_direct_write_schedule_iovec(dreq); @@ -870,53 +893,32 @@ static ssize_t nfs_direct_write_schedule_iovec(struct nfs_direct_req *dreq, NFS_I(inode)->write_io += iov_iter_count(iter); while (iov_iter_count(iter)) { - struct page **pagevec; - size_t bytes; - size_t pgbase; - unsigned npages, i; - - result = iov_iter_get_pages_alloc2(iter, &pagevec, - wsize, &pgbase); + result = nfs_direct_extract_pages(dreq, iter, wsize, &pos, &nfs_page_list); if (result < 0) break; - bytes = result; - npages = (result + pgbase + PAGE_SIZE - 1) / PAGE_SIZE; - for (i = 0; i < npages; i++) { - struct nfs_page *req; - unsigned int req_len = min_t(size_t, bytes, PAGE_SIZE - pgbase); - - req = nfs_page_create_from_page(dreq->ctx, pagevec[i], - pgbase, pos, req_len); - if (IS_ERR(req)) { - result = PTR_ERR(req); - break; - } - - if (desc.pg_error < 0) { - nfs_free_request(req); - result = desc.pg_error; - break; - } - - pgbase = 0; - bytes -= req_len; - requested_bytes += req_len; - pos += req_len; + while (!list_empty(&nfs_page_list)) { + struct nfs_page *req = nfs_list_entry(nfs_page_list.next); + size_t req_len = req->wb_bytes; + nfs_list_remove_request(req); if (defer) { nfs_mark_request_commit(req, NULL, &cinfo, 0); + requested_bytes += req_len; continue; } nfs_lock_request(req); - if (nfs_pageio_add_request(&desc, req)) + if (nfs_pageio_add_request(&desc, req)) { + requested_bytes += req_len; continue; + } /* Exit on hard errors */ if (desc.pg_error < 0 && desc.pg_error != -EAGAIN) { result = desc.pg_error; nfs_unlock_and_release_request(req); + nfs_release_request_list(&nfs_page_list); break; } @@ -927,11 +929,10 @@ static ssize_t nfs_direct_write_schedule_iovec(struct nfs_direct_req *dreq, spin_unlock(&dreq->lock); nfs_unlock_request(req); nfs_mark_request_commit(req, NULL, &cinfo, 0); + requested_bytes += req_len; desc.pg_error = 0; defer = true; } - nfs_direct_release_pages(pagevec, npages); - kvfree(pagevec); if (result < 0) break; } diff --git a/fs/nfs/filelayout/filelayout.c b/fs/nfs/filelayout/filelayout.c index 72e20b56fbc7..53808e42546a 100644 --- a/fs/nfs/filelayout/filelayout.c +++ b/fs/nfs/filelayout/filelayout.c @@ -55,12 +55,13 @@ static loff_t filelayout_get_dense_offset(struct nfs4_filelayout_segment *flseg, loff_t offset) { - u32 stripe_width = flseg->stripe_unit * flseg->dsaddr->stripe_count; + u64 stripe_width = (u64)flseg->stripe_unit * + flseg->dsaddr->stripe_count; u64 stripe_no; u32 rem; offset -= flseg->pattern_offset; - stripe_no = div_u64(offset, stripe_width); + stripe_no = div64_u64(offset, stripe_width); div_u64_rem(offset, flseg->stripe_unit, &rem); return stripe_no * flseg->stripe_unit + rem; @@ -186,7 +187,7 @@ static int filelayout_async_handle_error(struct rpc_task *task, dprintk("%s DS connection error %d\n", __func__, task->tk_status); nfs4_mark_deviceid_unavailable(devid); - pnfs_error_mark_layout_for_return(inode, lseg); + pnfs_error_mark_layout_for_return(inode, lseg, NULL); pnfs_set_lo_fail(lseg); rpc_wake_up(&tbl->slot_tbl_waitq); fallthrough; @@ -796,7 +797,7 @@ filelayout_pg_test(struct nfs_pageio_descriptor *pgio, struct nfs_page *prev, unsigned int size; u64 p_stripe, r_stripe; u32 stripe_offset; - u64 segment_offset = pgio->pg_lseg->pls_range.offset; + u64 pattern_offset = FILELAYOUT_LSEG(pgio->pg_lseg)->pattern_offset; u32 stripe_unit = FILELAYOUT_LSEG(pgio->pg_lseg)->stripe_unit; /* calls nfs_generic_pg_test */ @@ -808,8 +809,8 @@ filelayout_pg_test(struct nfs_pageio_descriptor *pgio, struct nfs_page *prev, /* see if req and prev are in the same stripe */ if (prev) { - p_stripe = (u64)req_offset(prev) - segment_offset; - r_stripe = (u64)req_offset(req) - segment_offset; + p_stripe = (u64)req_offset(prev) - pattern_offset; + r_stripe = (u64)req_offset(req) - pattern_offset; do_div(p_stripe, stripe_unit); do_div(r_stripe, stripe_unit); @@ -818,7 +819,7 @@ filelayout_pg_test(struct nfs_pageio_descriptor *pgio, struct nfs_page *prev, } /* calculate remaining bytes in the current stripe */ - div_u64_rem((u64)req_offset(req) - segment_offset, + div_u64_rem((u64)req_offset(req) - pattern_offset, stripe_unit, &stripe_offset); WARN_ON_ONCE(stripe_offset > stripe_unit); @@ -856,7 +857,7 @@ fl_pnfs_update_layout(struct inode *ino, status = filelayout_check_deviceid(lo, fl, gfp_flags); if (status) { - pnfs_error_mark_layout_for_return(ino, lseg); + pnfs_error_mark_layout_for_return(ino, lseg, NULL); pnfs_set_lo_fail(lseg); pnfs_put_lseg(lseg); lseg = NULL; diff --git a/fs/nfs/filelayout/filelayoutdev.c b/fs/nfs/filelayout/filelayoutdev.c index 88bc79ec3459..5995f32f07ec 100644 --- a/fs/nfs/filelayout/filelayoutdev.c +++ b/fs/nfs/filelayout/filelayoutdev.c @@ -177,27 +177,14 @@ nfs4_fl_alloc_deviceid_node(struct nfs_server *server, struct pnfs_device *pdev, trace_fl_getdevinfo(server, &pdev->dev_id, dsaddr->ds_list[i]->ds_remotestr); /* If DS was already in cache, free ds addrs */ - while (!list_empty(&dsaddrs)) { - da = list_first_entry(&dsaddrs, - struct nfs4_pnfs_ds_addr, - da_node); - list_del_init(&da->da_node); - kfree(da->da_remotestr); - kfree(da); - } + nfs4_pnfs_ds_addr_list_free(&dsaddrs); } folio_put(scratch); return dsaddr; out_err_drain_dsaddrs: - while (!list_empty(&dsaddrs)) { - da = list_first_entry(&dsaddrs, struct nfs4_pnfs_ds_addr, - da_node); - list_del_init(&da->da_node); - kfree(da->da_remotestr); - kfree(da); - } + nfs4_pnfs_ds_addr_list_free(&dsaddrs); out_err_free_deviceid: nfs4_fl_free_deviceid(dsaddr); /* stripe_indicies was part of dsaddr */ @@ -280,7 +267,7 @@ nfs4_fl_prepare_ds(struct pnfs_layout_segment *lseg, u32 ds_idx) goto out_test_devid; status = nfs4_pnfs_ds_connect(s, ds, devid, dataserver_timeo, - dataserver_retrans, 4, + dataserver_retrans, 0, 4, s->nfs_client->cl_minorversion, true); if (status) { nfs4_mark_deviceid_unavailable(devid); diff --git a/fs/nfs/flexfilelayout/flexfilelayout.c b/fs/nfs/flexfilelayout/flexfilelayout.c index 7fe8b91fa47c..94cc324b591f 100644 --- a/fs/nfs/flexfilelayout/flexfilelayout.c +++ b/fs/nfs/flexfilelayout/flexfilelayout.c @@ -44,10 +44,11 @@ static void ff_layout_read_record_layoutstats_done(struct rpc_task *task, static int ff_layout_mirror_prepare_stats(struct pnfs_layout_hdr *lo, struct nfs42_layoutstat_devinfo *devinfo, + struct nfs4_ff_layoutstat_priv *priv, int dev_limit, enum nfs4_ff_op_type type); static void ff_layout_encode_ff_layoutupdate(struct xdr_stream *xdr, const struct nfs42_layoutstat_devinfo *devinfo, - struct nfs4_ff_layout_ds_stripe *dss_info); + struct nfs4_ff_layoutstat_priv *priv); static struct pnfs_layout_hdr * ff_layout_alloc_layout_hdr(struct inode *inode, gfp_t gfp_flags) @@ -285,8 +286,8 @@ static struct nfs4_ff_layout_mirror *ff_layout_alloc_mirror(u32 dss_count, mirror->dss_count = dss_count; mirror->dss = - kzalloc_objs(struct nfs4_ff_layout_ds_stripe, dss_count, - gfp_flags); + kvzalloc_objs(struct nfs4_ff_layout_ds_stripe, dss_count, + gfp_flags); if (mirror->dss == NULL) { kfree(mirror); return NULL; @@ -312,10 +313,12 @@ static void ff_layout_free_mirror(struct nfs4_ff_layout_mirror *mirror) cred = rcu_access_pointer(mirror->dss[dss_id].rw_cred); put_cred(cred); nfs_close_local_fh(&mirror->dss[dss_id].nfl); - nfs4_ff_layout_put_deviceid(mirror->dss[dss_id].mirror_ds); + /* the last reference to the mirror is gone; no concurrency */ + nfs4_ff_layout_put_deviceid(rcu_dereference_protected( + mirror->dss[dss_id].mirror_ds, 1)); } - kfree(mirror->dss); + kvfree(mirror->dss); kfree(mirror); } @@ -857,33 +860,17 @@ nfs4_ff_layout_stat_io_end_write(struct rpc_task *task, spin_unlock(&mirror->lock); } -static void -ff_layout_mark_ds_unreachable(struct pnfs_layout_segment *lseg, u32 idx, u32 dss_id) -{ - struct nfs4_deviceid_node *devid = FF_LAYOUT_DEVID_NODE(lseg, idx, dss_id); - - if (devid) - nfs4_mark_deviceid_unavailable(devid); -} - -static void -ff_layout_mark_ds_reachable(struct pnfs_layout_segment *lseg, u32 idx, u32 dss_id) -{ - struct nfs4_deviceid_node *devid = FF_LAYOUT_DEVID_NODE(lseg, idx, dss_id); - - if (devid) - nfs4_mark_deviceid_available(devid); -} - -static struct nfs4_pnfs_ds * +static struct nfs4_ff_layout_ds * ff_layout_choose_ds_for_read(struct pnfs_layout_segment *lseg, u32 start_idx, u32 *best_idx, - u32 offset, u32 *dss_id, + u64 offset, u32 *dss_id, bool check_device) { struct nfs4_ff_layout_segment *fls = FF_LAYOUT_LSEG(lseg); struct nfs4_ff_layout_mirror *mirror; - struct nfs4_pnfs_ds *ds = ERR_PTR(-EAGAIN); + struct nfs4_ff_layout_ds *mirror_ds; + struct nfs4_ff_layout_ds *ret = ERR_PTR(-EAGAIN); + struct nfs4_pnfs_ds *ds; u32 idx; /* mirrors are initially sorted by efficiency */ @@ -893,70 +880,79 @@ ff_layout_choose_ds_for_read(struct pnfs_layout_segment *lseg, fls->stripe_unit, fls->mirror_array[idx]->dss_count, offset); - ds = nfs4_ff_layout_prepare_ds(lseg, mirror, *dss_id, false); - if (IS_ERR(ds)) + mirror_ds = ff_layout_get_mirror_ds(lseg->pls_layout, mirror, + *dss_id); + ds = nfs4_ff_layout_prepare_ds(lseg, mirror, mirror_ds, + *dss_id, OP_READ); + if (IS_ERR(ds)) { + nfs4_ff_layout_put_deviceid(mirror_ds); + ret = ERR_CAST(ds); continue; + } if (check_device && - nfs4_test_deviceid_unavailable(&mirror->dss[*dss_id].mirror_ds->id_node)) { + nfs4_test_deviceid_unavailable(&mirror_ds->id_node)) { + nfs4_ff_layout_put_deviceid(mirror_ds); // reinitialize the error state in case if this is the last iteration - ds = ERR_PTR(-EINVAL); + ret = ERR_PTR(-EINVAL); continue; } *best_idx = idx; - break; + return mirror_ds; } - return ds; + return ret; } -static struct nfs4_pnfs_ds * +static struct nfs4_ff_layout_ds * ff_layout_choose_any_ds_for_read(struct pnfs_layout_segment *lseg, u32 start_idx, u32 *best_idx, - u32 offset, u32 *dss_id) + u64 offset, u32 *dss_id) { return ff_layout_choose_ds_for_read(lseg, start_idx, best_idx, offset, dss_id, false); } -static struct nfs4_pnfs_ds * +static struct nfs4_ff_layout_ds * ff_layout_choose_valid_ds_for_read(struct pnfs_layout_segment *lseg, u32 start_idx, u32 *best_idx, - u32 offset, u32 *dss_id) + u64 offset, u32 *dss_id) { return ff_layout_choose_ds_for_read(lseg, start_idx, best_idx, offset, dss_id, true); } -static struct nfs4_pnfs_ds * +static struct nfs4_ff_layout_ds * ff_layout_choose_best_ds_for_read(struct pnfs_layout_segment *lseg, u32 start_idx, u32 *best_idx, - u32 offset, u32 *dss_id) + u64 offset, u32 *dss_id) { - struct nfs4_pnfs_ds *ds; + struct nfs4_ff_layout_ds *mirror_ds; - ds = ff_layout_choose_valid_ds_for_read(lseg, start_idx, best_idx, - offset, dss_id); - if (!IS_ERR(ds)) - return ds; + mirror_ds = ff_layout_choose_valid_ds_for_read(lseg, start_idx, + best_idx, offset, + dss_id); + if (!IS_ERR(mirror_ds)) + return mirror_ds; return ff_layout_choose_any_ds_for_read(lseg, start_idx, best_idx, offset, dss_id); } -static struct nfs4_pnfs_ds * +static struct nfs4_ff_layout_ds * ff_layout_get_ds_for_read(struct nfs_pageio_descriptor *pgio, u32 *best_idx, - u32 offset, + u64 offset, u32 *dss_id) { struct pnfs_layout_segment *lseg = pgio->pg_lseg; - struct nfs4_pnfs_ds *ds; + struct nfs4_ff_layout_ds *mirror_ds; - ds = ff_layout_choose_best_ds_for_read(lseg, pgio->pg_mirror_idx, - best_idx, offset, dss_id); - if (!IS_ERR(ds) || !pgio->pg_mirror_idx) - return ds; + mirror_ds = ff_layout_choose_best_ds_for_read(lseg, + pgio->pg_mirror_idx, + best_idx, offset, dss_id); + if (!IS_ERR(mirror_ds) || !pgio->pg_mirror_idx) + return mirror_ds; return ff_layout_choose_best_ds_for_read(lseg, 0, best_idx, offset, dss_id); } @@ -995,9 +991,8 @@ ff_layout_pg_test(struct nfs_pageio_descriptor *pgio, struct nfs_page *prev, { unsigned int size; u64 p_stripe, r_stripe; - u32 stripe_offset; - u64 segment_offset = pgio->pg_lseg->pls_range.offset; - u32 stripe_unit = FF_LAYOUT_LSEG(pgio->pg_lseg)->stripe_unit; + u64 stripe_offset; + u64 stripe_unit = FF_LAYOUT_LSEG(pgio->pg_lseg)->stripe_unit; /* calls nfs_generic_pg_test */ size = pnfs_generic_pg_test(pgio, prev, req); @@ -1008,23 +1003,23 @@ ff_layout_pg_test(struct nfs_pageio_descriptor *pgio, struct nfs_page *prev, /* see if req and prev are in the same stripe */ if (prev) { - p_stripe = (u64)req_offset(prev) - segment_offset; - r_stripe = (u64)req_offset(req) - segment_offset; - do_div(p_stripe, stripe_unit); - do_div(r_stripe, stripe_unit); + p_stripe = (u64)req_offset(prev); + r_stripe = (u64)req_offset(req); + p_stripe = div64_u64(p_stripe, stripe_unit); + r_stripe = div64_u64(r_stripe, stripe_unit); if (p_stripe != r_stripe) return 0; } /* calculate remaining bytes in the current stripe */ - div_u64_rem((u64)req_offset(req) - segment_offset, + div64_u64_rem((u64)req_offset(req), stripe_unit, &stripe_offset); WARN_ON_ONCE(stripe_offset > stripe_unit); if (stripe_offset >= stripe_unit) return 0; - return min(stripe_unit - (unsigned int)stripe_offset, size); + return min_t(u64, stripe_unit - stripe_offset, size); } static void @@ -1032,8 +1027,7 @@ ff_layout_pg_init_read(struct nfs_pageio_descriptor *pgio, struct nfs_page *req) { struct nfs_pgio_mirror *pgm; - struct nfs4_ff_layout_mirror *mirror; - struct nfs4_pnfs_ds *ds; + struct nfs4_ff_layout_ds *mirror_ds; u32 ds_idx, dss_id; if (NFS_SERVER(pgio->pg_inode)->flags & @@ -1055,9 +1049,9 @@ retry: /* Reset wb_nio, since getting layout segment was successful */ req->wb_nio = 0; - ds = ff_layout_get_ds_for_read(pgio, &ds_idx, - req_offset(req), &dss_id); - if (IS_ERR(ds)) { + mirror_ds = ff_layout_get_ds_for_read(pgio, &ds_idx, + req_offset(req), &dss_id); + if (IS_ERR(mirror_ds)) { if (!ff_layout_no_fallback_to_mds(pgio->pg_lseg)) goto out_mds; pnfs_generic_pg_cleanup(pgio); @@ -1066,9 +1060,9 @@ retry: goto retry; } - mirror = FF_LAYOUT_COMP(pgio->pg_lseg, ds_idx); pgm = &pgio->pg_mirrors[0]; - pgm->pg_bsize = mirror->dss[dss_id].mirror_ds->ds_versions[0].rsize; + pgm->pg_bsize = mirror_ds->ds_versions[0].rsize; + nfs4_ff_layout_put_deviceid(mirror_ds); pgio->pg_mirror_idx = ds_idx; return; @@ -1103,6 +1097,7 @@ ff_layout_pg_init_write(struct nfs_pageio_descriptor *pgio, struct nfs_page *req) { struct nfs4_ff_layout_mirror *mirror; + struct nfs4_ff_layout_ds *mirror_ds; struct nfs_pgio_mirror *pgm; struct nfs4_pnfs_ds *ds; u32 i, dss_id; @@ -1134,9 +1129,12 @@ retry: FF_LAYOUT_LSEG(pgio->pg_lseg)->stripe_unit, mirror->dss_count, req_offset(req)); + mirror_ds = ff_layout_get_mirror_ds(pgio->pg_lseg->pls_layout, + mirror, dss_id); ds = nfs4_ff_layout_prepare_ds(pgio->pg_lseg, mirror, - dss_id, true); + mirror_ds, dss_id, OP_WRITE); if (IS_ERR(ds)) { + nfs4_ff_layout_put_deviceid(mirror_ds); if (!ff_layout_no_fallback_to_mds(pgio->pg_lseg)) goto out_mds; pnfs_generic_pg_cleanup(pgio); @@ -1145,7 +1143,8 @@ retry: goto retry; } pgm = &pgio->pg_mirrors[i]; - pgm->pg_bsize = mirror->dss[dss_id].mirror_ds->ds_versions[0].wsize; + pgm->pg_bsize = mirror_ds->ds_versions[0].wsize; + nfs4_ff_layout_put_deviceid(mirror_ds); } if (NFS_SERVER(pgio->pg_inode)->flags & @@ -1277,14 +1276,16 @@ static void ff_layout_resend_pnfs_read(struct nfs_pgio_header *hdr) u32 idx = hdr->pgio_mirror_idx + 1; u32 new_idx = 0; u32 dss_id = 0; - struct nfs4_pnfs_ds *ds; + struct nfs4_ff_layout_ds *mirror_ds; - ds = ff_layout_choose_any_ds_for_read(hdr->lseg, idx, &new_idx, - hdr->args.offset, &dss_id); - if (IS_ERR(ds)) - pnfs_error_mark_layout_for_return(hdr->inode, hdr->lseg); - else + mirror_ds = ff_layout_choose_any_ds_for_read(hdr->lseg, idx, &new_idx, + hdr->args.offset, &dss_id); + if (IS_ERR(mirror_ds)) { + pnfs_error_mark_layout_for_return(hdr->inode, hdr->lseg, NULL); + } else { + nfs4_ff_layout_put_deviceid(mirror_ds); ff_layout_send_layouterror(hdr->lseg); + } pnfs_read_resend_pnfs(hdr, new_idx); } @@ -1293,7 +1294,7 @@ static void ff_layout_reset_read(struct nfs_pgio_header *hdr) struct rpc_task *task = &hdr->task; pnfs_layoutcommit_inode(hdr->inode, false); - pnfs_error_mark_layout_for_return(hdr->inode, hdr->lseg); + pnfs_error_mark_layout_for_return(hdr->inode, hdr->lseg, NULL); if (!test_and_set_bit(NFS_IOHDR_REDO, &hdr->flags)) { dprintk("%s Reset task %5u for i/o through MDS " @@ -1317,11 +1318,10 @@ static int ff_layout_async_handle_error_v4(struct rpc_task *task, struct nfs4_state *state, struct nfs_client *clp, struct pnfs_layout_segment *lseg, - u32 idx, u32 dss_id) + struct nfs4_deviceid_node *devid) { struct pnfs_layout_hdr *lo = lseg->pls_layout; struct inode *inode = lo->plh_inode; - struct nfs4_deviceid_node *devid = FF_LAYOUT_DEVID_NODE(lseg, idx, dss_id); struct nfs4_slot_table *tbl = nfs4_has_session(clp) ? &clp->cl_session->fc_slot_table : clp->cl_slot_tbl; @@ -1394,8 +1394,9 @@ static int ff_layout_async_handle_error_v4(struct rpc_task *task, case -ENODEV: dprintk("%s DS connection error %d\n", __func__, task->tk_status); - nfs4_delete_deviceid(devid->ld, devid->nfs_client, - &devid->deviceid); + if (devid) + nfs4_delete_deviceid(devid->ld, devid->nfs_client, + &devid->deviceid); rpc_wake_up(&tbl->slot_tbl_waitq); break; default: @@ -1419,9 +1420,8 @@ static int ff_layout_async_handle_error_v3(struct rpc_task *task, u32 op_status, struct nfs_client *clp, struct pnfs_layout_segment *lseg, - u32 idx, u32 dss_id) + struct nfs4_deviceid_node *devid) { - struct nfs4_deviceid_node *devid = FF_LAYOUT_DEVID_NODE(lseg, idx, dss_id); switch (op_status) { case NFS_OK: @@ -1467,8 +1467,9 @@ static int ff_layout_async_handle_error_v3(struct rpc_task *task, default: dprintk("%s DS connection error %d\n", __func__, task->tk_status); - nfs4_delete_deviceid(devid->ld, devid->nfs_client, - &devid->deviceid); + if (devid) + nfs4_delete_deviceid(devid->ld, devid->nfs_client, + &devid->deviceid); } out_reset_to_pnfs: /* FIXME: Need to prevent infinite looping here. */ @@ -1485,12 +1486,13 @@ static int ff_layout_async_handle_error(struct rpc_task *task, struct nfs4_state *state, struct nfs_client *clp, struct pnfs_layout_segment *lseg, - u32 idx, u32 dss_id) + struct nfs4_deviceid_node *devid) { int vers = clp->cl_nfs_mod->rpc_vers->number; if (task->tk_status >= 0) { - ff_layout_mark_ds_reachable(lseg, idx, dss_id); + if (devid) + nfs4_mark_deviceid_available(devid); return 0; } @@ -1501,10 +1503,10 @@ static int ff_layout_async_handle_error(struct rpc_task *task, switch (vers) { case 3: return ff_layout_async_handle_error_v3(task, op_status, clp, - lseg, idx, dss_id); + lseg, devid); case 4: return ff_layout_async_handle_error_v4(task, op_status, state, - clp, lseg, idx, dss_id); + clp, lseg, devid); default: /* should never happen */ WARN_ON_ONCE(1); @@ -1513,6 +1515,7 @@ static int ff_layout_async_handle_error(struct rpc_task *task, } static void ff_layout_io_track_ds_error(struct pnfs_layout_segment *lseg, + struct nfs4_deviceid_node *devid, u32 idx, u32 dss_id, u64 offset, u64 length, u32 *op_status, int opnum, int error) { @@ -1562,8 +1565,8 @@ static void ff_layout_io_track_ds_error(struct pnfs_layout_segment *lseg, mirror = FF_LAYOUT_COMP(lseg, idx); err = ff_layout_track_ds_error(FF_LAYOUT_FROM_HDR(lseg->pls_layout), - mirror, dss_id, offset, length, status, opnum, - nfs_io_gfp_mask()); + mirror, devid, dss_id, offset, length, + status, opnum, nfs_io_gfp_mask()); /* * I/O we cancelled ourselves to return a recalled or revoked layout @@ -1580,7 +1583,8 @@ static void ff_layout_io_track_ds_error(struct pnfs_layout_segment *lseg, case NFS4ERR_PERM: break; case NFS4ERR_NXIO: - ff_layout_mark_ds_unreachable(lseg, idx, dss_id); + if (devid) + nfs4_mark_deviceid_unavailable(devid); /* * Don't return the layout if this is a read and we still * have layouts to try @@ -1590,7 +1594,8 @@ static void ff_layout_io_track_ds_error(struct pnfs_layout_segment *lseg, fallthrough; default: pnfs_error_mark_layout_for_return(lseg->pls_layout->plh_inode, - lseg); + lseg, + &mirror->dss[dss_id].devid); } out: @@ -1609,7 +1614,7 @@ static int ff_layout_read_done_cb(struct rpc_task *task, int err; if (task->tk_status < 0) { - ff_layout_io_track_ds_error(hdr->lseg, + ff_layout_io_track_ds_error(hdr->lseg, hdr->ds_dev, hdr->pgio_mirror_idx, dss_id, hdr->args.offset, hdr->args.count, &hdr->res.op_status, OP_READ, @@ -1620,8 +1625,7 @@ static int ff_layout_read_done_cb(struct rpc_task *task, err = ff_layout_async_handle_error(task, hdr->res.op_status, hdr->args.context->state, hdr->ds_clp, hdr->lseg, - hdr->pgio_mirror_idx, - dss_id); + hdr->ds_dev); trace_nfs4_pnfs_read(hdr, err); clear_bit(NFS_IOHDR_RESEND_PNFS, &hdr->flags); @@ -1814,7 +1818,7 @@ static int ff_layout_write_done_cb(struct rpc_task *task, int err; if (task->tk_status < 0) { - ff_layout_io_track_ds_error(hdr->lseg, + ff_layout_io_track_ds_error(hdr->lseg, hdr->ds_dev, hdr->pgio_mirror_idx, dss_id, hdr->args.offset, hdr->args.count, &hdr->res.op_status, OP_WRITE, @@ -1825,8 +1829,7 @@ static int ff_layout_write_done_cb(struct rpc_task *task, err = ff_layout_async_handle_error(task, hdr->res.op_status, hdr->args.context->state, hdr->ds_clp, hdr->lseg, - hdr->pgio_mirror_idx, - dss_id); + hdr->ds_dev); trace_nfs4_pnfs_write(hdr, err); clear_bit(NFS_IOHDR_RESEND_PNFS, &hdr->flags); @@ -1868,7 +1871,7 @@ static int ff_layout_commit_done_cb(struct rpc_task *task, u32 dss_id = calc_dss_id_from_commit(data->lseg, data->ds_commit_index); if (task->tk_status < 0) { - ff_layout_io_track_ds_error(data->lseg, idx, dss_id, + ff_layout_io_track_ds_error(data->lseg, data->ds_dev, idx, dss_id, data->args.offset, data->args.count, &data->res.op_status, OP_COMMIT, task->tk_status); @@ -1876,8 +1879,8 @@ static int ff_layout_commit_done_cb(struct rpc_task *task, } err = ff_layout_async_handle_error(task, data->res.op_status, - NULL, data->ds_clp, data->lseg, idx, - dss_id); + NULL, data->ds_clp, data->lseg, + data->ds_dev); trace_nfs4_pnfs_commit_ds(data, err); switch (err) { @@ -2167,6 +2170,7 @@ ff_layout_read_pagelist(struct nfs_pgio_header *hdr) struct rpc_clnt *ds_clnt; struct nfsd_file *localio; struct nfs4_ff_layout_mirror *mirror; + struct nfs4_ff_layout_ds *mirror_ds; const struct cred *ds_cred; loff_t offset = hdr->args.offset; u32 idx = hdr->pgio_mirror_idx; @@ -2184,22 +2188,25 @@ ff_layout_read_pagelist(struct nfs_pgio_header *hdr) FF_LAYOUT_LSEG(lseg)->stripe_unit, mirror->dss_count, offset); - ds = nfs4_ff_layout_prepare_ds(lseg, mirror, dss_id, false); + mirror_ds = ff_layout_get_mirror_ds(lseg->pls_layout, mirror, dss_id); + ds = nfs4_ff_layout_prepare_ds(lseg, mirror, mirror_ds, dss_id, + OP_READ); if (IS_ERR(ds)) { ds_fatal_error = nfs_error_is_fatal(PTR_ERR(ds)); goto out_failed; } - ds_clnt = nfs4_ff_find_or_create_ds_client(mirror, ds->ds_clp, - hdr->inode, dss_id); + ds_clnt = nfs4_ff_find_or_create_ds_client(mirror_ds, ds->ds_clp, + hdr->inode); if (IS_ERR(ds_clnt)) goto out_failed; - ds_cred = ff_layout_get_ds_cred(mirror, &lseg->pls_range, hdr->cred, dss_id); + ds_cred = ff_layout_get_ds_cred(mirror, &lseg->pls_range, hdr->cred, + mirror_ds, dss_id); if (!ds_cred) goto out_failed; - vers = nfs4_ff_layout_ds_version(mirror, dss_id); + vers = nfs4_ff_layout_ds_version(mirror_ds); dprintk("%s USE DS: %s cl_count %d vers %d\n", __func__, ds->ds_remotestr, refcount_read(&ds->ds_clp->cl_count), vers); @@ -2211,7 +2218,8 @@ ff_layout_read_pagelist(struct nfs_pgio_header *hdr) if (fh) hdr->args.fh = fh; - nfs4_ff_layout_select_ds_stateid(mirror, dss_id, &hdr->args.stateid); + nfs4_ff_layout_select_ds_stateid(mirror, mirror_ds, dss_id, + &hdr->args.stateid); /* * Note that if we ever decide to split across DSes, @@ -2228,6 +2236,10 @@ ff_layout_read_pagelist(struct nfs_pgio_header *hdr) ff_layout_read_record_layoutstats_start(&hdr->task, hdr); } + /* Transfer the device node reference to the I/O; put on release */ + pnfs_put_ds_dev(hdr->ds_dev); + hdr->ds_dev = &mirror_ds->id_node; + /* Perform an asynchronous read to ds */ nfs_initiate_pgio(ds_clnt, hdr, ds_cred, ds->ds_clp->rpc_ops, vers == 3 ? &ff_layout_read_call_ops_v3 : @@ -2237,6 +2249,7 @@ ff_layout_read_pagelist(struct nfs_pgio_header *hdr) return PNFS_ATTEMPTED; out_failed: + nfs4_ff_layout_put_deviceid(mirror_ds); if (ff_layout_avoid_mds_available_ds(lseg) && !ds_fatal_error) return PNFS_TRY_AGAIN; if (ff_layout_no_fallback_to_mds(lseg)) { @@ -2244,7 +2257,8 @@ out_failed: * FF_FLAGS_NO_IO_THRU_MDS: force fresh LAYOUTGET, * never fall through to MDS I/O. */ - pnfs_error_mark_layout_for_return(hdr->inode, lseg); + pnfs_error_mark_layout_for_return(hdr->inode, lseg, + &mirror->dss[dss_id].devid); return PNFS_TRY_AGAIN; } trace_pnfs_mds_fallback_read_pagelist(hdr->inode, @@ -2262,6 +2276,7 @@ ff_layout_write_pagelist(struct nfs_pgio_header *hdr, int sync) struct rpc_clnt *ds_clnt; struct nfsd_file *localio; struct nfs4_ff_layout_mirror *mirror; + struct nfs4_ff_layout_ds *mirror_ds; const struct cred *ds_cred; loff_t offset = hdr->args.offset; int vers; @@ -2275,22 +2290,25 @@ ff_layout_write_pagelist(struct nfs_pgio_header *hdr, int sync) FF_LAYOUT_LSEG(lseg)->stripe_unit, mirror->dss_count, offset); - ds = nfs4_ff_layout_prepare_ds(lseg, mirror, dss_id, true); + mirror_ds = ff_layout_get_mirror_ds(lseg->pls_layout, mirror, dss_id); + ds = nfs4_ff_layout_prepare_ds(lseg, mirror, mirror_ds, dss_id, + OP_WRITE); if (IS_ERR(ds)) { ds_fatal_error = nfs_error_is_fatal(PTR_ERR(ds)); goto out_failed; } - ds_clnt = nfs4_ff_find_or_create_ds_client(mirror, ds->ds_clp, - hdr->inode, dss_id); + ds_clnt = nfs4_ff_find_or_create_ds_client(mirror_ds, ds->ds_clp, + hdr->inode); if (IS_ERR(ds_clnt)) goto out_failed; - ds_cred = ff_layout_get_ds_cred(mirror, &lseg->pls_range, hdr->cred, dss_id); + ds_cred = ff_layout_get_ds_cred(mirror, &lseg->pls_range, hdr->cred, + mirror_ds, dss_id); if (!ds_cred) goto out_failed; - vers = nfs4_ff_layout_ds_version(mirror, dss_id); + vers = nfs4_ff_layout_ds_version(mirror_ds); dprintk("%s ino %llu sync %d req %zu@%llu DS: %s cl_count %d vers %d\n", __func__, hdr->inode->i_ino, sync, (size_t) hdr->args.count, @@ -2305,7 +2323,8 @@ ff_layout_write_pagelist(struct nfs_pgio_header *hdr, int sync) if (fh) hdr->args.fh = fh; - nfs4_ff_layout_select_ds_stateid(mirror, dss_id, &hdr->args.stateid); + nfs4_ff_layout_select_ds_stateid(mirror, mirror_ds, dss_id, + &hdr->args.stateid); /* * Note that if we ever decide to split across DSes, @@ -2321,6 +2340,10 @@ ff_layout_write_pagelist(struct nfs_pgio_header *hdr, int sync) ff_layout_write_record_layoutstats_start(&hdr->task, hdr); } + /* Transfer the device node reference to the I/O; put on release */ + pnfs_put_ds_dev(hdr->ds_dev); + hdr->ds_dev = &mirror_ds->id_node; + /* Perform an asynchronous write */ nfs_initiate_pgio(ds_clnt, hdr, ds_cred, ds->ds_clp->rpc_ops, vers == 3 ? &ff_layout_write_call_ops_v3 : @@ -2330,6 +2353,7 @@ ff_layout_write_pagelist(struct nfs_pgio_header *hdr, int sync) return PNFS_ATTEMPTED; out_failed: + nfs4_ff_layout_put_deviceid(mirror_ds); if (ff_layout_avoid_mds_available_ds(lseg) && !ds_fatal_error) return PNFS_TRY_AGAIN; if (ff_layout_no_fallback_to_mds(lseg)) { @@ -2337,7 +2361,8 @@ out_failed: * FF_FLAGS_NO_IO_THRU_MDS: force fresh LAYOUTGET, * never fall through to MDS I/O. */ - pnfs_error_mark_layout_for_return(hdr->inode, lseg); + pnfs_error_mark_layout_for_return(hdr->inode, lseg, + &mirror->dss[dss_id].devid); return PNFS_TRY_AGAIN; } trace_pnfs_mds_fallback_write_pagelist(hdr->inode, @@ -2364,6 +2389,7 @@ static int ff_layout_initiate_commit(struct nfs_commit_data *data, int how) struct rpc_clnt *ds_clnt; struct nfsd_file *localio; struct nfs4_ff_layout_mirror *mirror; + struct nfs4_ff_layout_ds *mirror_ds = NULL; const struct cred *ds_cred; u32 idx, dss_id; int vers, ret; @@ -2376,20 +2402,23 @@ static int ff_layout_initiate_commit(struct nfs_commit_data *data, int how) idx = calc_mirror_idx_from_commit(lseg, data->ds_commit_index); mirror = FF_LAYOUT_COMP(lseg, idx); dss_id = calc_dss_id_from_commit(lseg, data->ds_commit_index); - ds = nfs4_ff_layout_prepare_ds(lseg, mirror, dss_id, true); + mirror_ds = ff_layout_get_mirror_ds(lseg->pls_layout, mirror, dss_id); + ds = nfs4_ff_layout_prepare_ds(lseg, mirror, mirror_ds, dss_id, + OP_COMMIT); if (IS_ERR(ds)) goto out_err; - ds_clnt = nfs4_ff_find_or_create_ds_client(mirror, ds->ds_clp, - data->inode, dss_id); + ds_clnt = nfs4_ff_find_or_create_ds_client(mirror_ds, ds->ds_clp, + data->inode); if (IS_ERR(ds_clnt)) goto out_err; - ds_cred = ff_layout_get_ds_cred(mirror, &lseg->pls_range, data->cred, dss_id); + ds_cred = ff_layout_get_ds_cred(mirror, &lseg->pls_range, data->cred, + mirror_ds, dss_id); if (!ds_cred) goto out_err; - vers = nfs4_ff_layout_ds_version(mirror, dss_id); + vers = nfs4_ff_layout_ds_version(mirror_ds); dprintk("%s ino %llu, how %d cl_count %d vers %d\n", __func__, data->inode->i_ino, how, refcount_read(&ds->ds_clp->cl_count), @@ -2410,6 +2439,10 @@ static int ff_layout_initiate_commit(struct nfs_commit_data *data, int how) ff_layout_commit_record_layoutstats_start(&data->task, data); } + /* Transfer the device node reference to the commit; put on release */ + pnfs_put_ds_dev(data->ds_dev); + data->ds_dev = &mirror_ds->id_node; + ret = nfs_initiate_commit(ds_clnt, data, ds->ds_clp->rpc_ops, vers == 3 ? &ff_layout_commit_call_ops_v3 : &ff_layout_commit_call_ops_v4, @@ -2417,6 +2450,7 @@ static int ff_layout_initiate_commit(struct nfs_commit_data *data, int how) put_cred(ds_cred); return ret; out_err: + nfs4_ff_layout_put_deviceid(mirror_ds); pnfs_generic_prepare_to_resend_writes(data); pnfs_generic_commit_release(data); return -EAGAIN; @@ -2459,7 +2493,8 @@ static bool ff_layout_match_io(const struct rpc_task *task, const void *data) return false; } -static void ff_layout_cancel_io(struct pnfs_layout_segment *lseg) +static void ff_layout_cancel_io(struct pnfs_layout_segment *lseg, + const struct nfs4_deviceid *devid) { struct nfs4_ff_layout_segment *flseg = FF_LAYOUT_LSEG(lseg); struct nfs4_ff_layout_mirror *mirror; @@ -2472,22 +2507,94 @@ static void ff_layout_cancel_io(struct pnfs_layout_segment *lseg) for (idx = 0; idx < flseg->mirror_array_cnt; idx++) { mirror = flseg->mirror_array[idx]; for (dss_id = 0; dss_id < mirror->dss_count; dss_id++) { - mirror_ds = mirror->dss[dss_id].mirror_ds; - if (IS_ERR_OR_NULL(mirror_ds)) + if (devid && memcmp(&mirror->dss[dss_id].devid, devid, + sizeof(*devid)) != 0) continue; - ds = mirror->dss[dss_id].mirror_ds->ds; - if (!ds) + rcu_read_lock(); + mirror_ds = rcu_dereference(mirror->dss[dss_id].mirror_ds); + if (IS_ERR_OR_NULL(mirror_ds) || + !atomic_inc_not_zero(&mirror_ds->id_node.ref)) { + rcu_read_unlock(); continue; + } + rcu_read_unlock(); + ds = mirror_ds->ds; + if (!ds) + goto next; ds_clp = ds->ds_clp; if (!ds_clp) - continue; + goto next; clnt = ds_clp->cl_rpcclient; if (!clnt) - continue; + goto next; if (!rpc_cancel_tasks(clnt, -ECANCELED, ff_layout_match_io, lseg)) - continue; + goto next; rpc_clnt_disconnect(clnt); +next: + nfs4_ff_layout_put_deviceid(mirror_ds); + } + } +} + +/* Called under @lo's inode i_lock. */ +static bool ff_layout_references_deviceid(struct pnfs_layout_hdr *lo, + const struct nfs4_deviceid *id) +{ + struct nfs4_flexfile_layout *flo = FF_LAYOUT_FROM_HDR(lo); + struct nfs4_ff_layout_mirror *mirror; + u32 dss_id; + + list_for_each_entry(mirror, &flo->mirrors, mirrors) + for (dss_id = 0; dss_id < mirror->dss_count; dss_id++) + if (memcmp(&mirror->dss[dss_id].devid, id, + sizeof(*id)) == 0) + return true; + return false; +} + +/* + * Un-pin every stripe node resolved from @id: in-flight I/O drains on the + * old node through its own reference, the next I/O re-resolves. + */ +static void ff_layout_reresolve_deviceid(struct pnfs_layout_hdr *lo, + const struct nfs4_deviceid *id, + bool immediate, + struct list_head *head) +{ + struct nfs4_flexfile_layout *flo = FF_LAYOUT_FROM_HDR(lo); + struct nfs4_ff_layout_mirror *mirror; + struct nfs4_ff_layout_ds *old; + struct nfs4_deviceid_put *put; + u32 dss_id; + + list_for_each_entry(mirror, &flo->mirrors, mirrors) { + for (dss_id = 0; dss_id < mirror->dss_count; dss_id++) { + if (memcmp(&mirror->dss[dss_id].devid, id, + sizeof(*id)) != 0) + continue; + /* Allocate before un-pinning: on failure the reference + * stays put rather than being dropped here, where the + * final put may not sleep. + */ + put = kzalloc_obj(*put, GFP_ATOMIC); + if (!put) + continue; + old = unrcu_pointer( + xchg(&mirror->dss[dss_id].mirror_ds, NULL)); + if (IS_ERR_OR_NULL(old)) { + kfree(put); + continue; + } + /* A node still hashed was fetched after the unhash + * and carries the new mapping; mark only the + * superseded ones. + */ + if (immediate && + hlist_unhashed_lockless(&old->id_node.node)) + nfs4_mark_deviceid_unavailable(&old->id_node); + put->dev = &old->id_node; + list_add(&put->node, head); } } } @@ -2700,7 +2807,7 @@ ff_layout_prepare_layoutreturn(struct nfs4_layoutreturn_args *args) spin_lock(&args->inode->i_lock); ff_args->num_dev = ff_layout_mirror_prepare_stats( - &ff_layout->generic_hdr, &ff_args->devinfo[0], + &ff_layout->generic_hdr, &ff_args->devinfo[0], &ff_args->priv[0], ARRAY_SIZE(ff_args->devinfo), NFS4_FF_OP_LAYOUTRETURN); spin_unlock(&args->inode->i_lock); @@ -2875,10 +2982,11 @@ ff_layout_encode_io_latency(struct xdr_stream *xdr, static void ff_layout_encode_ff_layoutupdate(struct xdr_stream *xdr, const struct nfs42_layoutstat_devinfo *devinfo, - struct nfs4_ff_layout_ds_stripe *dss_info) + struct nfs4_ff_layoutstat_priv *priv) { + struct nfs4_ff_layout_ds_stripe *dss_info = priv->dss_info; struct nfs4_pnfs_ds_addr *da; - struct nfs4_pnfs_ds *ds = dss_info->mirror_ds->ds; + struct nfs4_pnfs_ds *ds = priv->mirror_ds->ds; struct nfs_fh *fh = &dss_info->fh_versions[0]; __be32 *p; @@ -2925,10 +3033,10 @@ ff_layout_encode_layoutstats(struct xdr_stream *xdr, const void *args, static void ff_layout_free_layoutstats(struct nfs4_xdr_opaque_data *opaque) { - struct nfs4_ff_layout_ds_stripe *dss_info = opaque->data; - struct nfs4_ff_layout_mirror *mirror = dss_info->mirror; + struct nfs4_ff_layoutstat_priv *priv = opaque->data; - ff_layout_put_mirror(mirror); + nfs4_ff_layout_put_deviceid(priv->mirror_ds); + ff_layout_put_mirror(priv->dss_info->mirror); } static const struct nfs4_xdr_opaque_ops layoutstat_ops = { @@ -2939,20 +3047,23 @@ static const struct nfs4_xdr_opaque_ops layoutstat_ops = { static int ff_layout_mirror_prepare_stats(struct pnfs_layout_hdr *lo, struct nfs42_layoutstat_devinfo *devinfo, + struct nfs4_ff_layoutstat_priv *priv, int dev_limit, enum nfs4_ff_op_type type) { struct nfs4_flexfile_layout *ff_layout = FF_LAYOUT_FROM_HDR(lo); struct nfs4_ff_layout_mirror *mirror; struct nfs4_ff_layout_ds_stripe *dss_info; - struct nfs4_deviceid_node *dev; + struct nfs4_ff_layout_ds *mirror_ds; int i = 0, dss_id; + rcu_read_lock(); list_for_each_entry(mirror, &ff_layout->mirrors, mirrors) { for (dss_id = 0; dss_id < mirror->dss_count; ++dss_id) { dss_info = &mirror->dss[dss_id]; if (i >= dev_limit) break; - if (IS_ERR_OR_NULL(dss_info->mirror_ds)) + mirror_ds = rcu_dereference(dss_info->mirror_ds); + if (IS_ERR_OR_NULL(mirror_ds)) continue; if (!test_and_clear_bit(NFS4_FF_MIRROR_STAT_AVAIL, &mirror->flags) && @@ -2961,9 +3072,12 @@ ff_layout_mirror_prepare_stats(struct pnfs_layout_hdr *lo, /* mirror refcount put in cleanup_layoutstats */ if (!refcount_inc_not_zero(&mirror->ref)) continue; - dev = &dss_info->mirror_ds->id_node; + /* The pin holds a reference; it is exchanged out only + * under i_lock. Put in ff_layout_free_layoutstats(). + */ + atomic_inc(&mirror_ds->id_node.ref); memcpy(&devinfo->dev_id, - &dev->deviceid, + &mirror_ds->id_node.deviceid, NFS4_DEVICEID4_SIZE); devinfo->offset = 0; devinfo->length = NFS4_MAX_UINT64; @@ -2979,12 +3093,16 @@ ff_layout_mirror_prepare_stats(struct pnfs_layout_hdr *lo, spin_unlock(&mirror->lock); devinfo->layout_type = LAYOUT_FLEX_FILES; devinfo->ld_private.ops = &layoutstat_ops; - devinfo->ld_private.data = &mirror->dss[dss_id]; + priv->dss_info = dss_info; + priv->mirror_ds = mirror_ds; + devinfo->ld_private.data = priv; devinfo++; + priv++; i++; } } + rcu_read_unlock(); return i; } @@ -2992,21 +3110,28 @@ static int ff_layout_prepare_layoutstats(struct nfs42_layoutstat_args *args) { struct pnfs_layout_hdr *lo; struct nfs4_flexfile_layout *ff_layout; + struct nfs4_ff_layoutstat_priv *priv; const int dev_count = PNFS_LAYOUTSTATS_MAXDEV; - /* For now, send at most PNFS_LAYOUTSTATS_MAXDEV statistics */ - args->devinfo = kmalloc_objs(*args->devinfo, dev_count, - nfs_io_gfp_mask()); + /* + * For now, send at most PNFS_LAYOUTSTATS_MAXDEV statistics. + * The per-devinfo private entries are co-allocated after the + * devinfo array and freed along with it. + */ + args->devinfo = kmalloc(dev_count * (sizeof(*args->devinfo) + + sizeof(*priv)), + nfs_io_gfp_mask()); if (!args->devinfo) return -ENOMEM; + priv = (struct nfs4_ff_layoutstat_priv *)&args->devinfo[dev_count]; spin_lock(&args->inode->i_lock); lo = NFS_I(args->inode)->layout; if (lo && pnfs_layout_is_valid(lo)) { ff_layout = FF_LAYOUT_FROM_HDR(lo); args->num_dev = ff_layout_mirror_prepare_stats( - &ff_layout->generic_hdr, &args->devinfo[0], dev_count, - NFS4_FF_OP_LAYOUTSTATS); + &ff_layout->generic_hdr, &args->devinfo[0], priv, + dev_count, NFS4_FF_OP_LAYOUTSTATS); } else args->num_dev = 0; spin_unlock(&args->inode->i_lock); @@ -3055,6 +3180,8 @@ static struct pnfs_layoutdriver_type flexfilelayout_type = { .pg_write_ops = &ff_layout_pg_write_ops, .get_ds_info = ff_layout_get_ds_info, .free_deviceid_node = ff_layout_free_deviceid_node, + .reresolve_deviceid = ff_layout_reresolve_deviceid, + .layout_references_deviceid = ff_layout_references_deviceid, .read_pagelist = ff_layout_read_pagelist, .write_pagelist = ff_layout_write_pagelist, .alloc_deviceid_node = ff_layout_alloc_deviceid_node, diff --git a/fs/nfs/flexfilelayout/flexfilelayout.h b/fs/nfs/flexfilelayout/flexfilelayout.h index a5bd00f69e82..09c3cd6964fd 100644 --- a/fs/nfs/flexfilelayout/flexfilelayout.h +++ b/fs/nfs/flexfilelayout/flexfilelayout.h @@ -79,7 +79,7 @@ struct nfs4_ff_layout_ds_stripe { struct nfs4_ff_layout_mirror *mirror; struct nfs4_deviceid devid; u32 efficiency; - struct nfs4_ff_layout_ds *mirror_ds; + struct nfs4_ff_layout_ds __rcu *mirror_ds; u32 fh_versions_cnt; struct nfs_fh *fh_versions; nfs4_stateid stateid; @@ -124,9 +124,20 @@ struct nfs4_flexfile_layout { unsigned long flags; }; +/* + * Per-devinfo private data for a layoutstats/layoutreturn encode: the + * stripe the stats describe plus a reference on its device node so the + * node (and its DS addresses) stay valid until the XDR encode runs. + */ +struct nfs4_ff_layoutstat_priv { + struct nfs4_ff_layout_ds_stripe *dss_info; + struct nfs4_ff_layout_ds *mirror_ds; +}; + struct nfs4_flexfile_layoutreturn_args { struct list_head errors; struct nfs42_layoutstat_devinfo devinfo[FF_LAYOUTSTATS_MAXDEV]; + struct nfs4_ff_layoutstat_priv priv[FF_LAYOUTSTATS_MAXDEV]; unsigned int num_errors; unsigned int num_dev; struct page *pages[1]; @@ -162,20 +173,6 @@ FF_LAYOUT_COMP(struct pnfs_layout_segment *lseg, u32 idx) return NULL; } -static inline struct nfs4_deviceid_node * -FF_LAYOUT_DEVID_NODE(struct pnfs_layout_segment *lseg, u32 idx, u32 dss_id) -{ - struct nfs4_ff_layout_mirror *mirror = FF_LAYOUT_COMP(lseg, idx); - - if (mirror != NULL) { - struct nfs4_ff_layout_ds *mirror_ds = mirror->dss[dss_id].mirror_ds; - - if (!IS_ERR_OR_NULL(mirror_ds)) - return &mirror_ds->id_node; - } - return NULL; -} - static inline u32 FF_LAYOUT_MIRROR_COUNT(struct pnfs_layout_segment *lseg) { @@ -207,9 +204,9 @@ ff_layout_no_read_on_rw(struct pnfs_layout_segment *lseg) } static inline int -nfs4_ff_layout_ds_version(const struct nfs4_ff_layout_mirror *mirror, u32 dss_id) +nfs4_ff_layout_ds_version(const struct nfs4_ff_layout_ds *mirror_ds) { - return mirror->dss[dss_id].mirror_ds->ds_versions[0].version; + return mirror_ds->ds_versions[0].version; } static inline u32 @@ -220,7 +217,7 @@ nfs4_ff_layout_calc_dss_id(const u64 stripe_unit, const u32 dss_count, const lof if (dss_count == 1 || stripe_unit == 0) return 0; - do_div(tmp, stripe_unit); + tmp = div64_u64(tmp, stripe_unit); return do_div(tmp, dss_count); } @@ -232,6 +229,7 @@ void nfs4_ff_layout_put_deviceid(struct nfs4_ff_layout_ds *mirror_ds); void nfs4_ff_layout_free_deviceid(struct nfs4_ff_layout_ds *mirror_ds); int ff_layout_track_ds_error(struct nfs4_flexfile_layout *flo, struct nfs4_ff_layout_mirror *mirror, + const struct nfs4_deviceid_node *devid, u32 dss_id, u64 offset, u64 length, int status, enum nfs_opnum4 opnum, gfp_t gfp_flags); void ff_layout_send_layouterror(struct pnfs_layout_segment *lseg); @@ -245,23 +243,29 @@ struct nfs_fh * nfs4_ff_layout_select_ds_fh(struct nfs4_ff_layout_mirror *mirror, u32 dss_id); void nfs4_ff_layout_select_ds_stateid(const struct nfs4_ff_layout_mirror *mirror, + const struct nfs4_ff_layout_ds *mirror_ds, u32 dss_id, nfs4_stateid *stateid); +struct nfs4_ff_layout_ds * +ff_layout_get_mirror_ds(struct pnfs_layout_hdr *lo, + struct nfs4_ff_layout_mirror *mirror, + u32 dss_id); struct nfs4_pnfs_ds * nfs4_ff_layout_prepare_ds(struct pnfs_layout_segment *lseg, struct nfs4_ff_layout_mirror *mirror, + struct nfs4_ff_layout_ds *mirror_ds, u32 dss_id, - bool fail_return); + enum nfs_opnum4 opnum); struct rpc_clnt * -nfs4_ff_find_or_create_ds_client(struct nfs4_ff_layout_mirror *mirror, +nfs4_ff_find_or_create_ds_client(const struct nfs4_ff_layout_ds *mirror_ds, struct nfs_client *ds_clp, - struct inode *inode, - u32 dss_id); + struct inode *inode); const struct cred *ff_layout_get_ds_cred(struct nfs4_ff_layout_mirror *mirror, const struct pnfs_layout_range *range, const struct cred *mdscred, + const struct nfs4_ff_layout_ds *mirror_ds, u32 dss_id); bool ff_layout_avoid_mds_available_ds(struct pnfs_layout_segment *lseg); bool ff_layout_avoid_read_on_rw(struct pnfs_layout_segment *lseg); diff --git a/fs/nfs/flexfilelayout/flexfilelayoutdev.c b/fs/nfs/flexfilelayout/flexfilelayoutdev.c index 5c0216bd5fce..6165c41fcf62 100644 --- a/fs/nfs/flexfilelayout/flexfilelayoutdev.c +++ b/fs/nfs/flexfilelayout/flexfilelayoutdev.c @@ -20,6 +20,7 @@ static unsigned int dataserver_timeo = NFS_DEF_TCP_TIMEO; static unsigned int dataserver_retrans; +static unsigned int dataserver_nconnect; static bool ff_layout_has_available_ds(struct pnfs_layout_segment *lseg); @@ -159,26 +160,13 @@ nfs4_ff_alloc_deviceid_node(struct nfs_server *server, struct pnfs_device *pdev, goto out_err_drain_dsaddrs; /* If DS was already in cache, free ds addrs */ - while (!list_empty(&dsaddrs)) { - da = list_first_entry(&dsaddrs, - struct nfs4_pnfs_ds_addr, - da_node); - list_del_init(&da->da_node); - kfree(da->da_remotestr); - kfree(da); - } + nfs4_pnfs_ds_addr_list_free(&dsaddrs); folio_put(scratch); return new_ds; out_err_drain_dsaddrs: - while (!list_empty(&dsaddrs)) { - da = list_first_entry(&dsaddrs, struct nfs4_pnfs_ds_addr, - da_node); - list_del_init(&da->da_node); - kfree(da->da_remotestr); - kfree(da); - } + nfs4_pnfs_ds_addr_list_free(&dsaddrs); kfree(ds_versions); out_scratch: @@ -256,6 +244,7 @@ ff_layout_add_ds_error_locked(struct nfs4_flexfile_layout *flo, int ff_layout_track_ds_error(struct nfs4_flexfile_layout *flo, struct nfs4_ff_layout_mirror *mirror, + const struct nfs4_deviceid_node *devid, u32 dss_id, u64 offset, u64 length, int status, enum nfs_opnum4 opnum, gfp_t gfp_flags) { @@ -264,7 +253,7 @@ int ff_layout_track_ds_error(struct nfs4_flexfile_layout *flo, if (status == 0) return 0; - if (IS_ERR_OR_NULL(mirror->dss[dss_id].mirror_ds)) + if (devid == NULL) return -EINVAL; dserr = kmalloc_obj(*dserr, gfp_flags); @@ -277,8 +266,7 @@ int ff_layout_track_ds_error(struct nfs4_flexfile_layout *flo, dserr->status = status; dserr->opnum = opnum; nfs4_stateid_copy(&dserr->stateid, &mirror->dss[dss_id].stateid); - memcpy(&dserr->deviceid, &mirror->dss[dss_id].mirror_ds->id_node.deviceid, - NFS4_DEVICEID4_SIZE); + memcpy(&dserr->deviceid, &devid->deviceid, NFS4_DEVICEID4_SIZE); spin_lock(&flo->generic_hdr.plh_inode->i_lock); ff_layout_add_ds_error_locked(flo, dserr); @@ -317,50 +305,81 @@ nfs4_ff_layout_select_ds_fh(struct nfs4_ff_layout_mirror *mirror, u32 dss_id) void nfs4_ff_layout_select_ds_stateid(const struct nfs4_ff_layout_mirror *mirror, + const struct nfs4_ff_layout_ds *mirror_ds, u32 dss_id, nfs4_stateid *stateid) { - if (nfs4_ff_layout_ds_version(mirror, dss_id) == 4) + if (nfs4_ff_layout_ds_version(mirror_ds) == 4) nfs4_stateid_copy(stateid, &mirror->dss[dss_id].stateid); } -static bool -ff_layout_init_mirror_ds(struct pnfs_layout_hdr *lo, - struct nfs4_ff_layout_mirror *mirror, - u32 dss_id) +/* + * Resolve the stripe's deviceid on first use and pin the node on the + * mirror. Returns a node the caller must put, or an ERR_PTR. + */ +struct nfs4_ff_layout_ds * +ff_layout_get_mirror_ds(struct pnfs_layout_hdr *lo, + struct nfs4_ff_layout_mirror *mirror, + u32 dss_id) { + struct nfs4_ff_layout_ds *mirror_ds, *old; + struct nfs4_deviceid_node *node; + if (mirror == NULL) - goto outerr; - if (mirror->dss[dss_id].mirror_ds == NULL) { - struct nfs4_deviceid_node *node; - struct nfs4_ff_layout_ds *mirror_ds = ERR_PTR(-ENODEV); - - node = nfs4_find_get_deviceid(NFS_SERVER(lo->plh_inode), - &mirror->dss[dss_id].devid, lo->plh_lc_cred, - GFP_KERNEL); - if (node) - mirror_ds = FF_LAYOUT_MIRROR_DS(node); - - /* check for race with another call to this function */ - if (cmpxchg(&mirror->dss[dss_id].mirror_ds, NULL, mirror_ds) && - mirror_ds != ERR_PTR(-ENODEV)) - nfs4_put_deviceid_node(node); + return ERR_PTR(-ENODEV); + +retry: + rcu_read_lock(); + mirror_ds = rcu_dereference(mirror->dss[dss_id].mirror_ds); + if (mirror_ds && !IS_ERR(mirror_ds) && + atomic_inc_not_zero(&mirror_ds->id_node.ref)) { + rcu_read_unlock(); + return mirror_ds; + } + rcu_read_unlock(); + if (IS_ERR(mirror_ds)) + return mirror_ds; + if (mirror_ds != NULL) + /* raced with a reset; the field is being re-pointed */ + goto retry; + + node = nfs4_find_get_deviceid(NFS_SERVER(lo->plh_inode), + &mirror->dss[dss_id].devid, lo->plh_lc_cred, + GFP_KERNEL); + if (node) { + mirror_ds = FF_LAYOUT_MIRROR_DS(node); + /* + * Take the caller's reference before the pointer becomes + * visible below, so a concurrent reset of the installed + * pointer cannot drop the last reference under us. + */ + atomic_inc(&node->ref); + } else { + mirror_ds = ERR_PTR(-ENODEV); } - if (IS_ERR(mirror->dss[dss_id].mirror_ds)) - goto outerr; + /* check for race with another call to this function */ + old = unrcu_pointer(cmpxchg(&mirror->dss[dss_id].mirror_ds, + NULL, RCU_INITIALIZER(mirror_ds))); + if (old == NULL) + return mirror_ds; - return true; -outerr: - return false; + /* lost the race; use the winner's node instead */ + if (node) { + nfs4_put_deviceid_node(node); + nfs4_put_deviceid_node(node); + } + goto retry; } /** * nfs4_ff_layout_prepare_ds - prepare a DS connection for an RPC call * @lseg: the layout segment we're operating on * @mirror: layout mirror describing the DS to use + * @mirror_ds: referenced device node for the stripe, from + * ff_layout_get_mirror_ds() (may be an ERR_PTR) * @dss_id: DS stripe id to select stripe to use - * @fail_return: return layout on connect failure? + * @opnum: operation this connection is being prepared for * * Try to prepare a DS connection to accept an RPC call. This involves * selecting a mirror to use and connecting the client to it if it's not @@ -368,16 +387,19 @@ outerr: * * Since we only need a single functioning mirror to satisfy a read, we don't * want to return the layout if there is one. For writes though, any down - * mirror should result in a LAYOUTRETURN. @fail_return is how we distinguish - * between the two cases. + * mirror should result in a LAYOUTRETURN. @opnum is how we distinguish + * between the two cases. On failure, @opnum is also reported in the tracked + * device error so that the server can tell which class of I/O the client + * was unable to send to the mirror. * * Returns a pointer to a connected DS object on success or NULL on failure. */ struct nfs4_pnfs_ds * nfs4_ff_layout_prepare_ds(struct pnfs_layout_segment *lseg, struct nfs4_ff_layout_mirror *mirror, + struct nfs4_ff_layout_ds *mirror_ds, u32 dss_id, - bool fail_return) + enum nfs_opnum4 opnum) { struct nfs4_pnfs_ds *ds; struct inode *ino = lseg->pls_layout->plh_inode; @@ -385,10 +407,10 @@ nfs4_ff_layout_prepare_ds(struct pnfs_layout_segment *lseg, unsigned int max_payload; int status = -EAGAIN; - if (!ff_layout_init_mirror_ds(lseg->pls_layout, mirror, dss_id)) + if (IS_ERR_OR_NULL(mirror_ds)) goto noconnect; - ds = mirror->dss[dss_id].mirror_ds->ds; + ds = mirror_ds->ds; if (READ_ONCE(ds->ds_clp)) goto out; /* matching smp_wmb() in _nfs4_pnfs_v3/4_ds_connect */ @@ -397,11 +419,12 @@ nfs4_ff_layout_prepare_ds(struct pnfs_layout_segment *lseg, /* FIXME: For now we assume the server sent only one version of NFS * to use for the DS. */ - status = nfs4_pnfs_ds_connect(s, ds, &mirror->dss[dss_id].mirror_ds->id_node, + status = nfs4_pnfs_ds_connect(s, ds, &mirror_ds->id_node, dataserver_timeo, dataserver_retrans, - mirror->dss[dss_id].mirror_ds->ds_versions[0].version, - mirror->dss[dss_id].mirror_ds->ds_versions[0].minor_version, - mirror->dss[dss_id].mirror_ds->ds_versions[0].tightly_coupled); + dataserver_nconnect, + mirror_ds->ds_versions[0].version, + mirror_ds->ds_versions[0].minor_version, + mirror_ds->ds_versions[0].tightly_coupled); /* connect success, check rsize/wsize limit */ if (!status) { @@ -414,20 +437,24 @@ nfs4_ff_layout_prepare_ds(struct pnfs_layout_segment *lseg, max_payload = nfs_block_size(rpc_max_payload(ds->ds_clp->cl_rpcclient), NULL); - if (mirror->dss[dss_id].mirror_ds->ds_versions[0].rsize > max_payload) - mirror->dss[dss_id].mirror_ds->ds_versions[0].rsize = max_payload; - if (mirror->dss[dss_id].mirror_ds->ds_versions[0].wsize > max_payload) - mirror->dss[dss_id].mirror_ds->ds_versions[0].wsize = max_payload; + if (mirror_ds->ds_versions[0].rsize > max_payload) + mirror_ds->ds_versions[0].rsize = max_payload; + if (mirror_ds->ds_versions[0].wsize > max_payload) + mirror_ds->ds_versions[0].wsize = max_payload; goto out; } noconnect: ff_layout_track_ds_error(FF_LAYOUT_FROM_HDR(lseg->pls_layout), - mirror, dss_id, lseg->pls_range.offset, + mirror, + IS_ERR_OR_NULL(mirror_ds) ? + NULL : &mirror_ds->id_node, + dss_id, lseg->pls_range.offset, lseg->pls_range.length, NFS4ERR_NXIO, - OP_ILLEGAL, GFP_NOIO); + opnum, GFP_NOIO); ff_layout_send_layouterror(lseg); - if (fail_return || !ff_layout_has_available_ds(lseg)) - pnfs_error_mark_layout_for_return(ino, lseg); + if (opnum != OP_READ || !ff_layout_has_available_ds(lseg)) + pnfs_error_mark_layout_for_return(ino, lseg, + &mirror->dss[dss_id].devid); ds = ERR_PTR(status); out: return ds; @@ -437,11 +464,12 @@ const struct cred * ff_layout_get_ds_cred(struct nfs4_ff_layout_mirror *mirror, const struct pnfs_layout_range *range, const struct cred *mdscred, + const struct nfs4_ff_layout_ds *mirror_ds, u32 dss_id) { const struct cred *cred; - if (mirror && !mirror->dss[dss_id].mirror_ds->ds_versions[0].tightly_coupled) { + if (mirror && !mirror_ds->ds_versions[0].tightly_coupled) { cred = ff_layout_get_mirror_cred(mirror, range->iomode, dss_id); if (!cred) cred = get_cred(mdscred); @@ -453,20 +481,18 @@ ff_layout_get_ds_cred(struct nfs4_ff_layout_mirror *mirror, /** * nfs4_ff_find_or_create_ds_client - Find or create a DS rpc client - * @mirror: pointer to the mirror + * @mirror_ds: device node for the stripe * @ds_clp: nfs_client for the DS * @inode: pointer to inode - * @dss_id: DS stripe id * * Find or create a DS rpc client with th MDS server rpc client auth flavor * in the nfs_client cl_ds_clients list. */ struct rpc_clnt * -nfs4_ff_find_or_create_ds_client(struct nfs4_ff_layout_mirror *mirror, - struct nfs_client *ds_clp, struct inode *inode, - u32 dss_id) +nfs4_ff_find_or_create_ds_client(const struct nfs4_ff_layout_ds *mirror_ds, + struct nfs_client *ds_clp, struct inode *inode) { - switch (mirror->dss[dss_id].mirror_ds->ds_versions[0].version) { + switch (mirror_ds->ds_versions[0].version) { case 3: /* For NFSv3 DS, flavor is set when creating DS connections */ return ds_clp->cl_rpcclient; @@ -571,49 +597,60 @@ unsigned int ff_layout_fetch_ds_ioerr(struct pnfs_layout_hdr *lo, static bool ff_read_layout_has_available_ds(struct pnfs_layout_segment *lseg) { struct nfs4_ff_layout_mirror *mirror; - struct nfs4_deviceid_node *devid; + struct nfs4_ff_layout_ds *mirror_ds; + bool ret = false; u32 idx, dss_id; + rcu_read_lock(); for (idx = 0; idx < FF_LAYOUT_MIRROR_COUNT(lseg); idx++) { mirror = FF_LAYOUT_COMP(lseg, idx); if (!mirror) continue; for (dss_id = 0; dss_id < mirror->dss_count; dss_id++) { - if (!mirror->dss[dss_id].mirror_ds) - return true; - if (IS_ERR(mirror->dss[dss_id].mirror_ds)) + mirror_ds = rcu_dereference(mirror->dss[dss_id].mirror_ds); + if (!mirror_ds) { + ret = true; + goto out; + } + if (IS_ERR(mirror_ds)) continue; - devid = &mirror->dss[dss_id].mirror_ds->id_node; - if (!nfs4_test_deviceid_unavailable(devid)) - return true; + if (!nfs4_test_deviceid_unavailable(&mirror_ds->id_node)) { + ret = true; + goto out; + } } } - - return false; +out: + rcu_read_unlock(); + return ret; } static bool ff_rw_layout_has_available_ds(struct pnfs_layout_segment *lseg) { struct nfs4_ff_layout_mirror *mirror; - struct nfs4_deviceid_node *devid; + struct nfs4_ff_layout_ds *mirror_ds; + bool ret = false; u32 idx, dss_id; + rcu_read_lock(); for (idx = 0; idx < FF_LAYOUT_MIRROR_COUNT(lseg); idx++) { mirror = FF_LAYOUT_COMP(lseg, idx); if (!mirror) - return false; + goto out; for (dss_id = 0; dss_id < mirror->dss_count; dss_id++) { - if (IS_ERR(mirror->dss[dss_id].mirror_ds)) - return false; - if (!mirror->dss[dss_id].mirror_ds) + mirror_ds = rcu_dereference(mirror->dss[dss_id].mirror_ds); + if (IS_ERR(mirror_ds)) + goto out; + if (!mirror_ds) continue; - devid = &mirror->dss[dss_id].mirror_ds->id_node; - if (nfs4_test_deviceid_unavailable(devid)) - return false; + if (nfs4_test_deviceid_unavailable(&mirror_ds->id_node)) + goto out; } } - - return FF_LAYOUT_MIRROR_COUNT(lseg) != 0; + ret = FF_LAYOUT_MIRROR_COUNT(lseg) != 0; +out: + rcu_read_unlock(); + return ret; } static bool ff_layout_has_available_ds(struct pnfs_layout_segment *lseg) @@ -644,3 +681,8 @@ module_param(dataserver_timeo, uint, 0644); MODULE_PARM_DESC(dataserver_timeo, "The time (in tenths of a second) the " "NFSv4.1 client waits for a response from a " " data server before it retries an NFS request."); +module_param(dataserver_nconnect, uint, 0644); +MODULE_PARM_DESC(dataserver_nconnect, "The maximum number of connections " + "the NFSv4.1 client opens to each data server, " + "capping the value inherited from the MDS nconnect " + "mount option. 0 (default) applies no cap."); diff --git a/fs/nfs/inode.c b/fs/nfs/inode.c index 3022454f7698..3c9b2ec4e244 100644 --- a/fs/nfs/inode.c +++ b/fs/nfs/inode.c @@ -690,7 +690,7 @@ EXPORT_SYMBOL_GPL(nfs_update_delegated_mtime); #define NFS_VALID_ATTRS (ATTR_MODE|ATTR_UID|ATTR_GID|ATTR_SIZE|ATTR_ATIME|ATTR_ATIME_SET|ATTR_MTIME|ATTR_MTIME_SET|ATTR_FILE|ATTR_OPEN) int -nfs_setattr(struct mnt_idmap *idmap, struct dentry *dentry, +nfs_setattr(const struct mnt_idmap *idmap, struct dentry *dentry, struct iattr *attr) { struct inode *inode = d_inode(dentry); @@ -955,7 +955,7 @@ static u32 nfs_get_valid_attrmask(struct inode *inode) return reply_mask; } -int nfs_getattr(struct mnt_idmap *idmap, const struct path *path, +int nfs_getattr(const struct mnt_idmap *idmap, const struct path *path, struct kstat *stat, u32 request_mask, unsigned int query_flags) { struct inode *inode = d_inode(path->dentry); diff --git a/fs/nfs/internal.h b/fs/nfs/internal.h index abc81f5ae578..d6ea41a3f9b4 100644 --- a/fs/nfs/internal.h +++ b/fs/nfs/internal.h @@ -251,6 +251,7 @@ extern struct nfs_client *nfs4_set_ds_client(struct nfs_server *mds_srv, int ds_addrlen, int ds_proto, unsigned int ds_timeo, unsigned int ds_retrans, + unsigned int ds_nconnect, u32 minor_version, bool tightly_coupled); extern struct rpc_clnt *nfs4_find_or_create_ds_client(struct nfs_client *, @@ -260,7 +261,7 @@ extern void nfs4_session_limit_xasize(struct nfs_server *server); extern struct nfs_client *nfs3_set_ds_client(struct nfs_server *mds_srv, const struct sockaddr_storage *ds_addr, int ds_addrlen, int ds_proto, unsigned int ds_timeo, - unsigned int ds_retrans); + unsigned int ds_retrans, unsigned int ds_nconnect); #ifdef CONFIG_PROC_FS extern int __init nfs_fs_proc_init(void); extern void nfs_fs_proc_exit(void); @@ -396,18 +397,18 @@ extern unsigned long nfs_access_cache_scan(struct shrinker *shrink, struct shrink_control *sc); struct dentry *nfs_lookup(struct inode *, struct dentry *, unsigned int); void nfs_d_prune_case_insensitive_aliases(struct inode *inode); -int nfs_create(struct mnt_idmap *, struct inode *, struct dentry *, +int nfs_create(const struct mnt_idmap *, struct inode *, struct dentry *, umode_t); -struct dentry *nfs_mkdir(struct mnt_idmap *, struct inode *, struct dentry *, +struct dentry *nfs_mkdir(const struct mnt_idmap *, struct inode *, struct dentry *, umode_t); int nfs_rmdir(struct inode *, struct dentry *); int nfs_unlink(struct inode *, struct dentry *); -int nfs_symlink(struct mnt_idmap *, struct inode *, struct dentry *, +int nfs_symlink(const struct mnt_idmap *, struct inode *, struct dentry *, const char *); int nfs_link(struct dentry *, struct inode *, struct dentry *); -int nfs_mknod(struct mnt_idmap *, struct inode *, struct dentry *, umode_t, +int nfs_mknod(const struct mnt_idmap *, struct inode *, struct dentry *, umode_t, dev_t); -int nfs_rename(struct mnt_idmap *, struct inode *, struct dentry *, +int nfs_rename(const struct mnt_idmap *, struct inode *, struct dentry *, struct inode *, struct dentry *, unsigned int); #ifdef CONFIG_NFS_V4_2 diff --git a/fs/nfs/namespace.c b/fs/nfs/namespace.c index 6d0073c24771..c50d59c52c50 100644 --- a/fs/nfs/namespace.c +++ b/fs/nfs/namespace.c @@ -222,7 +222,7 @@ out_fc: } static int -nfs_namespace_getattr(struct mnt_idmap *idmap, +nfs_namespace_getattr(const struct mnt_idmap *idmap, const struct path *path, struct kstat *stat, u32 request_mask, unsigned int query_flags) { @@ -235,7 +235,7 @@ nfs_namespace_getattr(struct mnt_idmap *idmap, } static int -nfs_namespace_setattr(struct mnt_idmap *idmap, struct dentry *dentry, +nfs_namespace_setattr(const struct mnt_idmap *idmap, struct dentry *dentry, struct iattr *attr) { if (NFS_FH(d_inode(dentry))->size != 0) diff --git a/fs/nfs/netns.h b/fs/nfs/netns.h index 36658579100d..560fa95726b0 100644 --- a/fs/nfs/netns.h +++ b/fs/nfs/netns.h @@ -31,7 +31,10 @@ struct nfs_net { unsigned short nfs_callback_tcpport; unsigned short nfs_callback_tcpport6; int cb_users[NFS4_MAX_MINOR_VERSION + 1]; - struct list_head nfs4_data_server_cache; +#define NFS4_DS_CACHE_HASH_BITS 8 +#define NFS4_DS_CACHE_HASH_SIZE (1 << NFS4_DS_CACHE_HASH_BITS) + /* every entry is still in bucket 0 until the key is added */ + struct hlist_head nfs4_data_server_cache[NFS4_DS_CACHE_HASH_SIZE]; spinlock_t nfs4_data_server_lock; #endif /* CONFIG_NFS_V4 */ struct nfs_netns_client *nfs_client; diff --git a/fs/nfs/nfs3_fs.h b/fs/nfs/nfs3_fs.h index b333ea119ef5..ffcabadb3546 100644 --- a/fs/nfs/nfs3_fs.h +++ b/fs/nfs/nfs3_fs.h @@ -12,7 +12,7 @@ */ #ifdef CONFIG_NFS_V3_ACL extern struct posix_acl *nfs3_get_acl(struct inode *inode, int type, bool rcu); -extern int nfs3_set_acl(struct mnt_idmap *idmap, struct dentry *dentry, +extern int nfs3_set_acl(const struct mnt_idmap *idmap, struct dentry *dentry, struct posix_acl *acl, int type); extern int nfs3_proc_setacls(struct inode *inode, struct posix_acl *acl, struct posix_acl *dfacl); diff --git a/fs/nfs/nfs3acl.c b/fs/nfs/nfs3acl.c index a126eb31f62f..2549a1985b9a 100644 --- a/fs/nfs/nfs3acl.c +++ b/fs/nfs/nfs3acl.c @@ -254,7 +254,7 @@ int nfs3_proc_setacls(struct inode *inode, struct posix_acl *acl, } -int nfs3_set_acl(struct mnt_idmap *idmap, struct dentry *dentry, +int nfs3_set_acl(const struct mnt_idmap *idmap, struct dentry *dentry, struct posix_acl *acl, int type) { struct posix_acl *orig = acl, *dfacl = NULL, *alloc; diff --git a/fs/nfs/nfs3client.c b/fs/nfs/nfs3client.c index 5d97c1d38bb6..cf2f7be4b435 100644 --- a/fs/nfs/nfs3client.c +++ b/fs/nfs/nfs3client.c @@ -84,7 +84,8 @@ struct nfs_server *nfs3_clone_server(struct nfs_server *source, */ struct nfs_client *nfs3_set_ds_client(struct nfs_server *mds_srv, const struct sockaddr_storage *ds_addr, int ds_addrlen, - int ds_proto, unsigned int ds_timeo, unsigned int ds_retrans) + int ds_proto, unsigned int ds_timeo, unsigned int ds_retrans, + unsigned int ds_nconnect) { struct rpc_timeout ds_timeout; unsigned long connect_timeout = ds_timeo * (ds_retrans + 1) * HZ / 10; @@ -124,8 +125,12 @@ struct nfs_client *nfs3_set_ds_client(struct nfs_server *mds_srv, fallthrough; case XPRT_TRANSPORT_RDMA: case XPRT_TRANSPORT_TCP: - if (mds_clp->cl_nconnect > 1) + if (mds_clp->cl_nconnect > 1) { cl_init.nconnect = mds_clp->cl_nconnect; + if (ds_nconnect) + cl_init.nconnect = min(cl_init.nconnect, + ds_nconnect); + } } if (mds_srv->flags & NFS_MOUNT_NORESVPORT) diff --git a/fs/nfs/nfs4_fs.h b/fs/nfs/nfs4_fs.h index b48e5b87cb2a..76dae699d4d7 100644 --- a/fs/nfs/nfs4_fs.h +++ b/fs/nfs/nfs4_fs.h @@ -52,6 +52,7 @@ enum nfs4_client_state { NFS4CLNT_RECALL_ANY_LAYOUT_READ, NFS4CLNT_RECALL_ANY_LAYOUT_RW, NFS4CLNT_DELEGRETURN_DELAYED, + NFS4CLNT_DEVICEID_DELETE, }; #define NFS4_RENEW_TIMEOUT 0x01 @@ -493,6 +494,7 @@ int nfs41_discover_server_trunking(struct nfs_client *clp, struct nfs_client **, const struct cred *); extern void nfs4_schedule_session_recovery(struct nfs4_session *, int); extern void nfs41_notify_server(struct nfs_client *); +extern void nfs4_deviceid_delete_recover_run(struct nfs_client *clp); bool nfs4_check_serverowner_major_id(struct nfs41_server_owner *o1, struct nfs41_server_owner *o2); @@ -509,6 +511,7 @@ extern void nfs_inode_find_state_and_recover(struct inode *inode, const nfs4_stateid *stateid); extern int nfs4_state_mark_reclaim_nograce(struct nfs_client *, struct nfs4_state *); extern void nfs4_schedule_lease_recovery(struct nfs_client *); +extern void nfs4_reset_all_state(struct nfs_client *); extern int nfs4_wait_clnt_recover(struct nfs_client *clp); extern int nfs4_client_recover_expired_lease(struct nfs_client *clp); extern void nfs4_schedule_state_manager(struct nfs_client *); diff --git a/fs/nfs/nfs4client.c b/fs/nfs/nfs4client.c index b661f446ea49..fe779fb2ec72 100644 --- a/fs/nfs/nfs4client.c +++ b/fs/nfs/nfs4client.c @@ -217,6 +217,7 @@ struct nfs_client *nfs4_alloc_client(const struct nfs_client_initdata *cl_init) clp->cl_last_renewal = jiffies; init_waitqueue_head(&clp->cl_lock_waitq); INIT_LIST_HEAD(&clp->pending_cb_stateids); + INIT_LIST_HEAD(&clp->cl_deviceid_deletes); if (cl_init->minorversion != 0) __set_bit(NFS_CS_INFINITE_SLOTS, &clp->cl_flags); @@ -286,6 +287,7 @@ static void nfs4_shutdown_client(struct nfs_client *clp) nfs4_kill_renewd(clp); clp->cl_mvops->shutdown_client(clp); nfs4_destroy_callback(clp); + pnfs_deviceid_delete_queue_free(clp); if (__test_and_clear_bit(NFS_CS_IDMAP, &clp->cl_res_state)) nfs_idmap_delete(clp); @@ -792,7 +794,8 @@ static int nfs4_set_client(struct nfs_server *server, struct nfs_client *nfs4_set_ds_client(struct nfs_server *mds_srv, const struct sockaddr_storage *ds_addr, int ds_addrlen, int ds_proto, unsigned int ds_timeo, unsigned int ds_retrans, - u32 minor_version, bool tightly_coupled) + unsigned int ds_nconnect, u32 minor_version, + bool tightly_coupled) { struct rpc_timeout ds_timeout; struct nfs_client *mds_clp = mds_srv->nfs_client; @@ -830,6 +833,9 @@ struct nfs_client *nfs4_set_ds_client(struct nfs_server *mds_srv, case XPRT_TRANSPORT_TCP: if (mds_clp->cl_nconnect > 1) { cl_init.nconnect = mds_clp->cl_nconnect; + if (ds_nconnect) + cl_init.nconnect = min(cl_init.nconnect, + ds_nconnect); cl_init.max_connect = NFS_MAX_TRANSPORTS; } } diff --git a/fs/nfs/nfs4file.c b/fs/nfs/nfs4file.c index 6401f6363f75..9a434f5dda8d 100644 --- a/fs/nfs/nfs4file.c +++ b/fs/nfs/nfs4file.c @@ -402,6 +402,7 @@ static void __nfs42_ssc_close(struct file *filep) } static const struct nfs4_ssc_client_ops nfs4_ssc_clnt_ops_tbl = { + .owner = THIS_MODULE, .sco_open = __nfs42_ssc_open, .sco_close = __nfs42_ssc_close, }; diff --git a/fs/nfs/nfs4proc.c b/fs/nfs/nfs4proc.c index 04b1987115d5..beb659744760 100644 --- a/fs/nfs/nfs4proc.c +++ b/fs/nfs/nfs4proc.c @@ -7866,7 +7866,7 @@ int nfs4_lock_delegation_recall(struct file_lock *fl, struct nfs4_state *state, #define XATTR_NAME_NFSV4_ACL "system.nfs4_acl" static int nfs4_xattr_set_nfs4_acl(const struct xattr_handler *handler, - struct mnt_idmap *idmap, + const struct mnt_idmap *idmap, struct dentry *unused, struct inode *inode, const char *key, const void *buf, size_t buflen, int flags) @@ -7889,7 +7889,7 @@ static bool nfs4_xattr_list_nfs4_acl(struct dentry *dentry) #define XATTR_NAME_NFSV4_DACL "system.nfs4_dacl" static int nfs4_xattr_set_nfs4_dacl(const struct xattr_handler *handler, - struct mnt_idmap *idmap, + const struct mnt_idmap *idmap, struct dentry *unused, struct inode *inode, const char *key, const void *buf, size_t buflen, int flags) @@ -7912,7 +7912,7 @@ static bool nfs4_xattr_list_nfs4_dacl(struct dentry *dentry) #define XATTR_NAME_NFSV4_SACL "system.nfs4_sacl" static int nfs4_xattr_set_nfs4_sacl(const struct xattr_handler *handler, - struct mnt_idmap *idmap, + const struct mnt_idmap *idmap, struct dentry *unused, struct inode *inode, const char *key, const void *buf, size_t buflen, int flags) @@ -7935,7 +7935,7 @@ static bool nfs4_xattr_list_nfs4_sacl(struct dentry *dentry) #ifdef CONFIG_NFS_V4_SECURITY_LABEL static int nfs4_xattr_set_nfs4_label(const struct xattr_handler *handler, - struct mnt_idmap *idmap, + const struct mnt_idmap *idmap, struct dentry *unused, struct inode *inode, const char *key, const void *buf, size_t buflen, int flags) @@ -7965,7 +7965,7 @@ static const struct xattr_handler nfs4_xattr_nfs4_label_handler = { #ifdef CONFIG_NFS_V4_2 static int nfs4_xattr_set_nfs4_user(const struct xattr_handler *handler, - struct mnt_idmap *idmap, + const struct mnt_idmap *idmap, struct dentry *unused, struct inode *inode, const char *key, const void *buf, size_t buflen, int flags) @@ -9681,6 +9681,14 @@ nfs4_layoutget_handle_exception(struct rpc_task *task, status = -EOVERFLOW; goto out; /* + * NFS4ERR_TOOSMALL means the layout for the requested range + * exceeds what the client advertised in loga_maxcount (see + * RFC8881 section 18.43.3). + */ + case -ETOOSMALL: + status = -EMSGSIZE; + goto out; + /* * NFS4ERR_LAYOUTTRYLATER is a conflict with another client * (or clients) writing to the same RAID stripe except when * the minlength argument is 0 (see RFC5661 section 18.43.3). @@ -10498,6 +10506,134 @@ out_put_clp: return ret; } +/* + * GETDEVICEINFO surfacing the raw status; nfs4_get_device_info() + * swallows it. A device too large for one page fails with something + * other than -ENOENT, which still proves existence. + */ +static int nfs4_deviceid_validate(struct nfs_server *server, + const struct pnfs_layoutdriver_type *ld, + const struct nfs4_deviceid *id, const struct cred *cred) +{ + struct pnfs_device pdev; + struct page *page; + int status; + + page = alloc_page(GFP_KERNEL); + if (!page) + return -ENOMEM; + + memset(&pdev, 0, sizeof(pdev)); + memcpy(&pdev.dev_id, id, sizeof(pdev.dev_id)); + pdev.layout_type = ld->id; + pdev.pages = &page; + pdev.pglen = PAGE_SIZE; + pdev.maxcount = PAGE_SIZE - nfs41_maxgetdevinfo_overhead; + + status = nfs4_proc_getdeviceinfo(server, &pdev, cred); + __free_page(page); + return status; +} + +/* + * A DELETE for a deviceID we still hold layouts on implies the server + * revoked them: run the RFC 8881 Section 18.40.4 recovery. A layout the + * server still calls valid leaves the revocations unable to confirm the + * delete, so verify it with GETDEVICEINFO. + */ +static void nfs4_deviceid_delete_recover(struct nfs_client *clp, + const struct pnfs_layoutdriver_type *ld, + const struct nfs4_deviceid *id) +{ + LIST_HEAD(layouts); + struct nfs4_deviceid_ref *ref, *confirm = NULL; + bool revoked = false; + bool inconclusive = false; + int status; + + if (pnfs_layout_collect_deviceid_refs(clp, ld, id, &layouts)) { + /* Only a partial list -- an allocation failed, or an inode is + * being evicted. Leave the device cached and recover on a + * later notification. + */ + pnfs_layout_put_deviceid_refs(&layouts); + return; + } + + if (list_empty(&layouts)) { + nfs4_delete_deviceid(ld, clp, id); + return; + } + + list_for_each_entry(ref, &layouts, node) { + struct pnfs_layout_hdr *lo = ref->lo; + struct inode *inode = ref->inode; + bool invalidated = false; + LIST_HEAD(head); + + status = nfs41_test_stateid(NFS_SERVER(inode), &ref->stateid, + ref->cred); + switch (status) { + case NFS_OK: + case -NFS4ERR_OLD_STATEID: + if (!confirm) + confirm = ref; + break; + case -NFS4ERR_ADMIN_REVOKED: + case -NFS4ERR_DELEG_REVOKED: + case -NFS4ERR_EXPIRED: + case -NFS4ERR_BAD_STATEID: + spin_lock(&inode->i_lock); + if (pnfs_layout_is_valid(lo) && + nfs4_stateid_match_other(&ref->stateid, + &lo->plh_stateid)) { + pnfs_mark_layout_stateid_invalid(lo, &head); + revoked = true; + invalidated = true; + } + spin_unlock(&inode->i_lock); + pnfs_free_lseg_list(&head); + if (invalidated) + nfs_commit_inode(inode, 0); + nfs41_free_stateid(NFS_SERVER(inode), &ref->stateid, + ref->cred, true); + break; + default: + inconclusive = true; + break; + } + } + + if (confirm) { + status = nfs4_deviceid_validate(NFS_SERVER(confirm->inode), + ld, id, confirm->cred); + if (status == -ENOENT) { + /* Section 18.40.4 prescribes EXCHANGE_ID here; + * nfs4_schedule_lease_recovery() would only renew + * the existing lease. + */ + pr_warn_ratelimited("NFS: server %s deleted a deviceID referred to by a layout it still considers valid; re-establishing the client ID\n", + clp->cl_hostname); + nfs4_reset_all_state(clp); + nfs4_delete_deviceid(ld, clp, id); + } + } else if (revoked && !inconclusive) { + nfs4_delete_deviceid(ld, clp, id); + } + pnfs_layout_put_deviceid_refs(&layouts); +} + +void nfs4_deviceid_delete_recover_run(struct nfs_client *clp) +{ + struct nfs4_deviceid_delete *dd; + + while ((dd = pnfs_deviceid_delete_dequeue(clp)) != NULL) { + nfs4_deviceid_delete_recover(clp, dd->ld, &dd->id); + pnfs_put_layoutdriver(dd->ld); + kfree(dd); + } +} + static void nfs41_free_lock_state(struct nfs_server *server, struct nfs4_lock_state *lsp) { diff --git a/fs/nfs/nfs4state.c b/fs/nfs/nfs4state.c index a5dec0473e22..b5d6daa2c416 100644 --- a/fs/nfs/nfs4state.c +++ b/fs/nfs/nfs4state.c @@ -2354,7 +2354,7 @@ void nfs41_notify_server(struct nfs_client *clp) nfs4_schedule_state_manager(clp); } -static void nfs4_reset_all_state(struct nfs_client *clp) +void nfs4_reset_all_state(struct nfs_client *clp) { if (test_and_set_bit(NFS4CLNT_LEASE_EXPIRED, &clp->cl_state) == 0) { set_bit(NFS4CLNT_PURGE_STATE, &clp->cl_state); @@ -2669,6 +2669,11 @@ static void nfs4_state_manager(struct nfs_client *clp) set_bit(NFS4CLNT_RUN_MANAGER, &clp->cl_state); } nfs4_layoutreturn_any_run(clp); + if (test_and_clear_bit(NFS4CLNT_DEVICEID_DELETE, + &clp->cl_state)) { + nfs4_deviceid_delete_recover_run(clp); + set_bit(NFS4CLNT_RUN_MANAGER, &clp->cl_state); + } clear_bit(NFS4CLNT_RECALL_RUNNING, &clp->cl_state); } diff --git a/fs/nfs/nfs4xdr.c b/fs/nfs/nfs4xdr.c index fc049ce4ba8a..8b3d96b4a0f0 100644 --- a/fs/nfs/nfs4xdr.c +++ b/fs/nfs/nfs4xdr.c @@ -6187,7 +6187,7 @@ static int decode_layoutget(struct xdr_stream *xdr, struct rpc_rqst *req, dprintk("NFS: server cheating in layoutget reply: " "layout len %u > recvd %u\n", res->layoutp->len, recvd); - status = -EINVAL; + status = -EMSGSIZE; goto out; } diff --git a/fs/nfs/pagelist.c b/fs/nfs/pagelist.c index 7dd478ffc2fa..71f0ce2bc4ea 100644 --- a/fs/nfs/pagelist.c +++ b/fs/nfs/pagelist.c @@ -404,20 +404,28 @@ static struct nfs_page *nfs_page_create(struct nfs_lock_context *l_ctx, return req; } -static void nfs_page_assign_folio(struct nfs_page *req, struct folio *folio) +static void nfs_page_assign_folio(struct nfs_page *req, struct folio *folio, + bool pinned) { if (folio != NULL) { req->wb_folio = folio; - folio_get(folio); + if (pinned) + set_bit(PG_PINNED, &req->wb_flags); + else + folio_get(folio); set_bit(PG_FOLIO, &req->wb_flags); } } -static void nfs_page_assign_page(struct nfs_page *req, struct page *page) +static void nfs_page_assign_page(struct nfs_page *req, struct page *page, + bool pinned) { if (page != NULL) { req->wb_page = page; - get_page(page); + if (pinned) + set_bit(PG_PINNED, &req->wb_flags); + else + get_page(page); } } @@ -425,6 +433,7 @@ static void nfs_page_assign_page(struct nfs_page *req, struct page *page) * nfs_page_create_from_page - Create an NFS read/write request. * @ctx: open context to use * @page: page to write + * @pinned: true if page is pinned * @pgbase: starting offset within the page for the write * @offset: file offset for the write * @count: number of bytes to read/write @@ -435,6 +444,7 @@ static void nfs_page_assign_page(struct nfs_page *req, struct page *page) */ struct nfs_page *nfs_page_create_from_page(struct nfs_open_context *ctx, struct page *page, + bool pinned, unsigned int pgbase, loff_t offset, unsigned int count) { @@ -446,7 +456,9 @@ struct nfs_page *nfs_page_create_from_page(struct nfs_open_context *ctx, ret = nfs_page_create(l_ctx, pgbase, offset >> PAGE_SHIFT, offset_in_page(offset), count); if (!IS_ERR(ret)) { - nfs_page_assign_page(ret, page); + nfs_page_assign_page(ret, page, pinned); + if (pinned) + ret->wb_nr_pinned = 1; nfs_page_group_init(ret, NULL); } nfs_put_lock_context(l_ctx); @@ -457,6 +469,7 @@ struct nfs_page *nfs_page_create_from_page(struct nfs_open_context *ctx, * nfs_page_create_from_folio - Create an NFS read/write request. * @ctx: open context to use * @folio: folio to write + * @pinned: true if folio is pinned * @offset: starting offset within the folio for the write * @count: number of bytes to read/write * @@ -466,6 +479,7 @@ struct nfs_page *nfs_page_create_from_page(struct nfs_open_context *ctx, */ struct nfs_page *nfs_page_create_from_folio(struct nfs_open_context *ctx, struct folio *folio, + bool pinned, unsigned int offset, unsigned int count) { @@ -476,7 +490,10 @@ struct nfs_page *nfs_page_create_from_folio(struct nfs_open_context *ctx, return ERR_CAST(l_ctx); ret = nfs_page_create(l_ctx, offset, folio->index, offset, count); if (!IS_ERR(ret)) { - nfs_page_assign_folio(ret, folio); + nfs_page_assign_folio(ret, folio, pinned); + if (pinned) + ret->wb_nr_pinned = nfs_page_array_len(offset_in_page(offset), + count); nfs_page_group_init(ret, NULL); } nfs_put_lock_context(l_ctx); @@ -498,9 +515,11 @@ nfs_create_subreq(struct nfs_page *req, offset, count); if (!IS_ERR(ret)) { if (folio) - nfs_page_assign_folio(ret, folio); + nfs_page_assign_folio(ret, folio, + test_bit(PG_PINNED, &req->wb_flags)); else - nfs_page_assign_page(ret, page); + nfs_page_assign_page(ret, page, + test_bit(PG_PINNED, &req->wb_flags)); /* find the last request */ for (last = req->wb_head; last->wb_this_page != req->wb_head; @@ -552,11 +571,21 @@ static void nfs_clear_request(struct nfs_page *req) struct nfs_open_context *ctx; if (folio != NULL) { - folio_put(folio); + if (test_and_clear_bit(PG_PINNED, &req->wb_flags)) { + if (req->wb_nr_pinned > 0) + unpin_user_folio(folio, req->wb_nr_pinned); + } else { + folio_put(folio); + } req->wb_folio = NULL; clear_bit(PG_FOLIO, &req->wb_flags); } else if (page != NULL) { - put_page(page); + if (test_and_clear_bit(PG_PINNED, &req->wb_flags)) { + if (req->wb_nr_pinned > 0) + unpin_user_pages(&page, req->wb_nr_pinned); + } else { + put_page(page); + } req->wb_page = NULL; } if (l_ctx != NULL) { @@ -600,6 +629,23 @@ void nfs_release_request(struct nfs_page *req) EXPORT_SYMBOL_GPL(nfs_release_request); /* + * nfs_release_request_list - Release a list of NFS read/write requests + * @head: list of requests to release + * + * Removes each request from the list and drops it's refcount. + */ +void nfs_release_request_list(struct list_head *head) +{ + while (!list_empty(head)) { + struct nfs_page *req = nfs_list_entry(head->next); + + nfs_list_remove_request(req); + nfs_release_request(req); + } +} +EXPORT_SYMBOL_GPL(nfs_release_request_list); + +/* * nfs_generic_pg_test - determine if requests can be coalesced * @desc: pointer to descriptor * @prev: previous request in desc, or NULL diff --git a/fs/nfs/pnfs.c b/fs/nfs/pnfs.c index 4f9c0f639014..93a0852e1e8a 100644 --- a/fs/nfs/pnfs.c +++ b/fs/nfs/pnfs.c @@ -433,7 +433,7 @@ bool nfs4_layout_refresh_old_stateid(nfs4_stateid *dst, } /* Try to update the seqid to the most recent */ err = pnfs_mark_matching_lsegs_return(lo, &head, &range, 0, - true); + true, NULL); if (err != -EBUSY) { dst->seqid = lo->plh_stateid.seqid; *dst_range = range; @@ -487,7 +487,8 @@ static int pnfs_mark_layout_stateid_return(struct pnfs_layout_hdr *lo, .length = NFS4_MAX_UINT64, }; - return pnfs_mark_matching_lsegs_return(lo, lseg_list, &range, seq, true); + return pnfs_mark_matching_lsegs_return(lo, lseg_list, &range, seq, true, + NULL); } static int @@ -525,7 +526,7 @@ pnfs_layout_io_set_failed(struct pnfs_layout_hdr *lo, u32 iomode) spin_lock(&inode->i_lock); pnfs_layout_set_fail_bit(lo, pnfs_iomode_to_fail_bit(iomode)); - pnfs_mark_matching_lsegs_return(lo, &head, &range, 0, true); + pnfs_mark_matching_lsegs_return(lo, &head, &range, 0, true, NULL); spin_unlock(&inode->i_lock); pnfs_free_lseg_list(&head); dprintk("%s Setting layout IOMODE_%s fail bit\n", __func__, @@ -740,7 +741,7 @@ pnfs_mark_matching_lsegs_invalid(struct pnfs_layout_hdr *lo, if (mark_lseg_invalid(lseg, tmp_list)) continue; remaining++; - pnfs_lseg_cancel_io(server, lseg); + pnfs_lseg_cancel_io(server, lseg, NULL); } dprintk("%s:Return %i\n", __func__, remaining); return remaining; @@ -1167,11 +1168,12 @@ pnfs_alloc_init_layoutget_args(struct inode *ino, struct nfs_open_context *ctx, const nfs4_stateid *stateid, const struct pnfs_layout_range *range, - gfp_t gfp_flags) + size_t min_reply_sz, gfp_t gfp_flags) { struct nfs_server *server = pnfs_find_server(ino, ctx); size_t max_reply_sz = server->pnfs_curr_ld->max_layoutget_response; - size_t max_pages = max_response_pages(server); + size_t session_pages = max_response_pages(server); + size_t max_pages = session_pages; struct nfs4_layoutget *lgp; dprintk("--> %s\n", __func__); @@ -1186,6 +1188,19 @@ pnfs_alloc_init_layoutget_args(struct inode *ino, max_pages = npages; } + /* + * A previous LAYOUTGET on this layout or on this server did not + * fit the reply buffer: raise the layout driver's default up to + * the session's maximum response size. + */ + if (!min_reply_sz) + min_reply_sz = READ_ONCE(server->lg_reply_sz); + if (min_reply_sz) { + size_t npages = (min_reply_sz + PAGE_SIZE - 1) >> PAGE_SHIFT; + if (npages > max_pages) + max_pages = min(npages, session_pages); + } + lgp->args.layout.pages = nfs4_alloc_pages(max_pages, gfp_flags); if (!lgp->args.layout.pages) { kfree(lgp); @@ -1210,7 +1225,7 @@ pnfs_alloc_init_layoutget_args(struct inode *ino, lgp->args.minlength = i_size - range->offset; } } - lgp->args.maxcount = PNFS_LAYOUT_MAXSIZE; + lgp->args.maxcount = lgp->args.layout.pglen; pnfs_copy_range(&lgp->args.range, range); lgp->args.type = server->pnfs_curr_ld->id; lgp->args.inode = ino; @@ -1462,7 +1477,7 @@ _pnfs_return_layout(struct inode *ino) } valid_layout = pnfs_layout_is_valid(lo); pnfs_clear_layoutcommit(ino, &tmp_list); - pnfs_mark_matching_lsegs_return(lo, &tmp_list, &range, 0, true); + pnfs_mark_matching_lsegs_return(lo, &tmp_list, &range, 0, true, NULL); /* Don't send a LAYOUTRETURN if list was initially empty */ @@ -2145,6 +2160,7 @@ pnfs_update_layout(struct inode *ino, .inode = ino, }; unsigned long giveup = jiffies + (clp->cl_lease_time << 1); + size_t reply_sz = 0; bool first; if (!pnfs_enabled_sb(NFS_SERVER(ino))) { @@ -2304,7 +2320,8 @@ lookup_again: if (arg.length != NFS4_MAX_UINT64) arg.length = PAGE_ALIGN(arg.length); - lgp = pnfs_alloc_init_layoutget_args(ino, ctx, &stateid, &arg, gfp_flags); + lgp = pnfs_alloc_init_layoutget_args(ino, ctx, &stateid, &arg, reply_sz, + gfp_flags); if (!lgp) { lseg = ERR_PTR(-ENOMEM); trace_pnfs_update_layout(ino, pos, count, iomode, lo, NULL, @@ -2335,6 +2352,25 @@ lookup_again: lo, pnfs_iomode_to_fail_bit(iomode)); lseg = NULL; goto out_put_layout_hdr; + case -EMSGSIZE: { + /* + * The layout exceeded loga_maxcount (NFS4ERR_TOOSMALL): + * retry once with the reply buffer raised to the + * session's maximum response size before falling back + * to I/O through the MDS. + */ + size_t max = max_response_pages(server) << PAGE_SHIFT; + + if (reply_sz < max) { + reply_sz = max; + exception.retry = 1; + break; + } + pnfs_layout_set_fail_bit( + lo, pnfs_iomode_to_fail_bit(iomode)); + lseg = NULL; + goto out_put_layout_hdr; + } default: if (!nfs_error_is_fatal(PTR_ERR(lseg))) { pnfs_layout_clear_fail_bit(lo, pnfs_iomode_to_fail_bit(iomode)); @@ -2354,6 +2390,13 @@ lookup_again: goto lookup_again; } } else { + /* + * A LAYOUTGET that only succeeded with an escalated reply + * buffer: remember the size so that future LAYOUTGETs to + * this server skip the attempt at the driver's default. + */ + if (reply_sz) + WRITE_ONCE(server->lg_reply_sz, reply_sz); pnfs_layout_clear_fail_bit(lo, pnfs_iomode_to_fail_bit(iomode)); } @@ -2448,7 +2491,7 @@ static void _lgopen_prepare_attached(struct nfs4_opendata *data, lo = _pnfs_grab_empty_layout(ino, ctx); if (!lo) return; - lgp = pnfs_alloc_init_layoutget_args(ino, ctx, ¤t_stateid, &rng, + lgp = pnfs_alloc_init_layoutget_args(ino, ctx, ¤t_stateid, &rng, 0, nfs_io_gfp_mask()); if (!lgp) { clear_and_wake_up_bit(NFS_LAYOUT_FIRST_LAYOUTGET, &lo->plh_flags); @@ -2474,7 +2517,7 @@ static void _lgopen_prepare_floating(struct nfs4_opendata *data, }; struct nfs4_layoutget *lgp; - lgp = pnfs_alloc_init_layoutget_args(ino, ctx, ¤t_stateid, &rng, + lgp = pnfs_alloc_init_layoutget_args(ino, ctx, ¤t_stateid, &rng, 0, nfs_io_gfp_mask()); if (!lgp) return; @@ -2616,7 +2659,8 @@ pnfs_layout_process(struct nfs4_layoutget *lgp) .iomode = IOMODE_ANY, .length = NFS4_MAX_UINT64, }; - pnfs_mark_matching_lsegs_return(lo, &free_me, &range, 0, true); + pnfs_mark_matching_lsegs_return(lo, &free_me, &range, 0, true, + NULL); goto out_forget; } else { /* We have a completely new layout */ @@ -2649,6 +2693,7 @@ out_forget: * @return_range: describe layout segment ranges to be returned * @seq: stateid seqid to match * @cancel_io: signal io be cancelled + * @devid: only cancel io directed at this device (all devices if NULL) * * This function is mainly intended for use by layoutrecall. It attempts * to free the layout segment immediately, or else to mark it for return @@ -2663,7 +2708,8 @@ int pnfs_mark_matching_lsegs_return(struct pnfs_layout_hdr *lo, struct list_head *tmp_list, const struct pnfs_layout_range *return_range, - u32 seq, bool cancel_io) + u32 seq, bool cancel_io, + const struct nfs4_deviceid *devid) { struct pnfs_layout_segment *lseg, *next; struct nfs_server *server = NFS_SERVER(lo->plh_inode); @@ -2690,7 +2736,7 @@ pnfs_mark_matching_lsegs_return(struct pnfs_layout_hdr *lo, remaining++; set_bit(NFS_LSEG_LAYOUTRETURN, &lseg->pls_flags); if (cancel_io) - pnfs_lseg_cancel_io(server, lseg); + pnfs_lseg_cancel_io(server, lseg, devid); } if (remaining) { @@ -2708,7 +2754,8 @@ pnfs_mark_matching_lsegs_return(struct pnfs_layout_hdr *lo, static void pnfs_mark_layout_for_return(struct inode *inode, - const struct pnfs_layout_range *range) + const struct pnfs_layout_range *range, + const struct nfs4_deviceid *devid) { struct pnfs_layout_hdr *lo; bool return_now = false; @@ -2726,7 +2773,7 @@ pnfs_mark_layout_for_return(struct inode *inode, * for how it works. */ if (pnfs_mark_matching_lsegs_return(lo, &lo->plh_return_segs, range, 0, - true) != -EBUSY) { + true, devid) != -EBUSY) { const struct cred *cred; nfs4_stateid stateid; enum pnfs_iomode iomode; @@ -2743,7 +2790,8 @@ pnfs_mark_layout_for_return(struct inode *inode, } void pnfs_error_mark_layout_for_return(struct inode *inode, - struct pnfs_layout_segment *lseg) + struct pnfs_layout_segment *lseg, + const struct nfs4_deviceid *devid) { struct pnfs_layout_range range = { .iomode = lseg->pls_range.iomode, @@ -2751,7 +2799,7 @@ void pnfs_error_mark_layout_for_return(struct inode *inode, .length = NFS4_MAX_UINT64, }; - pnfs_mark_layout_for_return(inode, &range); + pnfs_mark_layout_for_return(inode, &range, devid); } EXPORT_SYMBOL_GPL(pnfs_error_mark_layout_for_return); @@ -2841,7 +2889,7 @@ restart: pnfs_get_layout_hdr(lo); pnfs_set_plh_return_info(lo, range->iomode, 0); if (pnfs_mark_matching_lsegs_return(lo, &lo->plh_return_segs, - range, 0, true) != 0 || + range, 0, true, NULL) != 0 || !pnfs_prepare_layoutreturn(lo, &stateid, &cred, &iomode)) { spin_unlock(&inode->i_lock); rcu_read_unlock(); @@ -2875,6 +2923,294 @@ pnfs_layout_return_unused_byclid(struct nfs_client *clp, &range); } +struct pnfs_reresolve_deviceid_args { + const struct pnfs_layoutdriver_type *ld; + const struct nfs4_deviceid *devid; + bool immediate; + struct list_head put_list; +}; + +static int pnfs_layout_reresolve_deviceid_byserver(struct nfs_server *server, + void *data) +{ + struct pnfs_reresolve_deviceid_args *args = data; + struct pnfs_layout_hdr *lo; + struct inode *inode; + + if (server->pnfs_curr_ld != args->ld) + return 0; + + rcu_read_lock(); + list_for_each_entry_rcu(lo, &server->layouts, plh_layouts) { + inode = lo->plh_inode; + if (!inode) + continue; + spin_lock(&inode->i_lock); + args->ld->reresolve_deviceid(lo, args->devid, args->immediate, + &args->put_list); + spin_unlock(&inode->i_lock); + } + rcu_read_unlock(); + return 0; +} + +/* + * Invoke @ld's reresolve_deviceid hook for @devid on every layout of @clp's + * servers, then drain the put_list once the locks are dropped. + */ +void +pnfs_layout_reresolve_deviceid_byclid(struct nfs_client *clp, + const struct pnfs_layoutdriver_type *ld, + const struct nfs4_deviceid *devid, + bool immediate) +{ + struct pnfs_reresolve_deviceid_args args = { + .ld = ld, + .devid = devid, + .immediate = immediate, + .put_list = LIST_HEAD_INIT(args.put_list), + }; + struct nfs4_deviceid_put *put, *tmp; + + if (!ld->reresolve_deviceid) + return; + + nfs_client_for_each_server(clp, + pnfs_layout_reresolve_deviceid_byserver, &args); + + list_for_each_entry_safe(put, tmp, &args.put_list, node) { + list_del(&put->node); + nfs4_put_deviceid_node(put->dev); + kfree(put); + } +} + +struct pnfs_deviceid_ref_args { + const struct pnfs_layoutdriver_type *ld; + const struct nfs4_deviceid *devid; + struct list_head *result; + bool found; +}; + +static int pnfs_layout_deviceid_referenced_byserver( + struct nfs_server *server, void *data) +{ + struct pnfs_deviceid_ref_args *args = data; + struct pnfs_layout_hdr *lo; + struct inode *inode; + + if (server->pnfs_curr_ld != args->ld) + return 0; + + rcu_read_lock(); + list_for_each_entry_rcu(lo, &server->layouts, plh_layouts) { + inode = lo->plh_inode; + if (!inode) + continue; + spin_lock(&inode->i_lock); + if (NFS_I(inode)->layout == lo && pnfs_layout_is_valid(lo) && + args->ld->layout_references_deviceid(lo, args->devid)) + args->found = true; + spin_unlock(&inode->i_lock); + if (args->found) + break; + } + rcu_read_unlock(); + return args->found; +} + +/* + * pnfs_layout_deviceid_referenced_byclid - does any live layout of + * @clp's servers using @ld still reference deviceid @devid? + */ +bool +pnfs_layout_deviceid_referenced_byclid(struct nfs_client *clp, + const struct pnfs_layoutdriver_type *ld, + const struct nfs4_deviceid *devid) +{ + struct pnfs_deviceid_ref_args args = { + .ld = ld, + .devid = devid, + }; + + if (!ld->layout_references_deviceid) + return false; + + nfs_client_for_each_server(clp, + pnfs_layout_deviceid_referenced_byserver, &args); + return args.found; +} + +static int pnfs_layout_collect_deviceid_refs_byserver( + struct nfs_server *server, void *data) +{ + struct pnfs_deviceid_ref_args *args = data; + struct nfs4_deviceid_ref *ref, *tmp; + struct pnfs_layout_hdr *lo; + struct inode *inode; + LIST_HEAD(putme); + bool matched; + int ret = 0; + + if (server->pnfs_curr_ld != args->ld) + return 0; + + rcu_read_lock(); + list_for_each_entry_rcu(lo, &server->layouts, plh_layouts) { + inode = lo->plh_inode; + if (!inode) + continue; + + spin_lock(&inode->i_lock); + matched = NFS_I(inode)->layout == lo && + pnfs_layout_is_valid(lo) && + args->ld->layout_references_deviceid(lo, args->devid); + if (!matched) { + spin_unlock(&inode->i_lock); + continue; + } + ref = kzalloc_obj(*ref, GFP_ATOMIC); + if (!ref) { + spin_unlock(&inode->i_lock); + ret = -ENOMEM; + break; + } + /* NFS_I()->layout == lo under i_lock means the refcount has + * not reached zero: pnfs_put_layout_hdr() decrements to zero + * and detaches in the same critical section. + */ + pnfs_get_layout_hdr(lo); + ref->lo = lo; + nfs4_stateid_copy(&ref->stateid, &lo->plh_stateid); + ref->cred = get_cred(lo->plh_lc_cred); + spin_unlock(&inode->i_lock); + + /* the pinned hdr does not hold the inode: grab it (and + * keep the superblock active) for use across RPCs + */ + ref->inode = nfs_igrab_and_active(inode); + if (!ref->inode) { + /* The layout may still name the deviceID, so report a + * partial list rather than silently shortening it. + * Defer the put: it can layoutreturn and sleep. + */ + list_add(&ref->node, &putme); + ret = -EAGAIN; + break; + } + list_add_tail(&ref->node, args->result); + } + rcu_read_unlock(); + + list_for_each_entry_safe(ref, tmp, &putme, node) { + list_del(&ref->node); + pnfs_put_layout_hdr(ref->lo); + put_cred(ref->cred); + kfree(ref); + } + return ret; +} + +/* + * Collect @clp's layouts referencing @devid onto @result as entries usable + * across sleeping RPCs; release with pnfs_layout_put_deviceid_refs(). + * A negative return means @result is only a partial set. + */ +int +pnfs_layout_collect_deviceid_refs(struct nfs_client *clp, + const struct pnfs_layoutdriver_type *ld, + const struct nfs4_deviceid *devid, + struct list_head *result) +{ + struct pnfs_deviceid_ref_args args = { + .ld = ld, + .devid = devid, + .result = result, + }; + + if (!ld->layout_references_deviceid) + return 0; + + return nfs_client_for_each_server(clp, + pnfs_layout_collect_deviceid_refs_byserver, &args); +} + +void +pnfs_layout_put_deviceid_refs(struct list_head *result) +{ + struct nfs4_deviceid_ref *ref, *tmp; + + list_for_each_entry_safe(ref, tmp, result, node) { + list_del(&ref->node); + put_cred(ref->cred); + pnfs_put_layout_hdr(ref->lo); + nfs_iput_and_deactive(ref->inode); + kfree(ref); + } +} + +/* + * Queue @id for the state manager's Section 18.40.4 recovery, + * dropping duplicates of an already-queued suspect. + */ +void pnfs_deviceid_delete_mark(struct nfs_client *clp, + const struct pnfs_layoutdriver_type *ld, + const struct nfs4_deviceid *id) +{ + struct nfs4_deviceid_delete *dd, *new; + + new = kzalloc_obj(*new, GFP_KERNEL); + if (!new) + return; /* lost notification; recovery waits for the next */ + new->ld = pnfs_find_layoutdriver(ld->id); + if (!new->ld) { + kfree(new); + return; + } + memcpy(&new->id, id, sizeof(new->id)); + + spin_lock(&clp->cl_lock); + list_for_each_entry(dd, &clp->cl_deviceid_deletes, list) { + if (dd->ld == new->ld && + !memcmp(&dd->id, &new->id, sizeof(dd->id))) { + spin_unlock(&clp->cl_lock); + pnfs_put_layoutdriver(new->ld); + kfree(new); + return; + } + } + list_add_tail(&new->list, &clp->cl_deviceid_deletes); + spin_unlock(&clp->cl_lock); + + set_bit(NFS4CLNT_DEVICEID_DELETE, &clp->cl_state); + nfs4_schedule_state_manager(clp); +} + +struct nfs4_deviceid_delete *pnfs_deviceid_delete_dequeue( + struct nfs_client *clp) +{ + struct nfs4_deviceid_delete *dd = NULL; + + spin_lock(&clp->cl_lock); + if (!list_empty(&clp->cl_deviceid_deletes)) { + dd = list_first_entry(&clp->cl_deviceid_deletes, + struct nfs4_deviceid_delete, list); + list_del(&dd->list); + } + spin_unlock(&clp->cl_lock); + return dd; +} + +void pnfs_deviceid_delete_queue_free(struct nfs_client *clp) +{ + struct nfs4_deviceid_delete *dd; + + while ((dd = pnfs_deviceid_delete_dequeue(clp)) != NULL) { + pnfs_put_layoutdriver(dd->ld); + kfree(dd); + } +} + /* Check if we have we have a valid layout but if there isn't an intersection * between the request and the pgio->pg_lseg, put this pgio->pg_lseg away. */ @@ -3103,6 +3439,7 @@ pnfs_do_write(struct nfs_pageio_descriptor *desc, static void pnfs_writehdr_free(struct nfs_pgio_header *hdr) { + pnfs_put_ds_dev(hdr->ds_dev); pnfs_put_lseg(hdr->lseg); nfs_pgio_header_free(hdr); } @@ -3248,6 +3585,7 @@ pnfs_do_read(struct nfs_pageio_descriptor *desc, struct nfs_pgio_header *hdr) static void pnfs_readhdr_free(struct nfs_pgio_header *hdr) { + pnfs_put_ds_dev(hdr->ds_dev); pnfs_put_lseg(hdr->lseg); nfs_pgio_header_free(hdr); } diff --git a/fs/nfs/pnfs.h b/fs/nfs/pnfs.h index 70d20f779678..2774d4adf4c3 100644 --- a/fs/nfs/pnfs.h +++ b/fs/nfs/pnfs.h @@ -57,7 +57,7 @@ struct nfs4_pnfs_ds_addr { }; struct nfs4_pnfs_ds { - struct list_head ds_node; /* nfs4_pnfs_dev_hlist dev_dslist */ + struct hlist_node ds_node; /* nfs_net nfs4_data_server_cache */ char *ds_remotestr; /* comma sep list of addrs */ struct list_head ds_addrs; const struct net *ds_net; @@ -171,6 +171,25 @@ struct pnfs_layoutdriver_type { struct nfs4_deviceid_node * (*alloc_deviceid_node) (struct nfs_server *server, struct pnfs_device *pdev, gfp_t gfp_flags); + /* + * Re-resolve @lo's references to the changed deviceid @id. Called + * under @lo's inode i_lock inside an RCU read-side critical section: + * must not sleep, allocations are GFP_ATOMIC. Rather than put the + * references it gives up (the final put can sleep), the hook + * allocates an nfs4_deviceid_put per reference and queues it on + * @put_list for the caller to put and free. On allocation failure + * it must leave the reference in place. + */ + void (*reresolve_deviceid)(struct pnfs_layout_hdr *lo, + const struct nfs4_deviceid *id, + bool immediate, + struct list_head *put_list); + /* + * Does @lo hold any reference to deviceid @id? Called under + * @lo's inode i_lock; must not sleep. + */ + bool (*layout_references_deviceid)(struct pnfs_layout_hdr *lo, + const struct nfs4_deviceid *id); int (*prepare_layoutreturn) (struct nfs4_layoutreturn_args *); @@ -178,7 +197,8 @@ struct pnfs_layoutdriver_type { int (*prepare_layoutcommit) (struct nfs4_layoutcommit_args *args); int (*prepare_layoutstats) (struct nfs42_layoutstat_args *args); - void (*cancel_io)(struct pnfs_layout_segment *lseg); + void (*cancel_io)(struct pnfs_layout_segment *lseg, + const struct nfs4_deviceid *devid); }; struct pnfs_commit_ops { @@ -301,7 +321,8 @@ int pnfs_mark_matching_lsegs_invalid(struct pnfs_layout_hdr *lo, int pnfs_mark_matching_lsegs_return(struct pnfs_layout_hdr *lo, struct list_head *tmp_list, const struct pnfs_layout_range *recall_range, - u32 seq, bool cancel_io); + u32 seq, bool cancel_io, + const struct nfs4_deviceid *devid); int pnfs_mark_layout_stateid_invalid(struct pnfs_layout_hdr *lo, struct list_head *lseg_list); bool pnfs_roc(struct inode *ino, struct nfs4_layoutreturn_args *args, @@ -351,9 +372,56 @@ int pnfs_read_done_resend_to_mds(struct nfs_pgio_header *); int pnfs_write_done_resend_to_mds(struct nfs_pgio_header *); struct nfs4_threshold *pnfs_mdsthreshold_alloc(void); void pnfs_error_mark_layout_for_return(struct inode *inode, - struct pnfs_layout_segment *lseg); + struct pnfs_layout_segment *lseg, + const struct nfs4_deviceid *devid); void pnfs_layout_return_unused_byclid(struct nfs_client *clp, enum pnfs_iomode iomode); +void pnfs_layout_reresolve_deviceid_byclid(struct nfs_client *clp, + const struct pnfs_layoutdriver_type *ld, + const struct nfs4_deviceid *devid, + bool immediate); +bool pnfs_layout_deviceid_referenced_byclid(struct nfs_client *clp, + const struct pnfs_layoutdriver_type *ld, + const struct nfs4_deviceid *devid); + +/* + * One live layout referencing a deviceID, collected for the + * CB_NOTIFY_DEVICEID DELETE recovery: the hdr is pinned, the inode + * igrab'd with its superblock active, and the layout stateid and + * cred snapshotted for TEST_STATEID. + */ +struct nfs4_deviceid_ref { + struct list_head node; + struct pnfs_layout_hdr *lo; + struct inode *inode; + nfs4_stateid stateid; + const struct cred *cred; +}; + +int pnfs_layout_collect_deviceid_refs(struct nfs_client *clp, + const struct pnfs_layoutdriver_type *ld, + const struct nfs4_deviceid *devid, + struct list_head *result); +void pnfs_layout_put_deviceid_refs(struct list_head *result); + +/* + * A CB_NOTIFY_DEVICEID DELETE naming a deviceID that live layouts + * still reference (RFC 8881 Section 18.40.4). Queued on + * nfs_client.cl_deviceid_deletes under cl_lock for the state manager + * to resolve; holds a layoutdriver reference. + */ +struct nfs4_deviceid_delete { + struct list_head list; + const struct pnfs_layoutdriver_type *ld; + struct nfs4_deviceid id; +}; + +void pnfs_deviceid_delete_mark(struct nfs_client *clp, + const struct pnfs_layoutdriver_type *ld, + const struct nfs4_deviceid *id); +struct nfs4_deviceid_delete *pnfs_deviceid_delete_dequeue( + struct nfs_client *clp); +void pnfs_deviceid_delete_queue_free(struct nfs_client *clp); int pnfs_layout_handle_reboot(struct nfs_client *clp); /* nfs4_deviceid_flags */ @@ -376,15 +444,32 @@ struct nfs4_deviceid_node { atomic_t ref; }; +/* One reference given up by reresolve_deviceid; nodes are shared, so a + * single pass can unpin the same node more than once. + */ +struct nfs4_deviceid_put { + struct list_head node; + struct nfs4_deviceid_node *dev; +}; + struct nfs4_deviceid_node * nfs4_find_get_deviceid(struct nfs_server *server, const struct nfs4_deviceid *id, const struct cred *cred, gfp_t gfp_mask); void nfs4_delete_deviceid(const struct pnfs_layoutdriver_type *, const struct nfs_client *, const struct nfs4_deviceid *); +void nfs4_deviceid_bump_change_epoch(struct nfs_client *clp); void nfs4_init_deviceid_node(struct nfs4_deviceid_node *, struct nfs_server *, const struct nfs4_deviceid *); bool nfs4_put_deviceid_node(struct nfs4_deviceid_node *); void nfs4_mark_deviceid_available(struct nfs4_deviceid_node *node); + +/* Put the device node reference carried by an in-flight I/O, if any */ +static inline void pnfs_put_ds_dev(struct nfs4_deviceid_node *dev) +{ + if (dev) + nfs4_put_deviceid_node(dev); +} + void nfs4_mark_deviceid_unavailable(struct nfs4_deviceid_node *node); bool nfs4_test_deviceid_unavailable(struct nfs4_deviceid_node *node); void nfs4_deviceid_purge_client(const struct nfs_client *); @@ -422,11 +507,13 @@ struct nfs4_pnfs_ds *nfs4_pnfs_ds_add(const struct net *net, void nfs4_pnfs_v3_ds_connect_unload(void); int nfs4_pnfs_ds_connect(struct nfs_server *mds_srv, struct nfs4_pnfs_ds *ds, struct nfs4_deviceid_node *devid, unsigned int timeo, - unsigned int retrans, u32 version, u32 minor_version, + unsigned int retrans, unsigned int nconnect, + u32 version, u32 minor_version, bool tightly_coupled); struct nfs4_pnfs_ds_addr *nfs4_decode_mp_ds_addr(struct net *net, struct xdr_stream *xdr, gfp_t gfp_flags); +void nfs4_pnfs_ds_addr_list_free(struct list_head *dsaddrs); void pnfs_layout_mark_request_commit(struct nfs_page *req, struct pnfs_layout_segment *lseg, struct nfs_commit_info *cinfo, @@ -690,10 +777,11 @@ pnfs_lseg_request_intersecting(struct pnfs_layout_segment *lseg, struct nfs_page } static inline void pnfs_lseg_cancel_io(struct nfs_server *server, - struct pnfs_layout_segment *lseg) + struct pnfs_layout_segment *lseg, + const struct nfs4_deviceid *devid) { if (server->pnfs_curr_ld->cancel_io) - server->pnfs_curr_ld->cancel_io(lseg); + server->pnfs_curr_ld->cancel_io(lseg, devid); } extern unsigned int layoutstats_timer; diff --git a/fs/nfs/pnfs_dev.c b/fs/nfs/pnfs_dev.c index 274abdd6d5f3..e8143977965c 100644 --- a/fs/nfs/pnfs_dev.c +++ b/fs/nfs/pnfs_dev.c @@ -40,8 +40,11 @@ /* * Device ID RCU cache. A device ID is unique per server and layout type. + * + * 256 buckets keeps the chains short at the 1024-or-more devices a + * striping deployment expects, for 2KB of BSS on 64-bit. */ -#define NFS4_DEVICE_ID_HASH_BITS 5 +#define NFS4_DEVICE_ID_HASH_BITS 8 #define NFS4_DEVICE_ID_HASH_SIZE (1 << NFS4_DEVICE_ID_HASH_BITS) #define NFS4_DEVICE_ID_HASH_MASK (NFS4_DEVICE_ID_HASH_SIZE - 1) @@ -181,6 +184,21 @@ __nfs4_find_get_deviceid(struct nfs_server *server, return d; } +/* + * Bumped before the stale entry is unhashed, so an insert serialised + * after the unhash by nfs4_deviceid_lock observes the new epoch. + */ +void +nfs4_deviceid_bump_change_epoch(struct nfs_client *clp) +{ + atomic_inc(&clp->cl_deviceid_change_epoch); +} + +/* Discarding a raced reply is an optimisation, not a correctness + * requirement, and the epoch moves at the server's rate: bound it. + */ +#define NFS4_DEVICEID_FETCH_RETRIES 3 + struct nfs4_deviceid_node * nfs4_find_get_deviceid(struct nfs_server *server, const struct nfs4_deviceid *id, const struct cred *cred, @@ -188,11 +206,14 @@ nfs4_find_get_deviceid(struct nfs_server *server, { long hash = nfs4_deviceid_hash(id); struct nfs4_deviceid_node *d, *new; + int epoch, tries = 0; +retry: d = __nfs4_find_get_deviceid(server, id, hash); if (d) goto found; + epoch = atomic_read(&server->nfs_client->cl_deviceid_change_epoch); new = nfs4_get_device_info(server, id, cred, gfp_mask); if (!new) { trace_nfs4_find_deviceid(server, id, -ENOENT); @@ -200,6 +221,13 @@ nfs4_find_get_deviceid(struct nfs_server *server, } spin_lock(&nfs4_deviceid_lock); + if (atomic_read(&server->nfs_client->cl_deviceid_change_epoch) != epoch && + ++tries <= NFS4_DEVICEID_FETCH_RETRIES) { + /* a mapping changed while we fetched; ours may be stale */ + spin_unlock(&nfs4_deviceid_lock); + server->pnfs_curr_ld->free_deviceid_node(new); + goto retry; + } d = __nfs4_find_get_deviceid(server, id, hash); if (d) { spin_unlock(&nfs4_deviceid_lock); diff --git a/fs/nfs/pnfs_nfs.c b/fs/nfs/pnfs_nfs.c index 93d63f75a355..7adb6f941cf2 100644 --- a/fs/nfs/pnfs_nfs.c +++ b/fs/nfs/pnfs_nfs.c @@ -15,6 +15,8 @@ #include "nfs4session.h" #include "internal.h" +#include <linux/hash.h> +#include <linux/jhash.h> #include "pnfs.h" #include "netns.h" #include "nfs4trace.h" @@ -55,6 +57,7 @@ void pnfs_generic_commit_release(void *calldata) struct nfs_commit_data *data = calldata; data->completion_ops->completion(data); + pnfs_put_ds_dev(data->ds_dev); pnfs_put_lseg(data->lseg); nfs_put_client(data->ds_clp); nfs_commitdata_release(data); @@ -576,8 +579,8 @@ same_sockaddr(struct sockaddr *addr1, struct sockaddr *addr2) } /* - * Checks if 'dsaddrs1' contains a subset of 'dsaddrs2'. If it does, - * declare a match. + * Checks if 'dsaddrs1' and 'dsaddrs2' hold the same set of addresses. + * If they do, declare a match. */ static bool _same_data_server_addrs_locked(const struct list_head *dsaddrs1, @@ -587,6 +590,10 @@ _same_data_server_addrs_locked(const struct list_head *dsaddrs1, struct sockaddr *sa1, *sa2; bool match = false; + if (list_count_nodes((struct list_head *)dsaddrs1) != + list_count_nodes((struct list_head *)dsaddrs2)) + return false; + list_for_each_entry(da1, dsaddrs1, da_node) { sa1 = (struct sockaddr *)&da1->da_addr; match = false; @@ -602,16 +609,58 @@ _same_data_server_addrs_locked(const struct list_head *dsaddrs1, return match; } +/* Hash family, address bytes, and port - as same_sockaddr() */ +static u32 +nfs4_ds_addr_hash(const struct sockaddr *sa) +{ + u32 h = sa->sa_family; + + switch (sa->sa_family) { + case AF_INET: { + const struct sockaddr_in *a = (const struct sockaddr_in *)sa; + + h = jhash(&a->sin_addr.s_addr, sizeof(a->sin_addr.s_addr), h); + h = jhash(&a->sin_port, sizeof(a->sin_port), h); + break; + } + case AF_INET6: { + const struct sockaddr_in6 *a = (const struct sockaddr_in6 *)sa; + + h = jhash(&a->sin6_addr, sizeof(a->sin6_addr), h); + h = jhash(&a->sin6_port, sizeof(a->sin6_port), h); + break; + } + } + return h; +} + /* - * Lookup DS by addresses and NFS version. nfs4_ds_cache_lock is held + * Bucket index for a DS cache key. Per-address hashes combine by + * addition so the multipath list order cannot change the bucket, + * matching the order-independent set comparison above. + */ +static u32 +nfs4_ds_cache_hash(const struct list_head *dsaddrs, u32 version) +{ + const struct nfs4_pnfs_ds_addr *da; + u32 h = 0; + + list_for_each_entry(da, dsaddrs, da_node) + h += nfs4_ds_addr_hash((const struct sockaddr *)&da->da_addr); + return hash_32(jhash_1word(version, h), NFS4_DS_CACHE_HASH_BITS); +} + +/* + * Lookup DS by addresses and NFS version. nfs4_data_server_lock is held */ static struct nfs4_pnfs_ds * _data_server_lookup_locked(const struct nfs_net *nn, const struct list_head *dsaddrs, u32 version) { struct nfs4_pnfs_ds *ds; + u32 bucket = nfs4_ds_cache_hash(dsaddrs, version); - list_for_each_entry(ds, &nn->nfs4_data_server_cache, ds_node) + hlist_for_each_entry(ds, &nn->nfs4_data_server_cache[bucket], ds_node) if (ds->ds_version == version && _same_data_server_addrs_locked(&ds->ds_addrs, dsaddrs)) return ds; @@ -633,23 +682,28 @@ static void nfs4_pnfs_ds_addr_free(struct nfs4_pnfs_ds_addr *da) kfree(da); } -static void destroy_ds(struct nfs4_pnfs_ds *ds) +void nfs4_pnfs_ds_addr_list_free(struct list_head *dsaddrs) { struct nfs4_pnfs_ds_addr *da; + while (!list_empty(dsaddrs)) { + da = list_first_entry(dsaddrs, struct nfs4_pnfs_ds_addr, + da_node); + list_del_init(&da->da_node); + nfs4_pnfs_ds_addr_free(da); + } +} +EXPORT_SYMBOL_GPL(nfs4_pnfs_ds_addr_list_free); + +static void destroy_ds(struct nfs4_pnfs_ds *ds) +{ dprintk("--> %s\n", __func__); ifdebug(FACILITY) print_ds(ds); nfs_put_client(ds->ds_clp); - while (!list_empty(&ds->ds_addrs)) { - da = list_first_entry(&ds->ds_addrs, - struct nfs4_pnfs_ds_addr, - da_node); - list_del_init(&da->da_node); - nfs4_pnfs_ds_addr_free(da); - } + nfs4_pnfs_ds_addr_list_free(&ds->ds_addrs); kfree(ds->ds_remotestr); kfree(ds); @@ -660,7 +714,7 @@ void nfs4_pnfs_ds_put(struct nfs4_pnfs_ds *ds) struct nfs_net *nn = net_generic(ds->ds_net, nfs_net_id); if (refcount_dec_and_lock(&ds->ds_count, &nn->nfs4_data_server_lock)) { - list_del_init(&ds->ds_node); + hlist_del_init(&ds->ds_node); spin_unlock(&nn->nfs4_data_server_lock); destroy_ds(ds); } @@ -726,6 +780,7 @@ nfs4_pnfs_ds_add(const struct net *net, struct list_head *dsaddrs, u32 version, { struct nfs_net *nn = net_generic(net, nfs_net_id); struct nfs4_pnfs_ds *tmp_ds, *ds = NULL; + struct hlist_head *bucket; char *remotestr; if (list_empty(dsaddrs)) { @@ -739,6 +794,8 @@ nfs4_pnfs_ds_add(const struct net *net, struct list_head *dsaddrs, u32 version, /* this is only used for debugging, so it's ok if its NULL */ remotestr = nfs4_pnfs_remotestr(dsaddrs, gfp_flags); + /* @dsaddrs is empty after the splice below. */ + bucket = &nn->nfs4_data_server_cache[nfs4_ds_cache_hash(dsaddrs, version)]; spin_lock(&nn->nfs4_data_server_lock); tmp_ds = _data_server_lookup_locked(nn, dsaddrs, version); @@ -747,11 +804,11 @@ nfs4_pnfs_ds_add(const struct net *net, struct list_head *dsaddrs, u32 version, list_splice_init(dsaddrs, &ds->ds_addrs); ds->ds_remotestr = remotestr; refcount_set(&ds->ds_count, 1); - INIT_LIST_HEAD(&ds->ds_node); + INIT_HLIST_NODE(&ds->ds_node); ds->ds_net = net; ds->ds_clp = NULL; ds->ds_version = version; - list_add(&ds->ds_node, &nn->nfs4_data_server_cache); + hlist_add_head(&ds->ds_node, bucket); dprintk("%s add new data server %s\n", __func__, ds->ds_remotestr); } else { @@ -787,7 +844,8 @@ static struct nfs_client *(*get_v3_ds_connect)( int ds_addrlen, int ds_proto, unsigned int ds_timeo, - unsigned int ds_retrans); + unsigned int ds_retrans, + unsigned int ds_nconnect); static bool load_v3_ds_connect(void) { @@ -810,7 +868,8 @@ void nfs4_pnfs_v3_ds_connect_unload(void) static int _nfs4_pnfs_v3_ds_connect(struct nfs_server *mds_srv, struct nfs4_pnfs_ds *ds, unsigned int timeo, - unsigned int retrans) + unsigned int retrans, + unsigned int nconnect) { struct nfs_client *clp = ERR_PTR(-EIO); struct nfs_client *mds_clp = mds_srv->nfs_client; @@ -862,7 +921,7 @@ static int _nfs4_pnfs_v3_ds_connect(struct nfs_server *mds_srv, ds_proto = XPRT_TRANSPORT_TCP_TLS; clp = get_v3_ds_connect(mds_srv, &da->da_addr, da->da_addrlen, - ds_proto, timeo, retrans); + ds_proto, timeo, retrans, nconnect); if (IS_ERR(clp)) continue; clp->cl_rpcclient->cl_softerr = 0; @@ -885,6 +944,7 @@ static int _nfs4_pnfs_v4_ds_connect(struct nfs_server *mds_srv, struct nfs4_pnfs_ds *ds, unsigned int timeo, unsigned int retrans, + unsigned int nconnect, u32 minor_version, bool tightly_coupled) { @@ -976,7 +1036,8 @@ static int _nfs4_pnfs_v4_ds_connect(struct nfs_server *mds_srv, clp = nfs4_set_ds_client(mds_srv, &da->da_addr, da->da_addrlen, ds_proto, - timeo, retrans, minor_version, + timeo, retrans, nconnect, + minor_version, tightly_coupled); if (IS_ERR(clp)) continue; @@ -1011,7 +1072,8 @@ out: */ int nfs4_pnfs_ds_connect(struct nfs_server *mds_srv, struct nfs4_pnfs_ds *ds, struct nfs4_deviceid_node *devid, unsigned int timeo, - unsigned int retrans, u32 version, u32 minor_version, + unsigned int retrans, unsigned int nconnect, + u32 version, u32 minor_version, bool tightly_coupled) { int err; @@ -1031,11 +1093,13 @@ int nfs4_pnfs_ds_connect(struct nfs_server *mds_srv, struct nfs4_pnfs_ds *ds, switch (version) { case 3: - err = _nfs4_pnfs_v3_ds_connect(mds_srv, ds, timeo, retrans); + err = _nfs4_pnfs_v3_ds_connect(mds_srv, ds, timeo, retrans, + nconnect); break; case 4: err = _nfs4_pnfs_v4_ds_connect(mds_srv, ds, timeo, retrans, - minor_version, tightly_coupled); + nconnect, minor_version, + tightly_coupled); break; default: dprintk("%s: unsupported DS version %d\n", __func__, version); diff --git a/fs/nfs/read.c b/fs/nfs/read.c index 2b70bd2b934b..e7497b029d6c 100644 --- a/fs/nfs/read.c +++ b/fs/nfs/read.c @@ -324,7 +324,7 @@ int nfs_read_add_folio(struct nfs_pageio_descriptor *pgio, aligned_len = min_t(unsigned int, ALIGN(len, rsize), fsize); - new = nfs_page_create_from_folio(ctx, folio, 0, aligned_len); + new = nfs_page_create_from_folio(ctx, folio, false, 0, aligned_len); if (IS_ERR(new)) { error = PTR_ERR(new); if (nfs_netfs_folio_unlock(folio)) diff --git a/fs/nfs/super.c b/fs/nfs/super.c index cb19f1540d98..23292680adbd 100644 --- a/fs/nfs/super.c +++ b/fs/nfs/super.c @@ -58,7 +58,6 @@ #include <linux/rcupdate.h> #include <linux/uaccess.h> -#include <linux/nfs_ssc.h> #include <uapi/linux/tls.h> @@ -92,12 +91,6 @@ const struct super_operations nfs_sops = { }; EXPORT_SYMBOL_GPL(nfs_sops); -#ifdef CONFIG_NFS_V4_2 -static const struct nfs_ssc_client_ops nfs_ssc_clnt_ops_tbl = { - .sco_sb_deactive = nfs_sb_deactive, -}; -#endif - #if IS_ENABLED(CONFIG_NFS_V4) static int __init register_nfs4_fs(void) { @@ -119,18 +112,6 @@ static void unregister_nfs4_fs(void) } #endif -#ifdef CONFIG_NFS_V4_2 -static void nfs_ssc_register_ops(void) -{ - nfs_ssc_register(&nfs_ssc_clnt_ops_tbl); -} - -static void nfs_ssc_unregister_ops(void) -{ - nfs_ssc_unregister(&nfs_ssc_clnt_ops_tbl); -} -#endif /* CONFIG_NFS_V4_2 */ - static struct shrinker *acl_shrinker; /* @@ -163,9 +144,6 @@ int __init register_nfs_fs(void) shrinker_register(acl_shrinker); -#ifdef CONFIG_NFS_V4_2 - nfs_ssc_register_ops(); -#endif return 0; error_3: nfs_unregister_sysctl(); @@ -185,9 +163,6 @@ void __exit unregister_nfs_fs(void) shrinker_free(acl_shrinker); nfs_unregister_sysctl(); unregister_nfs4_fs(); -#ifdef CONFIG_NFS_V4_2 - nfs_ssc_unregister_ops(); -#endif unregister_filesystem(&nfs_fs_type); } diff --git a/fs/nfs/unlink.c b/fs/nfs/unlink.c index b57cfaa4d516..c8d712204e64 100644 --- a/fs/nfs/unlink.c +++ b/fs/nfs/unlink.c @@ -67,6 +67,7 @@ static void nfs_async_unlink_release(void *calldata) struct super_block *sb = dentry->d_sb; up_read_non_owner(&NFS_I(d_inode(dentry->d_parent))->rmdir_sem); + d_lookup_acquire(dentry); d_lookup_done(dentry); nfs_free_unlinkdata(data); dput(dentry); @@ -159,6 +160,8 @@ static int nfs_call_unlink(struct dentry *dentry, struct inode *inode, struct nf return ret; } data->dentry = alias; + d_lookup_release(alias); + nfs_do_call_unlink(inode, data); return 1; } diff --git a/fs/nfs/write.c b/fs/nfs/write.c index b6967b528669..0b42d31f7a75 100644 --- a/fs/nfs/write.c +++ b/fs/nfs/write.c @@ -1087,7 +1087,7 @@ static struct nfs_page *nfs_setup_write_request(struct nfs_open_context *ctx, req = nfs_try_to_update_request(folio, offset, bytes); if (req != NULL) goto out; - req = nfs_page_create_from_folio(ctx, folio, offset, bytes); + req = nfs_page_create_from_folio(ctx, folio, false, offset, bytes); if (IS_ERR(req)) goto out; nfs_inode_add_request(req); diff --git a/fs/nfs_common/nfs_ssc.c b/fs/nfs_common/nfs_ssc.c index 832246b22c51..e521e3c836fe 100644 --- a/fs/nfs_common/nfs_ssc.c +++ b/fs/nfs_common/nfs_ssc.c @@ -10,82 +10,112 @@ #include <linux/module.h> #include <linux/fs.h> #include <linux/nfs_ssc.h> -#include "../nfs/nfs4_fs.h" +#include <linux/nfsd_ssc.h> +struct nfs_ssc_client_ops_tbl { + const struct nfs4_ssc_client_ops __rcu *ssc_nfs4_ops; +}; -struct nfs_ssc_client_ops_tbl nfs_ssc_client_tbl; -EXPORT_SYMBOL_GPL(nfs_ssc_client_tbl); +static struct nfs_ssc_client_ops_tbl nfs_ssc_client_tbl __read_mostly; -#ifdef CONFIG_NFS_V4_2 /** - * nfs42_ssc_register - install the NFS_V4 client ops in the nfs_ssc_client_tbl - * @ops: NFS_V4 ops to be installed + * nfsd42_ssc_open - Open a file to be used for server-to-server copy + * @ss_mnt: active mount point on which the source file resides + * @src_fh: file handle of the source file to be copied + * @stateid: stateid to use for COPY operation * - * Return values: - * None + * Caller must close the returned file using nfsd42_ssc_close(). + * + * Return: an open file, or an ERR_PTR on error */ -void nfs42_ssc_register(const struct nfs4_ssc_client_ops *ops) +struct file *nfsd42_ssc_open(struct vfsmount *ss_mnt, struct nfs_fh *src_fh, + nfs4_stateid *stateid) { - nfs_ssc_client_tbl.ssc_nfs4_ops = ops; + /* + * Built under CONFIG_NFS_V4_2_SSC_HELPER, which the NFS client + * enables on its own. The dispatch below is live only when the + * server also sets CONFIG_NFSD_V4_2_INTER_SSC; without it the + * source file cannot be opened, so callers get -EIO. + */ +#if IS_ENABLED(CONFIG_NFSD_V4_2_INTER_SSC) + const struct nfs4_ssc_client_ops *ops; + struct file *res; + + /* + * sco_open() sleeps and must not run inside an RCU read-side + * section. Pin the provider module so the open runs with the + * module held; try_module_get() fails once unregister begins, + * and the copy then gets -EIO. + */ + rcu_read_lock(); + ops = rcu_dereference(nfs_ssc_client_tbl.ssc_nfs4_ops); + if (ops && try_module_get(ops->owner)) { + rcu_read_unlock(); + res = ops->sco_open(ss_mnt, src_fh, stateid); + module_put(ops->owner); + return res; + } + rcu_read_unlock(); +#endif + + return ERR_PTR(-EIO); } -EXPORT_SYMBOL_GPL(nfs42_ssc_register); +EXPORT_SYMBOL_GPL(nfsd42_ssc_open); /** - * nfs42_ssc_unregister - uninstall the NFS_V4 client ops from - * the nfs_ssc_client_tbl - * @ops: ops to be uninstalled + * nfsd42_ssc_close - Close a file opened with nfsd42_ssc_open() + * @filp: struct file to be closed * - * Return values: - * None + * The real cleanup happens unconditionally in nfsd4_cleanup_inter_ssc(). + * The client ops table is read under RCU; nfs42_ssc_unregister() calls + * synchronize_rcu() so unregistration cannot complete while a close is + * in flight. */ -void nfs42_ssc_unregister(const struct nfs4_ssc_client_ops *ops) +void nfsd42_ssc_close(struct file *filp) { - if (nfs_ssc_client_tbl.ssc_nfs4_ops != ops) - return; + /* Live only under CONFIG_NFSD_V4_2_INTER_SSC; see nfsd42_ssc_open(). */ +#if IS_ENABLED(CONFIG_NFSD_V4_2_INTER_SSC) + const struct nfs4_ssc_client_ops *ops; - nfs_ssc_client_tbl.ssc_nfs4_ops = NULL; + rcu_read_lock(); + ops = rcu_dereference(nfs_ssc_client_tbl.ssc_nfs4_ops); + if (ops) + ops->sco_close(filp); + rcu_read_unlock(); +#endif } -EXPORT_SYMBOL_GPL(nfs42_ssc_unregister); -#endif /* CONFIG_NFS_V4_2 */ +EXPORT_SYMBOL_GPL(nfsd42_ssc_close); #ifdef CONFIG_NFS_V4_2 /** - * nfs_ssc_register - install the NFS_FS client ops in the nfs_ssc_client_tbl - * @ops: NFS_FS ops to be installed + * nfs42_ssc_register - install the NFS_V4 client ops in the nfs_ssc_client_tbl + * @ops: NFS_V4 ops to be installed * * Return values: * None */ -void nfs_ssc_register(const struct nfs_ssc_client_ops *ops) +void nfs42_ssc_register(const struct nfs4_ssc_client_ops *ops) { - nfs_ssc_client_tbl.ssc_nfs_ops = ops; + rcu_assign_pointer(nfs_ssc_client_tbl.ssc_nfs4_ops, ops); } -EXPORT_SYMBOL_GPL(nfs_ssc_register); +EXPORT_SYMBOL_GPL(nfs42_ssc_register); /** - * nfs_ssc_unregister - uninstall the NFS_FS client ops from + * nfs42_ssc_unregister - uninstall the NFS_V4 client ops from * the nfs_ssc_client_tbl * @ops: ops to be uninstalled * * Return values: * None */ -void nfs_ssc_unregister(const struct nfs_ssc_client_ops *ops) +void nfs42_ssc_unregister(const struct nfs4_ssc_client_ops *ops) { - if (nfs_ssc_client_tbl.ssc_nfs_ops != ops) + if (rcu_dereference_protected(nfs_ssc_client_tbl.ssc_nfs4_ops, + true) != ops) return; - nfs_ssc_client_tbl.ssc_nfs_ops = NULL; -} -EXPORT_SYMBOL_GPL(nfs_ssc_unregister); -#else -void nfs_ssc_register(const struct nfs_ssc_client_ops *ops) -{ + rcu_assign_pointer(nfs_ssc_client_tbl.ssc_nfs4_ops, NULL); + synchronize_rcu(); } -EXPORT_SYMBOL_GPL(nfs_ssc_register); - -void nfs_ssc_unregister(const struct nfs_ssc_client_ops *ops) -{ -} -EXPORT_SYMBOL_GPL(nfs_ssc_unregister); +EXPORT_SYMBOL_GPL(nfs42_ssc_unregister); #endif /* CONFIG_NFS_V4_2 */ diff --git a/fs/nfsd/blocklayout.c b/fs/nfsd/blocklayout.c index 5be7721c22c2..df02cf746479 100644 --- a/fs/nfsd/blocklayout.c +++ b/fs/nfsd/blocklayout.c @@ -9,6 +9,7 @@ #include <linux/nfsd/debug.h> +#include "nfserr.h" #include "blocklayoutxdr.h" #include "pnfs.h" #include "filecache.h" diff --git a/fs/nfsd/blocklayoutxdr.c b/fs/nfsd/blocklayoutxdr.c index f80dbc41fd5f..a6589f5c878a 100644 --- a/fs/nfsd/blocklayoutxdr.c +++ b/fs/nfsd/blocklayoutxdr.c @@ -8,11 +8,22 @@ #include <linux/nfs4.h> #include "nfsd.h" +#include "nfserr.h" #include "blocklayoutxdr.h" #include "vfs.h" #define NFSDDBG_FACILITY NFSDDBG_PNFS +static __be32 +nfsd4_decode_deviceid4(struct xdr_stream *xdr, struct nfsd4_deviceid *devid) +{ + __be32 *p = xdr_inline_decode(xdr, NFS4_DEVICEID4_SIZE); + + if (unlikely(!p)) + return nfserr_bad_xdr; + svcxdr_decode_deviceid4(p, devid); + return nfs_ok; +} /** * nfsd4_block_encode_layoutget - encode block/scsi layout extent array diff --git a/fs/nfsd/export.c b/fs/nfsd/export.c index b6e0c543e028..e5a0f1ababe6 100644 --- a/fs/nfsd/export.c +++ b/fs/nfsd/export.c @@ -21,6 +21,8 @@ #include <uapi/linux/nfsd_netlink.h> #include "nfsd.h" +#include "nfserr.h" +#include "nfs4ctl.h" #include "nfsfh.h" #include "netns.h" #include "pnfs.h" @@ -1890,21 +1892,19 @@ __be32 check_security_flavor(struct svc_export *exp, struct svc_rqst *rqstp, * check_nfsd_access - check if access to export is allowed. * @exp: svc_export that is being accessed. * @rqstp: svc_rqst attempting to access @exp. - * @may_bypass_gss: reduce strictness of authorization check * * Return values: * %nfs_ok if access is granted, or * %nfserr_wrongsec if access is denied */ -__be32 check_nfsd_access(struct svc_export *exp, struct svc_rqst *rqstp, - bool may_bypass_gss) +__be32 check_nfsd_access(struct svc_export *exp, struct svc_rqst *rqstp) { __be32 status; status = check_xprtsec_policy(exp, rqstp); if (status != nfs_ok) return status; - return check_security_flavor(exp, rqstp, may_bypass_gss); + return check_security_flavor(exp, rqstp, false); } /* diff --git a/fs/nfsd/export.h b/fs/nfsd/export.h index d2b09cd76145..117fb28db1e0 100644 --- a/fs/nfsd/export.h +++ b/fs/nfsd/export.h @@ -104,8 +104,7 @@ int nfsexp_flags(struct svc_cred *cred, struct svc_export *exp); __be32 check_xprtsec_policy(struct svc_export *exp, struct svc_rqst *rqstp); __be32 check_security_flavor(struct svc_export *exp, struct svc_rqst *rqstp, bool may_bypass_gss); -__be32 check_nfsd_access(struct svc_export *exp, struct svc_rqst *rqstp, - bool may_bypass_gss); +__be32 check_nfsd_access(struct svc_export *exp, struct svc_rqst *rqstp); /* * Function declarations diff --git a/fs/nfsd/filecache.c b/fs/nfsd/filecache.c index b9548eb17c77..3539149cc75f 100644 --- a/fs/nfsd/filecache.c +++ b/fs/nfsd/filecache.c @@ -43,6 +43,7 @@ #include "vfs.h" #include "nfsd.h" +#include "nfserr.h" #include "nfsfh.h" #include "netns.h" #include "filecache.h" diff --git a/fs/nfsd/flexfilelayout.c b/fs/nfsd/flexfilelayout.c index 6d531285ab43..0deb913493a3 100644 --- a/fs/nfsd/flexfilelayout.c +++ b/fs/nfsd/flexfilelayout.c @@ -13,7 +13,9 @@ #include <linux/sunrpc/addr.h> +#include "nfserr.h" #include "flexfilelayoutxdr.h" +#include "auth.h" #include "pnfs.h" #include "vfs.h" @@ -23,10 +25,10 @@ static __be32 nfsd4_ff_proc_layoutget(struct svc_rqst *rqstp, struct inode *inode, const struct svc_fh *fhp, struct nfsd4_layoutget *args) { + struct user_namespace *userns = nfsd_user_namespace(rqstp); struct nfsd4_layout_seg *seg = &args->lg_seg; u32 device_generation = 0; int error; - uid_t u; struct pnfs_ff_layout *fl; @@ -49,20 +51,22 @@ nfsd4_ff_proc_layoutget(struct svc_rqst *rqstp, struct inode *inode, fl->flags = FF_FLAGS_NO_LAYOUTCOMMIT | FF_FLAGS_NO_IO_THRU_MDS | FF_FLAGS_NO_READ_IO; - /* Do not allow a IOMODE_READ segment to have write pemissions */ - if (seg->iomode == IOMODE_READ) { - u = from_kuid(&init_user_ns, inode->i_uid) + 1; - fl->uid = make_kuid(&init_user_ns, u); - } else - fl->uid = inode->i_uid; - fl->gid = inode->i_gid; + fl->uid = from_kuid_munged(userns, inode->i_uid); + fl->gid = from_kgid_munged(userns, inode->i_gid); + + /* + * Do not allow an IOMODE_READ segment to have write permissions. + * The group is left intact so group-readable files stay readable; + * nfsd_setuser() squashes an unmapped uid to the export's anon ID. + */ + if (seg->iomode == IOMODE_READ) + fl->uid++; error = nfsd4_set_deviceid(&fl->deviceid, fhp, device_generation); if (error) goto out_error; - fl->fh.size = fhp->fh_handle.fh_size; - memcpy(fl->fh.data, &fhp->fh_handle.fh_raw, fl->fh.size); + fh_copy_shallow(&fl->fh, &fhp->fh_handle); /* Give whole file layout segments */ seg->offset = 0; diff --git a/fs/nfsd/flexfilelayoutxdr.c b/fs/nfsd/flexfilelayoutxdr.c index 374e52d3064a..e297100a2ac3 100644 --- a/fs/nfsd/flexfilelayoutxdr.c +++ b/fs/nfsd/flexfilelayoutxdr.c @@ -6,6 +6,7 @@ #include <linux/nfs4.h> #include "nfsd.h" +#include "nfserr.h" #include "flexfilelayoutxdr.h" #define NFSDDBG_FACILITY NFSDDBG_PNFS @@ -30,10 +31,10 @@ nfsd4_ff_encode_layoutget(struct xdr_stream *xdr, struct ff_idmap uid; struct ff_idmap gid; - fh_len = 4 + xdr_align_size(fl->fh.size); + fh_len = 4 + xdr_align_size(fl->fh.fh_size); - uid.len = sprintf(uid.buf, "%u", from_kuid(&init_user_ns, fl->uid)); - gid.len = sprintf(gid.buf, "%u", from_kgid(&init_user_ns, fl->gid)); + uid.len = sprintf(uid.buf, "%u", fl->uid); + gid.len = sprintf(gid.buf, "%u", fl->gid); /* data server entry: deviceid + efficiency + stateid + fh list + * user + group + flags + stats_collect_hint @@ -68,7 +69,7 @@ nfsd4_ff_encode_layoutget(struct xdr_stream *xdr, sizeof(stateid_opaque_t)); *p++ = cpu_to_be32(1); /* single file handle */ - p = xdr_encode_opaque(p, fl->fh.data, fl->fh.size); + p = xdr_encode_opaque(p, fl->fh.fh_raw, fl->fh.fh_size); p = xdr_encode_opaque(p, uid.buf, uid.len); p = xdr_encode_opaque(p, gid.buf, gid.len); diff --git a/fs/nfsd/flexfilelayoutxdr.h b/fs/nfsd/flexfilelayoutxdr.h index 6d5a1066a903..f7d1dd0708ec 100644 --- a/fs/nfsd/flexfilelayoutxdr.h +++ b/fs/nfsd/flexfilelayoutxdr.h @@ -6,6 +6,7 @@ #define _NFSD_FLEXFILELAYOUTXDR_H 1 #include <linux/inet.h> +#include "nfsfh.h" #include "xdr4.h" #define FF_FLAGS_NO_LAYOUTCOMMIT 1 @@ -35,11 +36,12 @@ struct pnfs_ff_device_addr { struct pnfs_ff_layout { u32 flags; u32 stats_collect_hint; - kuid_t uid; - kgid_t gid; + /* Values to encode; nfsd4_ff_proc_layoutget() has mapped these */ + u32 uid; + u32 gid; struct nfsd4_deviceid deviceid; stateid_t stateid; - struct nfs_fh fh; + struct knfsd_fh fh; }; __be32 nfsd4_ff_encode_getdeviceinfo(struct xdr_stream *xdr, diff --git a/fs/nfsd/localio.c b/fs/nfsd/localio.c index c458c01e9478..33b56d1b3f44 100644 --- a/fs/nfsd/localio.c +++ b/fs/nfsd/localio.c @@ -11,11 +11,9 @@ #include <linux/exportfs.h> #include <linux/sunrpc/svcauth.h> #include <linux/sunrpc/clnt.h> -#include <linux/nfs.h> #include <linux/nfs_common.h> +#include <linux/nfs_fh.h> #include <linux/nfslocalio.h> -#include <linux/nfs_fs.h> -#include <linux/nfs_xdr.h> #include <linux/string.h> #include "nfsd.h" @@ -55,7 +53,7 @@ nfsd_open_local_fh(struct net *net, struct auth_domain *dom, struct nfsd_file *localio; __be32 beres; - if (nfs_fh->size > NFS4_FHSIZE) + if (nfs_fh->size > NFS_MAXFHSIZE) return ERR_PTR(-EINVAL); if (!nfsd_net_try_get(net)) @@ -68,7 +66,7 @@ nfsd_open_local_fh(struct net *net, struct auth_domain *dom, return localio; /* nfs_fh -> svc_fh */ - fh_init(&fh, NFS4_FHSIZE); + fh_init(&fh, NFSD_FHSIZE_UNSPEC); fh.fh_handle.fh_size = nfs_fh->size; memcpy(fh.fh_handle.fh_raw, nfs_fh->data, nfs_fh->size); @@ -179,7 +177,7 @@ static bool localio_decode_uuidarg(struct svc_rqst *rqstp, struct localio_uuidarg *argp = rqstp->rq_argp; u8 uuid[UUID_SIZE]; - if (decode_opaque_fixed(xdr, uuid, UUID_SIZE)) + if (xdr_stream_decode_opaque_fixed(xdr, uuid, UUID_SIZE) < 0) return false; import_uuid(&argp->uuid, uuid); diff --git a/fs/nfsd/lockd.c b/fs/nfsd/lockd.c index 72a5b499839d..f24e45dc37a0 100644 --- a/fs/nfsd/lockd.c +++ b/fs/nfsd/lockd.c @@ -9,7 +9,9 @@ #include <linux/file.h> #include <linux/lockd/bind.h> +#include <linux/nfs_fh.h> #include "nfsd.h" +#include "nfserr.h" #include "vfs.h" #define NFSDDBG_FACILITY NFSDDBG_LOCKD @@ -33,8 +35,7 @@ static int nlm_fopen(struct svc_rqst *rqstp, struct nfs_fh *f, int access; struct svc_fh fh; - /* must initialize before using! but maxsize doesn't matter */ - fh_init(&fh,0); + fh_init(&fh, NFSD_FHSIZE_UNSPEC); fh.fh_handle.fh_size = f->size; memcpy(&fh.fh_handle.fh_raw, f->data, f->size); fh.fh_export = NULL; diff --git a/fs/nfsd/netns.h b/fs/nfsd/netns.h index 71eebfea020d..0ce7da20aba3 100644 --- a/fs/nfsd/netns.h +++ b/fs/nfsd/netns.h @@ -238,8 +238,21 @@ struct nfsd_net { int nfs4_max_clients; atomic_t nfsd_courtesy_clients; - struct shrinker *nfsd_client_shrinker; - struct work_struct nfsd_shrinker_work; + /* per-namespace; num_delegations in nfs4state.c is host-wide */ + atomic_long_t nfsd_delegations; + struct shrinker *nfsd_courtesy_shrinker; + struct shrinker *nfsd_deleg_shrinker; + struct work_struct nfsd_courtesy_work; + struct work_struct nfsd_deleg_work; + + /* courtesy scan requests the reaper has not retired yet */ + atomic_long_t nfsd_shrink_backlog; + + /* delegation scan requests the reaper has not retired yet */ + atomic_long_t nfsd_deleg_backlog; + + /* when deleg_reaper() last swept the client list */ + time64_t nfsd_last_recall_any; /* last time an admin-revoke happened for NFSv4.0 */ time64_t nfs40_last_revoke; diff --git a/fs/nfsd/nfs2acl.c b/fs/nfsd/nfs2acl.c index 190f5a001900..0a5c444fef99 100644 --- a/fs/nfsd/nfs2acl.c +++ b/fs/nfsd/nfs2acl.c @@ -6,9 +6,11 @@ */ #include "nfsd.h" +#include "nfserr.h" /* FIXME: nfsacl.h is a broken header */ #include <linux/nfsacl.h> #include <linux/gfp.h> +#include <linux/nfs3.h> #include "cache.h" #include "xdr3.h" #include "vfs.h" @@ -16,6 +18,48 @@ #define NFSDDBG_FACILITY NFSDDBG_PROC /* + * These maps are identical to the NFSv3 maps (nfs3proc.c). This enables + * the behavior of the two versions to diverge if needed. + */ +static const struct nfsd_access_map nfsd2_regaccess[] = { + { NFS3_ACCESS_READ, NFSD_MAY_READ }, + { NFS3_ACCESS_EXECUTE, NFSD_MAY_EXEC }, + { NFS3_ACCESS_MODIFY, NFSD_MAY_WRITE|NFSD_MAY_TRUNC }, + { NFS3_ACCESS_EXTEND, NFSD_MAY_WRITE }, + { 0, 0 } +}; + +static const struct nfsd_access_map nfsd2_diraccess[] = { + { NFS3_ACCESS_READ, NFSD_MAY_READ }, + { NFS3_ACCESS_LOOKUP, NFSD_MAY_EXEC }, + { NFS3_ACCESS_MODIFY, NFSD_MAY_EXEC|NFSD_MAY_WRITE|NFSD_MAY_TRUNC }, + { NFS3_ACCESS_EXTEND, NFSD_MAY_EXEC|NFSD_MAY_WRITE }, + { NFS3_ACCESS_DELETE, NFSD_MAY_REMOVE }, + { 0, 0 } +}; + +/* + * Some clients - Solaris 2.6 at least, make an access call to the NFS + * server to check for access for things like /dev/null (which really, + * NFSD doesn't care about). So NFSD provides simple access checking + * for those objects, looking mainly at mode bits, ignoring read-only + * filesystem checks. + */ +static const struct nfsd_access_map nfsd2_otheraccess[] = { + { NFS3_ACCESS_READ, NFSD_MAY_READ }, + { NFS3_ACCESS_EXECUTE, NFSD_MAY_EXEC }, + { NFS3_ACCESS_MODIFY, NFSD_MAY_WRITE|NFSD_MAY_LOCAL_ACCESS }, + { NFS3_ACCESS_EXTEND, NFSD_MAY_WRITE|NFSD_MAY_LOCAL_ACCESS }, + { 0, 0 } +}; + +static const struct nfsd_access_maps nfsd2_access_maps = { + .regular = nfsd2_regaccess, + .directory = nfsd2_diraccess, + .other = nfsd2_otheraccess, +}; + +/* * NULL call. */ static __be32 @@ -180,7 +224,9 @@ static __be32 nfsacld_proc_access(struct svc_rqst *rqstp) fh_copy(&resp->fh, &argp->fh); resp->access = argp->access; - resp->status = nfsd_access(rqstp, &resp->fh, &resp->access, NULL); + + resp->status = nfsd_access(rqstp, &resp->fh, &nfsd2_access_maps, + &resp->access, NULL); if (resp->status != nfs_ok) goto out; resp->status = fh_getattr(&resp->fh, &resp->stat); diff --git a/fs/nfsd/nfs3acl.c b/fs/nfsd/nfs3acl.c index 6b6b289db636..7183995182ab 100644 --- a/fs/nfsd/nfs3acl.c +++ b/fs/nfsd/nfs3acl.c @@ -6,6 +6,7 @@ */ #include "nfsd.h" +#include "nfserr.h" /* FIXME: nfsacl.h is a broken header */ #include <linux/nfsacl.h> #include <linux/gfp.h> diff --git a/fs/nfsd/nfs3proc.c b/fs/nfsd/nfs3proc.c index 0904d953d10e..60cd01b6a37d 100644 --- a/fs/nfsd/nfs3proc.c +++ b/fs/nfsd/nfs3proc.c @@ -9,10 +9,12 @@ #include <linux/ext2_fs.h> #include <linux/magic.h> #include <linux/namei.h> +#include <linux/nfs3.h> #include "cache.h" #include "xdr3.h" #include "vfs.h" +#include "nfserr.h" #include "filecache.h" #include "trace.h" @@ -48,6 +50,58 @@ static bool nfsd3_time_in_range(const struct iattr *iap) return true; } +static const struct nfsd_access_map nfsd3_regaccess[] = { + { NFS3_ACCESS_READ, NFSD_MAY_READ }, + { NFS3_ACCESS_EXECUTE, NFSD_MAY_EXEC }, + { NFS3_ACCESS_MODIFY, NFSD_MAY_WRITE|NFSD_MAY_TRUNC }, + { NFS3_ACCESS_EXTEND, NFSD_MAY_WRITE }, + { 0, 0 } +}; + +static const struct nfsd_access_map nfsd3_diraccess[] = { + { NFS3_ACCESS_READ, NFSD_MAY_READ }, + { NFS3_ACCESS_LOOKUP, NFSD_MAY_EXEC }, + { NFS3_ACCESS_MODIFY, NFSD_MAY_EXEC|NFSD_MAY_WRITE|NFSD_MAY_TRUNC }, + { NFS3_ACCESS_EXTEND, NFSD_MAY_EXEC|NFSD_MAY_WRITE }, + { NFS3_ACCESS_DELETE, NFSD_MAY_REMOVE }, + { 0, 0 } +}; + +/* + * Some clients - Solaris 2.6 at least, make an access call to the NFS + * server to check for access for things like /dev/null (which really, + * NFSD doesn't care about). So NFSD provides simple access checking + * for those objects, looking mainly at mode bits, ignoring read-only + * filesystem checks. + */ +static const struct nfsd_access_map nfsd3_otheraccess[] = { + { NFS3_ACCESS_READ, NFSD_MAY_READ }, + { NFS3_ACCESS_EXECUTE, NFSD_MAY_EXEC }, + { NFS3_ACCESS_MODIFY, NFSD_MAY_WRITE|NFSD_MAY_LOCAL_ACCESS }, + { NFS3_ACCESS_EXTEND, NFSD_MAY_WRITE|NFSD_MAY_LOCAL_ACCESS }, + { 0, 0 } +}; + +static const struct nfsd_access_maps nfsd3_access_maps = { + .regular = nfsd3_regaccess, + .directory = nfsd3_diraccess, + .other = nfsd3_otheraccess, +}; + +static int nfsd3_iocb_flags(enum nfs3_stable_how how) +{ + switch (how) { + case NFS_FILE_SYNC: + /* persist data and timestamps */ + return IOCB_DSYNC | IOCB_SYNC; + case NFS_DATA_SYNC: + /* persist data only */ + return IOCB_DSYNC; + default: + return 0; + } +} + static __be32 nfsd3_map_status(__be32 status) { switch (status) { @@ -171,7 +225,8 @@ nfsd3_proc_access(struct svc_rqst *rqstp) fh_copy(&resp->fh, &argp->fh); resp->access = argp->access; - resp->status = nfsd_access(rqstp, &resp->fh, &resp->access, NULL); + resp->status = nfsd_access(rqstp, &resp->fh, &nfsd3_access_maps, + &resp->access, NULL); resp->status = nfsd3_map_status(resp->status); return rpc_success; } @@ -260,7 +315,8 @@ nfsd3_proc_write(struct svc_rqst *rqstp) resp->committed = argp->stable; resp->status = nfsd_write(rqstp, &resp->fh, argp->offset, &argp->payload, &cnt, - resp->committed, resp->verf); + nfsd3_iocb_flags(resp->committed), + resp->verf); resp->count = cnt; resp->status = nfsd3_map_status(resp->status); return rpc_success; @@ -282,6 +338,7 @@ nfsd3_create_file(struct svc_rqst *rqstp, struct svc_fh *fhp, struct nfsd_attrs attrs = { .na_iattr = iap, }; + struct svc_export *exp; __u32 v_mtime, v_atime; struct inode *inode; __be32 status; @@ -320,7 +377,23 @@ nfsd3_create_file(struct svc_rqst *rqstp, struct svc_fh *fhp, goto out; } - status = fh_compose(resfhp, fhp->fh_export, child, fhp); + exp = exp_get(fhp->fh_export); + if (argp->createmode == NFS3_CREATE_UNCHECKED) { + /* + * If name is already in dcache we need to check for mountpoints + */ + if (d_is_reg(child) && + unlikely(nfsd_mountpoint(child, exp))) { + status = nfsd_cross_mnt(rqstp, &child, &exp); + if (status != nfs_ok) { + exp_put(exp); + goto out; + } + } + } + + status = fh_compose(resfhp, exp, child, fhp); + exp_put(exp); if (status != nfs_ok) goto out; diff --git a/fs/nfsd/nfs3xdr.c b/fs/nfsd/nfs3xdr.c index e481804bb120..a14e829e1c41 100644 --- a/fs/nfsd/nfs3xdr.c +++ b/fs/nfsd/nfs3xdr.c @@ -8,11 +8,13 @@ */ #include <linux/namei.h> +#include <linux/nfs3.h> #include <linux/sunrpc/svc_xprt.h> #include "xdr3.h" #include "auth.h" #include "netns.h" #include "vfs.h" +#include "nfserr.h" /* * Force construction of an empty post-op attr @@ -556,6 +558,8 @@ nfs3svc_decode_writeargs(struct svc_rqst *rqstp, struct xdr_stream *xdr) return false; if (xdr_stream_decode_u32(xdr, &args->stable) < 0) return false; + if (args->stable > NFS_FILE_SYNC) + return false; /* opaque data */ if (xdr_stream_decode_u32(xdr, &args->len) < 0) diff --git a/fs/nfsd/nfs4acl.c b/fs/nfsd/nfs4acl.c index 2c2f2fd89e87..94f6ad381ebe 100644 --- a/fs/nfsd/nfs4acl.c +++ b/fs/nfsd/nfs4acl.c @@ -40,6 +40,7 @@ #include "nfsfh.h" #include "nfsd.h" +#include "nfserr.h" #include "acl.h" #include "vfs.h" diff --git a/fs/nfsd/nfs4callback.c b/fs/nfsd/nfs4callback.c index 19dc337502ca..a6b31d3f2bf6 100644 --- a/fs/nfsd/nfs4callback.c +++ b/fs/nfsd/nfs4callback.c @@ -37,6 +37,7 @@ #include <linux/sunrpc/svc_xprt.h> #include <linux/slab.h> #include "nfsd.h" +#include "nfserr.h" #include "state.h" #include "netns.h" #include "stats.h" @@ -1529,12 +1530,14 @@ out: /** * nfsd41_cb_destroy_referring_call_list - release referring call info - * @cb: context of a callback that has completed + * @cb: context of callback to release referring calls from * * Callers who allocate referring calls using nfsd41_cb_referring_call() must * release those resources by calling nfsd41_cb_destroy_referring_call_list. * - * Caller serializes access to @cb. + * Caller serializes access to @cb. No CB_COMPOUND for @cb may be in + * flight, because encode_cb_sequence4args() walks this list as it + * encodes. */ void nfsd41_cb_destroy_referring_call_list(struct nfsd4_callback *cb) { @@ -1556,6 +1559,7 @@ void nfsd41_cb_destroy_referring_call_list(struct nfsd4_callback *cb) list_del(&rcl->__list); kfree(rcl); } + cb->cb_nr_referring_call_list = 0; } static void nfsd4_cb_prepare(struct rpc_task *task, void *calldata) diff --git a/fs/nfsd/nfs4ctl.h b/fs/nfsd/nfs4ctl.h new file mode 100644 index 000000000000..bcec4c4ef1d5 --- /dev/null +++ b/fs/nfsd/nfs4ctl.h @@ -0,0 +1,83 @@ +/* SPDX-License-Identifier: GPL-2.0 */ +/* + * Entry points by which the knfsd core drives the optional NFSv4 + * subsystem: state lifecycle, the laundromat workqueue, the recovery + * directory, junctions, the CLD notifier, and leases-net setup. + * + * Separated from nfsd.h so that the many translation units that + * include nfsd.h but call none of these -- among them the NFSv2 and + * NFSv3 paths -- do not have to parse them. The CONFIG_NFSD_V4=n + * stubs let the version-agnostic callers invoke the routines + * unconditionally. + */ + +#ifndef LINUX_NFSD_NFS4CTL_H +#define LINUX_NFSD_NFS4CTL_H + +#include <linux/stddef.h> +#include <linux/types.h> + +struct net; +struct inode; +struct dentry; +struct svc_rqst; +struct nfsd_net; + +#ifdef CONFIG_NFSD_V4 +extern unsigned long max_delegations; +int nfsd4_init_slabs(void); +void nfsd4_free_slabs(void); +int nfs4_state_start(void); +int nfs4_state_start_net(struct net *net); +void nfs4_state_shutdown(void); +void nfs4_state_shutdown_net(struct net *net); +int nfs4_reset_recoverydir(char *recdir); +char * nfs4_recoverydir(void); +bool nfsd4_spo_must_allow(struct svc_rqst *rqstp); +int nfsd4_create_laundry_wq(void); +void nfsd4_destroy_laundry_wq(void); +bool nfsd_wait_for_delegreturn(struct svc_rqst *rqstp, struct inode *inode); + +extern int nfsd4_is_junction(struct dentry *dentry); +extern int register_cld_notifier(void); +extern void unregister_cld_notifier(void); +#ifdef CONFIG_NFSD_V4_2_INTER_SSC +extern void nfsd4_ssc_init_umount_work(struct nfsd_net *nn); +#endif + +extern void nfsd4_init_leases_net(struct nfsd_net *nn); + +#else /* CONFIG_NFSD_V4 */ +static inline int nfsd4_init_slabs(void) { return 0; } +static inline void nfsd4_free_slabs(void) { } +static inline int nfs4_state_start(void) { return 0; } +static inline int nfs4_state_start_net(struct net *net) { return 0; } +static inline void nfs4_state_shutdown(void) { } +static inline void nfs4_state_shutdown_net(struct net *net) { } +static inline int nfs4_reset_recoverydir(char *recdir) { return 0; } +static inline char * nfs4_recoverydir(void) {return NULL; } +static inline bool nfsd4_spo_must_allow(struct svc_rqst *rqstp) +{ + return false; +} +static inline int nfsd4_create_laundry_wq(void) { return 0; }; +static inline void nfsd4_destroy_laundry_wq(void) {}; +static inline bool nfsd_wait_for_delegreturn(struct svc_rqst *rqstp, + struct inode *inode) +{ + return false; +} + +static inline int nfsd4_is_junction(struct dentry *dentry) +{ + return 0; +} + +static inline void nfsd4_init_leases_net(struct nfsd_net *nn) { }; + +#define register_cld_notifier() 0 +#define unregister_cld_notifier() do { } while(0) + +#endif /* CONFIG_NFSD_V4 */ + +#endif /* LINUX_NFSD_NFS4CTL_H */ diff --git a/fs/nfsd/nfs4idmap.c b/fs/nfsd/nfs4idmap.c index e9faf8b78f74..4e5297593963 100644 --- a/fs/nfsd/nfs4idmap.c +++ b/fs/nfsd/nfs4idmap.c @@ -41,6 +41,7 @@ #include "auth.h" #include "idmap.h" #include "nfsd.h" +#include "nfserr.h" #include "netns.h" #include "vfs.h" diff --git a/fs/nfsd/nfs4layouts.c b/fs/nfsd/nfs4layouts.c index 22bcb6d09f70..4187202f9acc 100644 --- a/fs/nfsd/nfs4layouts.c +++ b/fs/nfsd/nfs4layouts.c @@ -9,6 +9,7 @@ #include <linux/sched.h> #include <linux/sunrpc/addr.h> +#include "nfserr.h" #include "pnfs.h" #include "netns.h" #include "trace.h" diff --git a/fs/nfsd/nfs4proc.c b/fs/nfsd/nfs4proc.c index 50c07561e31f..bb74eef43938 100644 --- a/fs/nfsd/nfs4proc.c +++ b/fs/nfsd/nfs4proc.c @@ -38,19 +38,22 @@ #include <linux/slab.h> #include <linux/kthread.h> #include <linux/namei.h> +#include <linux/pagemap.h> #include <linux/sunrpc/addr.h> -#include <linux/nfs_ssc.h> +#include <linux/nfsd_ssc.h> #include "attr4.h" #include "idmap.h" #include "cache.h" #include "xdr4.h" +#include "nfs4ctl.h" #include "vfs.h" #include "current_stateid.h" #include "netns.h" #include "acl.h" #include "pnfs.h" +#include "nfserr.h" #include "trace.h" static bool inter_copy_offload_enable; @@ -69,6 +72,57 @@ MODULE_PARM_DESC(nfsd4_ssc_umount_timeout, #define NFSDDBG_FACILITY NFSDDBG_PROC +static const struct nfsd_access_map nfsd4_regaccess[] = { + { NFS4_ACCESS_READ, NFSD_MAY_READ }, + { NFS4_ACCESS_EXECUTE, NFSD_MAY_EXEC }, + { NFS4_ACCESS_MODIFY, NFSD_MAY_WRITE|NFSD_MAY_TRUNC }, + { NFS4_ACCESS_EXTEND, NFSD_MAY_WRITE }, + { NFS4_ACCESS_XAREAD, NFSD_MAY_READ }, + { NFS4_ACCESS_XAWRITE, NFSD_MAY_WRITE }, + { NFS4_ACCESS_XALIST, NFSD_MAY_READ }, + { 0, 0 } +}; + +static const struct nfsd_access_map nfsd4_diraccess[] = { + { NFS4_ACCESS_READ, NFSD_MAY_READ }, + { NFS4_ACCESS_LOOKUP, NFSD_MAY_EXEC }, + { NFS4_ACCESS_MODIFY, NFSD_MAY_EXEC|NFSD_MAY_WRITE|NFSD_MAY_TRUNC }, + { NFS4_ACCESS_EXTEND, NFSD_MAY_EXEC|NFSD_MAY_WRITE }, + { NFS4_ACCESS_DELETE, NFSD_MAY_REMOVE }, + { NFS4_ACCESS_XAREAD, NFSD_MAY_READ }, + { NFS4_ACCESS_XAWRITE, NFSD_MAY_WRITE }, + { NFS4_ACCESS_XALIST, NFSD_MAY_READ }, + { 0, 0 } +}; + +static const struct nfsd_access_map nfsd4_otheraccess[] = { + { NFS4_ACCESS_READ, NFSD_MAY_READ }, + { NFS4_ACCESS_EXECUTE, NFSD_MAY_EXEC }, + { NFS4_ACCESS_MODIFY, NFSD_MAY_WRITE|NFSD_MAY_LOCAL_ACCESS }, + { NFS4_ACCESS_EXTEND, NFSD_MAY_WRITE|NFSD_MAY_LOCAL_ACCESS }, + { 0, 0 } +}; + +static const struct nfsd_access_maps nfsd4_access_maps = { + .regular = nfsd4_regaccess, + .directory = nfsd4_diraccess, + .other = nfsd4_otheraccess, +}; + +static int nfsd4_iocb_flags(enum stable_how4 how) +{ + switch (how) { + case FILE_SYNC4: + /* persist data and timestamps */ + return IOCB_DSYNC | IOCB_SYNC; + case DATA_SYNC4: + /* persist data only */ + return IOCB_DSYNC; + default: + return 0; + } +} + static u32 nfsd_attrmask[] = { NFSD_WRITEABLE_ATTRS_WORD0, NFSD_WRITEABLE_ATTRS_WORD1, @@ -169,23 +223,17 @@ do_open_permission(struct svc_rqst *rqstp, struct svc_fh *current_fh, struct nfs return fh_verify(rqstp, current_fh, S_IFREG, accmode); } -static __be32 nfsd_check_obj_isreg(struct svc_fh *fh, u32 minor_version) +static int nfsd_check_obj_isreg(struct dentry *child) { - umode_t mode = d_inode(fh->fh_dentry)->i_mode; + umode_t mode = d_inode(child)->i_mode; if (S_ISREG(mode)) - return nfs_ok; + return 0; if (S_ISDIR(mode)) - return nfserr_isdir; + return -EISDIR; if (S_ISLNK(mode)) - return nfserr_symlink; - - /* RFC 7530 - 16.16.6 */ - if (minor_version == 0) - return nfserr_symlink; - else - return nfserr_wrong_type; - + return -ELOOP; + return -EFTYPE; } static void nfsd4_set_open_owner_reply_cache(struct nfsd4_compound_state *cstate, struct nfsd4_open *open, struct svc_fh *resfh) @@ -202,40 +250,50 @@ static inline bool nfsd4_create_is_exclusive(int createmode) createmode == NFS4_CREATE_EXCLUSIVE4_1; } -static __be32 -nfsd4_vfs_create(struct svc_fh *fhp, struct dentry **child, - struct nfsd4_open *open) +static struct file *do_lookup_open(struct path *parent, + struct qstr *name, + unsigned int oflags, + umode_t mode) { - struct file *filp; + struct file *filp = NULL; struct path path; - int oflags; + struct dentry *child; + int want_write_err = 0; - oflags = O_CREAT | O_LARGEFILE; - if (nfsd4_create_is_exclusive(open->op_createmode)) - oflags |= O_EXCL; + want_write_err = mnt_want_write(parent->mnt); - switch (open->op_share_access & NFS4_SHARE_ACCESS_BOTH) { - case NFS4_SHARE_ACCESS_WRITE: - oflags |= O_WRONLY; - break; - case NFS4_SHARE_ACCESS_BOTH: - oflags |= O_RDWR; - break; - default: - oflags |= O_RDONLY; + child = start_creating(&nop_mnt_idmap, parent->dentry, name); + if (IS_ERR(child)) { + filp = ERR_CAST(child); + goto out; } + path.mnt = parent->mnt; + path.dentry = child; - path.mnt = fhp->fh_export->ex_path.mnt; - path.dentry = *child; - filp = dentry_create(&path, oflags, open->op_iattr.ia_mode, - current_cred()); - *child = path.dentry; - - if (IS_ERR(filp)) - return nfserrno(PTR_ERR(filp)); + if (d_really_is_positive(child)) { + /* + * open the file so that we consistently have a valid + * op_filp and consequently a valid ->f_path.dentry. + */ + int err = nfsd_check_obj_isreg(child); - open->op_filp = filp; - return nfs_ok; + if (err) + filp = ERR_PTR(err); + else + filp = dentry_open(&path, oflags, current_cred()); + } else if (!(oflags & O_CREAT)) { + filp = ERR_PTR(-ENOENT); + } else if (want_write_err) { + filp = ERR_PTR(want_write_err); + } else { + filp = dentry_create(&path, oflags, mode, current_cred()); + child = path.dentry; + } + end_creating(child); +out: + if (!want_write_err) + mnt_drop_write(parent->mnt); + return filp; } /* @@ -254,11 +312,15 @@ nfsd4_create_file(struct svc_rqst *rqstp, struct svc_fh *fhp, .na_iattr = iap, .na_seclabel = &open->op_label, }; - struct dentry *parent, *child = ERR_PTR(-EINVAL); + int oflags = O_CREAT | O_LARGEFILE; + struct dentry *child = ERR_PTR(-EINVAL); + struct path parent = { + .mnt = fhp->fh_export->ex_path.mnt, + .dentry = fhp->fh_dentry, + }; __u32 v_mtime, v_atime; - struct inode *inode; - __be32 status; - int host_err; + __be32 status, create_status; + int want_write_err; if (name_is_dot_dotdot(open->op_fname, open->op_fnamelen)) return nfserr_exist; @@ -268,43 +330,69 @@ nfsd4_create_file(struct svc_rqst *rqstp, struct svc_fh *fhp, status = fh_verify(rqstp, fhp, S_IFDIR, NFSD_MAY_EXEC); if (status != nfs_ok) return status; - parent = fhp->fh_dentry; - inode = d_inode(parent); - - host_err = fh_want_write(fhp); - if (host_err) - return nfserrno(host_err); - if (open->op_acl) { - if (open->op_dpacl || open->op_pacl) { - status = nfserr_inval; - goto out; - } - if (is_create_with_attrs(open)) { - status = nfsd4_acl_to_attr(NF4REG, open->op_acl, - &attrs); - if (status) - goto out; + if (open->op_createmode == NFS4_CREATE_UNCHECKED) { + /* + * If name is already in dcache we need to check for mountpoints + */ + child = try_lookup_noperm(&QSTR_LEN(open->op_fname, + open->op_fnamelen), + parent.dentry); + if (child && !IS_ERR(child) && d_is_reg(child) && + unlikely(nfsd_mountpoint(child, fhp->fh_export))) { + struct svc_export *exp = exp_get(fhp->fh_export); + + status = nfsd_cross_mnt(rqstp, &child, &exp); + if (status == nfs_ok) + status = fh_compose(resfhp, exp, + child, fhp); + fh_fill_post_noop(fhp); + open->op_truncate = + (iap->ia_valid & ATTR_SIZE) && + !iap->ia_size; + dput(child); + exp_put(exp); + return status; } - } else if (is_create_with_attrs(open)) { - /* The dpacl and pacl will get released by nfsd_attrs_free(). */ - attrs.na_dpacl = open->op_dpacl; - attrs.na_pacl = open->op_pacl; - open->op_dpacl = NULL; - open->op_pacl = NULL; + if (!IS_ERR(child)) + dput(child); } - child = start_creating(&nop_mnt_idmap, parent, - &QSTR_LEN(open->op_fname, open->op_fnamelen)); - if (IS_ERR(child)) { - status = nfserrno(PTR_ERR(child)); - goto out; + if (!IS_POSIXACL(d_inode(parent.dentry))) + iap->ia_mode &= ~current_umask(); + + /* + * For the EXCLUSIVE modes we do our own uniqueness tests + * so don't want O_EXCL. + */ + if (open->op_createmode == NFS4_CREATE_GUARDED) + oflags |= O_EXCL; + + switch (open->op_share_access & NFS4_SHARE_ACCESS_BOTH) { + case NFS4_SHARE_ACCESS_WRITE: + oflags |= O_WRONLY; + break; + case NFS4_SHARE_ACCESS_BOTH: + oflags |= O_RDWR; + break; + default: + oflags |= O_RDONLY; } - if (d_really_is_negative(child)) { - status = fh_verify(rqstp, fhp, S_IFDIR, NFSD_MAY_CREATE); - if (status != nfs_ok) - goto out; + if (!is_create_with_attrs(open)) { + /* No attrs to check */ + } else if (open->op_acl) { + if (open->op_dpacl || open->op_pacl) { + /* Cannot specify both NFSv4 and Posix ACLs */ + return nfserr_inval; + } + status = nfsd4_acl_to_attr(NF4REG, open->op_acl, + &attrs); + if (status) + return status; + } else { + attrs.na_dpacl = posix_acl_dup(open->op_dpacl); + attrs.na_pacl = posix_acl_dup(open->op_pacl); } v_mtime = 0; @@ -322,24 +410,53 @@ nfsd4_create_file(struct svc_rqst *rqstp, struct svc_fh *fhp, */ v_mtime = verifier[0] & 0x7fffffff; v_atime = verifier[1] & 0x7fffffff; + + iap->ia_valid |= ATTR_MTIME | ATTR_ATIME | + ATTR_MTIME_SET|ATTR_ATIME_SET; + iap->ia_mtime.tv_sec = v_mtime; + iap->ia_atime.tv_sec = v_atime; + iap->ia_mtime.tv_nsec = 0; + iap->ia_atime.tv_nsec = 0; } - if (d_really_is_positive(child)) { - /* NFSv4 protocol requires change attributes even though - * no change happened. - */ - status = fh_fill_both_attrs(fhp); - if (status != nfs_ok) - goto out; + create_status = fh_verify(rqstp, fhp, S_IFDIR, NFSD_MAY_CREATE); + if (create_status) + /* Might still succeed if no create is needed */ + oflags &= ~O_CREAT; + + open->op_filp = do_lookup_open(&parent, + &QSTR_LEN(open->op_fname, + open->op_fnamelen), + oflags, + open->op_iattr.ia_mode); + if (IS_ERR(open->op_filp)) { + status = nfserrno(PTR_ERR(open->op_filp)); + open->op_filp = NULL; + if (status == nfserr_noent && create_status) + status = create_status; + goto out; + } - status = fh_compose(resfhp, fhp->fh_export, child, fhp); - if (status != nfs_ok) - goto out; + child = open->op_filp->f_path.dentry; + open->op_created = open->op_filp->f_mode & FMODE_CREATED; - switch (open->op_createmode) { - case NFS4_CREATE_UNCHECKED: - if (!d_is_reg(child)) - break; + status = fh_compose(resfhp, fhp->fh_export, child, fhp); + if (status != nfs_ok) + goto out; + + if (!open->op_created && + nfsd4_create_is_exclusive(open->op_createmode) && + inode_get_mtime_sec(d_inode(child)) == v_mtime && + inode_get_atime_sec(d_inode(child)) == v_atime && + d_inode(child)->i_size == 0) + open->op_created = true; + + if (!open->op_created) { + if (open->op_createmode == NFS4_CREATE_UNCHECKED) { + /* NFSv4 protocol requires change attributes + * even though no change happened. + */ + fh_fill_post_noop(fhp); /* * In NFSv4, we don't want to truncate the file @@ -347,63 +464,30 @@ nfsd4_create_file(struct svc_rqst *rqstp, struct svc_fh *fhp, * some other reason. Furthermore, if the size is * nonzero, we should ignore it according to spec! */ - open->op_truncate = (iap->ia_valid & ATTR_SIZE) && - !iap->ia_size; - break; - case NFS4_CREATE_GUARDED: + open->op_truncate = (d_is_reg(child) && + (iap->ia_valid & ATTR_SIZE) && + !iap->ia_size); + } else status = nfserr_exist; - break; - case NFS4_CREATE_EXCLUSIVE: - if (inode_get_mtime_sec(d_inode(child)) == v_mtime && - inode_get_atime_sec(d_inode(child)) == v_atime && - d_inode(child)->i_size == 0) { - open->op_created = true; - break; /* subtle */ - } - status = nfserr_exist; - break; - case NFS4_CREATE_EXCLUSIVE4_1: - if (inode_get_mtime_sec(d_inode(child)) == v_mtime && - inode_get_atime_sec(d_inode(child)) == v_atime && - d_inode(child)->i_size == 0) { - open->op_created = true; - goto set_attr; /* subtle */ - } - status = nfserr_exist; - } goto out; } - - if (!IS_POSIXACL(inode)) - iap->ia_mode &= ~current_umask(); - - status = fh_fill_pre_attrs(fhp); - if (status != nfs_ok) - goto out; - status = nfsd4_vfs_create(fhp, &child, open); - if (status != nfs_ok) - goto out; - open->op_created = true; + /* file was created */ fh_fill_post_attrs(fhp); - status = fh_compose(resfhp, fhp->fh_export, child, fhp); - if (status != nfs_ok) - goto out; - /* A newly created file already has a file size of zero. */ if ((iap->ia_valid & ATTR_SIZE) && (iap->ia_size == 0)) iap->ia_valid &= ~ATTR_SIZE; - if (nfsd4_create_is_exclusive(open->op_createmode)) { - iap->ia_valid = ATTR_MTIME | ATTR_ATIME | - ATTR_MTIME_SET|ATTR_ATIME_SET; - iap->ia_mtime.tv_sec = v_mtime; - iap->ia_atime.tv_sec = v_atime; - iap->ia_mtime.tv_nsec = 0; - iap->ia_atime.tv_nsec = 0; - } -set_attr: - status = nfsd_create_setattr(rqstp, fhp, resfhp, &attrs); + /* We will need write access to set the attrs */ + want_write_err = fh_want_write(fhp); + if (!want_write_err) { + status = nfsd_create_setattr(rqstp, fhp, + resfhp, &attrs); + fh_drop_write(fhp); + } else if (nfsd_attrs_valid(&attrs)) { + /* Needed write access */ + status = nfserrno(want_write_err); + } if (attrs.na_labelerr) open->op_bmval[2] &= ~FATTR4_WORD2_SECURITY_LABEL; @@ -414,9 +498,7 @@ set_attr: if (attrs.na_paclerr) open->op_bmval[2] &= ~FATTR4_WORD2_POSIX_ACCESS_ACL; out: - end_creating(child); nfsd_attrs_free(&attrs); - fh_drop_write(fhp); return status; } @@ -465,6 +547,9 @@ do_open_lookup(struct svc_rqst *rqstp, struct nfsd4_compound_state *cstate, stru fh_init(*resfh, NFS4_FHSIZE); open->op_truncate = false; + status = fh_fill_pre_attrs_unlocked(current_fh); + if (status) + goto out; if (open->op_create) { /* FIXME: check session persistence and pnfs flags. * The nfsv4.1 spec requires the following semantics: @@ -496,15 +581,15 @@ do_open_lookup(struct svc_rqst *rqstp, struct nfsd4_compound_state *cstate, stru } else { status = nfsd_lookup(rqstp, current_fh, open->op_fname, open->op_fnamelen, *resfh); - if (status == nfs_ok) - /* NFSv4 protocol requires change attributes even though - * no change happened. - */ - status = fh_fill_both_attrs(current_fh); + /* + * NFSv4 protocol requires change attributes even though + * no change happened. + */ + fh_fill_post_noop(current_fh); } if (status) goto out; - status = nfsd_check_obj_isreg(*resfh, cstate->minorversion); + status = nfserrno(nfsd_check_obj_isreg((*resfh)->fh_dentry)); if (status) goto out; @@ -516,6 +601,10 @@ do_open_lookup(struct svc_rqst *rqstp, struct nfsd4_compound_state *cstate, stru status = do_open_permission(rqstp, *resfh, open, accmode); set_change_info(&open->op_cinfo, current_fh); out: + if (status == nfserr_wrong_type && cstate->minorversion == 0) + /* RFC 7530 - 16.16.6 */ + return nfserr_symlink; + return status; } @@ -703,13 +792,6 @@ static __be32 nfsd4_open_omfg(struct svc_rqst *rqstp, struct nfsd4_compound_stat return nfsd4_open(rqstp, cstate, &op->u); } -static void -nfsd4_open_release(union nfsd4_op_u *u) -{ - posix_acl_release(u->open.op_dpacl); - posix_acl_release(u->open.op_pacl); -} - /* * filehandle-manipulating ops. */ @@ -788,17 +870,18 @@ nfsd4_access(struct svc_rqst *rqstp, struct nfsd4_compound_state *cstate, struct nfsd4_access *access = &u->access; u32 access_full; - access_full = NFS3_ACCESS_FULL; + access_full = NFS4_ACCESS_READ | NFS4_ACCESS_LOOKUP | + NFS4_ACCESS_MODIFY | NFS4_ACCESS_EXTEND | + NFS4_ACCESS_DELETE | NFS4_ACCESS_EXECUTE; if (cstate->minorversion >= 2) access_full |= NFS4_ACCESS_XALIST | NFS4_ACCESS_XAREAD | NFS4_ACCESS_XAWRITE; if (access->ac_req_access & ~access_full) return nfserr_inval; - access->ac_resp_access = access->ac_req_access; - return nfsd_access(rqstp, &cstate->current_fh, &access->ac_resp_access, - &access->ac_supported); + return nfsd_access(rqstp, &cstate->current_fh, &nfsd4_access_maps, + &access->ac_resp_access, &access->ac_supported); } static __be32 @@ -829,16 +912,13 @@ nfsd4_create(struct svc_rqst *rqstp, struct nfsd4_compound_state *cstate, struct nfsd_attrs attrs = { .na_iattr = &create->cr_iattr, .na_seclabel = &create->cr_label, - .na_dpacl = create->cr_dpacl, - .na_pacl = create->cr_pacl, + .na_dpacl = posix_acl_dup(create->cr_dpacl), + .na_pacl = posix_acl_dup(create->cr_pacl), }; struct svc_fh resfh; __be32 status; dev_t rdev; - create->cr_dpacl = NULL; - create->cr_pacl = NULL; - fh_init(&resfh, NFS4_FHSIZE); status = fh_verify(rqstp, &cstate->current_fh, S_IFDIR, NFSD_MAY_NOP); @@ -1248,8 +1328,8 @@ nfsd4_setattr(struct svc_rqst *rqstp, struct nfsd4_compound_state *cstate, struct nfsd_attrs attrs = { .na_iattr = &setattr->sa_iattr, .na_seclabel = &setattr->sa_label, - .na_pacl = setattr->sa_pacl, - .na_dpacl = setattr->sa_dpacl, + .na_pacl = posix_acl_dup(setattr->sa_pacl), + .na_dpacl = posix_acl_dup(setattr->sa_dpacl), }; bool save_no_wcc, deleg_attrs; struct nfs4_stid *st = NULL; @@ -1257,10 +1337,6 @@ nfsd4_setattr(struct svc_rqst *rqstp, struct nfsd4_compound_state *cstate, __be32 status = nfs_ok; int err; - /* Transfer ownership to attrs for cleanup via nfsd_attrs_free() */ - setattr->sa_pacl = NULL; - setattr->sa_dpacl = NULL; - deleg_attrs = setattr->sa_bmval[2] & (FATTR4_WORD2_TIME_DELEG_ACCESS | FATTR4_WORD2_TIME_DELEG_MODIFY); @@ -1377,7 +1453,7 @@ nfsd4_write(struct svc_rqst *rqstp, struct nfsd4_compound_state *cstate, write->wr_how_written = write->wr_stable_how; status = nfsd_vfs_write(rqstp, &cstate->current_fh, nf, write->wr_offset, &write->wr_payload, - &cnt, write->wr_how_written, + &cnt, nfsd4_iocb_flags(write->wr_how_written), (__be32 *)write->wr_verifier.data); nfsd_file_put(nf); @@ -1431,16 +1507,37 @@ nfsd4_clone(struct svc_rqst *rqstp, struct nfsd4_compound_state *cstate, { struct nfsd4_clone *clone = &u->clone; struct nfsd_file *src, *dst; + bool sync_failed = false; + errseq_t since; __be32 status; + int host_err; status = nfsd4_verify_copy(rqstp, cstate, &clone->cl_src_stateid, &src, &clone->cl_dst_stateid, &dst); if (status) goto out; - status = nfsd4_clone_file_range(rqstp, src, clone->cl_src_pos, - dst, clone->cl_dst_pos, clone->cl_count, - EX_ISSYNC(cstate->current_fh.fh_export)); + host_err = nfsd_clone_file_range(src->nf_file, clone->cl_src_pos, + dst->nf_file, clone->cl_dst_pos, + clone->cl_count, &since); + if (!host_err && EX_ISSYNC(cstate->current_fh.fh_export)) { + host_err = nfsd_clone_sync_range(src->nf_file, dst->nf_file, + clone->cl_dst_pos, + clone->cl_count, since); + sync_failed = host_err < 0; + } + if (host_err < 0) { + trace_nfsd_clone_file_range_err(rqstp, &cstate->save_fh, + clone->cl_src_pos, &cstate->current_fh, + clone->cl_dst_pos, clone->cl_count, host_err); + if (sync_failed) { + struct nfsd_net *nn = net_generic(dst->nf_net, + nfsd_net_id); + + nfsd_maybe_reset_write_verifier(nn, rqstp, host_err); + } + } + status = nfserrno(host_err); if (!status && (READ_ONCE(dst->nf_file->f_mode) & FMODE_NOCMTIME) != 0) nfsd_update_cmtime_attr(dst->nf_file, 0); @@ -1653,13 +1750,6 @@ void nfsd4_cancel_copy_by_sb(struct net *net, struct super_block *sb) #ifdef CONFIG_NFSD_V4_2_INTER_SSC -extern struct file *nfs42_ssc_open(struct vfsmount *ss_mnt, - struct nfs_fh *src_fh, - nfs4_stateid *stateid); -extern void nfs42_ssc_close(struct file *filep); - -extern void nfs_sb_deactive(struct super_block *sb); - #define NFSD42_INTERSSC_MOUNTOPS "vers=4.2,addr=%s,sec=sys" /* @@ -1882,7 +1972,7 @@ nfsd4_cleanup_inter_ssc(struct nfsd4_ssc_umount_item *nsui, struct file *filp, struct nfsd_net *nn = net_generic(dst->nf_net, nfsd_net_id); long timeout = msecs_to_jiffies(nfsd4_ssc_umount_timeout); - nfs42_ssc_close(filp); + nfsd42_ssc_close(filp); fput(filp); spin_lock(&nn->nfsd_ssc_lock); @@ -1914,12 +2004,6 @@ nfsd4_cleanup_inter_ssc(struct nfsd4_ssc_umount_item *nsui, struct file *filp, { } -static struct file *nfs42_ssc_open(struct vfsmount *ss_mnt, - struct nfs_fh *src_fh, - nfs4_stateid *stateid) -{ - return NULL; -} #endif /* CONFIG_NFSD_V4_2_INTER_SSC */ static __be32 @@ -1974,7 +2058,7 @@ static void nfsd4_init_copy_res(struct nfsd4_copy *copy, bool sync) { copy->cp_res.wr_stable_how = test_bit(NFSD4_COPY_F_COMMITTED, ©->cp_flags) ? - NFS_FILE_SYNC : NFS_UNSTABLE; + FILE_SYNC4 : UNSTABLE4; nfsd4_copy_set_sync(copy, sync); } @@ -2142,8 +2226,8 @@ static int nfsd4_do_async_copy(void *data) if (nfsd4_ssc_is_inter(copy)) { struct file *filp; - filp = nfs42_ssc_open(copy->ss_nsui->nsui_vfsmount, - ©->c_fh, ©->stateid); + filp = nfsd42_ssc_open(copy->ss_nsui->nsui_vfsmount, + ©->c_fh, ©->stateid); if (IS_ERR(filp)) { switch (PTR_ERR(filp)) { case -EBADF: @@ -3328,7 +3412,7 @@ nfsd4_proc_compound(struct svc_rqst *rqstp) if (current_fh->fh_export && need_wrongsec_check(rqstp)) - op->status = check_nfsd_access(current_fh->fh_export, rqstp, false); + op->status = check_nfsd_access(current_fh->fh_export, rqstp); } encode_op: if (op->status == nfserr_replay_me) { @@ -3842,7 +3926,6 @@ static const struct nfsd4_operation nfsd4_ops[] = { }, [OP_OPEN] = { .op_func = nfsd4_open, - .op_release = nfsd4_open_release, .op_flags = OP_HANDLES_WRONGSEC | OP_MODIFIES_SOMETHING, .op_name = "OP_OPEN", .op_rsize_bop = nfsd4_open_rsize, diff --git a/fs/nfsd/nfs4recover.c b/fs/nfsd/nfs4recover.c index d513971fb119..aee3a0b22d1c 100644 --- a/fs/nfsd/nfs4recover.c +++ b/fs/nfsd/nfs4recover.c @@ -47,6 +47,7 @@ #include <linux/nfsd/cld.h> #include "nfsd.h" +#include "nfs4ctl.h" #include "state.h" #include "vfs.h" #include "netns.h" diff --git a/fs/nfsd/nfs4state.c b/fs/nfsd/nfs4state.c index 9c4adf3110ae..1de6c6d757c3 100644 --- a/fs/nfsd/nfs4state.c +++ b/fs/nfsd/nfs4state.c @@ -45,13 +45,15 @@ #include <linux/string_helpers.h> #include <linux/fsnotify.h> #include <linux/rhashtable.h> -#include <linux/nfs_ssc.h> +#include <linux/nfsd_ssc.h> #include "xdr4.h" +#include "nfs4ctl.h" #include "xdr4cb.h" #include "vfs.h" #include "current_stateid.h" #include "stats.h" +#include "nfserr.h" #include "netns.h" #include "pnfs.h" @@ -91,7 +93,9 @@ static void nfs4_free_ol_stateid(struct nfs4_stid *stid); static void nfsd4_end_grace(struct nfsd_net *nn); static void _free_cpntf_state_locked(struct nfsd_net *nn, struct nfs4_cpntf_state *cps); static void nfsd4_file_hash_remove(struct nfs4_file *fi); -static void deleg_reaper(struct nfsd_net *nn); +static void deleg_reaper(struct nfsd_net *nn, unsigned long backlog); +static void nfsd4_drop_revoked_stid(struct nfs4_stid *s) + __releases(&s->sc_client->cl_lock); static const struct lease_manager_operations nfsd_lease_mng_ops; @@ -1157,14 +1161,18 @@ static struct nfs4_ol_stateid * nfs4_alloc_open_stateid(struct nfs4_client *clp) */ static void nfs4_free_deleg(struct nfs4_stid *stid) { + struct nfsd_net *nn = net_generic(stid->sc_client->net, nfsd_net_id); struct nfs4_delegation *dp = delegstateid(stid); WARN_ON_ONCE(!list_empty(&stid->sc_cp_list)); WARN_ON_ONCE(!list_empty(&dp->dl_perfile)); WARN_ON_ONCE(!list_empty(&dp->dl_perclnt)); WARN_ON_ONCE(!list_empty(&dp->dl_recall_lru)); + /* The list outlives one recall, so ->release() cannot free it. */ + nfsd41_cb_destroy_referring_call_list(&dp->dl_recall); kmem_cache_free(deleg_slab, stid); atomic_long_dec(&num_delegations); + atomic_long_dec(&nn->nfsd_delegations); } /* @@ -1249,6 +1257,7 @@ __alloc_init_deleg(struct nfs4_client *clp, struct nfs4_file *fp, struct nfs4_clnt_odstate *odstate, u32 dl_type, void (*sc_free)(struct nfs4_stid *)) { + struct nfsd_net *nn = net_generic(clp->net, nfsd_net_id); struct nfs4_delegation *dp; struct nfs4_stid *stid; long n; @@ -1257,6 +1266,7 @@ __alloc_init_deleg(struct nfs4_client *clp, struct nfs4_file *fp, return NULL; n = atomic_long_inc_return(&num_delegations); + atomic_long_inc(&nn->nfsd_delegations); if (n < 0 || n > max_delegations) goto out_dec; @@ -1279,6 +1289,9 @@ __alloc_init_deleg(struct nfs4_client *clp, struct nfs4_file *fp, dp->dl_type = dl_type; dp->dl_retries = 1; dp->dl_recalled = false; + dp->dl_recall_rejected = false; + dp->dl_recall_grant.valid = false; + dp->dl_recall_grant.retired_at_send = false; get_nfs4_file(fp); dp->dl_stid.sc_file = fp; nfsd4_init_cb(&dp->dl_recall, dp->dl_stid.sc_client, @@ -1286,6 +1299,7 @@ __alloc_init_deleg(struct nfs4_client *clp, struct nfs4_file *fp, return dp; out_dec: atomic_long_dec(&num_delegations); + atomic_long_dec(&nn->nfsd_delegations); return NULL; } @@ -1314,6 +1328,7 @@ static void nfs4_free_dir_deleg(struct nfs4_stid *stid) for (i = 0; i < ncn->ncn_evt_cnt; ++i) nfsd_notify_event_put(ncn->ncn_evt[i]); kfree(ncn->ncn_nf); + kfree(ncn->ncn_masks); for (i = 0; i < NOTIFY4_PAGE_ARRAY_SIZE; i++) { if (!ncn->ncn_pages[i]) break; @@ -1346,6 +1361,11 @@ alloc_init_dir_deleg(struct nfs4_client *clp, struct nfs4_file *fp) nfs4_put_stid(&dp->dl_stid); return NULL; } + ncn->ncn_masks = kcalloc(NOTIFY4_EVENT_QUEUE_SIZE, sizeof(*ncn->ncn_masks), GFP_KERNEL); + if (!ncn->ncn_masks) { + nfs4_put_stid(&dp->dl_stid); + return NULL; + } spin_lock_init(&ncn->ncn_lock); nfsd4_init_cb(&ncn->ncn_cb, dp->dl_stid.sc_client, &nfsd4_cb_notify_ops, NFSPROC4_CLNT_CB_NOTIFY); @@ -1512,6 +1532,7 @@ hash_delegation_locked(struct nfs4_delegation *dp, struct nfs4_file *fp) dp->dl_stid.sc_type = SC_TYPE_DELEG; list_add(&dp->dl_perfile, &fp->fi_delegations); list_add(&dp->dl_perclnt, &clp->cl_delegations); + clp->cl_deleg_count++; return 0; } @@ -1543,6 +1564,7 @@ unhash_delegation_locked(struct nfs4_delegation *dp, unsigned short statusmask) ++dp->dl_time; spin_lock(&fp->fi_lock); list_del_init(&dp->dl_perclnt); + dp->dl_stid.sc_client->cl_deleg_count--; list_del_init(&dp->dl_recall_lru); list_del_init(&dp->dl_perfile); spin_unlock(&fp->fi_lock); @@ -1563,27 +1585,22 @@ static void destroy_delegation(struct nfs4_delegation *dp) } /** - * revoke_delegation - perform nfs4 delegation structure cleanup - * @dp: pointer to the delegation + * revoke_delegation - dispose of a delegation the server has revoked + * @dp: delegation to dispose of + * + * The caller holds a reference on @dp, which this function consumes. + * On NFSv4.1 and newer, @dp's sc_status must already carry + * SC_STATUS_REVOKED or SC_STATUS_ADMIN_REVOKED. * - * This function assumes that it's called either from the administrative - * interface (nfsd4_revoke_states()) that's revoking a specific delegation - * stateid or it's called from a laundromat thread (nfsd4_landromat()) that - * determined that this specific state has expired and needs to be revoked - * (both mark state with the appropriate stid sc_status mode). It is also - * assumed that a reference was taken on the @dp state. This function - * consumes that reference. + * @dp is parked on the client's cl_revoked list to await a FREE_STATEID. + * Where none can arrive, @dp is destroyed here instead: FREE_STATEID has + * already freed it, or the client rejected the recall with + * NFS4ERR_BADHANDLE or NFS4ERR_BAD_STATEID and holds no record of the + * delegation. NFS4ERR_ADMIN_REVOKED still prompts one, so an + * administrative revoke waits on cl_revoked. * - * If this function finds that the @dp state is SC_STATUS_FREED it means - * that a FREE_STATEID operation for this stateid has been processed and - * we can proceed to removing it from recalled list. However, if @dp state - * isn't marked SC_STATUS_FREED, it means we need place it on the cl_revoked - * list and wait for the FREE_STATEID to arrive from the client. At the same - * time, we need to mark it as SC_STATUS_FREEABLE to indicate to the - * nfsd4_free_stateid() function that this stateid has already been added - * to the cl_revoked list and that nfsd4_free_stateid() is now responsible - * for removing it from the list. Inspection of where the delegation state - * in the revocation process is protected by the clp->cl_lock. + * Context: Takes and releases the client's cl_lock; may sleep after + * dropping it. */ static void revoke_delegation(struct nfs4_delegation *dp) { @@ -1601,6 +1618,19 @@ static void revoke_delegation(struct nfs4_delegation *dp) list_del_init(&dp->dl_recall_lru); goto out; } + if (dp->dl_recall_rejected && + !(dp->dl_stid.sc_status & SC_STATUS_ADMIN_REVOKED)) { + /* + * SC_STATUS_CLOSED, set under cl_lock, makes a racing + * FREE_STATEID bail out rather than drop this reference + * too. The put releases what cl_revoked would have held. + */ + dp->dl_stid.sc_status |= SC_STATUS_CLOSED; + spin_unlock(&clp->cl_lock); + nfs4_put_stid(&dp->dl_stid); + destroy_unhashed_deleg(dp); + return; + } list_add(&dp->dl_recall_lru, &clp->cl_revoked); dp->dl_stid.sc_status |= SC_STATUS_FREEABLE; out: @@ -2789,10 +2819,16 @@ void nfsd4_put_client(struct nfs4_client *clp) static void free_client(struct nfs4_client *clp) { - while (!list_empty(&clp->cl_sessions)) { + LIST_HEAD(reaplist); + + /* client_info_show() walks cl_sessions under cl_lock */ + spin_lock(&clp->cl_lock); + list_splice_init(&clp->cl_sessions, &reaplist); + spin_unlock(&clp->cl_lock); + while (!list_empty(&reaplist)) { struct nfsd4_session *ses; - ses = list_entry(clp->cl_sessions.next, struct nfsd4_session, - se_perclnt); + ses = list_entry(reaplist.next, struct nfsd4_session, + se_perclnt); list_del(&ses->se_perclnt); WARN_ON_ONCE(atomic_read(&ses->se_ref)); free_session(ses); @@ -2889,11 +2925,18 @@ __destroy_client(struct nfs4_client *clp) list_del_init(&dp->dl_recall_lru); destroy_unhashed_deleg(dp); } + /* + * A CB_RECALL reply can release revoked delegations concurrently: + * nfsd4_shutdown_callback() has not run yet. + */ + spin_lock(&clp->cl_lock); while (!list_empty(&clp->cl_revoked)) { dp = list_entry(clp->cl_revoked.next, struct nfs4_delegation, dl_recall_lru); - list_del_init(&dp->dl_recall_lru); - nfs4_put_stid(&dp->dl_stid); + /* this function drops ->cl_lock */ + nfsd4_drop_revoked_stid(&dp->dl_stid); + spin_lock(&clp->cl_lock); } + spin_unlock(&clp->cl_lock); while (!list_empty(&clp->cl_openowners)) { oo = list_entry(clp->cl_openowners.next, struct nfs4_openowner, oo_perclient); nfs4_get_stateowner(&oo->oo_owner); @@ -3737,14 +3780,9 @@ nfsd4_cb_notify_prepare(struct nfsd4_callback *cb) struct nfsd_notify_event *nne = events[i]; if (!error) { - u32 *maskp = (u32 *)xdr_reserve_space(&stream, sizeof(*maskp)); + u32 *maskp = &ncn->ncn_masks[i]; u8 *p; - if (!maskp) { - error = true; - goto put_event; - } - p = nfsd4_encode_notify_event(&stream, nne, dp, nf, maskp); if (!p) { pr_notice("Could not generate CB_NOTIFY from fsnotify mask 0x%x\n", @@ -3762,13 +3800,10 @@ put_event: nfsd_notify_event_put(nne); } if (!error && (dp->dl_notify_mask & BIT(NOTIFY4_CHANGE_DIR_ATTRS))) { - u32 *maskp = (u32 *)xdr_reserve_space(&stream, sizeof(*maskp)); + u32 *maskp = &ncn->ncn_masks[count]; u8 *p; - if (maskp) - p = nfsd4_encode_dir_attr_change(&stream, dp, nf); - else - p = ERR_PTR(-ENOBUFS); + p = nfsd4_encode_dir_attr_change(&stream, dp, nf); if (IS_ERR(p)) { /* @@ -4216,12 +4251,18 @@ nfsd4_set_ex_flags(struct nfs4_client *new, struct nfsd4_exchange_id *clid) static bool client_has_openowners(struct nfs4_client *clp) { struct nfs4_openowner *oo; + bool found = false; + spin_lock(&clp->cl_lock); list_for_each_entry(oo, &clp->cl_openowners, oo_perclient) { - if (!list_empty(&oo->oo_owner.so_stateids)) - return true; + if (!list_empty(&oo->oo_owner.so_stateids)) { + found = true; + break; + } } - return false; + spin_unlock(&clp->cl_lock); + + return found; } static bool client_has_state(struct nfs4_client *clp) @@ -5016,6 +5057,7 @@ __be32 nfsd4_sequence(struct svc_rqst *rqstp, struct nfsd4_compound_state *cstate, union nfsd4_op_u *u) { + struct nfsd4_compoundargs *args = rqstp->rq_argp; struct nfsd4_sequence *seq = &u->sequence; struct nfsd4_compoundres *resp = rqstp->rq_resp; struct xdr_stream *xdr = resp->xdr; @@ -5025,6 +5067,7 @@ nfsd4_sequence(struct svc_rqst *rqstp, struct nfsd4_compound_state *cstate, struct nfsd4_conn *conn; __be32 status; int buflen; + u32 maxlen, respsize; struct net *net = SVC_NET(rqstp); struct nfsd_net *nn = net_generic(net, nfsd_net_id); @@ -5102,7 +5145,22 @@ nfsd4_sequence(struct svc_rqst *rqstp, struct nfsd4_compound_state *cstate, session->se_fchannel.maxresp_sz; status = (seq->cachethis) ? nfserr_rep_too_big_to_cache : nfserr_rep_too_big; - if (xdr_restrict_buflen(xdr, buflen - rqstp->rq_auth_slack)) + if (buflen < rqstp->rq_auth_slack) + goto out_put_session; + maxlen = buflen - rqstp->rq_auth_slack; + + /* + * A SEQUENCE result too large for maxlen never reaches + * nfsd4_encode_sequence(), so cstate.data_offset stays zero and + * the reply cache overruns the slot. + */ + respsize = nfsd4_max_reply(rqstp, &args->ops[0]); + if (!nfsd4_last_compound_op(rqstp)) + respsize += COMPOUND_ERR_SLACK_SPACE; + if (xdr->buf->len + respsize > maxlen) + goto out_put_session; + + if (xdr_restrict_buflen(xdr, maxlen)) goto out_put_session; svc_reserve_auth(rqstp, buflen); @@ -5118,6 +5176,7 @@ nfsd4_sequence(struct svc_rqst *rqstp, struct nfsd4_compound_state *cstate, slot->sl_flags &= ~NFSD4_SLOT_CACHETHIS; cstate->slot = slot; + cstate->slot_owned = true; cstate->session = session; cstate->clp = clp; @@ -5192,7 +5251,11 @@ nfsd4_sequence_done(struct nfsd4_compoundres *resp) struct nfsd4_compound_state *cs = &resp->cstate; if (nfsd4_has_session(cs)) { - if (cs->status != nfserr_replay_cache) { + /* + * Only the request that claimed the slot may update its + * cached reply and clear NFSD4_SLOT_INUSE. + */ + if (cs->slot_owned) { nfsd4_store_cache_entry(resp); cs->slot->sl_flags &= ~NFSD4_SLOT_INUSE; } @@ -5526,26 +5589,110 @@ out: return -ENOMEM; } +#define NFSD_RECALL_ANY_COOLDOWN_SECS 5 + static unsigned long -nfsd4_state_shrinker_count(struct shrinker *shrink, struct shrink_control *sc) +nfsd4_courtesy_shrinker_count(struct shrinker *shrink, + struct shrink_control *sc) { struct nfsd_net *nn = shrink->private_data; - long count; + long backlog, count; count = atomic_read(&nn->nfsd_courtesy_clients); if (!count) - count = atomic_long_read(&num_delegations); - if (count) - queue_work(laundry_wq, &nn->nfsd_shrinker_work); - return (unsigned long)count; + return 0; + + queue_work(laundry_wq, &nn->nfsd_courtesy_work); + + /* Work already queued is not available to reclaim again. */ + backlog = atomic_long_read(&nn->nfsd_shrink_backlog); + return count > backlog ? count - backlog : 0; +} + +static unsigned long +nfsd4_deleg_shrinker_count(struct shrinker *shrink, struct shrink_control *sc) +{ + struct nfsd_net *nn = shrink->private_data; + time64_t elapsed; + long backlog, count; + + count = atomic_long_read(&nn->nfsd_delegations); + if (!count) + return 0; + + /* + * Delegations the last sweep reached stay unreclaimable until + * deleg_reaper()'s cooldown expires. CB_RECALL_ANY leaves the + * choice of delegations to the client, so there is no return + * to wait on instead. + */ + elapsed = ktime_get_boottime_seconds() - + READ_ONCE(nn->nfsd_last_recall_any); + if (elapsed < NFSD_RECALL_ANY_COOLDOWN_SECS) + return 0; + + /* + * Unlike the courtesy shrinker, this one queues no work. + * Nothing is recalled until a scan request arrives. Subtract + * the requests already recorded, or concurrent reclaimers + * each see the whole namespace and stack a scan on top of it. + */ + backlog = atomic_long_read(&nn->nfsd_deleg_backlog); + return count > backlog ? count - backlog : 0; } static unsigned long -nfsd4_state_shrinker_scan(struct shrinker *shrink, struct shrink_control *sc) +nfsd4_courtesy_shrinker_scan(struct shrinker *shrink, + struct shrink_control *sc) { + struct nfsd_net *nn = shrink->private_data; + + atomic_long_add(sc->nr_to_scan, &nn->nfsd_shrink_backlog); + queue_work(laundry_wq, &nn->nfsd_courtesy_work); + + /* + * The reaper runs from laundry_wq. Report no progress rather + * than claim memory that is not free yet. + */ return SHRINK_STOP; } +static unsigned long +nfsd4_deleg_shrinker_scan(struct shrinker *shrink, struct shrink_control *sc) +{ + struct nfsd_net *nn = shrink->private_data; + + atomic_long_add(sc->nr_to_scan, &nn->nfsd_deleg_backlog); + queue_work(laundry_wq, &nn->nfsd_deleg_work); + + /* + * The reaper sends CB_RECALL_ANY, so nothing is free when + * this returns. + */ + return SHRINK_STOP; +} + +static struct shrinker * +nfsd4_alloc_state_shrinker(struct nfsd_net *nn, const char *name, + unsigned long (*count)(struct shrinker *, + struct shrink_control *), + unsigned long (*scan)(struct shrinker *, + struct shrink_control *)) +{ + struct shrinker *shrink; + + shrink = shrinker_alloc(0, "%s:%s", name, nn->nfsd_name); + if (!shrink) + return NULL; + + shrink->count_objects = count; + shrink->scan_objects = scan; + shrink->private_data = nn; + + shrinker_register(shrink); + return shrink; +} + void nfsd4_init_leases_net(struct nfsd_net *nn) { @@ -6063,6 +6210,80 @@ bool nfsd_wait_for_delegreturn(struct svc_rqst *rqstp, struct inode *inode) return timeo > 0; } +/* + * gen_sessionid() composes a sessionid from the client's clientid and a + * sequence counter, so the sequence alone identifies the granting session. + */ +static void nfsd4_recall_grant_sessionid(const struct nfs4_delegation *dp, + struct nfsd4_sessionid *sid) +{ + sid->clientid = dp->dl_stid.sc_client->cl_clientid; + sid->sequence = dp->dl_recall_grant.sessionid_seq; + sid->reserved = 0; +} + +static bool nfsd4_recall_grant_slot_retired(struct nfs4_delegation *dp) +{ + struct nfs4_client *clp = dp->dl_stid.sc_client; + struct nfsd_net *nn = net_generic(clp->net, nfsd_net_id); + struct nfsd4_session *ses; + struct nfsd4_sessionid sid; + bool retired = false; + void *entry; + + if (!dp->dl_recall_grant.valid) + return false; + + nfsd4_recall_grant_sessionid(dp, &sid); + + /* + * A missing session does not prove the client saw the grant: a + * DESTROY_SESSION unhashes its own session before the reply to + * that compound is encoded. + */ + spin_lock(&nn->client_lock); + ses = __find_in_sessionid_hashtbl((struct nfs4_sessionid *)&sid, + clp->net); + entry = ses ? xa_load(&ses->se_slots, dp->dl_recall_grant.slotid) : NULL; + if (xa_is_value(entry)) { + /* + * A slot is freed only once the client has acknowledged + * the smaller slot table, which it cannot do while a + * request on that slot is outstanding. + */ + retired = true; + } else if (entry) { + struct nfsd4_slot *slot = entry; + + /* + * A reactivated slot was freed and rebuilt, so the same + * acknowledgment applies. The seqid test errs toward + * revoking: a rebuilt slot restarting at seqid 1 matches + * an old grant. + */ + retired = (slot->sl_flags & NFSD4_SLOT_REUSED) || + ((slot->sl_flags & NFSD4_SLOT_INITIALIZED) && + slot->sl_seqid != dp->dl_recall_grant.seqid); + } + spin_unlock(&nn->client_lock); + return retired; +} + +/* + * ->prepare does not run on every send: nfsd4_run_cb_work() skips it + * on a requeue, and a retry via rpc_restart_call_prepare() re-enters + * the RPC layer beneath it. The granting request does not change, so + * a send inherits a correct list. Retirement is the one transition + * the list has to follow. + */ +static void nfsd4_refresh_recall_grant(struct nfs4_delegation *dp) +{ + dp->dl_recall_grant.retired_at_send = + nfsd4_recall_grant_slot_retired(dp); + if (dp->dl_recall_grant.retired_at_send) + nfsd41_cb_destroy_referring_call_list(&dp->dl_recall); +} + static bool nfsd4_cb_recall_prepare(struct nfsd4_callback *cb) { struct nfs4_delegation *dp = cb_to_delegation(cb); @@ -6084,9 +6305,46 @@ static bool nfsd4_cb_recall_prepare(struct nfsd4_callback *cb) list_add_tail(&dp->dl_recall_lru, &nn->del_recall_lru); } spin_unlock(&nn->deleg_lock); + + nfsd4_refresh_recall_grant(dp); + + if (dp->dl_recall_grant.valid && !dp->dl_recall_grant.retired_at_send) { + struct nfsd4_sessionid sid; + + nfsd4_recall_grant_sessionid(dp, &sid); + nfsd41_cb_referring_call(&dp->dl_recall, + (struct nfs4_sessionid *)&sid, + dp->dl_recall_grant.slotid, + dp->dl_recall_grant.seqid); + } return true; } +/* + * cl_lock orders this against a laundromat reaping @dp: either + * revoke_delegation() observes dl_recall_rejected and destroys @dp, or + * it reached cl_revoked first and @dp is released here instead. + */ +static void nfsd4_deleg_recall_rejected(struct nfs4_delegation *dp) +{ + struct nfs4_client *clp = dp->dl_stid.sc_client; + + spin_lock(&clp->cl_lock); + if (dp->dl_stid.sc_status & (SC_STATUS_CLOSED | SC_STATUS_FREED | + SC_STATUS_ADMIN_REVOKED)) { + spin_unlock(&clp->cl_lock); + return; + } + if (dp->dl_stid.sc_status & SC_STATUS_FREEABLE) { + dp->dl_stid.sc_status |= SC_STATUS_CLOSED; + /* this function drops ->cl_lock */ + nfsd4_drop_revoked_stid(&dp->dl_stid); + return; + } + dp->dl_recall_rejected = true; + spin_unlock(&clp->cl_lock); +} + static int nfsd4_cb_recall_done(struct nfsd4_callback *cb, struct rpc_task *task) { @@ -6094,27 +6352,32 @@ static int nfsd4_cb_recall_done(struct nfsd4_callback *cb, trace_nfsd_cb_recall_done(&dp->dl_stid.sc_stateid, task); - if (dp->dl_stid.sc_status) - /* CLOSED or REVOKED */ - return 1; - switch (task->tk_status) { case 0: return 1; case -NFS4ERR_DELAY: + if (dp->dl_stid.sc_status) + /* CLOSED or REVOKED */ + return 1; rpc_delay(task, 2 * HZ); return 0; case -EBADHANDLE: case -NFS4ERR_BAD_STATEID: /* - * Race: client probably got cb_recall before open reply - * granting delegation. + * Retirement of the granting slot proves the client saw + * the grant. Trust the rejection only if the slot had + * retired when this recall was sent. */ - if (dp->dl_retries--) { + if (dp->dl_recall_grant.retired_at_send) { + nfsd4_deleg_recall_rejected(dp); + return 1; + } + if (!dp->dl_stid.sc_status && dp->dl_retries--) { + nfsd4_refresh_recall_grant(dp); rpc_delay(task, 2 * HZ); return 0; } - fallthrough; + return 1; default: return 1; } @@ -6708,9 +6971,25 @@ static bool nfsd4_want_deleg_timestamps(const struct nfsd4_open *open) return open->op_deleg_want & OPEN4_SHARE_ACCESS_WANT_DELEG_TIMESTAMPS; } +static void +nfs4_delegation_record_grant_slot(struct nfs4_delegation *dp, + const struct nfsd4_compound_state *cstate) +{ + const struct nfsd4_sessionid *sid; + + if (!cstate->session) + return; + sid = (struct nfsd4_sessionid *)cstate->session->se_sessionid.data; + dp->dl_recall_grant.sessionid_seq = sid->sequence; + dp->dl_recall_grant.slotid = cstate->slot->sl_index; + dp->dl_recall_grant.seqid = cstate->slot->sl_seqid; + dp->dl_recall_grant.valid = true; +} + static struct nfs4_delegation * -nfs4_set_delegation(struct nfsd4_open *open, struct nfs4_ol_stateid *stp, - struct svc_fh *parent) +nfs4_set_delegation(struct nfsd4_open *open, + const struct nfsd4_compound_state *cstate, + struct nfs4_ol_stateid *stp, struct svc_fh *parent) { bool deleg_ts = nfsd4_want_deleg_timestamps(open); struct nfs4_client *clp = stp->st_stid.sc_client; @@ -6800,6 +7079,14 @@ nfs4_set_delegation(struct nfsd4_open *open, struct nfs4_ol_stateid *stp, dp = alloc_init_deleg(clp, fp, odstate, dl_type); if (!dp) goto out_delegees; + + /* + * Record the granting slot before kernel_setlease() makes @dp + * visible to lease breakers. A conflicting open can drive + * CB_RECALL to completion from that point on. + */ + nfs4_delegation_record_grant_slot(dp, cstate); + if (stp->st_stid.sc_export) dp->dl_stid.sc_export = exp_get(stp->st_stid.sc_export); @@ -6964,6 +7251,7 @@ nfs4_open_delegation(struct svc_rqst *rqstp, struct nfsd4_open *open, struct nfs4_ol_stateid *stp, struct svc_fh *currentfh, struct svc_fh *fh) { + struct nfsd4_compoundres *resp = rqstp->rq_resp; struct nfs4_openowner *oo = openowner(stp->st_stateowner); bool deleg_ts = nfsd4_want_deleg_timestamps(open); struct nfs4_client *clp = stp->st_stid.sc_client; @@ -7000,7 +7288,7 @@ nfs4_open_delegation(struct svc_rqst *rqstp, struct nfsd4_open *open, default: goto out_no_deleg; } - dp = nfs4_set_delegation(open, stp, parent); + dp = nfs4_set_delegation(open, &resp->cstate, stp, parent); if (IS_ERR(dp)) goto out_no_deleg; @@ -7624,6 +7912,7 @@ nfs4_laundromat(struct nfsd_net *nn) struct nfs4_cpntf_state *cps; struct nfs4_client *clp; copy_stateid_t *cps_t; + long held, host, n; int i; if (clients_still_reclaiming(nn)) { @@ -7737,8 +8026,22 @@ nfs4_laundromat(struct nfsd_net *nn) /* service the server-to-server copy delayed unmount list */ nfsd4_ssc_expire_umount(nn); #endif - if (atomic_long_read(&num_delegations) >= max_delegations) - deleg_reaper(nn); + /* + * set_max_delegations() computes a zero max_delegations on a + * server with very little memory. @host is a divisor below. + */ + host = atomic_long_read(&num_delegations); + if (host && host >= max_delegations) { + /* + * max_delegations bounds the host, but the laundromat + * runs once per network namespace. Requesting the whole + * overage in each would multiply the request, so take + * only this namespace's share. + */ + held = atomic_long_read(&nn->nfsd_delegations); + n = host - max_delegations + 1; + deleg_reaper(nn, DIV64_U64_ROUND_UP((u64)n * held, host)); + } out: return max_t(time64_t, lt.new_timeo, NFSD_LAUNDROMAT_MINTIMEOUT); } @@ -7766,49 +8069,146 @@ courtesy_client_reaper(struct nfsd_net *nn) nfs4_process_client_reaplist(&reaplist); } +/* The two passes in deleg_reaper() must agree on which clients are asked. */ +static bool +deleg_reaper_eligible(const struct nfs4_client *clp, time64_t now) +{ + if (clp->cl_minorversion == 0) + return false; + if (clp->cl_state != NFSD4_ACTIVE) + return false; + if (atomic_read(&clp->cl_delegs_in_recall)) + return false; + if (test_bit(NFSD4_CALLBACK_RUNNING, &clp->cl_ra->ra_cb.cb_flags)) + return false; + if (now - clp->cl_ra_time < NFSD_RECALL_ANY_COOLDOWN_SECS) + return false; + if (clp->cl_cb_state != NFSD4_CB_UP) + return false; + return true; +} + static void -deleg_reaper(struct nfsd_net *nn) +deleg_reaper(struct nfsd_net *nn, unsigned long backlog) { struct list_head *pos, *next; struct nfs4_client *clp; + unsigned long remaining, share, total; + unsigned int count; + time64_t now; + + /* + * Recalling a delegation before it is needed costs the client + * an OPEN when it next touches the file. Leave + * nfsd_last_recall_any unstamped so the next sweep is not + * delayed. + */ + if (!backlog) + return; + now = ktime_get_boottime_seconds(); spin_lock(&nn->client_lock); - list_for_each_safe(pos, next, &nn->client_lru) { + + /* + * Only the clients this sweep asks contribute to the + * apportionment. Dividing the request among holders that are + * skipped under-serves it, and the shortfall goes nowhere: + * nfsd4_deleg_shrinker_worker() has already cleared + * nfsd_deleg_backlog. + */ + total = 0; + list_for_each(pos, &nn->client_lru) { clp = list_entry(pos, struct nfs4_client, cl_lru); - if (clp->cl_state != NFSD4_ACTIVE) + if (!deleg_reaper_eligible(clp, now)) continue; - if (list_empty(&clp->cl_delegations)) - continue; - if (atomic_read(&clp->cl_delegs_in_recall)) - continue; - if (ktime_get_boottime_seconds() - clp->cl_ra_time < 5) + /* + * This read races with hash_delegation_locked() and + * unhash_delegation_locked() on other CPUs. A stale + * count only skews the keep value; the next + * laundromat pass sees a more current one. + */ + total += data_race(READ_ONCE(clp->cl_deleg_count)); + } + if (!total) + goto out; + + /* + * Reclaim asks in batches and is not bound by what the count + * callback reported, so the backlog can exceed what these + * clients hold. Cap it to keep each share within the client's + * own count. + */ + backlog = min(backlog, total); + remaining = backlog; + + list_for_each_safe(pos, next, &nn->client_lru) { + clp = list_entry(pos, struct nfs4_client, cl_lru); + + if (!deleg_reaper_eligible(clp, now)) continue; - if (clp->cl_cb_state != NFSD4_CB_UP) + count = data_race(READ_ONCE(clp->cl_deleg_count)); + if (!count) continue; if (test_and_set_bit(NFSD4_CALLBACK_RUNNING, &clp->cl_ra->ra_cb.cb_flags)) continue; /* release in nfsd4_cb_recall_any_release */ kref_get(&clp->cl_nfsdfs.cl_ref); - clp->cl_ra_time = ktime_get_boottime_seconds(); - clp->cl_ra->ra_keep = 0; + clp->cl_ra_time = now; + /* + * Rounding up guarantees every holder gives up at least + * one. The round-up can overshoot @backlog, so stop + * once the request is met. client_lru is ordered by + * last renewal, so the least active clients are asked + * first. + */ + share = DIV64_U64_ROUND_UP((u64)backlog * count, total); + share = min(share, remaining); + remaining -= share; + clp->cl_ra->ra_keep = count - share; clp->cl_ra->ra_bmval[0] = BIT(RCA4_TYPE_MASK_RDATA_DLG) | - BIT(RCA4_TYPE_MASK_WDATA_DLG); + BIT(RCA4_TYPE_MASK_WDATA_DLG) | + BIT(RCA4_TYPE_MASK_DIR_DLG); trace_nfsd_cb_recall_any(clp->cl_ra); nfsd4_run_cb(&clp->cl_ra->ra_cb); + if (!remaining) + break; } +out: spin_unlock(&nn->client_lock); + + /* + * Stamp the sweep even when no recall went out. A sweep that + * found nothing eligible finds nothing on an immediate retry. + */ + WRITE_ONCE(nn->nfsd_last_recall_any, now); } static void -nfsd4_state_shrinker_worker(struct work_struct *work) +nfsd4_courtesy_shrinker_worker(struct work_struct *work) { struct nfsd_net *nn = container_of(work, struct nfsd_net, - nfsd_shrinker_work); + nfsd_courtesy_work); + long backlog; + /* + * Retire only the requests sampled here, so that requests + * arriving while the reaper runs are still discounted by + * nfsd4_courtesy_shrinker_count(). + */ + backlog = atomic_long_read(&nn->nfsd_shrink_backlog); courtesy_client_reaper(nn); - deleg_reaper(nn); + atomic_long_sub(backlog, &nn->nfsd_shrink_backlog); +} + +static void +nfsd4_deleg_shrinker_worker(struct work_struct *work) +{ + struct nfsd_net *nn = container_of(work, struct nfsd_net, + nfsd_deleg_work); + + deleg_reaper(nn, atomic_long_xchg(&nn->nfsd_deleg_backlog, 0)); } static inline __be32 nfs4_check_fh(struct svc_fh *fhp, struct nfs4_stid *stp) @@ -9770,21 +10170,31 @@ static int nfs4_state_create_net(struct net *net) INIT_DELAYED_WORK(&nn->laundromat_work, laundromat_main); /* Make sure this cannot run until client tracking is initialised */ disable_delayed_work(&nn->laundromat_work); - INIT_WORK(&nn->nfsd_shrinker_work, nfsd4_state_shrinker_worker); + INIT_WORK(&nn->nfsd_courtesy_work, nfsd4_courtesy_shrinker_worker); + INIT_WORK(&nn->nfsd_deleg_work, nfsd4_deleg_shrinker_worker); + atomic_long_set(&nn->nfsd_shrink_backlog, 0); + atomic_long_set(&nn->nfsd_deleg_backlog, 0); + nn->nfsd_last_recall_any = 0; get_net(net); - nn->nfsd_client_shrinker = shrinker_alloc(0, "nfsd-client"); - if (!nn->nfsd_client_shrinker) + nn->nfsd_courtesy_shrinker = + nfsd4_alloc_state_shrinker(nn, "nfsd-courtesy", + nfsd4_courtesy_shrinker_count, + nfsd4_courtesy_shrinker_scan); + if (!nn->nfsd_courtesy_shrinker) goto err_shrinker; - nn->nfsd_client_shrinker->scan_objects = nfsd4_state_shrinker_scan; - nn->nfsd_client_shrinker->count_objects = nfsd4_state_shrinker_count; - nn->nfsd_client_shrinker->private_data = nn; - - shrinker_register(nn->nfsd_client_shrinker); + nn->nfsd_deleg_shrinker = + nfsd4_alloc_state_shrinker(nn, "nfsd-delegation", + nfsd4_deleg_shrinker_count, + nfsd4_deleg_shrinker_scan); + if (!nn->nfsd_deleg_shrinker) + goto err_deleg_shrinker; return 0; +err_deleg_shrinker: + shrinker_free(nn->nfsd_courtesy_shrinker); err_shrinker: put_net(net); kfree(nn->sessionid_hashtbl); @@ -9885,8 +10295,10 @@ nfs4_state_shutdown_net(struct net *net) struct list_head *pos, *next, reaplist; struct nfsd_net *nn = net_generic(net, nfsd_net_id); - shrinker_free(nn->nfsd_client_shrinker); - cancel_work_sync(&nn->nfsd_shrinker_work); + shrinker_free(nn->nfsd_courtesy_shrinker); + shrinker_free(nn->nfsd_deleg_shrinker); + cancel_work_sync(&nn->nfsd_courtesy_work); + cancel_work_sync(&nn->nfsd_deleg_work); disable_delayed_work_sync(&nn->laundromat_work); locks_end_grace(&nn->nfsd4_manager); @@ -10294,6 +10706,7 @@ nfsd_get_dir_deleg(struct nfsd4_compound_state *cstate, dp = alloc_init_dir_deleg(clp, fp); if (!dp) goto out_delegees; + nfs4_delegation_record_grant_slot(dp, cstate); if (cstate->current_fh.fh_export) dp->dl_stid.sc_export = exp_get(cstate->current_fh.fh_export); diff --git a/fs/nfsd/nfs4xdr.c b/fs/nfsd/nfs4xdr.c index 606ddcb085c0..00ddaac499c6 100644 --- a/fs/nfsd/nfs4xdr.c +++ b/fs/nfsd/nfs4xdr.c @@ -54,6 +54,7 @@ #include "xdr4.h" #include "vfs.h" #include "state.h" +#include "nfserr.h" #include "cache.h" #include "netns.h" #include "pnfs.h" @@ -113,27 +114,27 @@ static int zero_clientid(clientid_t *clid) return (clid->cl_boot == 0) && (clid->cl_id == 0); } -/** - * svcxdr_tmpalloc - allocate memory to be freed after compound processing - * @argp: NFSv4 compound argument structure - * @len: length of buffer to allocate - * - * Allocates a buffer of size @len to be freed when processing the compound - * operation described in @argp finishes. - */ static void * -svcxdr_tmpalloc(struct nfsd4_compoundargs *argp, size_t len) +svcxdr_tmpalloc_release(struct nfsd4_compoundargs *argp, size_t len, + void (*release)(void *)) { struct svcxdr_tmpbuf *tb; tb = kmalloc_flex(*tb, buf, len); if (!tb) return NULL; + tb->release = release; tb->next = argp->to_free; argp->to_free = tb; return tb->buf; } +static void * +svcxdr_tmpalloc(struct nfsd4_compoundargs *argp, size_t len) +{ + return svcxdr_tmpalloc_release(argp, len, NULL); +} + /* * For xdr strings that need to be passed to other kernel api's * as null-terminated strings. @@ -441,10 +442,16 @@ nfsd4_decode_posixace4(struct nfsd4_compoundargs *argp, return status; } +static void svcxdr_release_pacl(void *p) +{ + posix_acl_release(*(struct posix_acl **)p); +} + static noinline __be32 nfsd4_decode_posixacl(struct nfsd4_compoundargs *argp, struct posix_acl **acl) { struct posix_acl_entry *ace; + struct posix_acl **slot; __be32 status; u32 count; @@ -484,6 +491,15 @@ nfsd4_decode_posixacl(struct nfsd4_compoundargs *argp, struct posix_acl **acl) if (count >= 3) sort_pacl_range(*acl, 0, count - 1); + slot = svcxdr_tmpalloc_release(argp, sizeof(*slot), + svcxdr_release_pacl); + if (!slot) { + posix_acl_release(*acl); + *acl = NULL; + return nfserr_jukebox; + } + *slot = *acl; + return nfs_ok; } @@ -676,7 +692,6 @@ nfsd4_decode_fattr4(struct nfsd4_compoundargs *argp, u32 *bmval, u32 bmlen, status = nfsd4_decode_posixacl(argp, &pacl); if (status) { - posix_acl_release(*dpaclp); *dpaclp = NULL; return status; } @@ -686,12 +701,8 @@ nfsd4_decode_fattr4(struct nfsd4_compoundargs *argp, u32 *bmval, u32 bmlen, /* request sanity: did attrlist4 contain the expected number of words? */ if (attrlist4_count != xdr_stream_pos(argp->xdr) - starting_pos) { -#ifdef CONFIG_NFSD_V4_POSIX_ACLS - posix_acl_release(*dpaclp); - posix_acl_release(*paclp); *dpaclp = NULL; *paclp = NULL; -#endif return nfserr_bad_xdr; } @@ -1606,7 +1617,7 @@ nfsd4_decode_write(struct nfsd4_compoundargs *argp, union nfsd4_op_u *u) return nfserr_bad_xdr; if (xdr_stream_decode_u32(argp->xdr, &write->wr_stable_how) < 0) return nfserr_bad_xdr; - if (write->wr_stable_how > NFS_FILE_SYNC) + if (write->wr_stable_how > FILE_SYNC4) return nfserr_bad_xdr; if (xdr_stream_decode_u32(argp->xdr, &write->wr_buflen) < 0) return nfserr_bad_xdr; @@ -1920,6 +1931,17 @@ nfsd4_decode_get_dir_delegation(struct nfsd4_compoundargs *argp, #ifdef CONFIG_NFSD_PNFS static __be32 +nfsd4_decode_deviceid4(struct xdr_stream *xdr, struct nfsd4_deviceid *devid) +{ + __be32 *p = xdr_inline_decode(xdr, NFS4_DEVICEID4_SIZE); + + if (unlikely(!p)) + return nfserr_bad_xdr; + svcxdr_decode_deviceid4(p, devid); + return nfs_ok; +} + +static __be32 nfsd4_decode_getdeviceinfo(struct nfsd4_compoundargs *argp, union nfsd4_op_u *u) { @@ -2733,6 +2755,87 @@ nfsd4_decode_compound(struct nfsd4_compoundargs *argp) return true; } +static __always_inline __be32 +nfsd4_encode_bool(struct xdr_stream *xdr, bool val) +{ + __be32 *p = xdr_reserve_space(xdr, XDR_UNIT); + + if (unlikely(p == NULL)) + return nfserr_resource; + *p = val ? xdr_one : xdr_zero; + return nfs_ok; +} + +static __always_inline __be32 +nfsd4_encode_uint32_t(struct xdr_stream *xdr, u32 val) +{ + __be32 *p = xdr_reserve_space(xdr, XDR_UNIT); + + if (unlikely(p == NULL)) + return nfserr_resource; + *p = cpu_to_be32(val); + return nfs_ok; +} + +#define nfsd4_encode_aceflag4(x, v) nfsd4_encode_uint32_t(x, v) +#define nfsd4_encode_acemask4(x, v) nfsd4_encode_uint32_t(x, v) +#define nfsd4_encode_acetype4(x, v) nfsd4_encode_uint32_t(x, v) +#define nfsd4_encode_count4(x, v) nfsd4_encode_uint32_t(x, v) +#define nfsd4_encode_mode4(x, v) nfsd4_encode_uint32_t(x, v) +#define nfsd4_encode_nfs_lease4(x, v) nfsd4_encode_uint32_t(x, v) +#define nfsd4_encode_qop4(x, v) nfsd4_encode_uint32_t(x, v) +#define nfsd4_encode_sequenceid4(x, v) nfsd4_encode_uint32_t(x, v) +#define nfsd4_encode_slotid4(x, v) nfsd4_encode_uint32_t(x, v) + +static __always_inline __be32 +nfsd4_encode_uint64_t(struct xdr_stream *xdr, u64 val) +{ + __be32 *p = xdr_reserve_space(xdr, XDR_UNIT * 2); + + if (unlikely(p == NULL)) + return nfserr_resource; + put_unaligned_be64(val, p); + return nfs_ok; +} + +#define nfsd4_encode_changeid4(x, v) nfsd4_encode_uint64_t(x, v) +#define nfsd4_encode_nfs_cookie4(x, v) nfsd4_encode_uint64_t(x, v) +#define nfsd4_encode_length4(x, v) nfsd4_encode_uint64_t(x, v) +#define nfsd4_encode_offset4(x, v) nfsd4_encode_uint64_t(x, v) + +static __always_inline __be32 +nfsd4_encode_opaque_fixed(struct xdr_stream *xdr, const void *data, + size_t size) +{ + __be32 *p = xdr_reserve_space(xdr, xdr_align_size(size)); + size_t pad = xdr_pad_size(size); + + if (unlikely(p == NULL)) + return nfserr_resource; + memcpy(p, data, size); + if (pad) + memset((char *)p + size, 0, pad); + return nfs_ok; +} + +static __always_inline __be32 +nfsd4_encode_opaque(struct xdr_stream *xdr, const void *data, size_t size) +{ + size_t pad = xdr_pad_size(size); + __be32 *p; + + p = xdr_reserve_space(xdr, XDR_UNIT + xdr_align_size(size)); + if (unlikely(p == NULL)) + return nfserr_resource; + *p++ = cpu_to_be32(size); + memcpy(p, data, size); + if (pad) + memset((char *)p + size, 0, pad); + return nfs_ok; +} + +#define nfsd4_encode_component4(x, d, s) nfsd4_encode_opaque(x, d, s) + static __be32 nfsd4_encode_nfs_fh4(struct xdr_stream *xdr, const struct knfsd_fh *fh_handle) { @@ -3110,7 +3213,7 @@ static __be32 fattr_handle_absent_fs(u32 *bmval0, u32 *bmval1, u32 *bmval2, u32 *bmval1 & ~WORD1_ABSENT_FS_ATTRS) { if (*bmval0 & FATTR4_WORD0_RDATTR_ERROR || *bmval0 & FATTR4_WORD0_FS_LOCATIONS) - *rdattr_err = NFSERR_MOVED; + *rdattr_err = NFS4ERR_MOVED; else return nfserr_moved; } @@ -4285,21 +4388,16 @@ out: static bool nfsd4_setup_notify_entry4(struct notify_entry4 *ne, struct xdr_stream *xdr, struct dentry *dentry, struct nfs4_delegation *dp, - struct nfsd_file *nf, char *name, u32 namelen) + struct nfsd_file *nf, char *name, u32 namelen, + u32 *attrmask) { struct path path = nf->nf_file->f_path; struct nfsd4_fattr_args args = { }; const u32 *reqmask; - uint32_t *attrmask; __be32 status; bool parent; int ret; - /* Reserve space for attrmask */ - attrmask = xdr_reserve_space(xdr, 3 * sizeof(uint32_t)); - if (!attrmask) - return false; - ne->ne_file.data = name; ne->ne_file.len = namelen; ne->ne_attrs.attrmask.element = attrmask; @@ -4383,6 +4481,7 @@ u8 *nfsd4_encode_notify_event(struct xdr_stream *xdr, struct nfsd_notify_event * struct nfs4_delegation *dp, struct nfsd_file *nf, u32 *notify_mask) { + u32 attrmask[3][3] = { }; u8 *p = NULL; *notify_mask = 0; @@ -4391,7 +4490,8 @@ u8 *nfsd4_encode_notify_event(struct xdr_stream *xdr, struct nfsd_notify_event * struct notify_remove4 nr = { }; if (!nfsd4_setup_notify_entry4(&nr.nrm_old_entry, xdr, nne->ne_dentry, dp, - nf, nne->ne_name, nne->ne_namelen)) + nf, nne->ne_name, nne->ne_namelen, + attrmask[0])) goto out_err; p = (u8 *)xdr->p; if (!xdrgen_encode_notify_remove4(xdr, &nr)) @@ -4402,14 +4502,16 @@ u8 *nfsd4_encode_notify_event(struct xdr_stream *xdr, struct nfsd_notify_event * struct notify_remove4 old = { }; if (!nfsd4_setup_notify_entry4(&na.nad_new_entry, xdr, nne->ne_dentry, dp, - nf, nne->ne_name, nne->ne_namelen)) + nf, nne->ne_name, nne->ne_namelen, + attrmask[0])) goto out_err; /* If a file was overwritten, report it in nad_old_entry */ if (nne->ne_target) { if (!nfsd4_setup_notify_entry4(&old.nrm_old_entry, xdr, NULL, dp, nf, - nne->ne_name, nne->ne_namelen)) + nne->ne_name, nne->ne_namelen, + attrmask[1])) goto out_err; na.nad_old_entry.count = 1; na.nad_old_entry.element = &old; @@ -4428,19 +4530,19 @@ u8 *nfsd4_encode_notify_event(struct xdr_stream *xdr, struct nfsd_notify_event * /* Don't send any attributes in the old_entry since they're the same in new */ if (!nfsd4_setup_notify_entry4(&nr.nrn_old_entry.nrm_old_entry, xdr, NULL, dp, nf, nne->ne_name, - nne->ne_namelen)) + nne->ne_namelen, attrmask[0])) goto out_err; if (!nfsd4_setup_notify_entry4(&nr.nrn_new_entry.nad_new_entry, xdr, nne->ne_dentry, dp, nf, newname, - nne->ne_newnamelen)) + nne->ne_newnamelen, attrmask[1])) goto out_err; /* If a file was overwritten, report it in nad_old_entry */ if (nne->ne_target) { if (!nfsd4_setup_notify_entry4(&old.nrm_old_entry, xdr, NULL, dp, nf, newname, - nne->ne_newnamelen)) + nne->ne_newnamelen, attrmask[2])) goto out_err; nr.nrn_new_entry.nad_old_entry.count = 1; nr.nrn_new_entry.nad_old_entry.element = &old; @@ -4476,11 +4578,12 @@ u8 *nfsd4_encode_dir_attr_change(struct xdr_stream *xdr, struct nfs4_delegation { struct dentry *dentry = nf->nf_file->f_path.dentry; struct notify_attr4 na = { }; + u32 attrmask[3] = { }; u8 *p; /* RFC 8881 s10.4.3: ne_file must be a zero-length string for dir attrs */ if (!nfsd4_setup_notify_entry4(&na.na_changed_entry, xdr, - dentry, dp, nf, "", 0)) + dentry, dp, nf, "", 0, attrmask)) return ERR_PTR(-ENOBUFS); /* No requested attributes to report; omit the event */ @@ -4574,8 +4677,6 @@ nfsd4_encode_entry4_fattr(struct nfsd4_readdir *cd, const char *name, * directly from the mountpoint dentry. */ if (nfsd_mountpoint(dentry, exp)) { - int err; - if (!(exp->ex_flags & NFSEXP_V4ROOT) && !attributes_need_mount(cd->rd_bmval)) { ignore_crossmnt = 1; @@ -4586,12 +4687,7 @@ nfsd4_encode_entry4_fattr(struct nfsd4_readdir *cd, const char *name, * Different "."/".." handling? Something else? * At least, add a comment here to explain.... */ - err = nfsd_cross_mnt(cd->rd_rqstp, &dentry, &exp); - if (err) { - nfserr = nfserrno(err); - goto out_put; - } - nfserr = check_nfsd_access(exp, cd->rd_rqstp, false); + nfserr = nfsd_cross_mnt(cd->rd_rqstp, &dentry, &exp); if (nfserr) goto out_put; crossed = true; @@ -6637,11 +6733,20 @@ nfsd4_encode_operation(struct nfsd4_compoundres *resp, struct nfsd4_op *op) unsigned int op_status_offset; nfsd4_enc encoder; - if (xdr_stream_encode_u32(xdr, op->opnum) != XDR_UNIT) + /* + * nfsd4_proc_compound() stops the COMPOUND early only + * when op->status is set, so a header that cannot be + * encoded has to report the failure here. + */ + if (xdr_stream_encode_u32(xdr, op->opnum) != XDR_UNIT) { + op->status = nfsd4_check_resp_size(resp, XDR_UNIT * 2); goto release; + } op_status_offset = xdr->buf->len; - if (!xdr_reserve_space(xdr, XDR_UNIT)) + if (!xdr_reserve_space(xdr, XDR_UNIT)) { + op->status = nfsd4_check_resp_size(resp, XDR_UNIT); goto release; + } if (op->opnum == OP_ILLEGAL) goto status; @@ -6751,7 +6856,10 @@ void nfsd4_release_compoundargs(struct svc_rqst *rqstp) } while (args->to_free) { struct svcxdr_tmpbuf *tb = args->to_free; + args->to_free = tb->next; + if (tb->release) + tb->release(tb->buf); kfree(tb); } } diff --git a/fs/nfsd/nfscache.c b/fs/nfsd/nfscache.c index c7db532c8523..80364b91331a 100644 --- a/fs/nfsd/nfscache.c +++ b/fs/nfsd/nfscache.c @@ -19,6 +19,7 @@ #include <net/checksum.h> #include "nfsd.h" +#include "nfserr.h" #include "netns.h" #include "stats.h" #include "cache.h" diff --git a/fs/nfsd/nfsctl.c b/fs/nfsd/nfsctl.c index 5abb2d4274c9..3a6945b3bd8b 100644 --- a/fs/nfsd/nfsctl.c +++ b/fs/nfsd/nfsctl.c @@ -20,9 +20,12 @@ #include <linux/module.h> #include <linux/fsnotify.h> #include <linux/nfslocalio.h> +#include <linux/nfs3.h> #include "idmap.h" #include "nfsd.h" +#include "nfserr.h" +#include "nfs4ctl.h" #include "netns.h" #include "stats.h" #include "cache.h" @@ -477,7 +480,7 @@ static ssize_t write_pool_threads(struct file *file, char *buf, size_t size) char *mesg = buf; int i; int rv; - int len; + size_t len; int npools; int *nthreads; struct net *net = netns(file); @@ -531,9 +534,13 @@ static ssize_t write_pool_threads(struct file *file, char *buf, size_t size) mesg = buf; size = SIMPLE_TRANSACTION_LIMIT; - for (i = 0; i < npools && size > 0; i++) { - snprintf(mesg, size, "%d%c", nthreads[i], (i == npools-1 ? '\n' : ' ')); - len = strlen(mesg); + for (i = 0; i < npools; i++) { + len = snprintf(mesg, size, "%d%c", nthreads[i], + (i == npools - 1 ? '\n' : ' ')); + if (len >= size) { + rv = -ENAMETOOLONG; + goto out_free; + } size -= len; mesg += len; } @@ -1581,14 +1588,29 @@ int nfsd_nl_rpc_status_get_dumpit(struct sk_buff *skb, rqstp->rq_proc == NFSPROC4_COMPOUND) { /* NFSv4 compound */ struct nfsd4_compoundargs *args; + struct nfsd4_op *ops; + u32 opcnt; int j; args = rqstp->rq_argp; - genl_rqstp.rq_opcnt = min_t(u32, args->opcnt, + opcnt = READ_ONCE(args->opcnt); + ops = READ_ONCE(args->ops); + + /* + * Finish the seqcount retry before + * dereferencing ops. An unchanged counter means + * opcnt and ops came from the same COMPOUND, + * where opcnt cannot exceed what ops holds. + */ + smp_rmb(); + if (READ_ONCE(rqstp->rq_status_counter) != + status_counter) + continue; + + genl_rqstp.rq_opcnt = min_t(u32, opcnt, ARRAY_SIZE(genl_rqstp.rq_opnum)); for (j = 0; j < genl_rqstp.rq_opcnt; j++) - genl_rqstp.rq_opnum[j] = - args->ops[j].opnum; + genl_rqstp.rq_opnum[j] = ops[j].opnum; } #endif /* CONFIG_NFSD_V4 */ diff --git a/fs/nfsd/nfsd.h b/fs/nfsd/nfsd.h index 76a69d9a4e73..a145294c59c8 100644 --- a/fs/nfsd/nfsd.h +++ b/fs/nfsd/nfsd.h @@ -14,8 +14,6 @@ #include <linux/nfs.h> #include <linux/nfs2.h> -#include <linux/nfs3.h> -#include <linux/nfs4.h> #include <linux/sunrpc/svc.h> #include <linux/sunrpc/svc_xprt.h> @@ -147,49 +145,7 @@ extern u64 nfsd_io_cache_write __read_mostly; extern int nfsd_max_blksize; -static inline int nfsd_v4client(struct svc_rqst *rq) -{ - return rq && rq->rq_prog == NFS_PROGRAM && rq->rq_vers == 4; -} - -/* - * NFSv4 State - */ -#ifdef CONFIG_NFSD_V4 -extern unsigned long max_delegations; -int nfsd4_init_slabs(void); -void nfsd4_free_slabs(void); -int nfs4_state_start(void); -int nfs4_state_start_net(struct net *net); -void nfs4_state_shutdown(void); -void nfs4_state_shutdown_net(struct net *net); -int nfs4_reset_recoverydir(char *recdir); -char * nfs4_recoverydir(void); -bool nfsd4_spo_must_allow(struct svc_rqst *rqstp); -int nfsd4_create_laundry_wq(void); -void nfsd4_destroy_laundry_wq(void); -bool nfsd_wait_for_delegreturn(struct svc_rqst *rqstp, struct inode *inode); -#else -static inline int nfsd4_init_slabs(void) { return 0; } -static inline void nfsd4_free_slabs(void) { } -static inline int nfs4_state_start(void) { return 0; } -static inline int nfs4_state_start_net(struct net *net) { return 0; } -static inline void nfs4_state_shutdown(void) { } -static inline void nfs4_state_shutdown_net(struct net *net) { } -static inline int nfs4_reset_recoverydir(char *recdir) { return 0; } -static inline char * nfs4_recoverydir(void) {return NULL; } -static inline bool nfsd4_spo_must_allow(struct svc_rqst *rqstp) -{ - return false; -} -static inline int nfsd4_create_laundry_wq(void) { return 0; }; -static inline void nfsd4_destroy_laundry_wq(void) {}; -static inline bool nfsd_wait_for_delegreturn(struct svc_rqst *rqstp, - struct inode *inode) -{ - return false; -} -#endif +bool nfsd_v4client(struct svc_rqst *rqstp); /* * lockd binding @@ -198,195 +154,4 @@ void nfsd_lockd_init(void); void nfsd_lockd_shutdown(void); -/* - * These macros provide pre-xdr'ed values for faster operation. - */ -#define nfs_ok cpu_to_be32(NFS_OK) -#define nfserr_perm cpu_to_be32(NFSERR_PERM) -#define nfserr_noent cpu_to_be32(NFSERR_NOENT) -#define nfserr_io cpu_to_be32(NFSERR_IO) -#define nfserr_nxio cpu_to_be32(NFSERR_NXIO) -#define nfserr_acces cpu_to_be32(NFSERR_ACCES) -#define nfserr_exist cpu_to_be32(NFSERR_EXIST) -#define nfserr_xdev cpu_to_be32(NFSERR_XDEV) -#define nfserr_nodev cpu_to_be32(NFSERR_NODEV) -#define nfserr_notdir cpu_to_be32(NFSERR_NOTDIR) -#define nfserr_isdir cpu_to_be32(NFSERR_ISDIR) -#define nfserr_inval cpu_to_be32(NFSERR_INVAL) -#define nfserr_fbig cpu_to_be32(NFSERR_FBIG) -#define nfserr_nospc cpu_to_be32(NFSERR_NOSPC) -#define nfserr_rofs cpu_to_be32(NFSERR_ROFS) -#define nfserr_mlink cpu_to_be32(NFSERR_MLINK) -#define nfserr_nametoolong cpu_to_be32(NFSERR_NAMETOOLONG) -#define nfserr_notempty cpu_to_be32(NFSERR_NOTEMPTY) -#define nfserr_dquot cpu_to_be32(NFSERR_DQUOT) -#define nfserr_stale cpu_to_be32(NFSERR_STALE) -#define nfserr_remote cpu_to_be32(NFSERR_REMOTE) -#define nfserr_wflush cpu_to_be32(NFSERR_WFLUSH) -#define nfserr_badhandle cpu_to_be32(NFSERR_BADHANDLE) -#define nfserr_notsync cpu_to_be32(NFSERR_NOT_SYNC) -#define nfserr_badcookie cpu_to_be32(NFSERR_BAD_COOKIE) -#define nfserr_notsupp cpu_to_be32(NFSERR_NOTSUPP) -#define nfserr_toosmall cpu_to_be32(NFSERR_TOOSMALL) -#define nfserr_serverfault cpu_to_be32(NFSERR_SERVERFAULT) -#define nfserr_badtype cpu_to_be32(NFSERR_BADTYPE) -#define nfserr_jukebox cpu_to_be32(NFSERR_JUKEBOX) -#define nfserr_denied cpu_to_be32(NFSERR_DENIED) -#define nfserr_deadlock cpu_to_be32(NFSERR_DEADLOCK) -#define nfserr_expired cpu_to_be32(NFSERR_EXPIRED) -#define nfserr_bad_cookie cpu_to_be32(NFSERR_BAD_COOKIE) -#define nfserr_same cpu_to_be32(NFSERR_SAME) -#define nfserr_clid_inuse cpu_to_be32(NFSERR_CLID_INUSE) -#define nfserr_stale_clientid cpu_to_be32(NFSERR_STALE_CLIENTID) -#define nfserr_resource cpu_to_be32(NFSERR_RESOURCE) -#define nfserr_moved cpu_to_be32(NFSERR_MOVED) -#define nfserr_nofilehandle cpu_to_be32(NFSERR_NOFILEHANDLE) -#define nfserr_minor_vers_mismatch cpu_to_be32(NFSERR_MINOR_VERS_MISMATCH) -#define nfserr_share_denied cpu_to_be32(NFSERR_SHARE_DENIED) -#define nfserr_stale_stateid cpu_to_be32(NFSERR_STALE_STATEID) -#define nfserr_old_stateid cpu_to_be32(NFSERR_OLD_STATEID) -#define nfserr_bad_stateid cpu_to_be32(NFSERR_BAD_STATEID) -#define nfserr_bad_seqid cpu_to_be32(NFSERR_BAD_SEQID) -#define nfserr_symlink cpu_to_be32(NFSERR_SYMLINK) -#define nfserr_not_same cpu_to_be32(NFSERR_NOT_SAME) -#define nfserr_lock_range cpu_to_be32(NFSERR_LOCK_RANGE) -#define nfserr_restorefh cpu_to_be32(NFSERR_RESTOREFH) -#define nfserr_attrnotsupp cpu_to_be32(NFSERR_ATTRNOTSUPP) -#define nfserr_bad_xdr cpu_to_be32(NFSERR_BAD_XDR) -#define nfserr_openmode cpu_to_be32(NFSERR_OPENMODE) -#define nfserr_badowner cpu_to_be32(NFSERR_BADOWNER) -#define nfserr_locks_held cpu_to_be32(NFSERR_LOCKS_HELD) -#define nfserr_op_illegal cpu_to_be32(NFSERR_OP_ILLEGAL) -#define nfserr_grace cpu_to_be32(NFSERR_GRACE) -#define nfserr_no_grace cpu_to_be32(NFSERR_NO_GRACE) -#define nfserr_reclaim_bad cpu_to_be32(NFSERR_RECLAIM_BAD) -#define nfserr_badname cpu_to_be32(NFSERR_BADNAME) -#define nfserr_admin_revoked cpu_to_be32(NFS4ERR_ADMIN_REVOKED) -#define nfserr_cb_path_down cpu_to_be32(NFSERR_CB_PATH_DOWN) -#define nfserr_locked cpu_to_be32(NFSERR_LOCKED) -#define nfserr_wrongsec cpu_to_be32(NFSERR_WRONGSEC) -#define nfserr_delay cpu_to_be32(NFS4ERR_DELAY) -#define nfserr_badiomode cpu_to_be32(NFS4ERR_BADIOMODE) -#define nfserr_badlayout cpu_to_be32(NFS4ERR_BADLAYOUT) -#define nfserr_bad_session_digest cpu_to_be32(NFS4ERR_BAD_SESSION_DIGEST) -#define nfserr_badsession cpu_to_be32(NFS4ERR_BADSESSION) -#define nfserr_badslot cpu_to_be32(NFS4ERR_BADSLOT) -#define nfserr_complete_already cpu_to_be32(NFS4ERR_COMPLETE_ALREADY) -#define nfserr_conn_not_bound_to_session cpu_to_be32(NFS4ERR_CONN_NOT_BOUND_TO_SESSION) -#define nfserr_deleg_already_wanted cpu_to_be32(NFS4ERR_DELEG_ALREADY_WANTED) -#define nfserr_back_chan_busy cpu_to_be32(NFS4ERR_BACK_CHAN_BUSY) -#define nfserr_layouttrylater cpu_to_be32(NFS4ERR_LAYOUTTRYLATER) -#define nfserr_layoutunavailable cpu_to_be32(NFS4ERR_LAYOUTUNAVAILABLE) -#define nfserr_nomatching_layout cpu_to_be32(NFS4ERR_NOMATCHING_LAYOUT) -#define nfserr_recallconflict cpu_to_be32(NFS4ERR_RECALLCONFLICT) -#define nfserr_unknown_layouttype cpu_to_be32(NFS4ERR_UNKNOWN_LAYOUTTYPE) -#define nfserr_seq_misordered cpu_to_be32(NFS4ERR_SEQ_MISORDERED) -#define nfserr_sequence_pos cpu_to_be32(NFS4ERR_SEQUENCE_POS) -#define nfserr_req_too_big cpu_to_be32(NFS4ERR_REQ_TOO_BIG) -#define nfserr_rep_too_big cpu_to_be32(NFS4ERR_REP_TOO_BIG) -#define nfserr_rep_too_big_to_cache cpu_to_be32(NFS4ERR_REP_TOO_BIG_TO_CACHE) -#define nfserr_retry_uncached_rep cpu_to_be32(NFS4ERR_RETRY_UNCACHED_REP) -#define nfserr_unsafe_compound cpu_to_be32(NFS4ERR_UNSAFE_COMPOUND) -#define nfserr_too_many_ops cpu_to_be32(NFS4ERR_TOO_MANY_OPS) -#define nfserr_op_not_in_session cpu_to_be32(NFS4ERR_OP_NOT_IN_SESSION) -#define nfserr_hash_alg_unsupp cpu_to_be32(NFS4ERR_HASH_ALG_UNSUPP) -#define nfserr_clientid_busy cpu_to_be32(NFS4ERR_CLIENTID_BUSY) -#define nfserr_pnfs_io_hole cpu_to_be32(NFS4ERR_PNFS_IO_HOLE) -#define nfserr_seq_false_retry cpu_to_be32(NFS4ERR_SEQ_FALSE_RETRY) -#define nfserr_bad_high_slot cpu_to_be32(NFS4ERR_BAD_HIGH_SLOT) -#define nfserr_deadsession cpu_to_be32(NFS4ERR_DEADSESSION) -#define nfserr_encr_alg_unsupp cpu_to_be32(NFS4ERR_ENCR_ALG_UNSUPP) -#define nfserr_pnfs_no_layout cpu_to_be32(NFS4ERR_PNFS_NO_LAYOUT) -#define nfserr_not_only_op cpu_to_be32(NFS4ERR_NOT_ONLY_OP) -#define nfserr_wrong_cred cpu_to_be32(NFS4ERR_WRONG_CRED) -#define nfserr_wrong_type cpu_to_be32(NFS4ERR_WRONG_TYPE) -#define nfserr_dirdeleg_unavail cpu_to_be32(NFS4ERR_DIRDELEG_UNAVAIL) -#define nfserr_reject_deleg cpu_to_be32(NFS4ERR_REJECT_DELEG) -#define nfserr_returnconflict cpu_to_be32(NFS4ERR_RETURNCONFLICT) -#define nfserr_deleg_revoked cpu_to_be32(NFS4ERR_DELEG_REVOKED) -#define nfserr_partner_notsupp cpu_to_be32(NFS4ERR_PARTNER_NOTSUPP) -#define nfserr_partner_no_auth cpu_to_be32(NFS4ERR_PARTNER_NO_AUTH) -#define nfserr_union_notsupp cpu_to_be32(NFS4ERR_UNION_NOTSUPP) -#define nfserr_offload_denied cpu_to_be32(NFS4ERR_OFFLOAD_DENIED) -#define nfserr_wrong_lfs cpu_to_be32(NFS4ERR_WRONG_LFS) -#define nfserr_badlabel cpu_to_be32(NFS4ERR_BADLABEL) -#define nfserr_file_open cpu_to_be32(NFS4ERR_FILE_OPEN) -#define nfserr_xattr2big cpu_to_be32(NFS4ERR_XATTR2BIG) -#define nfserr_noxattr cpu_to_be32(NFS4ERR_NOXATTR) - -/* - * Error codes for internal use. These are based at an impossible - * nfsstat4 value so that, once converted to be32, they cannot conflict - * with any value defined by the protocol (compare the nlm__int__* codes - * in fs/lockd/lockd.h). - */ -enum { -/* end-of-file indicator in readdir */ - NFSERR_EOF = 30000, -#define nfserr_eof cpu_to_be32(NFSERR_EOF) - -/* replay detected */ - NFSERR_REPLAY_ME, -#define nfserr_replay_me cpu_to_be32(NFSERR_REPLAY_ME) - -/* nfs41 replay detected */ - NFSERR_REPLAY_CACHE, -#define nfserr_replay_cache cpu_to_be32(NFSERR_REPLAY_CACHE) - -/* symlink found where dir expected - handled differently to - * other symlink found errors by NFSv3. - */ - NFSERR_SYMLINK_NOT_DIR, -#define nfserr_symlink_not_dir cpu_to_be32(NFSERR_SYMLINK_NOT_DIR) -}; - -#ifdef CONFIG_NFSD_V4 - -/* before processing a COMPOUND operation, we have to check that there - * is enough space in the buffer for XDR encode to succeed. otherwise, - * we might process an operation with side effects, and be unable to - * tell the client that the operation succeeded. - * - * COMPOUND_SLACK_SPACE - this is the minimum bytes of buffer space - * needed to encode an "ordinary" _successful_ operation. (GETATTR, - * READ, READDIR, and READLINK have their own buffer checks.) if we - * fall below this level, we fail the next operation with NFS4ERR_RESOURCE. - * - * COMPOUND_ERR_SLACK_SPACE - this is the minimum bytes of buffer space - * needed to encode an operation which has failed with NFS4ERR_RESOURCE. - * care is taken to ensure that we never fall below this level for any - * reason. - */ -#define COMPOUND_SLACK_SPACE 140 /* OP_GETFH */ -#define COMPOUND_ERR_SLACK_SPACE 16 /* OP_SETATTR */ - -#define NFSD_LAUNDROMAT_MINTIMEOUT 1 /* seconds */ -#define NFSD_COURTESY_CLIENT_TIMEOUT (24 * 60 * 60) /* seconds */ -#define NFSD_CLIENT_MAX_TRIM_PER_RUN 128 -#define NFS4_CLIENTS_PER_GB 1024 -#define NFSD_DELEGRETURN_TIMEOUT (HZ / 34) /* 30ms */ -#define NFSD_CB_GETATTR_TIMEOUT NFSD_DELEGRETURN_TIMEOUT - -extern int nfsd4_is_junction(struct dentry *dentry); -extern int register_cld_notifier(void); -extern void unregister_cld_notifier(void); -#ifdef CONFIG_NFSD_V4_2_INTER_SSC -extern void nfsd4_ssc_init_umount_work(struct nfsd_net *nn); -#endif - -extern void nfsd4_init_leases_net(struct nfsd_net *nn); - -#else /* CONFIG_NFSD_V4 */ -static inline int nfsd4_is_junction(struct dentry *dentry) -{ - return 0; -} - -static inline void nfsd4_init_leases_net(struct nfsd_net *nn) { }; - -#define register_cld_notifier() 0 -#define unregister_cld_notifier() do { } while(0) - -#endif /* CONFIG_NFSD_V4 */ - #endif /* LINUX_NFSD_NFSD_H */ diff --git a/fs/nfsd/nfserr.h b/fs/nfsd/nfserr.h new file mode 100644 index 000000000000..d1c434b834b2 --- /dev/null +++ b/fs/nfsd/nfserr.h @@ -0,0 +1,159 @@ +/* SPDX-License-Identifier: GPL-2.0 */ +/* + * Pre-xdr'ed nfsd error values and nfsd-internal error codes. + * + * Separated from nfsd.h so that nfsd.h itself does not have to pull + * in <linux/nfs4.h>: the NFS4ERR_* values used below are the only + * reason that include was needed. + */ + +#ifndef LINUX_NFSD_NFSERR_H +#define LINUX_NFSD_NFSERR_H + +#include <linux/nfs.h> +#include <linux/nfs3.h> +#include <linux/nfs4.h> + +/* + * These macros provide pre-xdr'ed values for faster operation. + */ +#define nfs_ok cpu_to_be32(NFS_OK) +#define nfserr_perm cpu_to_be32(NFSERR_PERM) +#define nfserr_noent cpu_to_be32(NFSERR_NOENT) +#define nfserr_io cpu_to_be32(NFSERR_IO) +#define nfserr_nxio cpu_to_be32(NFSERR_NXIO) +#define nfserr_acces cpu_to_be32(NFSERR_ACCES) +#define nfserr_exist cpu_to_be32(NFSERR_EXIST) +#define nfserr_xdev cpu_to_be32(NFS3ERR_XDEV) +#define nfserr_nodev cpu_to_be32(NFSERR_NODEV) +#define nfserr_notdir cpu_to_be32(NFSERR_NOTDIR) +#define nfserr_isdir cpu_to_be32(NFSERR_ISDIR) +#define nfserr_inval cpu_to_be32(NFS3ERR_INVAL) +#define nfserr_fbig cpu_to_be32(NFSERR_FBIG) +#define nfserr_nospc cpu_to_be32(NFSERR_NOSPC) +#define nfserr_rofs cpu_to_be32(NFSERR_ROFS) +#define nfserr_mlink cpu_to_be32(NFS3ERR_MLINK) +#define nfserr_nametoolong cpu_to_be32(NFSERR_NAMETOOLONG) +#define nfserr_notempty cpu_to_be32(NFSERR_NOTEMPTY) +#define nfserr_dquot cpu_to_be32(NFSERR_DQUOT) +#define nfserr_stale cpu_to_be32(NFSERR_STALE) +#define nfserr_remote cpu_to_be32(NFS3ERR_REMOTE) +#define nfserr_wflush cpu_to_be32(NFSERR_WFLUSH) +#define nfserr_badhandle cpu_to_be32(NFS3ERR_BADHANDLE) +#define nfserr_notsync cpu_to_be32(NFS3ERR_NOT_SYNC) +#define nfserr_badcookie cpu_to_be32(NFS3ERR_BAD_COOKIE) +#define nfserr_notsupp cpu_to_be32(NFS3ERR_NOTSUPP) +#define nfserr_toosmall cpu_to_be32(NFS3ERR_TOOSMALL) +#define nfserr_serverfault cpu_to_be32(NFS3ERR_SERVERFAULT) +#define nfserr_badtype cpu_to_be32(NFS3ERR_BADTYPE) +#define nfserr_jukebox cpu_to_be32(NFS3ERR_JUKEBOX) +#define nfserr_delay cpu_to_be32(NFS4ERR_DELAY) +#define nfserr_same cpu_to_be32(NFS4ERR_SAME) +#define nfserr_denied cpu_to_be32(NFS4ERR_DENIED) +#define nfserr_deadlock cpu_to_be32(NFS4ERR_DEADLOCK) +#define nfserr_expired cpu_to_be32(NFS4ERR_EXPIRED) +#define nfserr_bad_cookie cpu_to_be32(NFS4ERR_BAD_COOKIE) +#define nfserr_clid_inuse cpu_to_be32(NFS4ERR_CLID_INUSE) +#define nfserr_stale_clientid cpu_to_be32(NFS4ERR_STALE_CLIENTID) +#define nfserr_resource cpu_to_be32(NFS4ERR_RESOURCE) +#define nfserr_moved cpu_to_be32(NFS4ERR_MOVED) +#define nfserr_nofilehandle cpu_to_be32(NFS4ERR_NOFILEHANDLE) +#define nfserr_minor_vers_mismatch cpu_to_be32(NFS4ERR_MINOR_VERS_MISMATCH) +#define nfserr_share_denied cpu_to_be32(NFS4ERR_SHARE_DENIED) +#define nfserr_stale_stateid cpu_to_be32(NFS4ERR_STALE_STATEID) +#define nfserr_old_stateid cpu_to_be32(NFS4ERR_OLD_STATEID) +#define nfserr_bad_stateid cpu_to_be32(NFS4ERR_BAD_STATEID) +#define nfserr_bad_seqid cpu_to_be32(NFS4ERR_BAD_SEQID) +#define nfserr_symlink cpu_to_be32(NFS4ERR_SYMLINK) +#define nfserr_not_same cpu_to_be32(NFS4ERR_NOT_SAME) +#define nfserr_lock_range cpu_to_be32(NFS4ERR_LOCK_RANGE) +#define nfserr_restorefh cpu_to_be32(NFS4ERR_RESTOREFH) +#define nfserr_attrnotsupp cpu_to_be32(NFS4ERR_ATTRNOTSUPP) +#define nfserr_bad_xdr cpu_to_be32(NFS4ERR_BADXDR) +#define nfserr_openmode cpu_to_be32(NFS4ERR_OPENMODE) +#define nfserr_badowner cpu_to_be32(NFS4ERR_BADOWNER) +#define nfserr_locks_held cpu_to_be32(NFS4ERR_LOCKS_HELD) +#define nfserr_op_illegal cpu_to_be32(NFS4ERR_OP_ILLEGAL) +#define nfserr_grace cpu_to_be32(NFS4ERR_GRACE) +#define nfserr_no_grace cpu_to_be32(NFS4ERR_NO_GRACE) +#define nfserr_reclaim_bad cpu_to_be32(NFS4ERR_RECLAIM_BAD) +#define nfserr_badname cpu_to_be32(NFS4ERR_BADNAME) +#define nfserr_admin_revoked cpu_to_be32(NFS4ERR_ADMIN_REVOKED) +#define nfserr_cb_path_down cpu_to_be32(NFS4ERR_CB_PATH_DOWN) +#define nfserr_locked cpu_to_be32(NFS4ERR_LOCKED) +#define nfserr_wrongsec cpu_to_be32(NFS4ERR_WRONGSEC) +#define nfserr_badiomode cpu_to_be32(NFS4ERR_BADIOMODE) +#define nfserr_badlayout cpu_to_be32(NFS4ERR_BADLAYOUT) +#define nfserr_bad_session_digest cpu_to_be32(NFS4ERR_BAD_SESSION_DIGEST) +#define nfserr_badsession cpu_to_be32(NFS4ERR_BADSESSION) +#define nfserr_badslot cpu_to_be32(NFS4ERR_BADSLOT) +#define nfserr_complete_already cpu_to_be32(NFS4ERR_COMPLETE_ALREADY) +#define nfserr_conn_not_bound_to_session cpu_to_be32(NFS4ERR_CONN_NOT_BOUND_TO_SESSION) +#define nfserr_deleg_already_wanted cpu_to_be32(NFS4ERR_DELEG_ALREADY_WANTED) +#define nfserr_back_chan_busy cpu_to_be32(NFS4ERR_BACK_CHAN_BUSY) +#define nfserr_layouttrylater cpu_to_be32(NFS4ERR_LAYOUTTRYLATER) +#define nfserr_layoutunavailable cpu_to_be32(NFS4ERR_LAYOUTUNAVAILABLE) +#define nfserr_nomatching_layout cpu_to_be32(NFS4ERR_NOMATCHING_LAYOUT) +#define nfserr_recallconflict cpu_to_be32(NFS4ERR_RECALLCONFLICT) +#define nfserr_unknown_layouttype cpu_to_be32(NFS4ERR_UNKNOWN_LAYOUTTYPE) +#define nfserr_seq_misordered cpu_to_be32(NFS4ERR_SEQ_MISORDERED) +#define nfserr_sequence_pos cpu_to_be32(NFS4ERR_SEQUENCE_POS) +#define nfserr_req_too_big cpu_to_be32(NFS4ERR_REQ_TOO_BIG) +#define nfserr_rep_too_big cpu_to_be32(NFS4ERR_REP_TOO_BIG) +#define nfserr_rep_too_big_to_cache cpu_to_be32(NFS4ERR_REP_TOO_BIG_TO_CACHE) +#define nfserr_retry_uncached_rep cpu_to_be32(NFS4ERR_RETRY_UNCACHED_REP) +#define nfserr_unsafe_compound cpu_to_be32(NFS4ERR_UNSAFE_COMPOUND) +#define nfserr_too_many_ops cpu_to_be32(NFS4ERR_TOO_MANY_OPS) +#define nfserr_op_not_in_session cpu_to_be32(NFS4ERR_OP_NOT_IN_SESSION) +#define nfserr_hash_alg_unsupp cpu_to_be32(NFS4ERR_HASH_ALG_UNSUPP) +#define nfserr_clientid_busy cpu_to_be32(NFS4ERR_CLIENTID_BUSY) +#define nfserr_pnfs_io_hole cpu_to_be32(NFS4ERR_PNFS_IO_HOLE) +#define nfserr_seq_false_retry cpu_to_be32(NFS4ERR_SEQ_FALSE_RETRY) +#define nfserr_bad_high_slot cpu_to_be32(NFS4ERR_BAD_HIGH_SLOT) +#define nfserr_deadsession cpu_to_be32(NFS4ERR_DEADSESSION) +#define nfserr_encr_alg_unsupp cpu_to_be32(NFS4ERR_ENCR_ALG_UNSUPP) +#define nfserr_pnfs_no_layout cpu_to_be32(NFS4ERR_PNFS_NO_LAYOUT) +#define nfserr_not_only_op cpu_to_be32(NFS4ERR_NOT_ONLY_OP) +#define nfserr_wrong_cred cpu_to_be32(NFS4ERR_WRONG_CRED) +#define nfserr_wrong_type cpu_to_be32(NFS4ERR_WRONG_TYPE) +#define nfserr_dirdeleg_unavail cpu_to_be32(NFS4ERR_DIRDELEG_UNAVAIL) +#define nfserr_reject_deleg cpu_to_be32(NFS4ERR_REJECT_DELEG) +#define nfserr_returnconflict cpu_to_be32(NFS4ERR_RETURNCONFLICT) +#define nfserr_deleg_revoked cpu_to_be32(NFS4ERR_DELEG_REVOKED) +#define nfserr_partner_notsupp cpu_to_be32(NFS4ERR_PARTNER_NOTSUPP) +#define nfserr_partner_no_auth cpu_to_be32(NFS4ERR_PARTNER_NO_AUTH) +#define nfserr_union_notsupp cpu_to_be32(NFS4ERR_UNION_NOTSUPP) +#define nfserr_offload_denied cpu_to_be32(NFS4ERR_OFFLOAD_DENIED) +#define nfserr_wrong_lfs cpu_to_be32(NFS4ERR_WRONG_LFS) +#define nfserr_badlabel cpu_to_be32(NFS4ERR_BADLABEL) +#define nfserr_file_open cpu_to_be32(NFS4ERR_FILE_OPEN) +#define nfserr_xattr2big cpu_to_be32(NFS4ERR_XATTR2BIG) +#define nfserr_noxattr cpu_to_be32(NFS4ERR_NOXATTR) + +/* + * Error codes for internal use. These are based at an impossible + * nfsstat4 value so that, once converted to be32, they cannot conflict + * with any value defined by the protocol (compare the nlm__int__* codes + * in fs/lockd/lockd.h). + */ +enum { +/* end-of-file indicator in readdir */ + NFSERR_EOF = 30000, +#define nfserr_eof cpu_to_be32(NFSERR_EOF) + +/* replay detected */ + NFSERR_REPLAY_ME, +#define nfserr_replay_me cpu_to_be32(NFSERR_REPLAY_ME) + +/* nfs41 replay detected */ + NFSERR_REPLAY_CACHE, +#define nfserr_replay_cache cpu_to_be32(NFSERR_REPLAY_CACHE) + +/* symlink found where dir expected - handled differently to + * other symlink found errors by NFSv3. + */ + NFSERR_SYMLINK_NOT_DIR, +#define nfserr_symlink_not_dir cpu_to_be32(NFSERR_SYMLINK_NOT_DIR) +}; + +#endif /* LINUX_NFSD_NFSERR_H */ diff --git a/fs/nfsd/nfsfh.c b/fs/nfsd/nfsfh.c index c7c60c35bdfc..2bd6907f443f 100644 --- a/fs/nfsd/nfsfh.c +++ b/fs/nfsd/nfsfh.c @@ -9,10 +9,12 @@ */ #include <linux/exportfs.h> +#include <linux/nfs3.h> #include <linux/sunrpc/svcauth_gss.h> #include <crypto/utils.h> #include "nfsd.h" +#include "nfserr.h" #include "netns.h" #include "stats.h" #include "vfs.h" @@ -334,6 +336,8 @@ static __be32 nfsd_set_fh_dentry(struct svc_rqst *rqstp, struct net *net, } switch (fhp->fh_maxsize) { + case NFSD_FHSIZE_UNSPEC: + break; case NFS4_FHSIZE: if (dentry->d_sb->s_export_op->flags & EXPORT_OP_NOATOMIC_ATTR) fhp->fh_no_atomic_attr = true; @@ -782,35 +786,54 @@ __be32 fh_getattr(const struct svc_fh *fhp, struct kstat *stat) AT_STATX_SYNC_AS_STAT)); } -/** - * fh_fill_pre_attrs - Fill in pre-op attributes - * @fhp: file handle to be updated - * - */ -__be32 __must_check fh_fill_pre_attrs(struct svc_fh *fhp) +static __be32 __must_check __fh_fill_pre_attrs(struct svc_fh *fhp) { bool v4 = (fhp->fh_maxsize == NFS4_FHSIZE); - struct kstat stat; __be32 err; if (fhp->fh_no_wcc || fhp->fh_pre_saved) return nfs_ok; - err = fh_getattr(fhp, &stat); + err = fh_getattr(fhp, &fhp->fh_post_attr); if (err) return err; if (v4) - fhp->fh_pre_change = nfsd4_change_attribute(&stat); + fhp->fh_pre_change = fhp->fh_post_change = + nfsd4_change_attribute(&fhp->fh_post_attr); - fhp->fh_pre_mtime = stat.mtime; - fhp->fh_pre_ctime = stat.ctime; - fhp->fh_pre_size = stat.size; + fhp->fh_pre_mtime = fhp->fh_post_attr.mtime; + fhp->fh_pre_ctime = fhp->fh_post_attr.ctime; + fhp->fh_pre_size = fhp->fh_post_attr.size; fhp->fh_pre_saved = true; return nfs_ok; } /** + * fh_fill_pre_attrs - Fill in pre-op attributes + * @fhp: file handle to be updated + * + * Post-op attrs are filled and pre-op attrs are copied + * from there. The post-op attrs can later be replaced by + * fh_fill_post_attrs() or activated by fh_fill_post_noop(). + * + * The inode must be locked. + * + * Returns: error from vfs_getattr() which must be checked. + */ +__be32 __must_check fh_fill_pre_attrs(struct svc_fh *fhp) +{ + lockdep_assert_held_write(&fhp->fh_dentry->d_inode->i_rwsem); + return __fh_fill_pre_attrs(fhp); +} + +__be32 __must_check fh_fill_pre_attrs_unlocked(struct svc_fh *fhp) +{ + fhp->fh_no_atomic_attr = true; + return __fh_fill_pre_attrs(fhp); +} + +/** * fh_fill_post_attrs - Fill in post-op attributes * @fhp: file handle to be updated * @@ -826,6 +849,9 @@ __be32 fh_fill_post_attrs(struct svc_fh *fhp) if (fhp->fh_post_saved) printk("nfsd: inode locked twice during operation.\n"); + if (!fhp->fh_no_atomic_attr) + lockdep_assert_held_write(&fhp->fh_dentry->d_inode->i_rwsem); + err = fh_getattr(fhp, &fhp->fh_post_attr); if (err) return err; @@ -837,29 +863,6 @@ __be32 fh_fill_post_attrs(struct svc_fh *fhp) return nfs_ok; } -/** - * fh_fill_both_attrs - Fill pre-op and post-op attributes - * @fhp: file handle to be updated - * - * This is used when the directory wasn't changed, but wcc attributes - * are needed anyway. - */ -__be32 __must_check fh_fill_both_attrs(struct svc_fh *fhp) -{ - __be32 err; - - err = fh_fill_post_attrs(fhp); - if (err) - return err; - - fhp->fh_pre_change = fhp->fh_post_change; - fhp->fh_pre_mtime = fhp->fh_post_attr.mtime; - fhp->fh_pre_ctime = fhp->fh_post_attr.ctime; - fhp->fh_pre_size = fhp->fh_post_attr.size; - fhp->fh_pre_saved = true; - return nfs_ok; -} - /* * Release a file handle. */ diff --git a/fs/nfsd/nfsfh.h b/fs/nfsd/nfsfh.h index cdeb5eea65a8..7d8e3f015307 100644 --- a/fs/nfsd/nfsfh.h +++ b/fs/nfsd/nfsfh.h @@ -246,6 +246,20 @@ fh_copy_shallow(struct knfsd_fh *dst, const struct knfsd_fh *src) memcpy(&dst->fh_raw, &src->fh_raw, src->fh_size); } +#define NFSD_FHSIZE_UNSPEC 0 + +/** + * fh_init - Prepare a file handle for fh_compose() or fh_verify() + * @fhp: File handle to initialize + * @maxsize: Largest file handle, in bytes, to build in @fhp + * + * @maxsize bounds the handle fh_compose() may build: NFS_FHSIZE, + * NFS3_FHSIZE, and NFS4_FHSIZE additionally select version-specific + * handling in fh_verify(). Callers that only verify an incoming + * handle pass NFSD_FHSIZE_UNSPEC, which cannot be composed. + * + * Return: @fhp + */ static __inline__ struct svc_fh * fh_init(struct svc_fh *fhp, int maxsize) { @@ -337,6 +351,18 @@ static inline void fh_clear_pre_post_attrs(struct svc_fh *fhp) u64 nfsd4_change_attribute(const struct kstat *stat); __be32 __must_check fh_fill_pre_attrs(struct svc_fh *fhp); +__be32 __must_check fh_fill_pre_attrs_unlocked(struct svc_fh *fhp); __be32 fh_fill_post_attrs(struct svc_fh *fhp); -__be32 __must_check fh_fill_both_attrs(struct svc_fh *fhp); + +/** + * fh_fill_post_noop - Copy pre attrs to post attrs + * @fhp: file handle to be updated + * + * This is used when the directory wasn't changed, but wcc attributes + * are needed anyway. + */ +static inline void fh_fill_post_noop(struct svc_fh *fhp) +{ + fhp->fh_post_saved = true; +} #endif /* _LINUX_NFSD_NFSFH_H */ diff --git a/fs/nfsd/nfsproc.c b/fs/nfsd/nfsproc.c index e2b5f8a241be..09d360839825 100644 --- a/fs/nfsd/nfsproc.c +++ b/fs/nfsd/nfsproc.c @@ -10,6 +10,7 @@ #include "cache.h" #include "xdr.h" #include "vfs.h" +#include "nfserr.h" #include "trace.h" #define NFSDDBG_FACILITY NFSDDBG_PROC @@ -39,6 +40,21 @@ static __be32 nfsd_map_status(__be32 status) return status; } +/* + * Because NFSv2 does not have an NFSERR_SYMLINK, Solaris returns + * NFSERR_ISDIR when the target of a READ or WRITE is any object + * that is not a regular file. + */ +static __be32 nfsd_map_io_status(__be32 status) +{ + switch (status) { + case nfserr_symlink: + case nfserr_wrong_type: + return nfserr_isdir; + } + return nfsd_map_status(status); +} + static __be32 nfsd_proc_null(struct svc_rqst *rqstp) { @@ -237,7 +253,7 @@ nfsd_proc_read(struct svc_rqst *rqstp) resp->status = fh_getattr(&resp->fh, &resp->stat); else if (resp->status == nfserr_jukebox) set_bit(RQ_DROPME, &rqstp->rq_flags); - resp->status = nfsd_map_status(resp->status); + resp->status = nfsd_map_io_status(resp->status); return rpc_success; } @@ -265,12 +281,12 @@ nfsd_proc_write(struct svc_rqst *rqstp) fh_copy(&resp->fh, &argp->fh); resp->status = nfsd_write(rqstp, &resp->fh, argp->offset, - &argp->payload, &cnt, NFS_DATA_SYNC, NULL); + &argp->payload, &cnt, IOCB_DSYNC, NULL); if (resp->status == nfs_ok) resp->status = fh_getattr(&resp->fh, &resp->stat); else if (resp->status == nfserr_jukebox) set_bit(RQ_DROPME, &rqstp->rq_flags); - resp->status = nfsd_map_status(resp->status); + resp->status = nfsd_map_io_status(resp->status); return rpc_success; } @@ -291,6 +307,7 @@ nfsd_proc_create(struct svc_rqst *rqstp) struct nfsd_attrs attrs = { .na_iattr = attr, }; + struct svc_export *exp; struct inode *inode; struct dentry *dchild; int type, mode; @@ -319,8 +336,22 @@ nfsd_proc_create(struct svc_rqst *rqstp) resp->status = nfserrno(PTR_ERR(dchild)); goto out_write; } + /* + * If name exists we need to check for mountpoints + */ + exp = exp_get(dirfhp->fh_export); + if (d_is_reg(dchild) && + unlikely(nfsd_mountpoint(dchild, exp))) { + resp->status = nfsd_cross_mnt(rqstp, &dchild, &exp); + if (resp->status != nfs_ok) { + exp_put(exp); + goto out_unlock; + } + } + fh_init(newfhp, NFS_FHSIZE); - resp->status = fh_compose(newfhp, dirfhp->fh_export, dchild, dirfhp); + resp->status = fh_compose(newfhp, exp, dchild, dirfhp); + exp_put(exp); if (!resp->status && d_really_is_negative(dchild)) resp->status = nfserr_noent; if (resp->status) { diff --git a/fs/nfsd/nfssvc.c b/fs/nfsd/nfssvc.c index 2edf716ea022..c04ef9d180ce 100644 --- a/fs/nfsd/nfssvc.c +++ b/fs/nfsd/nfssvc.c @@ -25,7 +25,10 @@ #include <net/addrconf.h> #include <net/ipv6.h> #include <net/net_namespace.h> + #include "nfsd.h" +#include "nfserr.h" +#include "nfs4ctl.h" #include "cache.h" #include "vfs.h" #include "netns.h" @@ -204,6 +207,11 @@ int nfsd_minorversion(struct nfsd_net *nn, u32 minorversion, enum vers_op change return 0; } +bool nfsd_v4client(struct svc_rqst *rqstp) +{ + return rqstp && rqstp->rq_prog == NFS_PROGRAM && rqstp->rq_vers == 4; +} + bool nfsd_net_try_get(struct net *net) __must_hold(rcu) { struct nfsd_net *nn = net_generic(net, nfsd_net_id); diff --git a/fs/nfsd/nfsxdr.c b/fs/nfsd/nfsxdr.c index 019f0cc971a7..0961c13d6ab1 100644 --- a/fs/nfsd/nfsxdr.c +++ b/fs/nfsd/nfsxdr.c @@ -8,6 +8,7 @@ #include <linux/filelock.h> #include "vfs.h" +#include "nfserr.h" #include "xdr.h" #include "auth.h" diff --git a/fs/nfsd/state.h b/fs/nfsd/state.h index 2d00a411c663..cd9294f024bb 100644 --- a/fs/nfsd/state.h +++ b/fs/nfsd/state.h @@ -45,6 +45,25 @@ #include "nfsfh.h" #include "nfsd.h" +/* + * Before processing a COMPOUND operation, we have to check that there + * is enough space in the buffer for XDR encode to succeed. otherwise, + * we might process an operation with side effects, and be unable to + * tell the client that the operation succeeded. + * + * COMPOUND_ERR_SLACK_SPACE - this is the minimum bytes of buffer space + * needed to encode an operation which has failed with NFS4ERR_RESOURCE. + * care is taken to ensure that we never fall below this level for any + * reason. + */ +#define COMPOUND_ERR_SLACK_SPACE 16 /* OP_SETATTR */ + +#define NFSD_LAUNDROMAT_MINTIMEOUT 1 /* seconds */ +#define NFSD_CLIENT_MAX_TRIM_PER_RUN 128 +#define NFS4_CLIENTS_PER_GB 1024 +#define NFSD_DELEGRETURN_TIMEOUT (HZ / 34) /* 30ms */ +#define NFSD_CB_GETATTR_TIMEOUT NFSD_DELEGRETURN_TIMEOUT + typedef struct { u32 cl_boot; u32 cl_id; @@ -126,10 +145,13 @@ struct nfs4_stid { #define SC_TYPE_COPY BIT(4) unsigned short sc_type; -/* nn->deleg_lock protects sc_status for delegation stateids. - * ->cl_lock protects sc_status for open and lock stateids. - * ->st_mutex also protect sc_status for open stateids. - * ->ls_lock protects sc_status for layout stateids. +/* + * nn->deleg_lock protects sc_status for hashed delegation stateids. + * ->cl_lock protects the bits set as one is disposed of + * (SC_STATUS_CLOSED, SC_STATUS_FREEABLE, SC_STATUS_FREED) and + * sc_status for open and lock stateids. ->st_mutex also protects + * sc_status for open stateids. ->ls_lock protects sc_status for + * layout stateids. */ /* * For an open stateid kept around *only* to process close replays. @@ -249,6 +271,7 @@ struct nfsd4_cb_notify { struct nfsd_notify_event *ncn_evt[NOTIFY4_EVENT_QUEUE_SIZE]; // list of events struct page *ncn_pages[NOTIFY4_PAGE_ARRAY_SIZE]; // for encoding struct notify4 *ncn_nf; // array of notify4's to be sent + u32 *ncn_masks; // host-order notify_mask backing for ncn_nf[] bool ncn_encode_err; // did encoding fail? struct nfsd4_callback ncn_cb; // notify4 callback }; @@ -270,7 +293,8 @@ struct nfsd4_cb_notify { * If the server attempts to recall a delegation and the client doesn't do so * before a timeout, the server may also revoke the delegation. In that case, * the object will either be destroyed (v4.0) or moved to a per-client list of - * revoked delegations (v4.1+). + * revoked delegations (v4.1+). A v4.1+ client that rejects the recall holds + * no record of the delegation, so the object is destroyed rather than listed. * * This object is a superset of the nfs4_stid. */ @@ -286,9 +310,19 @@ struct nfs4_delegation { int dl_retries; struct nfsd4_callback dl_recall; bool dl_recalled; + bool dl_recall_rejected; bool dl_written; bool dl_setattr; + /* Forward-channel slot that carried the granting request */ + struct { + u32 sessionid_seq; + u32 slotid; + u32 seqid; + bool valid; + bool retired_at_send; + } dl_recall_grant; + union { /* for CB_GETATTR */ struct nfs4_cb_fattr dl_cb_fattr; @@ -599,6 +633,8 @@ struct nfs4_client { unsigned int cl_state; atomic_t cl_delegs_in_recall; + /* Length of cl_delegations, updated under nn->deleg_lock */ + unsigned int cl_deleg_count; struct nfsd4_cb_recall_any *cl_ra; time64_t cl_ra_time; diff --git a/fs/nfsd/trace.h b/fs/nfsd/trace.h index 7d7a1483109a..2ae7f150a72c 100644 --- a/fs/nfsd/trace.h +++ b/fs/nfsd/trace.h @@ -13,7 +13,6 @@ #include <linux/sunrpc/xprt.h> #include <trace/misc/fs.h> #include <trace/misc/fsnotify.h> -#include <trace/misc/nfs.h> #include <trace/misc/sunrpc.h> #include "export.h" @@ -790,6 +789,20 @@ TRACE_EVENT(nfsd_stateowner_replay, __entry->opnum, __entry->status) ); +#define show_nfs4_seq4_status(x) \ + __print_flags(x, "|", \ + { SEQ4_STATUS_CB_PATH_DOWN, "CB_PATH_DOWN" }, \ + { SEQ4_STATUS_CB_GSS_CONTEXTS_EXPIRING, "CB_GSS_CONTEXTS_EXPIRING" }, \ + { SEQ4_STATUS_CB_GSS_CONTEXTS_EXPIRED, "CB_GSS_CONTEXTS_EXPIRED" }, \ + { SEQ4_STATUS_EXPIRED_ALL_STATE_REVOKED, "EXPIRED_ALL_STATE_REVOKED" }, \ + { SEQ4_STATUS_EXPIRED_SOME_STATE_REVOKED, "EXPIRED_SOME_STATE_REVOKED" }, \ + { SEQ4_STATUS_ADMIN_STATE_REVOKED, "ADMIN_STATE_REVOKED" }, \ + { SEQ4_STATUS_RECALLABLE_STATE_REVOKED, "RECALLABLE_STATE_REVOKED" }, \ + { SEQ4_STATUS_LEASE_MOVED, "LEASE_MOVED" }, \ + { SEQ4_STATUS_RESTART_RECLAIM_NEEDED, "RESTART_RECLAIM_NEEDED" }, \ + { SEQ4_STATUS_CB_PATH_DOWN_SESSION, "CB_PATH_DOWN_SESSION" }, \ + { SEQ4_STATUS_BACKCHANNEL_FAULT, "BACKCHANNEL_FAULT" }) + TRACE_EVENT_CONDITION(nfsd_seq4_status, TP_PROTO( const struct svc_rqst *rqstp, @@ -1694,6 +1707,14 @@ TRACE_EVENT(nfsd_cb_setup_err, /* Not a real opcode, but there is no 0 operation. */ #define _CB_NULL 0 +TRACE_DEFINE_ENUM(OP_CB_GETATTR); +TRACE_DEFINE_ENUM(OP_CB_RECALL); +TRACE_DEFINE_ENUM(OP_CB_LAYOUTRECALL); +TRACE_DEFINE_ENUM(OP_CB_RECALL_ANY); +TRACE_DEFINE_ENUM(OP_CB_NOTIFY); +TRACE_DEFINE_ENUM(OP_CB_NOTIFY_LOCK); +TRACE_DEFINE_ENUM(OP_CB_OFFLOAD); + #define show_nfsd_cb_opcode(val) \ __print_symbolic(val, \ { _CB_NULL, "CB_NULL" }, \ @@ -1918,6 +1939,28 @@ TRACE_EVENT(nfsd_cb_offload, __entry->fh_hash, __entry->count, __entry->status) ); +TRACE_DEFINE_ENUM(RCA4_TYPE_MASK_RDATA_DLG); +TRACE_DEFINE_ENUM(RCA4_TYPE_MASK_WDATA_DLG); +TRACE_DEFINE_ENUM(RCA4_TYPE_MASK_DIR_DLG); +TRACE_DEFINE_ENUM(RCA4_TYPE_MASK_FILE_LAYOUT); +TRACE_DEFINE_ENUM(RCA4_TYPE_MASK_BLK_LAYOUT); +TRACE_DEFINE_ENUM(RCA4_TYPE_MASK_OBJ_LAYOUT_MIN); +TRACE_DEFINE_ENUM(RCA4_TYPE_MASK_OBJ_LAYOUT_MAX); +TRACE_DEFINE_ENUM(RCA4_TYPE_MASK_OTHER_LAYOUT_MIN); +TRACE_DEFINE_ENUM(RCA4_TYPE_MASK_OTHER_LAYOUT_MAX); + +#define show_rca_mask(x) \ + __print_flags(x, "|", \ + { BIT(RCA4_TYPE_MASK_RDATA_DLG), "RDATA_DLG" }, \ + { BIT(RCA4_TYPE_MASK_WDATA_DLG), "WDATA_DLG" }, \ + { BIT(RCA4_TYPE_MASK_DIR_DLG), "DIR_DLG" }, \ + { BIT(RCA4_TYPE_MASK_FILE_LAYOUT), "FILE_LAYOUT" }, \ + { BIT(RCA4_TYPE_MASK_BLK_LAYOUT), "BLK_LAYOUT" }, \ + { BIT(RCA4_TYPE_MASK_OBJ_LAYOUT_MIN), "OBJ_LAYOUT_MIN" }, \ + { BIT(RCA4_TYPE_MASK_OBJ_LAYOUT_MAX), "OBJ_LAYOUT_MAX" }, \ + { BIT(RCA4_TYPE_MASK_OTHER_LAYOUT_MIN), "OTHER_LAYOUT_MIN" }, \ + { BIT(RCA4_TYPE_MASK_OTHER_LAYOUT_MAX), "OTHER_LAYOUT_MAX" }) + TRACE_EVENT(nfsd_cb_recall_any, TP_PROTO( const struct nfsd4_cb_recall_any *ra diff --git a/fs/nfsd/vfs.c b/fs/nfsd/vfs.c index 8923a9910a08..4789f2ec2078 100644 --- a/fs/nfsd/vfs.c +++ b/fs/nfsd/vfs.c @@ -34,15 +34,14 @@ #include <linux/sunrpc/xdr.h> #include <linux/fileattr.h> -#include "xdr3.h" - #ifdef CONFIG_NFSD_V4 #include "acl.h" #include "idmap.h" -#include "xdr4.h" #endif /* CONFIG_NFSD_V4 */ #include "nfsd.h" +#include "nfserr.h" +#include "nfs4ctl.h" #include "netns.h" #include "stats.h" #include "vfs.h" @@ -63,7 +62,7 @@ u64 nfsd_io_cache_write __read_mostly = NFSD_IO_BUFFERED; * it's an error we don't expect, log it once and return nfserr_io. */ __be32 -nfserrno (int errno) +nfserrno(int errno) { static struct { __be32 nfserr; @@ -107,6 +106,8 @@ nfserrno (int errno) { nfserr_perm, -ENOKEY }, { nfserr_no_grace, -ENOGRACE}, { nfserr_io, -EBADMSG }, + { nfserr_symlink, -ELOOP }, + { nfserr_wrong_type, -EFTYPE }, }; int i; @@ -118,15 +119,15 @@ nfserrno (int errno) return nfserr_io; } -/* - * Called from nfsd_lookup and encode_dirent. Check if we have crossed +/* + * Called from nfsd_lookup and encode_dirent. Check if we have crossed * a mount point. - * Returns -EAGAIN or -ETIMEDOUT leaving *dpp and *expp unchanged, + * Returns an nfs error leaving *dpp and *expp unchanged, * or nfs_ok having possibly changed *dpp and *expp */ -int -nfsd_cross_mnt(struct svc_rqst *rqstp, struct dentry **dpp, - struct svc_export **expp) +__be32 +nfsd_cross_mnt(struct svc_rqst *rqstp, struct dentry **dpp, + struct svc_export **expp) { struct svc_export *exp = *expp, *exp2 = NULL; struct dentry *dentry = *dpp; @@ -134,6 +135,7 @@ nfsd_cross_mnt(struct svc_rqst *rqstp, struct dentry **dpp, .dentry = dget(dentry)}; unsigned int follow_flags = 0; int err = 0; + __be32 nfserr = nfs_ok; if (exp->ex_flags & NFSEXP_CROSSMOUNT) follow_flags = LOOKUP_AUTOMOUNT; @@ -163,23 +165,28 @@ nfsd_cross_mnt(struct svc_rqst *rqstp, struct dentry **dpp, err = 0; } else if (nfsd_v4client(rqstp) || (exp->ex_flags & NFSEXP_CROSSMOUNT) || EX_NOHIDE(exp2)) { - /* successfully crossed mount point */ - /* - * This is subtle: path.dentry is *not* on path.mnt - * at this point. The only reason we are safe is that - * original mnt is pinned down by exp, so we should - * put path *before* putting exp - */ - *dpp = path.dentry; - path.dentry = dentry; - *expp = exp2; - exp2 = exp; + nfserr = check_nfsd_access(exp, rqstp); + if (nfserr == nfs_ok) { + /* successfully crossed mount point */ + /* + * This is subtle: path.dentry is *not* on path.mnt + * at this point. The only reason we are safe is that + * original mnt is pinned down by exp, so we should + * put path *before* putting exp + */ + *dpp = path.dentry; + path.dentry = dentry; + *expp = exp2; + exp2 = exp; + } } out: path_put(&path); if (exp2) exp_put(exp2); - return err; + if (nfserr) + return nfserr; + return nfserrno(err); } static void follow_to_parent(struct path *path) @@ -277,10 +284,12 @@ nfsd_lookup_dentry(struct svc_rqst *rqstp, struct svc_fh *fhp, if (IS_ERR(dentry)) goto out_nfserr; if (nfsd_mountpoint(dentry, exp)) { - host_err = nfsd_cross_mnt(rqstp, &dentry, &exp); - if (host_err) { + __be32 nfserr = nfsd_cross_mnt(rqstp, &dentry, &exp); + + if (nfserr) { dput(dentry); - goto out_nfserr; + exp_put(exp); + return nfserr; } } } @@ -327,9 +336,6 @@ nfsd_lookup(struct svc_rqst *rqstp, struct svc_fh *fhp, const char *name, err = nfsd_lookup_dentry(rqstp, fhp, name, len, &exp, &dentry); if (err) return err; - err = check_nfsd_access(exp, rqstp, false); - if (err) - goto out; /* * Note: we compose the file handle now, but as the * dentry may be negative, it may need to be updated. @@ -337,15 +343,26 @@ nfsd_lookup(struct svc_rqst *rqstp, struct svc_fh *fhp, const char *name, err = fh_compose(resfh, exp, dentry, fhp); if (!err && d_really_is_negative(dentry)) err = nfserr_noent; -out: + dput(dentry); exp_put(exp); return err; } -static void -commit_reset_write_verifier(struct nfsd_net *nn, struct svc_rqst *rqstp, - int err) +/** + * nfsd_maybe_reset_write_verifier - Reset the write verifier after an I/O error + * @nn: nfsd namespace holding the write verifier + * @rqstp: RPC transaction context + * @err: errno reported by the failed operation + * + * A write verifier reset tells clients that unstable data the server has + * already acknowledged might have been lost. Client response is to resend + * in-flight dirty data. + * + * Context: Process context. + */ +void nfsd_maybe_reset_write_verifier(struct nfsd_net *nn, + struct svc_rqst *rqstp, int err) { switch (err) { case -EAGAIN: @@ -682,56 +699,67 @@ int nfsd4_is_junction(struct dentry *dentry) return 1; } -static struct nfsd4_compound_state *nfsd4_get_cstate(struct svc_rqst *rqstp) +/** + * nfsd_clone_file_range - Clone a range of one file into another + * @src: file the range is cloned from + * @src_pos: offset in @src where the source range begins + * @dst: file the range is cloned into + * @dst_pos: offset in @dst where the destination range begins + * @count: length of the range, or zero to clone through end-of-file + * @since: receives @dst's writeback error state, sampled before the clone + * + * A caller that has to place the cloned data on durable storage passes + * @since to nfsd_clone_sync_range() once this call succeeds. Sampling + * happens here because a writeback error raised by the clone's own + * dirty pages has to fall inside the sampled interval. + * + * Context: Process context. + * Return: zero on success, or a negative errno + */ +int nfsd_clone_file_range(struct file *src, u64 src_pos, struct file *dst, + u64 dst_pos, u64 count, errseq_t *since) { - return &((struct nfsd4_compoundres *)rqstp->rq_resp)->cstate; + loff_t cloned; + + *since = READ_ONCE(dst->f_wb_err); + cloned = vfs_clone_file_range(src, src_pos, dst, dst_pos, count, 0); + if (cloned < 0) + return cloned; + if (count && cloned != count) + return -EINVAL; + return 0; } -__be32 nfsd4_clone_file_range(struct svc_rqst *rqstp, - struct nfsd_file *nf_src, u64 src_pos, - struct nfsd_file *nf_dst, u64 dst_pos, - u64 count, bool sync) +/** + * nfsd_clone_sync_range - Commit a cloned range to durable storage + * @src: file the range was cloned from, whose metadata is committed too + * @dst: file the range was cloned into + * @dst_pos: offset in @dst where the cloned range begins + * @count: length of the range, or zero if the clone ran to end-of-file + * @since: @dst's writeback error state as sampled by + * nfsd_clone_file_range() + * + * Context: Process context. + * Return: zero on success, or a negative errno + */ +int nfsd_clone_sync_range(struct file *src, struct file *dst, u64 dst_pos, + u64 count, errseq_t since) { - struct file *src = nf_src->nf_file; - struct file *dst = nf_dst->nf_file; - errseq_t since; - loff_t cloned; - __be32 ret = 0; + loff_t dst_end = count ? dst_pos + count - 1 : LLONG_MAX; + int status; - since = READ_ONCE(dst->f_wb_err); - cloned = vfs_clone_file_range(src, src_pos, dst, dst_pos, count, 0); - if (cloned < 0) { - ret = nfserrno(cloned); - goto out_err; - } - if (count && cloned != count) { - ret = nfserrno(-EINVAL); - goto out_err; - } - if (sync) { - loff_t dst_end = count ? dst_pos + count - 1 : LLONG_MAX; - int status = vfs_fsync_range(dst, dst_pos, dst_end, 0); - - if (!status) - status = filemap_check_wb_err(dst->f_mapping, since); - if (!status) - status = commit_inode_metadata(file_inode(src)); - if (status < 0) { - struct nfsd_net *nn = net_generic(nf_dst->nf_net, - nfsd_net_id); - - trace_nfsd_clone_file_range_err(rqstp, - &nfsd4_get_cstate(rqstp)->save_fh, - src_pos, - &nfsd4_get_cstate(rqstp)->current_fh, - dst_pos, - count, status); - commit_reset_write_verifier(nn, rqstp, status); - ret = nfserrno(status); - } + status = vfs_fsync_range(dst, dst_pos, dst_end, 0); + if (!status) + status = filemap_check_wb_err(dst->f_mapping, since); + if (!status) { + /* + * A reflink marks extents shared in the source inode too, + * so the source's metadata has to reach durable storage + * even though its data is untouched. + */ + status = commit_inode_metadata(file_inode(src)); } -out_err: - return ret; + return status; } ssize_t nfsd_copy_file_range(struct file *src, u64 src_pos, struct file *dst, @@ -773,64 +801,21 @@ __be32 nfsd4_vfs_fallocate(struct svc_rqst *rqstp, struct svc_fh *fhp, } #endif /* defined(CONFIG_NFSD_V4) */ -/* - * Check server access rights to a file system object +/** + * nfsd_access - Check caller's access rights to a file system object + * @rqstp: RPC transaction context + * @fhp: target NFS filehandle + * @maps: tables mapping on-the-wire access bits to NFSD_MAY flags + * @access: requested access bits on entry, permitted bits on return + * @supported: optional output of the access bits the server supports + * + * Return: nfs_ok on success, otherwise an nfserr status code */ -struct accessmap { - u32 access; - int how; -}; -static struct accessmap nfs3_regaccess[] = { - { NFS3_ACCESS_READ, NFSD_MAY_READ }, - { NFS3_ACCESS_EXECUTE, NFSD_MAY_EXEC }, - { NFS3_ACCESS_MODIFY, NFSD_MAY_WRITE|NFSD_MAY_TRUNC }, - { NFS3_ACCESS_EXTEND, NFSD_MAY_WRITE }, - -#ifdef CONFIG_NFSD_V4 - { NFS4_ACCESS_XAREAD, NFSD_MAY_READ }, - { NFS4_ACCESS_XAWRITE, NFSD_MAY_WRITE }, - { NFS4_ACCESS_XALIST, NFSD_MAY_READ }, -#endif - - { 0, 0 } -}; - -static struct accessmap nfs3_diraccess[] = { - { NFS3_ACCESS_READ, NFSD_MAY_READ }, - { NFS3_ACCESS_LOOKUP, NFSD_MAY_EXEC }, - { NFS3_ACCESS_MODIFY, NFSD_MAY_EXEC|NFSD_MAY_WRITE|NFSD_MAY_TRUNC}, - { NFS3_ACCESS_EXTEND, NFSD_MAY_EXEC|NFSD_MAY_WRITE }, - { NFS3_ACCESS_DELETE, NFSD_MAY_REMOVE }, - -#ifdef CONFIG_NFSD_V4 - { NFS4_ACCESS_XAREAD, NFSD_MAY_READ }, - { NFS4_ACCESS_XAWRITE, NFSD_MAY_WRITE }, - { NFS4_ACCESS_XALIST, NFSD_MAY_READ }, -#endif - - { 0, 0 } -}; - -static struct accessmap nfs3_anyaccess[] = { - /* Some clients - Solaris 2.6 at least, make an access call - * to the server to check for access for things like /dev/null - * (which really, the server doesn't care about). So - * We provide simple access checking for them, looking - * mainly at mode bits, and we make sure to ignore read-only - * filesystem checks - */ - { NFS3_ACCESS_READ, NFSD_MAY_READ }, - { NFS3_ACCESS_EXECUTE, NFSD_MAY_EXEC }, - { NFS3_ACCESS_MODIFY, NFSD_MAY_WRITE|NFSD_MAY_LOCAL_ACCESS }, - { NFS3_ACCESS_EXTEND, NFSD_MAY_WRITE|NFSD_MAY_LOCAL_ACCESS }, - - { 0, 0 } -}; - -__be32 -nfsd_access(struct svc_rqst *rqstp, struct svc_fh *fhp, u32 *access, u32 *supported) +__be32 nfsd_access(struct svc_rqst *rqstp, struct svc_fh *fhp, + const struct nfsd_access_maps *maps, + u32 *access, u32 *supported) { - struct accessmap *map; + const struct nfsd_access_map *map; struct svc_export *export; struct dentry *dentry; u32 query, result = 0, sresult = 0; @@ -844,12 +829,11 @@ nfsd_access(struct svc_rqst *rqstp, struct svc_fh *fhp, u32 *access, u32 *suppor dentry = fhp->fh_dentry; if (d_is_reg(dentry)) - map = nfs3_regaccess; + map = maps->regular; else if (d_is_dir(dentry)) - map = nfs3_diraccess; + map = maps->directory; else - map = nfs3_anyaccess; - + map = maps->other; query = *access; for (; map->access; map++) { @@ -859,7 +843,7 @@ nfsd_access(struct svc_rqst *rqstp, struct svc_fh *fhp, u32 *access, u32 *suppor sresult |= map->access; err2 = nfsd_permission(&rqstp->rq_cred, export, - dentry, map->how); + dentry, map->may); switch (err2) { case nfs_ok: result |= map->access; @@ -1062,7 +1046,6 @@ static __be32 nfsd_finish_read(struct svc_rqst *rqstp, struct svc_fh *fhp, nfsd_stats_io_read_add(nn, fhp->fh_export, host_err); *eof = nfsd_eof_on_read(file, offset, host_err, *count); *count = host_err; - fsnotify_access(file); trace_nfsd_read_io_done(rqstp, fhp, offset, *count); return 0; } else { @@ -1087,19 +1070,11 @@ __be32 nfsd_splice_read(struct svc_rqst *rqstp, struct svc_fh *fhp, struct file *file, loff_t offset, unsigned long *count, u32 *eof) { - struct splice_desc sd = { - .len = 0, - .total_len = *count, - .pos = offset, - .u.data = rqstp, - }; ssize_t host_err; trace_nfsd_read_splice(rqstp, fhp, offset, *count); - host_err = rw_verify_area(READ, file, &offset, *count); - if (!host_err) - host_err = splice_direct_to_actor(file, &sd, - nfsd_direct_splice_actor); + host_err = vfs_splice_to_actor(file, offset, *count, + nfsd_direct_splice_actor, rqstp); return nfsd_finish_read(rqstp, fhp, file, offset, count, eof, host_err); } @@ -1424,7 +1399,7 @@ nfsd_direct_write(struct svc_rqst *rqstp, struct svc_fh *fhp, * @offset: Byte offset of start * @payload: xdr_buf containing the write payload * @cnt: IN: number of bytes to write, OUT: number of bytes actually written - * @stable: An NFS stable_how value + * @iocb_flags: VFS IOCB_* flags expressing the requested write stability * @verf: NFS WRITE verifier * * Upon return, caller must invoke fh_put on @fhp. @@ -1436,7 +1411,7 @@ __be32 nfsd_vfs_write(struct svc_rqst *rqstp, struct svc_fh *fhp, struct nfsd_file *nf, loff_t offset, const struct xdr_buf *payload, unsigned long *cnt, - int stable, __be32 *verf) + int iocb_flags, __be32 *verf) { struct nfsd_net *nn = net_generic(SVC_NET(rqstp), nfsd_net_id); struct file *file = nf->nf_file; @@ -1473,21 +1448,11 @@ nfsd_vfs_write(struct svc_rqst *rqstp, struct svc_fh *fhp, exp = fhp->fh_export; if (!EX_ISSYNC(exp)) - stable = NFS_UNSTABLE; + iocb_flags = 0; init_sync_kiocb(&kiocb, file); kiocb.ki_pos = offset; - if (likely(!fhp->fh_use_wgather)) { - switch (stable) { - case NFS_FILE_SYNC: - /* persist data and timestamps */ - kiocb.ki_flags |= IOCB_DSYNC | IOCB_SYNC; - break; - case NFS_DATA_SYNC: - /* persist data only */ - kiocb.ki_flags |= IOCB_DSYNC; - break; - } - } + if (likely(!fhp->fh_use_wgather)) + kiocb.ki_flags |= iocb_flags; nvecs = xdr_buf_to_bvec(rqstp->rq_bvec, rqstp->rq_maxpages, payload); if (nvecs < 0) { @@ -1517,21 +1482,21 @@ nfsd_vfs_write(struct svc_rqst *rqstp, struct svc_fh *fhp, break; } if (host_err < 0) { - commit_reset_write_verifier(nn, rqstp, host_err); + nfsd_maybe_reset_write_verifier(nn, rqstp, host_err); goto out_nfserr; } nfsd_stats_io_write_add(nn, exp, *cnt); fsnotify_modify(file); host_err = filemap_check_wb_err(file->f_mapping, since); if (host_err < 0) { - commit_reset_write_verifier(nn, rqstp, host_err); + nfsd_maybe_reset_write_verifier(nn, rqstp, host_err); goto out_nfserr; } - if (stable && fhp->fh_use_wgather) { + if (iocb_flags && fhp->fh_use_wgather) { host_err = wait_for_concurrent_writes(file); if (host_err < 0) - commit_reset_write_verifier(nn, rqstp, host_err); + nfsd_maybe_reset_write_verifier(nn, rqstp, host_err); } out_nfserr: @@ -1619,7 +1584,7 @@ __be32 nfsd_read(struct svc_rqst *rqstp, struct svc_fh *fhp, * @offset: Byte offset of start * @payload: xdr_buf containing the write payload * @cnt: IN: number of bytes to write, OUT: number of bytes actually written - * @stable: An NFS stable_how value + * @iocb_flags: VFS IOCB_* flags expressing the requested write stability * @verf: NFS WRITE verifier * * Upon return, caller must invoke fh_put on @fhp. @@ -1629,8 +1594,8 @@ __be32 nfsd_read(struct svc_rqst *rqstp, struct svc_fh *fhp, */ __be32 nfsd_write(struct svc_rqst *rqstp, struct svc_fh *fhp, loff_t offset, - const struct xdr_buf *payload, unsigned long *cnt, int stable, - __be32 *verf) + const struct xdr_buf *payload, unsigned long *cnt, + int iocb_flags, __be32 *verf) { struct nfsd_file *nf; __be32 err; @@ -1642,7 +1607,7 @@ nfsd_write(struct svc_rqst *rqstp, struct svc_fh *fhp, loff_t offset, goto out; err = nfsd_vfs_write(rqstp, fhp, nf, offset, payload, cnt, - stable, verf); + iocb_flags, verf); nfsd_file_put(nf); out: trace_nfsd_write_done(rqstp, fhp, offset, *cnt); @@ -1707,14 +1672,14 @@ nfsd_commit(struct svc_rqst *rqstp, struct svc_fh *fhp, struct nfsd_file *nf, err2 = filemap_check_wb_err(nf->nf_file->f_mapping, since); if (err2 < 0) - commit_reset_write_verifier(nn, rqstp, err2); + nfsd_maybe_reset_write_verifier(nn, rqstp, err2); err = nfserrno(err2); break; case -EINVAL: err = nfserr_notsupp; break; default: - commit_reset_write_verifier(nn, rqstp, err2); + nfsd_maybe_reset_write_verifier(nn, rqstp, err2); err = nfserrno(err2); } } else diff --git a/fs/nfsd/vfs.h b/fs/nfsd/vfs.h index 4af2ff9e9dfe..f0cb184643f2 100644 --- a/fs/nfsd/vfs.h +++ b/fs/nfsd/vfs.h @@ -37,6 +37,17 @@ #define NFSD_MAY_CREATE (NFSD_MAY_EXEC|NFSD_MAY_WRITE) #define NFSD_MAY_REMOVE (NFSD_MAY_EXEC|NFSD_MAY_WRITE|NFSD_MAY_TRUNC) +struct nfsd_access_map { + u32 access; + int may; +}; + +struct nfsd_access_maps { + const struct nfsd_access_map *regular; + const struct nfsd_access_map *directory; + const struct nfsd_access_map *other; +}; + struct nfsd_file; /* @@ -75,9 +86,14 @@ static inline bool nfsd_attrs_valid(struct nfsd_attrs *attrs) attrs->na_pacl || attrs->na_dpacl); } +struct nfsd_net; + __be32 nfserrno (int errno); -int nfsd_cross_mnt(struct svc_rqst *rqstp, struct dentry **dpp, - struct svc_export **expp); +void nfsd_maybe_reset_write_verifier(struct nfsd_net *nn, + struct svc_rqst *rqstp, + int err); +__be32 nfsd_cross_mnt(struct svc_rqst *rqstp, struct dentry **dpp, + struct svc_export **expp); __be32 nfsd_lookup(struct svc_rqst *, struct svc_fh *, const char *, unsigned int, struct svc_fh *); __be32 nfsd_lookup_dentry(struct svc_rqst *, struct svc_fh *, @@ -89,10 +105,11 @@ int nfsd_mountpoint(struct dentry *, struct svc_export *); #ifdef CONFIG_NFSD_V4 __be32 nfsd4_vfs_fallocate(struct svc_rqst *, struct svc_fh *, struct file *, loff_t, loff_t, int); -__be32 nfsd4_clone_file_range(struct svc_rqst *rqstp, - struct nfsd_file *nf_src, u64 src_pos, - struct nfsd_file *nf_dst, u64 dst_pos, - u64 count, bool sync); +int nfsd_clone_file_range(struct file *src, u64 src_pos, + struct file *dst, u64 dst_pos, + u64 count, errseq_t *since); +int nfsd_clone_sync_range(struct file *src, struct file *dst, + u64 dst_pos, u64 count, errseq_t since); #endif /* CONFIG_NFSD_V4 */ __be32 nfsd_create_locked(struct svc_rqst *, struct svc_fh *, struct nfsd_attrs *attrs, int type, dev_t rdev, @@ -100,7 +117,9 @@ __be32 nfsd_create_locked(struct svc_rqst *, struct svc_fh *, __be32 nfsd_create(struct svc_rqst *, struct svc_fh *, char *name, int len, struct nfsd_attrs *attrs, int type, dev_t rdev, struct svc_fh *res); -__be32 nfsd_access(struct svc_rqst *, struct svc_fh *, u32 *, u32 *); +__be32 nfsd_access(struct svc_rqst *rqstp, struct svc_fh *fhp, + const struct nfsd_access_maps *maps, + u32 *access, u32 *supported); __be32 nfsd_create_setattr(struct svc_rqst *rqstp, struct svc_fh *fhp, struct svc_fh *resfhp, struct nfsd_attrs *iap); __be32 nfsd_commit(struct svc_rqst *rqst, struct svc_fh *fhp, @@ -135,11 +154,13 @@ __be32 nfsd_read(struct svc_rqst *rqstp, struct svc_fh *fhp, u32 *eof); __be32 nfsd_write(struct svc_rqst *rqstp, struct svc_fh *fhp, loff_t offset, const struct xdr_buf *payload, - unsigned long *cnt, int stable, __be32 *verf); + unsigned long *cnt, int iocb_flags, + __be32 *verf); __be32 nfsd_vfs_write(struct svc_rqst *rqstp, struct svc_fh *fhp, struct nfsd_file *nf, loff_t offset, const struct xdr_buf *payload, - unsigned long *cnt, int stable, __be32 *verf); + unsigned long *cnt, int iocb_flags, + __be32 *verf); __be32 nfsd_readlink(struct svc_rqst *, struct svc_fh *, char *, int *); __be32 nfsd_symlink(struct svc_rqst *, struct svc_fh *, diff --git a/fs/nfsd/xdr.h b/fs/nfsd/xdr.h index df540c940cef..52660b9f8dde 100644 --- a/fs/nfsd/xdr.h +++ b/fs/nfsd/xdr.h @@ -1,8 +1,8 @@ /* SPDX-License-Identifier: GPL-2.0 */ /* XDR types for nfsd. This is mainly a typing exercise. */ -#ifndef LINUX_NFSD_H -#define LINUX_NFSD_H +#ifndef _LINUX_NFSD_XDR_H +#define _LINUX_NFSD_XDR_H #include <linux/vfs.h> #include "nfsd.h" @@ -175,4 +175,4 @@ bool svcxdr_encode_stat(struct xdr_stream *xdr, __be32 status); bool svcxdr_encode_fattr(struct svc_rqst *rqstp, struct xdr_stream *xdr, const struct svc_fh *fhp, const struct kstat *stat); -#endif /* LINUX_NFSD_H */ +#endif /* _LINUX_NFSD_XDR_H */ diff --git a/fs/nfsd/xdr3.h b/fs/nfsd/xdr3.h index 344203874b4c..cad875d14231 100644 --- a/fs/nfsd/xdr3.h +++ b/fs/nfsd/xdr3.h @@ -39,7 +39,7 @@ struct nfsd3_writeargs { svc_fh fh; __u64 offset; __u32 count; - int stable; + __u32 stable; __u32 len; struct xdr_buf payload; }; diff --git a/fs/nfsd/xdr4.h b/fs/nfsd/xdr4.h index c7eda5bc833b..ec198993429e 100644 --- a/fs/nfsd/xdr4.h +++ b/fs/nfsd/xdr4.h @@ -37,6 +37,8 @@ #ifndef _LINUX_NFSD_XDR4_H #define _LINUX_NFSD_XDR4_H +#include <linux/nfs_fh.h> + #include "state.h" #include "vfs.h" @@ -50,134 +52,6 @@ #define HAS_CSTATE_FLAG(c, f) ((c)->sid_flags & (f)) #define CLEAR_CSTATE_FLAG(c, f) ((c)->sid_flags &= ~(f)) -/** - * nfsd4_encode_bool - Encode an XDR bool type result - * @xdr: target XDR stream - * @val: boolean value to encode - * - * Return values: - * %nfs_ok: @val encoded; @xdr advanced to next position - * %nfserr_resource: stream buffer space exhausted - */ -static __always_inline __be32 -nfsd4_encode_bool(struct xdr_stream *xdr, bool val) -{ - __be32 *p = xdr_reserve_space(xdr, XDR_UNIT); - - if (unlikely(p == NULL)) - return nfserr_resource; - *p = val ? xdr_one : xdr_zero; - return nfs_ok; -} - -/** - * nfsd4_encode_uint32_t - Encode an XDR uint32_t type result - * @xdr: target XDR stream - * @val: integer value to encode - * - * Return values: - * %nfs_ok: @val encoded; @xdr advanced to next position - * %nfserr_resource: stream buffer space exhausted - */ -static __always_inline __be32 -nfsd4_encode_uint32_t(struct xdr_stream *xdr, u32 val) -{ - __be32 *p = xdr_reserve_space(xdr, XDR_UNIT); - - if (unlikely(p == NULL)) - return nfserr_resource; - *p = cpu_to_be32(val); - return nfs_ok; -} - -#define nfsd4_encode_aceflag4(x, v) nfsd4_encode_uint32_t(x, v) -#define nfsd4_encode_acemask4(x, v) nfsd4_encode_uint32_t(x, v) -#define nfsd4_encode_acetype4(x, v) nfsd4_encode_uint32_t(x, v) -#define nfsd4_encode_count4(x, v) nfsd4_encode_uint32_t(x, v) -#define nfsd4_encode_mode4(x, v) nfsd4_encode_uint32_t(x, v) -#define nfsd4_encode_nfs_lease4(x, v) nfsd4_encode_uint32_t(x, v) -#define nfsd4_encode_qop4(x, v) nfsd4_encode_uint32_t(x, v) -#define nfsd4_encode_sequenceid4(x, v) nfsd4_encode_uint32_t(x, v) -#define nfsd4_encode_slotid4(x, v) nfsd4_encode_uint32_t(x, v) - -/** - * nfsd4_encode_uint64_t - Encode an XDR uint64_t type result - * @xdr: target XDR stream - * @val: integer value to encode - * - * Return values: - * %nfs_ok: @val encoded; @xdr advanced to next position - * %nfserr_resource: stream buffer space exhausted - */ -static __always_inline __be32 -nfsd4_encode_uint64_t(struct xdr_stream *xdr, u64 val) -{ - __be32 *p = xdr_reserve_space(xdr, XDR_UNIT * 2); - - if (unlikely(p == NULL)) - return nfserr_resource; - put_unaligned_be64(val, p); - return nfs_ok; -} - -#define nfsd4_encode_changeid4(x, v) nfsd4_encode_uint64_t(x, v) -#define nfsd4_encode_nfs_cookie4(x, v) nfsd4_encode_uint64_t(x, v) -#define nfsd4_encode_length4(x, v) nfsd4_encode_uint64_t(x, v) -#define nfsd4_encode_offset4(x, v) nfsd4_encode_uint64_t(x, v) - -/** - * nfsd4_encode_opaque_fixed - Encode a fixed-length XDR opaque type result - * @xdr: target XDR stream - * @data: pointer to data - * @size: length of data in bytes - * - * Return values: - * %nfs_ok: @data encoded; @xdr advanced to next position - * %nfserr_resource: stream buffer space exhausted - */ -static __always_inline __be32 -nfsd4_encode_opaque_fixed(struct xdr_stream *xdr, const void *data, - size_t size) -{ - __be32 *p = xdr_reserve_space(xdr, xdr_align_size(size)); - size_t pad = xdr_pad_size(size); - - if (unlikely(p == NULL)) - return nfserr_resource; - memcpy(p, data, size); - if (pad) - memset((char *)p + size, 0, pad); - return nfs_ok; -} - -/** - * nfsd4_encode_opaque - Encode a variable-length XDR opaque type result - * @xdr: target XDR stream - * @data: pointer to data - * @size: length of data in bytes - * - * Return values: - * %nfs_ok: @data encoded; @xdr advanced to next position - * %nfserr_resource: stream buffer space exhausted - */ -static __always_inline __be32 -nfsd4_encode_opaque(struct xdr_stream *xdr, const void *data, size_t size) -{ - size_t pad = xdr_pad_size(size); - __be32 *p; - - p = xdr_reserve_space(xdr, XDR_UNIT + xdr_align_size(size)); - if (unlikely(p == NULL)) - return nfserr_resource; - *p++ = cpu_to_be32(size); - memcpy(p, data, size); - if (pad) - memset((char *)p + size, 0, pad); - return nfs_ok; -} - -#define nfsd4_encode_component4(x, d, s) nfsd4_encode_opaque(x, d, s) - struct nfsd4_compound_state { struct svc_fh current_fh; struct svc_fh save_fh; @@ -188,6 +62,7 @@ struct nfsd4_compound_state { struct nfsd4_slot *slot; int data_offset; bool spo_must_allowed; + bool slot_owned; size_t iovlen; u32 minorversion; __be32 status; @@ -642,17 +517,6 @@ svcxdr_decode_deviceid4(__be32 *p, struct nfsd4_deviceid *devid) return p; } -static inline __be32 -nfsd4_decode_deviceid4(struct xdr_stream *xdr, struct nfsd4_deviceid *devid) -{ - __be32 *p = xdr_inline_decode(xdr, NFS4_DEVICEID4_SIZE); - - if (unlikely(!p)) - return nfserr_bad_xdr; - svcxdr_decode_deviceid4(p, devid); - return nfs_ok; -} - struct nfsd4_layout_seg { u32 iomode; u64 offset; @@ -736,6 +600,19 @@ struct nfsd4_cb_offload { u32 co_referring_seqno; }; +struct nfsd4_ssc_umount_item { + struct list_head nsui_list; + bool nsui_busy; + /* + * nsui_refcnt inited to 2, 1 on list and 1 for consumer. Entry + * is removed when refcnt drops to 1 and nsui_expire expires. + */ + refcount_t nsui_refcnt; + unsigned long nsui_expire; + struct vfsmount *nsui_vfsmount; + char nsui_ipaddr[RPC_MAX_ADDRBUFLEN + 1]; +}; + struct nfsd4_copy { /* request */ stateid_t cp_src_stateid; @@ -924,6 +801,7 @@ bool nfsd4_cache_this_op(struct nfsd4_op *); */ struct svcxdr_tmpbuf { struct svcxdr_tmpbuf *next; + void (*release)(void *buf); char buf[]; }; diff --git a/fs/nfsd/xdr4cb.h b/fs/nfsd/xdr4cb.h index b06d0170d7c4..838f8629821f 100644 --- a/fs/nfsd/xdr4cb.h +++ b/fs/nfsd/xdr4cb.h @@ -6,29 +6,30 @@ #define cb_compound_enc_hdr_sz 4 #define cb_compound_dec_hdr_sz (3 + (NFS4_MAXTAGLEN >> 2)) #define sessionid_sz (NFS4_MAX_SESSIONID_LEN >> 2) +#define op_enc_sz 1 #define enc_referring_call4_sz (1 + 1) #define enc_referring_call_list4_sz (sessionid_sz + 1 + \ enc_referring_call4_sz) -#define cb_sequence_enc_sz (sessionid_sz + 4 + \ - enc_referring_call_list4_sz) +#define cb_sequence_enc_sz (op_enc_sz + sessionid_sz + 4 + \ + 1 + enc_referring_call_list4_sz) #define cb_sequence_dec_sz (op_dec_sz + sessionid_sz + 4) -#define op_enc_sz 1 #define op_dec_sz 2 #define enc_nfs4_fh_sz (1 + (NFS4_FHSIZE >> 2)) #define enc_stateid_sz (NFS4_STATEID_SIZE >> 2) #define NFS4_enc_cb_recall_sz (cb_compound_enc_hdr_sz + \ cb_sequence_enc_sz + \ - 1 + enc_stateid_sz + \ - enc_nfs4_fh_sz) + op_enc_sz + enc_stateid_sz + \ + 1 + enc_nfs4_fh_sz) #define NFS4_dec_cb_recall_sz (cb_compound_dec_hdr_sz + \ cb_sequence_dec_sz + \ op_dec_sz) #define NFS4_enc_cb_layout_sz (cb_compound_enc_hdr_sz + \ cb_sequence_enc_sz + \ - 1 + 3 + \ - enc_nfs4_fh_sz + 4) + op_enc_sz + 3 + 1 + \ + enc_nfs4_fh_sz + 4 + \ + enc_stateid_sz) #define NFS4_dec_cb_layout_sz (cb_compound_dec_hdr_sz + \ cb_sequence_dec_sz + \ op_dec_sz) @@ -47,7 +48,7 @@ #define NFS4_enc_cb_notify_lock_sz (cb_compound_enc_hdr_sz + \ cb_sequence_enc_sz + \ - 2 + 1 + \ + op_enc_sz + 2 + 1 + \ XDR_QUADLEN(NFS4_OPAQUE_LIMIT) + \ enc_nfs4_fh_sz) #define NFS4_dec_cb_notify_lock_sz (cb_compound_dec_hdr_sz + \ @@ -57,6 +58,7 @@ XDR_QUADLEN(NFS4_VERIFIER_SIZE)) #define NFS4_enc_cb_offload_sz (cb_compound_enc_hdr_sz + \ cb_sequence_enc_sz + \ + op_enc_sz + \ enc_nfs4_fh_sz + \ enc_stateid_sz + \ enc_cb_offload_info_sz) @@ -65,7 +67,7 @@ op_dec_sz) #define NFS4_enc_cb_recall_any_sz (cb_compound_enc_hdr_sz + \ cb_sequence_enc_sz + \ - 1 + 1 + 1) + op_enc_sz + 1 + 1 + 1) #define NFS4_dec_cb_recall_any_sz (cb_compound_dec_hdr_sz + \ cb_sequence_dec_sz + \ op_dec_sz) diff --git a/fs/nilfs2/inode.c b/fs/nilfs2/inode.c index 34e6096069ad..8953719c4950 100644 --- a/fs/nilfs2/inode.c +++ b/fs/nilfs2/inode.c @@ -904,7 +904,7 @@ void nilfs_evict_inode(struct inode *inode) */ } -int nilfs_setattr(struct mnt_idmap *idmap, struct dentry *dentry, +int nilfs_setattr(const struct mnt_idmap *idmap, struct dentry *dentry, struct iattr *iattr) { struct nilfs_transaction_info ti; @@ -943,7 +943,7 @@ out_err: return err; } -int nilfs_permission(struct mnt_idmap *idmap, struct inode *inode, +int nilfs_permission(const struct mnt_idmap *idmap, struct inode *inode, int mask) { struct nilfs_root *root = NILFS_I(inode)->i_root; diff --git a/fs/nilfs2/ioctl.c b/fs/nilfs2/ioctl.c index 01a04080ef70..f56a79767c69 100644 --- a/fs/nilfs2/ioctl.c +++ b/fs/nilfs2/ioctl.c @@ -135,7 +135,7 @@ int nilfs_fileattr_get(struct dentry *dentry, struct file_kattr *fa) * * Return: 0 on success, or a negative error code on failure. */ -int nilfs_fileattr_set(struct mnt_idmap *idmap, +int nilfs_fileattr_set(const struct mnt_idmap *idmap, struct dentry *dentry, struct file_kattr *fa) { struct inode *inode = d_inode(dentry); diff --git a/fs/nilfs2/namei.c b/fs/nilfs2/namei.c index e037e0c6e31a..d0ae24f37854 100644 --- a/fs/nilfs2/namei.c +++ b/fs/nilfs2/namei.c @@ -85,7 +85,7 @@ nilfs_lookup(struct inode *dir, struct dentry *dentry, unsigned int flags) * If the create succeeds, we fill in the inode information * with d_instantiate(). */ -static int nilfs_create(struct mnt_idmap *idmap, struct inode *dir, +static int nilfs_create(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *dentry, umode_t mode) { struct inode *inode; @@ -113,7 +113,7 @@ static int nilfs_create(struct mnt_idmap *idmap, struct inode *dir, } static int -nilfs_mknod(struct mnt_idmap *idmap, struct inode *dir, +nilfs_mknod(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *dentry, umode_t mode, dev_t rdev) { struct inode *inode; @@ -138,7 +138,7 @@ nilfs_mknod(struct mnt_idmap *idmap, struct inode *dir, return err; } -static int nilfs_symlink(struct mnt_idmap *idmap, struct inode *dir, +static int nilfs_symlink(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *dentry, const char *symname) { struct nilfs_transaction_info ti; @@ -218,7 +218,7 @@ static int nilfs_link(struct dentry *old_dentry, struct inode *dir, return err; } -static struct dentry *nilfs_mkdir(struct mnt_idmap *idmap, struct inode *dir, +static struct dentry *nilfs_mkdir(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *dentry, umode_t mode) { struct inode *inode; @@ -358,7 +358,7 @@ static int nilfs_rmdir(struct inode *dir, struct dentry *dentry) return err; } -static int nilfs_rename(struct mnt_idmap *idmap, +static int nilfs_rename(const struct mnt_idmap *idmap, struct inode *old_dir, struct dentry *old_dentry, struct inode *new_dir, struct dentry *new_dentry, unsigned int flags) diff --git a/fs/nilfs2/nilfs.h b/fs/nilfs2/nilfs.h index 4fc42d3787a4..28d35b2fa7f7 100644 --- a/fs/nilfs2/nilfs.h +++ b/fs/nilfs2/nilfs.h @@ -270,7 +270,7 @@ extern int nilfs_sync_file(struct file *, loff_t, loff_t, int); /* ioctl.c */ int nilfs_fileattr_get(struct dentry *dentry, struct file_kattr *m); -int nilfs_fileattr_set(struct mnt_idmap *idmap, +int nilfs_fileattr_set(const struct mnt_idmap *idmap, struct dentry *dentry, struct file_kattr *fa); long nilfs_ioctl(struct file *, unsigned int, unsigned long); long nilfs_compat_ioctl(struct file *file, unsigned int cmd, unsigned long arg); @@ -299,10 +299,10 @@ struct inode *nilfs_iget_for_shadow(struct inode *inode); extern void nilfs_update_inode(struct inode *, struct buffer_head *, int); extern void nilfs_truncate(struct inode *); extern void nilfs_evict_inode(struct inode *); -extern int nilfs_setattr(struct mnt_idmap *, struct dentry *, +extern int nilfs_setattr(const struct mnt_idmap *, struct dentry *, struct iattr *); extern void nilfs_write_failed(struct address_space *mapping, loff_t to); -int nilfs_permission(struct mnt_idmap *idmap, struct inode *inode, +int nilfs_permission(const struct mnt_idmap *idmap, struct inode *inode, int mask); int nilfs_load_inode_block(struct inode *inode, struct buffer_head **pbh); extern int nilfs_inode_dirty(struct inode *); diff --git a/fs/nls/nls_iso8859-14.c b/fs/nls/nls_iso8859-14.c index c789eccb8a69..60b9400f915a 100644 --- a/fs/nls/nls_iso8859-14.c +++ b/fs/nls/nls_iso8859-14.c @@ -138,24 +138,23 @@ static const unsigned char page00[256] = { }; static const unsigned char page01[256] = { - 0x00, 0x00, 0xa1, 0xa2, 0x00, 0x00, 0x00, 0x00, /* 0x00-0x07 */ - 0x00, 0x00, 0xa6, 0xab, 0x00, 0x00, 0x00, 0x00, /* 0x08-0x0f */ + 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, /* 0x00-0x07 */ + 0x00, 0x00, 0xa4, 0xa5, 0x00, 0x00, 0x00, 0x00, /* 0x08-0x0f */ 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, /* 0x10-0x17 */ - 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0xb0, 0xb1, /* 0x18-0x1f */ + 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, /* 0x18-0x1f */ 0xb2, 0xb3, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, /* 0x20-0x27 */ 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, /* 0x28-0x2f */ 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, /* 0x30-0x37 */ 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, /* 0x38-0x3f */ - 0xb4, 0xb5, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, /* 0x40-0x47 */ + 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, /* 0x40-0x47 */ 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, /* 0x48-0x4f */ - 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0xb7, 0xb9, /* 0x50-0x57 */ + 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, /* 0x50-0x57 */ 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, /* 0x58-0x5f */ - 0xbb, 0xbf, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, /* 0x60-0x67 */ - 0x00, 0x00, 0xd7, 0xf7, 0x00, 0x00, 0x00, 0x00, /* 0x68-0x6f */ + 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, /* 0x60-0x67 */ + 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, /* 0x68-0x6f */ 0x00, 0x00, 0x00, 0x00, 0xd0, 0xf0, 0xde, 0xfe, /* 0x70-0x77 */ 0xaf, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, /* 0x78-0x7f */ - - 0xa8, 0xb8, 0xaa, 0xba, 0xbd, 0xbe, 0x00, 0x00, /* 0x80-0x87 */ + 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, /* 0x80-0x87 */ 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, /* 0x88-0x8f */ 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, /* 0x90-0x97 */ 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, /* 0x98-0x9f */ @@ -169,7 +168,7 @@ static const unsigned char page01[256] = { 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, /* 0xd8-0xdf */ 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, /* 0xe0-0xe7 */ 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, /* 0xe8-0xef */ - 0x00, 0x00, 0xac, 0xbc, 0x00, 0x00, 0x00, 0x00, /* 0xf0-0xf7 */ + 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, /* 0xf0-0xf7 */ 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, /* 0xf8-0xff */ }; @@ -178,7 +177,7 @@ static const unsigned char page1e[256] = { 0x00, 0x00, 0xa6, 0xab, 0x00, 0x00, 0x00, 0x00, /* 0x08-0x0f */ 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, /* 0x10-0x17 */ 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0xb0, 0xb1, /* 0x18-0x1f */ - 0xb2, 0xb3, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, /* 0x20-0x27 */ + 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, /* 0x20-0x27 */ 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, /* 0x28-0x2f */ 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, /* 0x30-0x37 */ 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, /* 0x38-0x3f */ @@ -188,9 +187,8 @@ static const unsigned char page1e[256] = { 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, /* 0x58-0x5f */ 0xbb, 0xbf, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, /* 0x60-0x67 */ 0x00, 0x00, 0xd7, 0xf7, 0x00, 0x00, 0x00, 0x00, /* 0x68-0x6f */ - 0x00, 0x00, 0x00, 0x00, 0xd0, 0xf0, 0xde, 0xfe, /* 0x70-0x77 */ - 0xaf, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, /* 0x78-0x7f */ - + 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, /* 0x70-0x77 */ + 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, /* 0x78-0x7f */ 0xa8, 0xb8, 0xaa, 0xba, 0xbd, 0xbe, 0x00, 0x00, /* 0x80-0x87 */ 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, /* 0x88-0x8f */ 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, 0x00, /* 0x90-0x97 */ diff --git a/fs/notify/fanotify/fanotify_user.c b/fs/notify/fanotify/fanotify_user.c index a32c6634d592..63c9759fc3b0 100644 --- a/fs/notify/fanotify/fanotify_user.c +++ b/fs/notify/fanotify/fanotify_user.c @@ -1837,7 +1837,7 @@ static int fanotify_events_supported(struct fsnotify_group *group, /* * mount and sb marks are not allowed on kernel internal pseudo fs, * like pipe_mnt, because that would subscribe to events on all the - * anonynous pipes in the system. + * anonymous pipes in the system. * * SB_NOUSER covers all of the internal pseudo fs whose objects are not * exposed to user's mount namespace, but there are other SB_KERNMOUNT diff --git a/fs/nsfs.c b/fs/nsfs.c index c3b6ae76594a..56ea0bb9ef3a 100644 --- a/fs/nsfs.c +++ b/fs/nsfs.c @@ -348,8 +348,8 @@ static long ns_ioctl(struct file *filp, unsigned int ioctl, return ret; FD_PREPARE(fdf, O_CLOEXEC, dentry_open(&path, O_RDONLY, current_cred())); - if (fdf.err) - return fdf.err; + if (fdf->fd < 0) + return fdf->fd; /* * If @uinfo is passed return all information about the * mount namespace as well. diff --git a/fs/ntfs/attrib.c b/fs/ntfs/attrib.c index c949ff765075..5439c12f9280 100644 --- a/fs/ntfs/attrib.c +++ b/fs/ntfs/attrib.c @@ -103,6 +103,17 @@ int ntfs_map_runlist_nolock(struct ntfs_inode *ni, s64 vcn, struct ntfs_attr_sea base_ni = ni; else base_ni = ni->ext.base_ntfs_ino; + /* + * ntfs_read_inode_mount() builds $MFT's runlist itself, so nothing + * should reach here for $MFT. A crafted image can: the read that + * gets here already holds the $MFT folio lock it would wait on. + */ + if (unlikely(NVolMftBootstrap(ni->vol) && + base_ni == NTFS_I(ni->vol->mft_ino))) { + ntfs_error(ni->vol->sb, + "$MFT needs its own extent records to describe itself; cannot mount."); + return -EIO; + } if (!ctx) { ctx_is_temporary = ctx_needs_reset = true; m = map_mft_record(base_ni); @@ -317,7 +328,7 @@ int ntfs_map_runlist(struct ntfs_inode *ni, s64 vcn) struct runlist_element *ntfs_attr_vcn_to_rl(struct ntfs_inode *ni, s64 vcn, s64 *lcn) { struct runlist_element *rl = ni->runlist.rl; - int err; + int err = 0; bool is_retry = false; if (!rl) { @@ -335,12 +346,34 @@ remap_rl: if (*lcn <= LCN_RL_NOT_MAPPED && is_retry == false) { is_retry = true; - if (!ntfs_map_runlist_nolock(ni, vcn, NULL)) { + err = ntfs_map_runlist_nolock(ni, vcn, NULL); + if (!err) { rl = ni->runlist.rl; goto remap_rl; } } + /* + * Neither the runlist nor the retry mapped @vcn, e.g. because the + * extent mft record holding it is corrupt or because the mapping + * pairs end too soon. ntfs_map_runlist_nolock() reports the latter + * as -ENOENT, as @vcn lies past the extent it found. Below the + * allocated size, callers would treat LCN_RL_NOT_MAPPED or LCN_ENOENT + * as a hole, so fail instead. At or beyond it nothing is mapped: the + * runlist ends there with LCN_ENOENT, or with LCN_RL_NOT_MAPPED if + * only a later extent has been mapped, so return that end as it is. + */ + if (*lcn <= LCN_RL_NOT_MAPPED) { + unsigned long flags; + s64 allocated_size; + + read_lock_irqsave(&ni->size_lock, flags); + allocated_size = ni->allocated_size; + read_unlock_irqrestore(&ni->size_lock, flags); + if ((s64)ntfs_cluster_to_bytes(ni->vol, vcn) < allocated_size) + return ERR_PTR(err == -ENOMEM ? -ENOMEM : -EIO); + } + return rl; } @@ -909,7 +942,7 @@ static int ntfs_attr_find(const __le32 type, const __le16 *name, rc = ntfs_collate_names(name, name_len, (__le16 *)((u8 *)a + le16_to_cpu(a->name_offset)), - a->name_length, 1, IGNORE_CASE, + a->name_length, true, IGNORE_CASE, upcase, upcase_len); /* * If @name collates before a->name, there is no @@ -922,7 +955,7 @@ static int ntfs_attr_find(const __le32 type, const __le16 *name, continue; rc = ntfs_collate_names(name, name_len, (__le16 *)((u8 *)a + le16_to_cpu(a->name_offset)), - a->name_length, 1, CASE_SENSITIVE, + a->name_length, true, CASE_SENSITIVE, upcase, upcase_len); if (rc == -1) return -ENOENT; @@ -1313,7 +1346,7 @@ find_attr_list_attr: register int rc; rc = ntfs_collate_names(name, name_len, al_name, - al_name_len, 1, IGNORE_CASE, + al_name_len, true, IGNORE_CASE, vol->upcase, vol->upcase_len); /* * If @name collates before al_name, there is no @@ -1326,7 +1359,7 @@ find_attr_list_attr: continue; rc = ntfs_collate_names(name, name_len, al_name, - al_name_len, 1, CASE_SENSITIVE, + al_name_len, true, CASE_SENSITIVE, vol->upcase, vol->upcase_len); if (rc == -1) goto not_found; @@ -1775,7 +1808,7 @@ int ntfs_attr_size_bounds_check(const struct ntfs_volume *vol, const __le32 type * $ATTRIBUTE_LIST has a maximum size of 256kiB, but this is not * listed in $AttrDef. */ - if (unlikely(type == AT_ATTRIBUTE_LIST && size > 256 * 1024)) + if (unlikely(type == AT_ATTRIBUTE_LIST && size > NTFS_MAX_ATTR_LIST_SIZE)) return -ERANGE; /* Get the $AttrDef entry for the attribute @type. */ ad = ntfs_attr_find_in_attrdef(vol, type); @@ -4234,12 +4267,14 @@ static int ntfs_attr_make_resident(struct ntfs_inode *ni, struct ntfs_attr_searc * ntfs_non_resident_attr_shrink - shrink a non-resident, open ntfs attribute * @ni: non-resident ntfs attribute to shrink * @newsize: new size (in bytes) to which to shrink the attribute + * @pagecache_truncated: page cache was already truncated to @newsize * * Reduce the size of a non-resident, open ntfs attribute @na to @newsize bytes. */ static int ntfs_non_resident_attr_shrink(struct ntfs_inode *ni, const s64 newsize, - struct ntfs_inode *locked_ni) + struct ntfs_inode *locked_ni, + bool pagecache_truncated) { struct ntfs_volume *vol; struct ntfs_attr_search_ctx *ctx; @@ -4389,7 +4424,8 @@ static int ntfs_non_resident_attr_shrink(struct ntfs_inode *ni, * later writeback map a vcn past the new allocation, which fails with * -ENOENT and loses the write. */ - truncate_inode_pages(VFS_I(ni)->i_mapping, newsize); + if (!pagecache_truncated) + truncate_inode_pages(VFS_I(ni)->i_mapping, newsize); /* Update data size in the index. */ if (ni->type == AT_DATA && ni->name == AT_UNNAMED) @@ -4450,6 +4486,7 @@ static int ntfs_non_resident_attr_expand(struct ntfs_inode *ni, const s64 newsiz struct ntfs_inode *base_ni; struct super_block *sb = ni->vol->sb; size_t new_rl_count; + unsigned long flags; ntfs_debug("Inode 0x%llx, attr 0x%x, new size %lld old size %lld\n", (unsigned long long)ni->mft_no, ni->type, @@ -4680,11 +4717,20 @@ rollback: if (err2) ntfs_debug("Leaking clusters"); - /* Now, truncate the runlist itself. */ + /* + * Now, truncate the runlist itself. Restore allocated_size before + * dropping the lock: ntfs_attr_vcn_to_rl() fails a lookup below the + * allocated size that falls past the end of the runlist. + */ if (ni != locked_ni) down_write(&ni->runlist.lock); err2 = ntfs_rl_truncate_nolock(vol, &ni->runlist, ntfs_bytes_to_cluster(vol, org_alloc_size)); + if (!err2) { + write_lock_irqsave(&ni->size_lock, flags); + ni->allocated_size = org_alloc_size; + write_unlock_irqrestore(&ni->size_lock, flags); + } if (ni != locked_ni) up_write(&ni->runlist.lock); if (err2) { @@ -4696,8 +4742,6 @@ rollback: ni->runlist.rl = NULL; ntfs_error(sb, "Couldn't truncate runlist. Rollback failed"); } else { - /* Prepare to mapping pairs update. */ - ni->allocated_size = org_alloc_size; /* Restore mapping pairs. */ if (ni != locked_ni) down_read(&ni->runlist.lock); @@ -4991,7 +5035,7 @@ int __ntfs_attr_truncate_vfs(struct ntfs_inode *ni, const s64 newsize, up_write(&ni->runlist.lock); } else err = ntfs_non_resident_attr_shrink( - ni, newsize, NULL); + ni, newsize, NULL, true); } else err = ntfs_resident_attr_resize(ni, newsize, 0, NVolDisableSparse(ni->vol) ? @@ -5104,7 +5148,7 @@ int ntfs_attr_truncate_i_locked(struct ntfs_inode *ni, const s64 newsize, ni, newsize, 0, holes, locked_ni); else err = ntfs_non_resident_attr_shrink( - ni, newsize, locked_ni); + ni, newsize, locked_ni, false); } else err = ntfs_resident_attr_resize(ni, newsize, 0, holes); ntfs_debug("Return status %d\n", err); @@ -5855,7 +5899,7 @@ int ntfs_attr_fallocate(struct ntfs_inode *ni, loff_t start, loff_t byte_len, bo goto out; } - if (signal_pending(current)) + if (fatal_signal_pending(current)) goto signal_out; vcn += alloc_cnt; @@ -5876,7 +5920,7 @@ int ntfs_attr_fallocate(struct ntfs_inode *ni, loff_t start, loff_t byte_len, bo try_alloc_cnt, &balloc, false, false); up_write(&ni->runlist.lock); mutex_unlock(&ni->mrec_lock); - if (err || signal_pending(current)) + if (err || fatal_signal_pending(current)) goto signal_out; vcn += alloc_cnt; diff --git a/fs/ntfs/attrlist.c b/fs/ntfs/attrlist.c index bb191953dcb1..3660e7fd24b1 100644 --- a/fs/ntfs/attrlist.c +++ b/fs/ntfs/attrlist.c @@ -14,8 +14,6 @@ #include "attrlist.h" #include "lcnalloc.h" -#define NTFS_MAX_ATTR_LIST_SIZE (256 * 1024) - /* * ntfs_attrlist_need - check whether inode need attribute list * @ni: opened ntfs inode for which perform check @@ -170,14 +168,18 @@ static int ntfs_attrlist_repack(struct inode *attr_vi, return 0; restore_old_runlist: + /* + * Restore allocated_size before dropping the runlist lock: + * ntfs_attr_vcn_to_rl() fails a lookup below the allocated size that + * falls past the end of the runlist. + */ down_write(&attr_ni->runlist.lock); attr_ni->runlist.rl = old_rl; attr_ni->runlist.count = old_rl_count; - up_write(&attr_ni->runlist.lock); - write_lock_irqsave(&attr_ni->size_lock, flags); attr_ni->allocated_size = old_alloc_size; write_unlock_irqrestore(&attr_ni->size_lock, flags); + up_write(&attr_ni->runlist.lock); restore_err = ntfs_attr_update_mapping_pairs_locked( attr_ni, 0, locked_ni); @@ -335,8 +337,8 @@ int ntfs_attrlist_entry_add(struct ntfs_inode *ni, struct attr_record *attr) ni_mrec = map_mft_record(ni); if (IS_ERR(ni_mrec)) { - ntfs_debug("Invalid arguments.\n"); - return -EIO; + ntfs_debug("Failed to map mft record.\n"); + return PTR_ERR(ni_mrec); } mref = MK_LE_MREF(ni->mft_no, le16_to_cpu(ni_mrec->sequence_number)); diff --git a/fs/ntfs/bitmap.c b/fs/ntfs/bitmap.c index 5a4457551306..912fdcca8e01 100644 --- a/fs/ntfs/bitmap.c +++ b/fs/ntfs/bitmap.c @@ -20,7 +20,7 @@ int ntfs_trim_fs(struct ntfs_volume *vol, struct fstrim_range *range) struct folio *folio; unsigned long *bitmap; char *kaddr; - u64 end, trimmed = 0, start_buf, end_buf, end_cluster; + u64 end, trimmed = 0, start_buf, end_buf, end_cluster, page_cluster; u64 start_cluster = ntfs_bytes_to_cluster(vol, range->start); u32 dq = bdev_discard_granularity(vol->sb->s_bdev); int ret = 0; @@ -45,8 +45,8 @@ int ntfs_trim_fs(struct ntfs_volume *vol, struct fstrim_range *range) return -ENOMEM; buf_clusters = PAGE_SIZE * 8; - start_index = start_cluster >> 15; - end_index = (end_cluster + buf_clusters - 1) >> 15; + start_index = start_cluster / buf_clusters; + end_index = DIV_ROUND_UP_ULL(end_cluster, buf_clusters); for (index = start_index; index < end_index; index++) { folio = ntfs_get_locked_folio(vol->lcnbmp_ino->i_mapping, @@ -59,19 +59,20 @@ int ntfs_trim_fs(struct ntfs_volume *vol, struct fstrim_range *range) kaddr = kmap_local_folio(folio, 0); bitmap = (unsigned long *)kaddr; - start_buf = max_t(u64, index * buf_clusters, start_cluster); - end_buf = min_t(u64, (index + 1) * buf_clusters, end_cluster); + page_cluster = (u64)index * buf_clusters; + start_buf = max_t(u64, page_cluster, start_cluster); + end_buf = min_t(u64, page_cluster + buf_clusters, end_cluster); end = start_buf; while (end < end_buf) { u64 aligned_start, aligned_end, aligned_count; - u64 start = find_next_zero_bit(bitmap, end_buf - start_buf, - end - start_buf) + start_buf; + u64 start = find_next_zero_bit(bitmap, end_buf - page_cluster, + end - page_cluster) + page_cluster; if (start >= end_buf) break; - end = find_next_bit(bitmap, end_buf - start_buf, - start - start_buf) + start_buf; + end = find_next_bit(bitmap, end_buf - page_cluster, + start - page_cluster) + page_cluster; aligned_start = ALIGN(ntfs_cluster_to_bytes(vol, start), dq); aligned_end = ALIGN_DOWN(ntfs_cluster_to_bytes(vol, end), dq); diff --git a/fs/ntfs/collate.c b/fs/ntfs/collate.c index 77e34038902d..5417288c5825 100644 --- a/fs/ntfs/collate.c +++ b/fs/ntfs/collate.c @@ -73,7 +73,7 @@ static int ntfs_collate_ntofs_ulongs(struct ntfs_volume *vol, if (data1_len != data2_len || data1_len & 3) { ntfs_error(vol->sb, "data1_len or data2_len not valid\n"); - return -1; + return -EINVAL; } len = data1_len; @@ -100,11 +100,11 @@ static int ntfs_collate_file_name(struct ntfs_volume *vol, { int rc; - rc = ntfs_file_compare_values(data1, data2, -EINVAL, - IGNORE_CASE, vol->upcase, vol->upcase_len); + rc = ntfs_file_compare_values(data1, data2, + true, IGNORE_CASE, vol->upcase, vol->upcase_len); if (!rc) rc = ntfs_file_compare_values(data1, data2, - -EINVAL, CASE_SENSITIVE, vol->upcase, vol->upcase_len); + true, CASE_SENSITIVE, vol->upcase, vol->upcase_len); return rc; } diff --git a/fs/ntfs/dir.c b/fs/ntfs/dir.c index df60138f9b2d..b95173f068cb 100644 --- a/fs/ntfs/dir.c +++ b/fs/ntfs/dir.c @@ -238,7 +238,7 @@ found_it: */ rc = ntfs_collate_names(uname, uname_len, (__le16 *)&ie->key.file_name.file_name, - ie->key.file_name.file_name_length, 1, + ie->key.file_name.file_name_length, false, IGNORE_CASE, vol->upcase, vol->upcase_len); /* * If uname collates before the name of the current entry, there @@ -257,7 +257,7 @@ found_it: */ rc = ntfs_collate_names(uname, uname_len, (__le16 *)&ie->key.file_name.file_name, - ie->key.file_name.file_name_length, 1, + ie->key.file_name.file_name_length, false, CASE_SENSITIVE, vol->upcase, vol->upcase_len); if (rc == -1) break; @@ -331,7 +331,6 @@ descend_into_child_node: } memcpy_from_folio(kaddr, folio, 0, PAGE_SIZE); - post_read_mst_fixup((struct ntfs_record *)kaddr, PAGE_SIZE); folio_unlock(folio); folio_put(folio); fast_descend_into_child_node: @@ -352,6 +351,14 @@ fast_descend_into_child_node: vcn, dir_ni->mft_no); goto unm_err_out; } + err = post_read_mst_fixup((struct ntfs_record *)ia, + dir_ni->itype.index.block_size); + if (err) { + ntfs_error(sb, + "MST fixup failed for index block vcn %lld in directory inode 0x%llx.", + vcn, dir_ni->mft_no); + goto unm_err_out; + } err = ntfs_index_block_inconsistent(vol, ia, dir_ni->itype.index.block_size, vcn, COLLATION_FILE_NAME, @@ -474,7 +481,7 @@ found_it2: */ rc = ntfs_collate_names(uname, uname_len, (__le16 *)&ie->key.file_name.file_name, - ie->key.file_name.file_name_length, 1, + ie->key.file_name.file_name_length, false, IGNORE_CASE, vol->upcase, vol->upcase_len); /* * If uname collates before the name of the current entry, there @@ -493,7 +500,7 @@ found_it2: */ rc = ntfs_collate_names(uname, uname_len, (__le16 *)&ie->key.file_name.file_name, - ie->key.file_name.file_name_length, 1, + ie->key.file_name.file_name_length, false, CASE_SENSITIVE, vol->upcase, vol->upcase_len); if (rc == -1) break; @@ -628,11 +635,11 @@ static inline int ntfs_filldir(struct ntfs_volume *vol, } mref = MREF_LE(ie->data.dir.indexed_file); - if (ie->key.file_name.file_attributes & + if (ie->key.file_name.file_attributes & FILE_ATTR_REPARSE_POINT) + dt_type = ntfs_reparse_tag_dt_types(vol, mref); + else if (ie->key.file_name.file_attributes & FILE_ATTR_DUP_FILE_NAME_INDEX_PRESENT) dt_type = DT_DIR; - else if (ie->key.file_name.file_attributes & FILE_ATTR_REPARSE_POINT) - dt_type = ntfs_reparse_tag_dt_types(vol, mref); else dt_type = DT_REG; diff --git a/fs/ntfs/ea.c b/fs/ntfs/ea.c index b4fcfbe2da4c..ddc201e3e2aa 100644 --- a/fs/ntfs/ea.c +++ b/fs/ntfs/ea.c @@ -875,7 +875,7 @@ static int ntfs_validate_fattr(struct ntfs_inode *ni, __le32 fattr) } static int ntfs_setxattr(const struct xattr_handler *handler, - struct mnt_idmap *idmap, struct dentry *unused, + const struct mnt_idmap *idmap, struct dentry *unused, struct inode *inode, const char *name, const void *value, size_t size, int flags) { @@ -977,7 +977,7 @@ const struct xattr_handler * const ntfs_xattr_handlers[] = { // clang-format on #ifdef CONFIG_NTFS_FS_POSIX_ACL -struct posix_acl *ntfs_get_acl(struct mnt_idmap *idmap, struct dentry *dentry, +struct posix_acl *ntfs_get_acl(const struct mnt_idmap *idmap, struct dentry *dentry, int type) { struct inode *inode = d_inode(dentry); @@ -1021,7 +1021,7 @@ struct posix_acl *ntfs_get_acl(struct mnt_idmap *idmap, struct dentry *dentry, return acl; } -static noinline int ntfs_set_acl_ex(struct mnt_idmap *idmap, +static noinline int ntfs_set_acl_ex(const struct mnt_idmap *idmap, struct inode *inode, struct posix_acl *acl, int type, bool init_acl) { @@ -1103,13 +1103,13 @@ out: return err; } -int ntfs_set_acl(struct mnt_idmap *idmap, struct dentry *dentry, +int ntfs_set_acl(const struct mnt_idmap *idmap, struct dentry *dentry, struct posix_acl *acl, int type) { return ntfs_set_acl_ex(idmap, d_inode(dentry), acl, type, false); } -int ntfs_init_acl(struct mnt_idmap *idmap, struct inode *inode, +int ntfs_init_acl(const struct mnt_idmap *idmap, struct inode *inode, struct inode *dir) { struct posix_acl *default_acl, *acl; diff --git a/fs/ntfs/ea.h b/fs/ntfs/ea.h index acb39c2a6fbc..690fafe181fb 100644 --- a/fs/ntfs/ea.h +++ b/fs/ntfs/ea.h @@ -17,11 +17,11 @@ int ntfs_ea_set_wsl_inode(struct inode *inode, dev_t rdev, __le16 *ea_size, ssize_t ntfs_listxattr(struct dentry *dentry, char *buffer, size_t size); #ifdef CONFIG_NTFS_FS_POSIX_ACL -struct posix_acl *ntfs_get_acl(struct mnt_idmap *idmap, struct dentry *dentry, +struct posix_acl *ntfs_get_acl(const struct mnt_idmap *idmap, struct dentry *dentry, int type); -int ntfs_set_acl(struct mnt_idmap *idmap, struct dentry *dentry, +int ntfs_set_acl(const struct mnt_idmap *idmap, struct dentry *dentry, struct posix_acl *acl, int type); -int ntfs_init_acl(struct mnt_idmap *idmap, struct inode *inode, +int ntfs_init_acl(const struct mnt_idmap *idmap, struct inode *inode, struct inode *dir); #else #define ntfs_get_acl NULL diff --git a/fs/ntfs/file.c b/fs/ntfs/file.c index 007d1614b9ac..bb8641103d26 100644 --- a/fs/ntfs/file.c +++ b/fs/ntfs/file.c @@ -252,6 +252,22 @@ static int ntfs_file_fsync(struct file *filp, loff_t start, loff_t end, return ret; } +static void ntfs_pagecache_extend(struct inode *vi, loff_t from, loff_t to) +{ + pagecache_isize_extended(vi, from, to); + + /* + * NTFS initialized_size is byte-granular, so a partial old-EOF page + * must stay write-protected until page_mkwrite() updates it. The + * generic helper skips this when the block size is at least a page, + * the extension ends before the rounded block boundary, or that + * boundary is page-aligned. + */ + if (from < to && from & (PAGE_SIZE - 1)) + unmap_mapping_range(vi->i_mapping, round_down(from, PAGE_SIZE), + PAGE_SIZE, 0); +} + static int ntfs_setattr_size(struct inode *vi, struct iattr *attr) { struct ntfs_inode *ni = NTFS_I(vi); @@ -281,7 +297,7 @@ static int ntfs_setattr_size(struct inode *vi, struct iattr *attr) if (attr->ia_size > old_size) { truncate_pagecache(vi, old_size); i_size_write(vi, attr->ia_size); - pagecache_isize_extended(vi, old_size, attr->ia_size); + ntfs_pagecache_extend(vi, old_size, attr->ia_size); } else { truncate_setsize(vi, attr->ia_size); } @@ -302,7 +318,7 @@ static int ntfs_setattr_size(struct inode *vi, struct iattr *attr) * NOTE: Changes in inode size are not supported yet for compressed or * encrypted files. */ -int ntfs_setattr(struct mnt_idmap *idmap, struct dentry *dentry, +int ntfs_setattr(const struct mnt_idmap *idmap, struct dentry *dentry, struct iattr *attr) { struct inode *vi = d_inode(dentry); @@ -325,8 +341,7 @@ int ntfs_setattr(struct mnt_idmap *idmap, struct dentry *dentry, goto out; } - if (!(vol->vol_flags & VOLUME_IS_DIRTY)) - ntfs_set_volume_flags(vol, VOLUME_IS_DIRTY); + ntfs_set_volume_flags(vol, VOLUME_IS_DIRTY); if (ia_valid & ATTR_SIZE) { err = ntfs_setattr_size(vi, attr); @@ -372,7 +387,7 @@ out: return err; } -int ntfs_getattr(struct mnt_idmap *idmap, const struct path *path, +int ntfs_getattr(const struct mnt_idmap *idmap, const struct path *path, struct kstat *stat, unsigned int request_mask, unsigned int query_flags) { @@ -620,8 +635,14 @@ static ssize_t ntfs_file_write_iter(struct kiocb *iocb, struct iov_iter *from) goto out_lock; } - if (!(vol->vol_flags & VOLUME_IS_DIRTY)) - ntfs_set_volume_flags(vol, VOLUME_IS_DIRTY); + /* + * The volume must be marked dirty before the modification is made, + * without an unlocked check of the in-memory flag: the dirty bit + * is only cleared at the quiescent transitions, under the same + * $Volume mrec_lock this call takes, so an unlocked skip could + * lose the set to one of them. + */ + ntfs_set_volume_flags(vol, VOLUME_IS_DIRTY); pos = iocb->ki_pos; count = ret; @@ -867,7 +888,7 @@ static int ntfs_ioctl_fitrim(struct ntfs_volume *vol, unsigned long arg) if (range.len < vol->cluster_size) return -EINVAL; - range.minlen = max_t(u32, range.minlen, bdev_discard_granularity(dev)); + range.minlen = max_t(u64, range.minlen, bdev_discard_granularity(dev)); err = ntfs_trim_fs(vol, &range); if (err < 0) @@ -1131,7 +1152,7 @@ static long ntfs_fallocate(struct file *file, int mode, loff_t offset, loff_t le struct ntfs_inode *ni = NTFS_I(vi); struct ntfs_volume *vol = ni->vol; int err = 0; - loff_t old_size; + loff_t old_size, new_size; if (mode & ~(NTFS_FALLOC_FL_SUPPORTED)) return -EOPNOTSUPP; @@ -1153,19 +1174,16 @@ static long ntfs_fallocate(struct file *file, int mode, loff_t offset, loff_t le return err; } - if (!(vol->vol_flags & VOLUME_IS_DIRTY)) { - err = ntfs_set_volume_flags(vol, VOLUME_IS_DIRTY); - if (err) - return err; - } - - old_size = i_size_read(vi); + err = ntfs_set_volume_flags(vol, VOLUME_IS_DIRTY); + if (err) + return err; inode_lock(vi); if (NInoCompressed(ni) || NInoEncrypted(ni) || NInoWofCompressed(ni)) { inode_unlock(vi); return -EOPNOTSUPP; } + old_size = i_size_read(vi); inode_dio_wait(vi); /* Take invalidate_lock for all fallocate operations to prevent races */ @@ -1194,10 +1212,12 @@ static long ntfs_fallocate(struct file *file, int mode, loff_t offset, loff_t le err = file_modified(file); out: - if (!err && mode == 0 && NInoNonResident(ni) && - offset > old_size) { - truncate_pagecache(vi, old_size); - pagecache_isize_extended(vi, old_size, offset); + if (!err && mode == 0 && NInoNonResident(ni)) { + new_size = i_size_read(vi); + if (new_size > old_size) { + truncate_pagecache(vi, old_size); + ntfs_pagecache_extend(vi, old_size, new_size); + } } filemap_invalidate_unlock(vi->i_mapping); diff --git a/fs/ntfs/index.c b/fs/ntfs/index.c index 580998990bc9..4410df0a8f31 100644 --- a/fs/ntfs/index.c +++ b/fs/ntfs/index.c @@ -852,7 +852,7 @@ int ntfs_index_lookup(const void *key, const u32 key_len, struct ntfs_index_cont if (ni->vol->cluster_size <= icx->block_size) icx->vcn_size_bits = ni->vol->cluster_size_bits; else - icx->vcn_size_bits = ni->vol->sector_size_bits; + icx->vcn_size_bits = NTFS_BLOCK_SIZE_BITS; icx->cr = ir->collation_rule; if (!ntfs_is_collation_rule_supported(icx->cr)) { diff --git a/fs/ntfs/index.h b/fs/ntfs/index.h index 9a03f53bba47..a898e348f2c9 100644 --- a/fs/ntfs/index.h +++ b/fs/ntfs/index.h @@ -37,7 +37,7 @@ * (maximum is MAX_PARENT_VCN) * @ib_dirty: true if the current index block (@ia/@ib) was modified * @block_size: size of index blocks in bytes (from $INDEX_ROOT or $Boot) - * @vcn_size_bits: log2(cluster size) + * @vcn_size_bits: log2 of the VCN unit size * @sync_write: true if synchronous writeback is requested for this context * * @idx_ni is the index inode this context belongs to. diff --git a/fs/ntfs/inode.c b/fs/ntfs/inode.c index a777de8a80c7..b666cdb3b102 100644 --- a/fs/ntfs/inode.c +++ b/fs/ntfs/inode.c @@ -651,6 +651,25 @@ void ntfs_set_vfs_operations(struct inode *inode, mode_t mode, dev_t dev) } } +static bool ntfs_non_resident_sizes_inconsistent(struct inode *vi, + const struct attr_record *a) +{ + s64 allocated_size = le64_to_cpu(a->data.non_resident.allocated_size); + s64 data_size = le64_to_cpu(a->data.non_resident.data_size); + s64 initialized_size = le64_to_cpu(a->data.non_resident.initialized_size); + + if (initialized_size >= 0 && initialized_size <= data_size && + data_size <= allocated_size && + !ntfs_bytes_to_cluster_off(NTFS_I(vi)->vol, allocated_size)) + return false; + + ntfs_error(vi->i_sb, + "Attribute 0x%x of inode 0x%llx is corrupt (initialized size %lld, data size %lld, allocated size %lld).", + le32_to_cpu(a->type), NTFS_I(vi)->mft_no, initialized_size, + data_size, allocated_size); + return true; +} + /* * ntfs_read_locked_inode - read an inode from its device * @vi: inode to read @@ -805,6 +824,8 @@ static int ntfs_read_locked_inode(struct inode *vi) goto unm_err_out; } } else { + s64 attr_list_size; + if (vi->i_ino == FILE_MFT) goto skip_attr_list_load; ntfs_debug("Attribute list found in inode 0x%llx.", ni->mft_no); @@ -827,11 +848,14 @@ static int ntfs_read_locked_inode(struct inode *vi) ni->mft_no); } /* Now allocate memory for the attribute list. */ - ni->attr_list_size = (u32)ntfs_attr_size(a); - if (!ni->attr_list_size) { - ntfs_error(vi->i_sb, "Attr_list_size is zero"); + attr_list_size = ntfs_attr_size(a); + if (attr_list_size <= 0 || + attr_list_size > NTFS_MAX_ATTR_LIST_SIZE) { + ntfs_error(vi->i_sb, "Invalid attribute list size %lld (mft_no 0x%llx).", + (long long)attr_list_size, ni->mft_no); goto unm_err_out; } + ni->attr_list_size = (u32)attr_list_size; ni->attr_list = kvzalloc(ni->attr_list_size, GFP_NOFS); if (!ni->attr_list) { ntfs_error(vi->i_sb, @@ -1021,8 +1045,8 @@ view_index_meta: ni->itype.index.vcn_size = vol->cluster_size; ni->itype.index.vcn_size_bits = vol->cluster_size_bits; } else { - ni->itype.index.vcn_size = vol->sector_size; - ni->itype.index.vcn_size_bits = vol->sector_size_bits; + ni->itype.index.vcn_size = NTFS_BLOCK_SIZE; + ni->itype.index.vcn_size_bits = NTFS_BLOCK_SIZE_BITS; } /* Setup the index allocation attribute, even if not present. */ @@ -1143,14 +1167,21 @@ view_index_meta: goto unm_err_out; } + /* + * Windows stores the standard compression unit in + * sparse attributes even when the data is not compressed. + */ if (NInoSparse(ni) && a->data.non_resident.compression_unit && a->data.non_resident.compression_unit != - vol->sparse_compression_unit) { + vol->sparse_compression_unit && + a->data.non_resident.compression_unit != + STANDARD_COMPRESSION_UNIT) { ntfs_error(vi->i_sb, - "Found non-standard compression unit (%u instead of 0 or %d). Cannot handle this.", + "Found non-standard compression unit (%u instead of 0, %d, or %d). Cannot handle this.", a->data.non_resident.compression_unit, - vol->sparse_compression_unit); + vol->sparse_compression_unit, + STANDARD_COMPRESSION_UNIT); err = -EOPNOTSUPP; goto unm_err_out; } @@ -1179,6 +1210,8 @@ view_index_meta: "First extent of $DATA attribute has non zero lowest_vcn."); goto unm_err_out; } + if (ntfs_non_resident_sizes_inconsistent(vi, a)) + goto unm_err_out; vi->i_size = ni->data_size = le64_to_cpu(a->data.non_resident.data_size); ni->initialized_size = le64_to_cpu(a->data.non_resident.initialized_size); ni->allocated_size = le64_to_cpu(a->data.non_resident.allocated_size); @@ -1441,6 +1474,8 @@ static int ntfs_read_locked_attr_inode(struct inode *base_vi, struct inode *vi) ntfs_error(vi->i_sb, "First extent of attribute has non-zero lowest_vcn."); goto unm_err_out; } + if (ntfs_non_resident_sizes_inconsistent(vi, a)) + goto unm_err_out; vi->i_size = ni->data_size = le64_to_cpu(a->data.non_resident.data_size); ni->initialized_size = le64_to_cpu(a->data.non_resident.initialized_size); ni->allocated_size = le64_to_cpu(a->data.non_resident.allocated_size); @@ -1612,8 +1647,8 @@ static int ntfs_read_locked_index_inode(struct inode *base_vi, struct inode *vi) ni->itype.index.vcn_size = vol->cluster_size; ni->itype.index.vcn_size_bits = vol->cluster_size_bits; } else { - ni->itype.index.vcn_size = vol->sector_size; - ni->itype.index.vcn_size_bits = vol->sector_size_bits; + ni->itype.index.vcn_size = NTFS_BLOCK_SIZE; + ni->itype.index.vcn_size_bits = NTFS_BLOCK_SIZE_BITS; } /* Find index allocation attribute. */ @@ -1670,6 +1705,8 @@ static int ntfs_read_locked_index_inode(struct inode *base_vi, struct inode *vi) "First extent of $INDEX_ALLOCATION attribute has non zero lowest_vcn."); goto unm_err_out; } + if (ntfs_non_resident_sizes_inconsistent(vi, a)) + goto unm_err_out; vi->i_size = ni->data_size = le64_to_cpu(a->data.non_resident.data_size); ni->initialized_size = le64_to_cpu(a->data.non_resident.initialized_size); ni->allocated_size = le64_to_cpu(a->data.non_resident.allocated_size); @@ -1955,6 +1992,7 @@ int ntfs_read_inode_mount(struct inode *vi) } else /* if (!err) */ { struct attr_list_entry *al_entry, *next_al_entry; u8 *al_end; + s64 attr_list_size; static const char *es = " Not allowed. $MFT is corrupt. You should run chkdsk."; ntfs_debug("Attribute list attribute found in $MFT."); @@ -1978,11 +2016,14 @@ int ntfs_read_inode_mount(struct inode *vi) "Resident attribute list attribute in $MFT system file is marked encrypted/sparse which is not true. However, Windows allows this and chkdsk does not detect or correct it so we will just ignore the invalid flags and pretend they are not set."); } /* Now allocate memory for the attribute list. */ - ni->attr_list_size = (u32)ntfs_attr_size(a); - if (!ni->attr_list_size) { - ntfs_error(sb, "Attr_list_size is zero"); + attr_list_size = ntfs_attr_size(a); + if (attr_list_size <= 0 || + attr_list_size > NTFS_MAX_ATTR_LIST_SIZE) { + ntfs_error(sb, "Invalid attribute list size %lld (mft_no 0x%llx).%s", + (long long)attr_list_size, ni->mft_no, es); goto put_err_out; } + ni->attr_list_size = (u32)attr_list_size; ni->attr_list = kvzalloc(round_up(ni->attr_list_size, SECTOR_SIZE), GFP_NOFS); if (!ni->attr_list) { @@ -1999,6 +2040,8 @@ int ntfs_read_inode_mount(struct inode *vi) "Attribute list has non zero lowest_vcn. $MFT is corrupt. You should run chkdsk."); goto put_err_out; } + if (ntfs_non_resident_sizes_inconsistent(vi, a)) + goto put_err_out; rl = ntfs_mapping_pairs_decompress(vol, a, NULL, &new_rl_count); if (IS_ERR(rl)) { @@ -2071,6 +2114,11 @@ int ntfs_read_inode_mount(struct inode *vi) /* Now load all attribute extents. */ a = NULL; next_vcn = last_vcn = highest_vcn = 0; + /* + * Reading one of $MFT's own extent records in this loop can re-enter + * ntfs_map_runlist_nolock() for $MFT; see the check there. + */ + NVolSetMftBootstrap(vol); while (!(err = ntfs_attr_lookup(AT_DATA, NULL, 0, 0, next_vcn, NULL, 0, ctx))) { struct runlist_element *nrl; @@ -2123,6 +2171,16 @@ int ntfs_read_inode_mount(struct inode *vi) ni->initialized_size = le64_to_cpu(a->data.non_resident.initialized_size); ni->allocated_size = le64_to_cpu(a->data.non_resident.allocated_size); /* + * Records between allocated_size and data_size are not + * on disk, and would be read as zeros. + */ + if (vi->i_size > ni->allocated_size) { + ntfs_error(sb, + "$MFT data size %lld exceeds its allocated size %lld. $MFT is corrupt. Run chkdsk.", + vi->i_size, ni->allocated_size); + goto put_err_out; + } + /* * Verify the number of mft records does not exceed * 2^32 - 1. */ @@ -2153,6 +2211,7 @@ int ntfs_read_inode_mount(struct inode *vi) err = ntfs_read_locked_inode(vi); if (err) { ntfs_error(sb, "ntfs_read_inode() of $MFT failed.\n"); + NVolClearMftBootstrap(vol); ntfs_attr_put_search_ctx(ctx); /* Revert to the safe super operations. */ kfree(m); @@ -2186,6 +2245,7 @@ int ntfs_read_inode_mount(struct inode *vi) goto put_err_out; } } + NVolClearMftBootstrap(vol); if (err != -ENOENT) { ntfs_error(sb, "Failed to lookup $MFT/$DATA attribute extent. Run chkdsk.\n"); goto put_err_out; @@ -2220,6 +2280,8 @@ em_put_err_out: put_err_out: ntfs_attr_put_search_ctx(ctx); err_out: + /* Also reached from inside the $DATA loop. */ + NVolClearMftBootstrap(vol); ntfs_error(sb, "Failed. Marking inode as bad."); kfree(m); return -1; @@ -2915,8 +2977,7 @@ err_out: * abort if deprotection or checks fail. * * Finally attach the ntfs inode to its base inode @base_ni and return a - * pointer to the ntfs_inode structure on success or NULL on error, with errno - * set to the error code. + * pointer to the ntfs_inode structure on success or ERR_PTR() on error. * * Note, extent inodes are never closed directly. They are automatically * disposed off by the closing of the base inode. @@ -2932,7 +2993,7 @@ static struct ntfs_inode *ntfs_extent_inode_open(struct ntfs_inode *base_ni, struct super_block *sb; if (!base_ni) - return NULL; + return ERR_PTR(-EINVAL); sb = base_ni->vol->sb; ntfs_debug("Opening extent inode %llu (base mft record %llu).\n", @@ -2951,7 +3012,7 @@ static struct ntfs_inode *ntfs_extent_inode_open(struct ntfs_inode *base_ni, if (IS_ERR(ni_mrec)) { ntfs_error(sb, "failed to map mft record for %llu", ni->mft_no); - goto out; + return ERR_CAST(ni_mrec); } /* Verify the sequence number if given. */ seq_no = MSEQNO_LE(mref); @@ -2960,7 +3021,7 @@ static struct ntfs_inode *ntfs_extent_inode_open(struct ntfs_inode *base_ni, ntfs_error(sb, "Found stale extent mft reference mft=%llu", ni->mft_no); unmap_mft_record(ni); - goto out; + return ERR_PTR(-EIO); } unmap_mft_record(ni); goto out; @@ -2969,7 +3030,7 @@ static struct ntfs_inode *ntfs_extent_inode_open(struct ntfs_inode *base_ni, /* Wasn't there, we need to load the extent inode. */ ni = ntfs_new_extent_inode(base_ni->vol->sb, mft_no); if (!ni) - goto out; + return ERR_PTR(-ENOMEM); ni->seq_no = (u16)MSEQNO_LE(mref); ni->nr_extents = -1; @@ -2979,8 +3040,10 @@ static struct ntfs_inode *ntfs_extent_inode_open(struct ntfs_inode *base_ni, i = (base_ni->nr_extents + 4) * sizeof(struct ntfs_inode *); extent_nis = kvzalloc(i, GFP_NOFS); - if (!extent_nis) - goto err_out; + if (!extent_nis) { + ntfs_destroy_ext_inode(ni); + return ERR_PTR(-ENOMEM); + } if (base_ni->nr_extents) { memcpy(extent_nis, base_ni->ext.extent_ntfs_inos, i - 4 * sizeof(struct ntfs_inode *)); @@ -2993,10 +3056,6 @@ static struct ntfs_inode *ntfs_extent_inode_open(struct ntfs_inode *base_ni, out: ntfs_debug("\n"); return ni; -err_out: - ntfs_destroy_ext_inode(ni); - ni = NULL; - goto out; } /* @@ -3034,9 +3093,12 @@ int ntfs_inode_attach_all_extents(struct ntfs_inode *ni) while ((u8 *)ale < ni->attr_list + ni->attr_list_size) { if (ni->mft_no != MREF_LE(ale->mft_reference) && prev_attached != MREF_LE(ale->mft_reference)) { - if (!ntfs_extent_inode_open(ni, ale->mft_reference)) { + struct ntfs_inode *ext_ni; + + ext_ni = ntfs_extent_inode_open(ni, ale->mft_reference); + if (IS_ERR(ext_ni)) { ntfs_debug("Couldn't attach extent inode.\n"); - return -1; + return PTR_ERR(ext_ni); } prev_attached = MREF_LE(ale->mft_reference); } @@ -3629,9 +3691,10 @@ static inline int ntfs_enlarge_attribute(struct inode *vi, s64 pos, s64 count, return -EOPNOTSUPP; if (pos + count > ni->data_size) { - if (ntfs_attr_truncate(ni, pos + count)) { + ret = ntfs_attr_truncate(ni, pos + count); + if (ret) { ntfs_debug("Failed to truncate attribute"); - return -1; + return ret; } ntfs_attr_reinit_search_ctx(ctx); @@ -3719,7 +3782,7 @@ static s64 __ntfs_inode_non_resident_attr_pwrite(struct inode *vi, index = pos >> PAGE_SHIFT; while (count) { - if (count == PAGE_SIZE) { + if (count == PAGE_SIZE && !offset_in_page(pos)) { folio = __filemap_get_folio(vi->i_mapping, index, FGP_CREAT | FGP_LOCK, mapping_gfp_mask(mapping)); @@ -3739,7 +3802,9 @@ static s64 __ntfs_inode_non_resident_attr_pwrite(struct inode *vi, folio_lock(folio); } - if (count == PAGE_SIZE) { + folio_wait_writeback(folio); + + if (count == PAGE_SIZE && !offset_in_page(pos)) { offset = 0; attr_len = count; } else { @@ -3758,7 +3823,9 @@ static s64 __ntfs_inode_non_resident_attr_pwrite(struct inode *vi, struct runlist_element *rl; int bio_err; - lcn_count = max_t(s64, 1, ntfs_bytes_to_cluster(vol, attr_len)); + lcn_count = max_t(s64, 1, + ntfs_bytes_to_cluster(vol, attr_len + + vol->cluster_size - 1)); vcn = ntfs_pidx_to_cluster(vol, folio->index); do { diff --git a/fs/ntfs/inode.h b/fs/ntfs/inode.h index ff61bd402df0..45a396c97846 100644 --- a/fs/ntfs/inode.h +++ b/fs/ntfs/inode.h @@ -330,9 +330,9 @@ int ntfs_read_inode_mount(struct inode *vi); int ntfs_show_options(struct seq_file *sf, struct dentry *root); int ntfs_truncate_vfs(struct inode *vi, loff_t new_size, loff_t i_size); -int ntfs_setattr(struct mnt_idmap *idmap, struct dentry *dentry, +int ntfs_setattr(const struct mnt_idmap *idmap, struct dentry *dentry, struct iattr *attr); -int ntfs_getattr(struct mnt_idmap *idmap, const struct path *path, +int ntfs_getattr(const struct mnt_idmap *idmap, const struct path *path, struct kstat *stat, unsigned int request_mask, unsigned int query_flags); diff --git a/fs/ntfs/iomap.c b/fs/ntfs/iomap.c index c812d7f19b36..b4475963e57a 100644 --- a/fs/ntfs/iomap.c +++ b/fs/ntfs/iomap.c @@ -394,7 +394,7 @@ static int ntfs_write_simple_iomap_begin_non_resident(struct inode *inode, loff_ loff_t vcn_ofs, rl_length; struct runlist_element *rl, *rlc; bool is_retry = false; - int err = 0; + int err = 0, map_err = 0; s64 vcn, lcn; s64 max_clu_count = ntfs_bytes_to_cluster(vol, round_up(length, vol->cluster_size)); @@ -429,12 +429,24 @@ remap_rl: if (lcn <= LCN_RL_NOT_MAPPED && is_retry == false) { is_retry = true; - if (!ntfs_map_runlist_nolock(ni, vcn, NULL)) { + map_err = ntfs_map_runlist_nolock(ni, vcn, NULL); + if (!map_err) { rl = ni->runlist.rl; goto remap_rl; } } + /* + * As in ntfs_attr_vcn_to_rl(): a runlist fragment that could not be + * mapped is not a hole. Treating it as one would put a delalloc + * extent over clusters that are allocated on disk but unknown to us. + */ + if (lcn == LCN_RL_NOT_MAPPED) { + up_write(&ni->runlist.lock); + mutex_unlock(&ni->mrec_lock); + return map_err == -ENOMEM ? -ENOMEM : -EIO; + } + max_clu_count = min(max_clu_count, rl->length - (vcn - rl->vcn)); if (max_clu_count == 0) { ntfs_error(inode->i_sb, diff --git a/fs/ntfs/layout.h b/fs/ntfs/layout.h index 9438fd9b668e..2de83d4ca40b 100644 --- a/fs/ntfs/layout.h +++ b/fs/ntfs/layout.h @@ -811,14 +811,17 @@ enum { * on XP SP2+. * @data.non_resident.reserved: 5 bytes for 8-byte alignment. * @data.non_resident.allocated_size: - * Allocated disk space in bytes. - * For compressed: logical allocated size. + * Allocated size in bytes, a multiple of + * the cluster size. For compressed and + * sparse attributes holes count as + * allocated; the clusters actually in use + * are in compressed_size. * @data.non_resident.data_size: Logical attribute value size in bytes. - * Can be larger than allocated_size if - * compressed/sparse. + * Never larger than allocated_size, also + * when compressed/sparse. * @data.non_resident.initialized_size: * Initialized portion size in bytes. - * Usually equals data_size. + * Usually equals data_size, never larger. * @data.non_resident.compressed_size: * Compressed on-disk size in bytes. * Only present when compressed or sparse. @@ -1880,8 +1883,8 @@ struct index_header { * @index_block_size: Size of each index block in bytes * (in $INDEX_ALLOCATION). * @clusters_per_index_block: - * Clusters per index block (or log2(bytes) - * if < cluster). + * Clusters per index block, or 512-byte units when + * the index block is smaller than a cluster. * Power of 2; used for encoding block size. * @reserved: 3 bytes reserved/alignment (zero). * @index: Index header for root entries (entries follow @@ -1921,7 +1924,7 @@ struct index_root { * @lsn: Log sequence number of last modification. * @index_block_vcn: VCN of this index block. * Units: clusters if cluster_size <= index_block_size; - * sectors otherwise. + * 512-byte blocks otherwise. * @index: Index header describing entries in this block. * * When creating the index block, we place the update sequence array at this diff --git a/fs/ntfs/mft.c b/fs/ntfs/mft.c index 4b7449495375..e01e367a588d 100644 --- a/fs/ntfs/mft.c +++ b/fs/ntfs/mft.c @@ -10,7 +10,10 @@ #include <linux/writeback.h> #include <linux/bio.h> +#include <linux/blkdev.h> +#include <linux/completion.h> #include <linux/iomap.h> +#include <linux/math64.h> #include "bitmap.h" #include "lcnalloc.h" @@ -429,81 +432,152 @@ void __mark_mft_record_dirty(struct ntfs_inode *ni) __mark_inode_dirty(VFS_I(base_ni), I_DIRTY_DATASYNC); } -/* - * ntfs_bio_end_io - bio completion callback for MFT record writes - * - * Decrements the folio reference count that was incremented before - * submit_bio(). This prevents a race condition where umount could - * evict the inode and release the folio while I/O is still in flight, - * potentially causing data corruption or use-after-free. - */ -static void ntfs_bio_end_io(struct bio *bio) +struct ntfs_mft_io_unit { + sector_t sector; + unsigned int folio_ofs; + unsigned int len; +}; + +struct ntfs_mft_write_ctx { + struct folio *folio; + struct address_space *mapping; + struct ntfs_volume *vol; + struct completion *done; + int error; + bool writeback_started; + struct bio bio; +}; + +static struct bio_set ntfs_mft_bioset; + +int ntfs_mft_bioset_init(void) +{ + return bioset_init(&ntfs_mft_bioset, BIO_POOL_SIZE, + offsetof(struct ntfs_mft_write_ctx, bio), + BIOSET_NEED_BVECS); +} + +void ntfs_mft_bioset_exit(void) +{ + bioset_exit(&ntfs_mft_bioset); +} + +static void ntfs_mft_end_io(struct bio *bio) +{ + struct ntfs_mft_write_ctx *ctx = + container_of(bio, struct ntfs_mft_write_ctx, bio); + struct folio *folio = ctx->folio; + struct completion *done = ctx->done; + int err; + + if (bio->bi_status) + err = blk_status_to_errno(bio->bi_status); + else + err = ctx->error; + if (err) { + mapping_set_error(ctx->mapping, err); + NVolSetErrors(ctx->vol); + ntfs_error(ctx->vol->sb, "I/O error while writing MFT: %d", + err); + } + + if (!done) + bio_put(bio); + folio_end_writeback(folio); + folio_put(folio); + if (done) + complete(done); +} + +static void ntfs_start_mft_writeback(struct ntfs_mft_write_ctx *ctx) +{ + if (ctx->writeback_started) + return; + + inode_attach_wb(ctx->mapping->host, NULL); + folio_get(ctx->folio); + folio_start_writeback(ctx->folio); + ctx->writeback_started = true; +} + +static struct bio *ntfs_alloc_mft_parent_bio(struct ntfs_volume *vol, + struct folio *folio, + struct completion *done) { - if (bio->bi_private) - folio_put((struct folio *)bio->bi_private); - bio_put(bio); + struct ntfs_mft_write_ctx *ctx; + struct bio *parent; + + parent = bio_alloc_bioset(vol->sb->s_bdev, 1, REQ_OP_WRITE, GFP_NOIO, + &ntfs_mft_bioset); + if (!parent) + return NULL; + + ctx = container_of(parent, struct ntfs_mft_write_ctx, bio); + ctx->folio = folio; + ctx->mapping = folio->mapping; + ctx->vol = vol; + ctx->done = done; + ctx->error = 0; + ctx->writeback_started = false; + parent->bi_end_io = ntfs_mft_end_io; + return parent; } /* - * ntfs_sync_mft_mirror - synchronize an mft record to the mft mirror - * @vol: ntfs volume on which the mft record to synchronize resides - * @mft_no: mft record number of mft record to synchronize - * @m: mapped, mst protected (extent) mft record to synchronize - * - * Write the mapped, mst protected (extent) mft record @m with mft record - * number @mft_no to the mft mirror ($MFTMirr) of the ntfs volume @vol. - * - * On success return 0. On error return -errno and set the volume errors flag - * in the ntfs volume @vol. - * - * NOTE: We always perform synchronous i/o. + * Write one MFT I/O unit to $MFTMirr. The source folio contains the MST + * protected image that will be written to $MFT. */ -int ntfs_sync_mft_mirror(struct ntfs_volume *vol, const u64 mft_no, - struct mft_record *m) +static int ntfs_sync_mft_mirror_unit(struct ntfs_volume *vol, + struct folio *source, u64 mirror_file_ofs, + const struct ntfs_mft_io_unit *unit) { - u8 *kmirr; - struct folio *folio; - unsigned int folio_ofs; - int err = 0; - struct bio *bio; + struct folio *mirror; + struct bio_vec bvec; + struct bio bio; + u64 mirror_size; + unsigned int mirror_ofs; + u8 *src, *dst; + int err; - ntfs_debug("Entering for inode 0x%llx.", mft_no); + if (unlikely(!vol->mftmirr_ino)) + return -EIO; - if (unlikely(!vol->mftmirr_ino)) { - /* This could happen during umount... */ + mirror_size = (u64)vol->mftmirr_size * vol->mft_record_size; + if (mirror_file_ofs >= mirror_size || + unit->len > mirror_size - mirror_file_ofs) + return -EIO; + if (unit->folio_ofs + unit->len > folio_size(source)) + return -EIO; + + mirror = read_mapping_folio(vol->mftmirr_ino->i_mapping, + mirror_file_ofs >> PAGE_SHIFT, NULL); + if (IS_ERR(mirror)) + return PTR_ERR(mirror); + + folio_lock(mirror); + if (folio_test_writeback(mirror)) + folio_wait_writeback(mirror); + mirror_ofs = mirror_file_ofs - folio_pos(mirror); + if (mirror_ofs + unit->len > folio_size(mirror)) { err = -EIO; - goto err_out; - } - /* Get the page containing the mirror copy of the mft record @m. */ - folio = read_mapping_folio(vol->mftmirr_ino->i_mapping, - NTFS_MFT_NR_TO_PIDX(vol, mft_no), NULL); - if (IS_ERR(folio)) { - ntfs_error(vol->sb, "Failed to map mft mirror page."); - err = PTR_ERR(folio); - goto err_out; + goto out_unlock; } - folio_lock(folio); - folio_clear_uptodate(folio); - /* Offset of the mft mirror record inside the page. */ - folio_ofs = NTFS_MFT_NR_TO_POFS(vol, mft_no); - /* The address in the page of the mirror copy of the mft record @m. */ - kmirr = kmap_local_folio(folio, 0) + folio_ofs; - /* Copy the mst protected mft record to the mirror. */ - memcpy(kmirr, m, vol->mft_record_size); - kunmap_local(kmirr); - - bio = bio_alloc(vol->sb->s_bdev, 1, REQ_OP_WRITE, GFP_NOIO); - bio->bi_iter.bi_sector = - ntfs_bytes_to_bio_sector(NTFS_CLU_TO_B(vol, vol->mftmirr_lcn) + - ((u64)folio->index << PAGE_SHIFT) + - folio_ofs); - - if (bio_add_folio(bio, folio, vol->mft_record_size, folio_ofs)) - err = submit_bio_wait(bio); - else + folio_clear_uptodate(mirror); + src = kmap_local_folio(source, unit->folio_ofs); + dst = kmap_local_folio(mirror, mirror_ofs); + memcpy(dst, src, unit->len); + kunmap_local(dst); + kunmap_local(src); + + bio_init(&bio, vol->sb->s_bdev, &bvec, 1, REQ_OP_WRITE); + bio.bi_iter.bi_sector = ntfs_bytes_to_bio_sector( + NTFS_CLU_TO_B(vol, vol->mftmirr_lcn) + mirror_file_ofs); + if (!bio_add_folio(&bio, mirror, unit->len, mirror_ofs)) err = -EIO; - bio_put(bio); + else + err = submit_bio_wait(&bio); + bio_uninit(&bio); /* * The in-memory mirror is now valid because we just memcpy()'d the @@ -512,23 +586,79 @@ int ntfs_sync_mft_mirror(struct ntfs_volume *vol, const u64 mft_no, * the stale on-disk mirror and overwrite this copy. The error is * propagated to the caller via @err. */ - folio_mark_uptodate(folio); + folio_mark_uptodate(mirror); - folio_unlock(folio); - folio_put(folio); - if (likely(!err)) { - ntfs_debug("Done."); - } else { - ntfs_error(vol->sb, "I/O error while writing mft mirror record 0x%llx!", mft_no); -err_out: - ntfs_error(vol->sb, - "Failed to synchronize $MFTMirr (error code %i). Volume will be left marked dirty on umount. Run chkdsk on the partition after umounting to correct this.", - err); - NVolSetErrors(vol); - } +out_unlock: + folio_unlock(mirror); + folio_put(mirror); return err; } +static int ntfs_sync_mft_mirror_record(struct ntfs_volume *vol, + struct folio *source, const u64 mft_no) +{ + u64 mirror_file_ofs = (u64)mft_no * vol->mft_record_size; + struct ntfs_mft_io_unit unit = { + .folio_ofs = NTFS_MFT_NR_TO_POFS(vol, mft_no), + .len = vol->mft_record_size, + }; + + mirror_file_ofs = + round_down(mirror_file_ofs, (u64)vol->mft_io_unit_size); + unit.folio_ofs = round_down(unit.folio_ofs, vol->mft_io_unit_size); + unit.len = vol->mft_io_unit_size; + + return ntfs_sync_mft_mirror_unit(vol, source, mirror_file_ofs, &unit); +} + +static int ntfs_prepare_mft_record_io_units(struct ntfs_inode *ni, + struct ntfs_mft_io_unit units[2]) +{ + struct ntfs_volume *vol = ni->vol; + u64 record_byte = (u64)ni->mft_no * vol->mft_record_size; + u64 disk_byte; + unsigned int nr_units = ni->mft_lcn_count; + unsigned int cluster_ofs; + unsigned int i; + + if (!nr_units || nr_units > ARRAY_SIZE(ni->mft_lcn)) + return -EIO; + for (i = 0; i < nr_units; i++) + if (ni->mft_lcn[i] < 0) + return -EIO; + + cluster_ofs = ntfs_bytes_to_cluster_off(vol, record_byte); + if (vol->mft_io_unit_size > vol->mft_record_size) { + cluster_ofs = round_down(cluster_ofs, vol->mft_io_unit_size); + disk_byte = NTFS_CLU_TO_B(vol, ni->mft_lcn[0]) + cluster_ofs; + units[0] = (struct ntfs_mft_io_unit){ + .sector = ntfs_bytes_to_bio_sector(disk_byte), + .folio_ofs = round_down(ni->folio_ofs, + vol->mft_io_unit_size), + .len = vol->mft_io_unit_size, + }; + return 1; + } + + disk_byte = NTFS_CLU_TO_B(vol, ni->mft_lcn[0]) + cluster_ofs; + units[0] = (struct ntfs_mft_io_unit){ + .sector = ntfs_bytes_to_bio_sector(disk_byte), + .folio_ofs = ni->folio_ofs, + .len = vol->mft_record_size, + }; + if (nr_units == 1) + return 1; + + units[0].len = vol->cluster_size; + disk_byte = NTFS_CLU_TO_B(vol, ni->mft_lcn[1]); + units[1] = (struct ntfs_mft_io_unit){ + .sector = ntfs_bytes_to_bio_sector(disk_byte), + .folio_ofs = ni->folio_ofs + vol->cluster_size, + .len = vol->mft_record_size - vol->cluster_size, + }; + return 2; +} + /* * write_mft_record_nolock - write out a mapped (extent) mft record * @ni: ntfs inode describing the mapped (extent) mft record @@ -541,25 +671,35 @@ err_out: * * We only write the mft record if the ntfs inode @ni is dirty. * - * On success, clean the mft record and return 0. - * On error (specifically ENOMEM), we redirty the record so it can be retried. - * For other errors, we mark the volume with errors. + * On success, clean the mft record and return 0. On ENOMEM, redirty the + * record so it can be retried. Asynchronous callers return success after + * redirtying while synchronous callers receive the error. For other errors, + * mark the volume with errors. + * + * If @sync is false, PG_writeback keeps the folio stable and serializes later + * writers until the I/O completes. */ int write_mft_record_nolock(struct ntfs_inode *ni, struct mft_record *m, int sync) { struct ntfs_volume *vol = ni->vol; struct folio *folio = ni->folio; - int err = 0, i = 0; + DECLARE_COMPLETION_ONSTACK(done); + struct ntfs_mft_io_unit units[2]; + struct ntfs_mft_write_ctx *ctx; + struct bio *parent, *child = NULL; + unsigned int nr_units; + int err = 0; u8 *kaddr; struct mft_record *fixup_m; - struct bio *bio; - unsigned int offset = 0, folio_size; ntfs_debug("Entering for inode 0x%llx.", ni->mft_no); WARN_ON(NInoAttr(ni)); WARN_ON(!folio_test_locked(folio)); + if (folio_test_writeback(folio)) + folio_wait_writeback(folio); + /* * If the struct ntfs_inode is clean no need to do anything. If it is dirty, * mark it as clean now so that it can be redirtied later on if needed. @@ -569,6 +709,11 @@ int write_mft_record_nolock(struct ntfs_inode *ni, struct mft_record *m, int syn if (!NInoTestClearDirty(ni)) goto done; + err = ntfs_prepare_mft_record_io_units(ni, units); + if (err < 0) + goto err_out; + nr_units = err; + kaddr = kmap_local_folio(folio, 0); fixup_m = (struct mft_record *)(kaddr + ni->folio_ofs); memcpy(fixup_m, m, vol->mft_record_size); @@ -580,68 +725,74 @@ int write_mft_record_nolock(struct ntfs_inode *ni, struct mft_record *m, int syn goto unmap_err_out; } - folio_size = vol->mft_record_size / ni->mft_lcn_count; - while (i < ni->mft_lcn_count) { - unsigned int clu_off; - - clu_off = (unsigned int)((s64)ni->mft_no * vol->mft_record_size + offset) & - vol->cluster_size_mask; - - bio = bio_alloc(vol->sb->s_bdev, 1, REQ_OP_WRITE, GFP_NOIO); - bio->bi_iter.bi_sector = - ntfs_bytes_to_bio_sector(NTFS_CLU_TO_B(vol, ni->mft_lcn[i]) + - clu_off); + parent = ntfs_alloc_mft_parent_bio(vol, folio, sync ? &done : NULL); + if (!parent) { + err = -ENOMEM; + goto unmap_err_out; + } + ctx = container_of(parent, struct ntfs_mft_write_ctx, bio); - if (!bio_add_folio(bio, folio, folio_size, - ni->folio_ofs + offset)) { - err = -EIO; - goto put_bio_out; + parent->bi_iter.bi_sector = units[0].sector; + if (!bio_add_folio(parent, folio, units[0].len, units[0].folio_ofs)) { + err = -EIO; + goto put_bios_out; + } + + if (nr_units == 2) { + if (bio_end_sector(parent) != units[1].sector || + !bio_add_folio(parent, folio, units[1].len, + units[1].folio_ofs)) { + child = bio_alloc(vol->sb->s_bdev, 1, REQ_OP_WRITE, + GFP_NOIO); + if (!child) { + err = -ENOMEM; + goto put_bios_out; + } + child->bi_iter.bi_sector = units[1].sector; + if (!bio_add_folio(child, folio, units[1].len, + units[1].folio_ofs)) { + err = -EIO; + goto put_bios_out; + } } + } - /* Synchronize the mft mirror now if not @sync. */ - if (!sync && ni->mft_no < vol->mftmirr_size) { - int sub_err = ntfs_sync_mft_mirror(vol, ni->mft_no, - fixup_m); - if (unlikely(sub_err) && !err) - err = sub_err; - } + if (ni->mft_no < vol->mftmirr_size) { + err = ntfs_sync_mft_mirror_record(vol, folio, ni->mft_no); + if (err) + ctx->error = err; + } - if (sync) { - int sub_err = submit_bio_wait(bio); + kunmap_local(kaddr); - bio_put(bio); - if (unlikely(sub_err) && !err) - err = sub_err; - } else { - folio_get(folio); - bio->bi_private = folio; - bio->bi_end_io = ntfs_bio_end_io; - submit_bio(bio); - } - offset += vol->cluster_size; - i++; + ntfs_start_mft_writeback(ctx); + if (child) { + bio_chain(child, parent); + submit_bio(child); } + submit_bio(parent); - /* If @sync, now synchronize the mft mirror. */ - if (sync && ni->mft_no < vol->mftmirr_size) { - int sub_err = ntfs_sync_mft_mirror(vol, ni->mft_no, fixup_m); - - if (unlikely(sub_err) && !err) - err = sub_err; + if (sync) { + wait_for_completion(&done); + if (!err && parent->bi_status) + err = blk_status_to_errno(parent->bi_status); + bio_put(parent); } - kunmap_local(kaddr); + if (unlikely(err)) { /* I/O error during writing. This is really bad! */ ntfs_error(vol->sb, "I/O error while writing mft record 0x%llx! Marking base inode as bad. You should unmount the volume and run chkdsk.", ni->mft_no); - goto err_out; + return err; } done: ntfs_debug("Done."); return 0; -put_bio_out: - bio_put(bio); +put_bios_out: + if (child) + bio_put(child); + bio_put(parent); unmap_err_out: kunmap_local(kaddr); err_out: @@ -655,7 +806,8 @@ err_out: ntfs_error(vol->sb, "Not enough memory to write mft record. Redirtying so the write is retried later."); mark_mft_record_dirty(ni); - err = 0; + if (!sync) + err = 0; } else NVolSetErrors(vol); return err; @@ -714,7 +866,8 @@ static int ntfs_test_inode_wb(struct inode *vi, u64 ino, void *data) * it in @ref_vi instead of calling iput() directly. The caller must call * iput() on @ref_vi after releasing the folio lock. * - * Return 'true' if the mft record may be written out and 'false' if not. + * Return 'true' if the mft record may be written out by this path and 'false' + * if direct inode writeback owns it. * * The caller has locked the page and cleared the uptodate flag on it which * means that we can safely write out any dirty mft records that do not have @@ -811,9 +964,10 @@ static bool ntfs_may_write_mft_record(struct ntfs_volume *vol, const u64 mft_no, mft_no); /* * The write has to occur while we hold the mft record lock so - * return the locked ntfs inode. + * return the locked ntfs inode and its vfs inode reference. */ *locked_ni = ni; + *ref_vi = vi; return true; } ntfs_debug("Inode 0x%llx is not in icache.", mft_no); @@ -1198,6 +1352,10 @@ static int ntfs_mft_bitmap_extend_allocation_nolock(struct ntfs_volume *vol) size_t new_rl_count; ntfs_debug("Extending mft bitmap allocation."); + /* The initial bitmap scan must finish before we lock or change a folio. */ + if (!NVolFreeClusterKnown(vol)) + wait_event(vol->free_waitq, NVolFreeClusterKnown(vol)); + mft_ni = NTFS_I(vol->mft_ino); mftbmp_ni = NTFS_I(vol->mftbmp_ino); /* @@ -1223,6 +1381,11 @@ static int ntfs_mft_bitmap_extend_allocation_nolock(struct ntfs_volume *vol) lcn = rl->lcn + rl->length; ntfs_debug("Last lcn of mft bitmap attribute is 0x%llx.", (long long)lcn); + /* There is no adjacent cluster if the last run ends at the volume end. */ + if (lcn >= vol->nr_clusters) { + lcn = -1; + goto alloc_cluster; + } /* * Attempt to get the cluster following the last allocated cluster by * hand as it may be in the MFT zone so the allocator would not give it @@ -1241,10 +1404,18 @@ static int ntfs_mft_bitmap_extend_allocation_nolock(struct ntfs_volume *vol) folio_lock(folio); b = (u8 *)kmap_local_folio(folio, 0) + (ll & ~PAGE_MASK); tb = 1 << (lcn & 7ull); - if (*b != 0xff && !(*b & tb)) { + /* + * A page skipped by the initial scan has no free bits recorded. + * Honor that and the space reserved for delayed allocation. + */ + if (*b != 0xff && !(*b & tb) && + vol->lcn_empty_bits_per_page[ll >> PAGE_SHIFT] && + ntfs_available_clusters_count(vol, 1) > 0) { /* Next cluster is free, allocate it. */ *b |= tb; folio_mark_dirty(folio); + ntfs_dec_free_clusters(vol, 1); + ntfs_set_lcn_empty_bits(vol, ll >> PAGE_SHIFT, 1, 1); folio_unlock(folio); kunmap_local(b); folio_put(folio); @@ -1259,6 +1430,7 @@ static int ntfs_mft_bitmap_extend_allocation_nolock(struct ntfs_volume *vol) kunmap_local(b); folio_put(folio); up_write(&vol->lcnbmp_lock); +alloc_cluster: /* Allocate a cluster from the DATA_ZONE. */ rl2 = ntfs_cluster_alloc(vol, rl[1].vcn, 1, lcn, DATA_ZONE, true, false, false); @@ -1268,6 +1440,8 @@ static int ntfs_mft_bitmap_extend_allocation_nolock(struct ntfs_volume *vol) "Failed to allocate a cluster for the mft bitmap."); return PTR_ERR(rl2); } + /* The adjacent cluster may have become available while unlocked. */ + status.added_cluster = rl2->lcn == lcn; rl = ntfs_runlists_merge(&mftbmp_ni->runlist, rl2, 0, &new_rl_count); if (IS_ERR(rl)) { up_write(&mftbmp_ni->runlist.lock); @@ -1282,8 +1456,8 @@ static int ntfs_mft_bitmap_extend_allocation_nolock(struct ntfs_volume *vol) } mftbmp_ni->runlist.rl = rl; mftbmp_ni->runlist.count = new_rl_count; - status.added_run = 1; - ntfs_debug("Adding one run to mft bitmap."); + status.added_run = !status.added_cluster; + ntfs_debug("Allocated one cluster for mft bitmap."); /* Find the last run in the new runlist. */ for (; rl[1].length; rl++) ; @@ -2017,6 +2191,8 @@ static int ntfs_mft_record_format(const struct ntfs_volume *vol, const s64 mft_n return PTR_ERR(folio); } folio_lock(folio); + if (folio_test_writeback(folio)) + folio_wait_writeback(folio); folio_clear_uptodate(folio); m = (struct mft_record *)((u8 *)kmap_local_folio(folio, 0) + ofs); err = ntfs_mft_record_layout(vol, mft_no, m); @@ -2534,6 +2710,8 @@ mft_rec_already_initialized: goto undo_mftbmp_alloc; } folio_lock(folio); + if (folio_test_writeback(folio)) + folio_wait_writeback(folio); folio_clear_uptodate(folio); m = (struct mft_record *)((u8 *)kmap_local_folio(folio, 0) + ofs); /* If we just formatted the mft record no need to do it again. */ @@ -2804,7 +2982,7 @@ int ntfs_mft_record_free(struct ntfs_volume *vol, struct ntfs_inode *ni) * record to be freed is guaranteed to do it already. */ NInoSetDirty(ni); - err = write_mft_record(ni, ni_mrec, 0); + err = write_mft_record(ni, ni_mrec, 1); if (err) goto sync_rollback; @@ -2854,19 +3032,150 @@ sync_rollback: return err; } -static s64 lcn_from_index(struct ntfs_volume *vol, struct ntfs_inode *ni, - unsigned long index) +static bool ntfs_add_mft_io_unit(struct bio *bio, struct folio *folio, + const struct ntfs_mft_io_unit *unit) { - s64 vcn; - s64 lcn; + if (!bio->bi_iter.bi_size) + bio->bi_iter.bi_sector = unit->sector; + else if (bio_end_sector(bio) != unit->sector) + return false; + + return bio_add_folio(bio, folio, unit->len, unit->folio_ofs); +} + +static void ntfs_release_mft_write_refs(struct ntfs_inode **locked_nis, + unsigned int nr_locked_nis, + struct inode **ref_inos, + unsigned int nr_ref_inos) +{ + while (nr_locked_nis-- > 0) { + struct ntfs_inode *tni = locked_nis[nr_locked_nis]; + + mutex_unlock(&tni->mrec_lock); + atomic_dec(&tni->count); + } + + while (nr_ref_inos-- > 0) + iput(ref_inos[nr_ref_inos]); +} - vcn = ntfs_pidx_to_cluster(vol, index); +static int ntfs_map_mft_io_locked(struct ntfs_inode *ni, u64 folio_byte, + struct ntfs_mft_io_unit *unit) +{ + struct ntfs_volume *vol = ni->vol; + u64 file_ofs = folio_byte + unit->folio_ofs; + s64 vcn, lcn; + + lockdep_assert_held(&ni->runlist.lock); + if (!ni->runlist.rl) + return -EIO; - down_read(&ni->runlist.lock); - lcn = ntfs_attr_vcn_to_lcn_nolock(ni, vcn, false); + vcn = ntfs_bytes_to_cluster(vol, file_ofs); + /* + * $MFT is fully mapped at mount time and extensions merge already + * mapped runs under the write lock. Do not remap here: the generic + * helper may drop this read lock while upgrading it. + */ + lcn = ntfs_rl_vcn_to_lcn(ni->runlist.rl, vcn); + if (lcn == LCN_RL_NOT_MAPPED) + return -EAGAIN; + if (lcn < 0) + return -EIO; + if (ntfs_bytes_to_cluster_off(vol, file_ofs) + unit->len > + vol->cluster_size) + return -EIO; + + unit->sector = ntfs_bytes_to_bio_sector( + NTFS_CLU_TO_B(vol, lcn) + + ntfs_bytes_to_cluster_off(vol, file_ofs)); + return 0; +} + +static int ntfs_map_mft_io_for_folio(struct ntfs_inode *ni, u64 folio_byte, + struct ntfs_mft_io_unit *unit) +{ + int err; + + if (!down_read_trylock(&ni->runlist.lock)) + return -EAGAIN; + err = ntfs_map_mft_io_locked(ni, folio_byte, unit); up_read(&ni->runlist.lock); + return err; +} + +static int ntfs_prepare_mft_folio_units(struct ntfs_inode *ni, u64 folio_byte, + u64 file_limit, + const unsigned long *record_selected, + struct ntfs_mft_io_unit *units, + unsigned int *nr_units, + unsigned int max_units, bool *defer) +{ + struct ntfs_volume *vol = ni->vol; + u64 folio_end = folio_byte + PAGE_SIZE; + u64 unit_byte = folio_byte; + + while (unit_byte < folio_end && unit_byte < file_limit) { + struct ntfs_mft_io_unit unit; + u64 record_byte; + u64 unit_end; + u64 io_unit_end; + u64 cluster_end; + bool selected = false; + int err; + + io_unit_end = + round_down(unit_byte, (u64)vol->mft_io_unit_size) + + vol->mft_io_unit_size; + cluster_end = ntfs_cluster_to_bytes( + vol, ntfs_bytes_to_cluster(vol, unit_byte) + 1); + unit_end = min3(io_unit_end, cluster_end, folio_end); + + record_byte = round_down(unit_byte, (u64)vol->mft_record_size); + while (record_byte < min(unit_end, file_limit)) { + unsigned int record; + + record = div_u64(record_byte - folio_byte, + vol->mft_record_size); + if (test_bit(record, record_selected)) { + selected = true; + break; + } + record_byte += vol->mft_record_size; + } + /* + * Direct inode writeback owns unselected records. Their folio + * images are stable while this folio is locked, so a containing + * unit can preserve them when writing a selected record. + */ + if (!selected) + goto next; + + unit = (struct ntfs_mft_io_unit){ + .folio_ofs = unit_byte - folio_byte, + .len = unit_end - unit_byte, + }; + err = ntfs_map_mft_io_for_folio(ni, folio_byte, &unit); + if (err == -EAGAIN) { + *defer = true; + goto next; + } + if (err) + return err; + if (*nr_units >= max_units) + return -EOVERFLOW; + units[(*nr_units)++] = unit; +next: + unit_byte = unit_end; + } + return 0; +} - return lcn; +static void ntfs_mft_write_error(struct ntfs_volume *vol, + struct address_space *mapping, int err) +{ + mapping_set_error(mapping, err); + NVolSetErrors(vol); + ntfs_error(vol->sb, "Error while writing MFT folio: %d", err); } /* @@ -2875,9 +3184,8 @@ static s64 lcn_from_index(struct ntfs_volume *vol, struct ntfs_inode *ni, * @wbc: Writeback control structure * * This function is called as part of the address_space_operations - * .writepages implementation for the $MFT inode (or $MFTMirr). - * It handles writing one folio (normally 4KiB page) worth of MFT records - * to the underlying block device. + * .writepages implementation for the $MFT inode. Despite its historical + * name, it handles one folio containing MFT records. * * Return: 0 on success, or -errno on error. */ @@ -2887,195 +3195,190 @@ static int ntfs_write_mft_block(struct folio *folio, struct writeback_control *w struct inode *vi = mapping->host; struct ntfs_inode *ni = NTFS_I(vi); struct ntfs_volume *vol = ni->vol; - u8 *kaddr; - struct ntfs_inode **locked_nis __free(kfree) = kmalloc_objs(struct ntfs_inode *, - PAGE_SIZE / NTFS_BLOCK_SIZE, - GFP_NOFS); - int nr_locked_nis = 0, err = 0, mft_ofs, prev_mft_ofs; - struct inode **ref_inos __free(kfree) = kmalloc_objs(struct inode *, - PAGE_SIZE / NTFS_BLOCK_SIZE, - GFP_NOFS); - int nr_ref_inos = 0; - struct bio *bio = NULL; - u64 mft_no; - struct ntfs_inode *tni; - s64 lcn; - s64 vcn = ntfs_pidx_to_cluster(vol, folio->index); - s64 end_vcn = ntfs_bytes_to_cluster(vol, ni->allocated_size); - unsigned int folio_sz; - loff_t i_size = i_size_read(vi); + struct ntfs_inode **locked_nis __free(kfree) = + kmalloc_objs(struct ntfs_inode *, PAGE_SIZE / NTFS_BLOCK_SIZE, + GFP_NOFS); + struct inode **ref_inos __free(kfree) = + kmalloc_objs(struct inode *, PAGE_SIZE / NTFS_BLOCK_SIZE, GFP_NOFS); + struct ntfs_mft_io_unit *units __free(kfree) = NULL; + struct bio *parent = NULL, *child = NULL; + struct ntfs_mft_write_ctx *ctx = NULL; + u8 *kaddr = NULL; + DECLARE_BITMAP(record_selected, PAGE_SIZE / NTFS_BLOCK_SIZE) = {}; + u64 folio_byte, file_limit, folio_end, mirror_size; + unsigned int nr_records, nr_units = 0, max_units; + unsigned int nr_locked_nis = 0, nr_ref_inos = 0; + unsigned int record, unit_idx; + unsigned long flags; + loff_t i_size; + s64 allocated_size; + bool defer = false, redirty = false; + int err = 0; - ntfs_debug("Entering for inode 0x%llx, attribute type 0x%x, folio index 0x%lx.", - ni->mft_no, ni->type, folio->index); + read_lock_irqsave(&ni->size_lock, flags); + i_size = i_size_read(vi); + allocated_size = ni->allocated_size; + read_unlock_irqrestore(&ni->size_lock, flags); - if (!locked_nis || !ref_inos) { - folio_redirty_for_writepage(wbc, folio); - folio_unlock(folio); - return -ENOMEM; + ntfs_debug("Entering for inode 0x%llx, folio index 0x%lx.", ni->mft_no, + folio->index); + WARN_ON(!folio_test_locked(folio)); + + folio_byte = folio_pos(folio); + folio_end = folio_byte + PAGE_SIZE; + mirror_size = (u64)vol->mftmirr_size * vol->mft_record_size; + + if (i_size <= folio_byte) + folio_zero_segment(folio, 0, PAGE_SIZE); + else if ((u64)i_size < folio_end) + folio_zero_segment(folio, i_size - folio_byte, PAGE_SIZE); + + file_limit = min_t(u64, i_size, allocated_size); + nr_records = PAGE_SIZE / vol->mft_record_size; + max_units = PAGE_SIZE / min(vol->cluster_size, vol->mft_record_size); + units = kmalloc_array(max_units, sizeof(*units), GFP_NOFS); + if (!locked_nis || !ref_inos || !units) { + err = -ENOMEM; + goto out_noio; } - /* We have to zero every time due to mmap-at-end-of-file. */ - if (folio->index >= (i_size >> folio_shift(folio))) - /* The page straddles i_size. */ - folio_zero_segment(folio, - offset_in_folio(folio, i_size), - folio_size(folio)); + kaddr = kmap_local_folio(folio, 0); + folio_clear_uptodate(folio); + + for (record = 0; record < nr_records; record++) { + struct ntfs_inode *tni = NULL; + struct inode *ref_vi = NULL; + u64 record_byte = + folio_byte + (u64)record * vol->mft_record_size; + u64 mft_no; + + if (record_byte >= file_limit) + continue; + mft_no = record_byte >> vol->mft_record_size_bits; + if (!ntfs_may_write_mft_record( + vol, mft_no, + (struct mft_record *)(kaddr + + record * vol->mft_record_size), + &tni, &ref_vi)) { + if (ref_vi) + ref_inos[nr_ref_inos++] = ref_vi; + continue; + } + if (ref_vi) + ref_inos[nr_ref_inos++] = ref_vi; + if (tni) { + locked_nis[nr_locked_nis++] = tni; + if (tni->nr_extents < 0 && + tni->ext.base_ntfs_ino == NTFS_I(vol->mft_ino)) + continue; + } + __set_bit(record, record_selected); + } - lcn = lcn_from_index(vol, ni, folio->index); - if (lcn <= LCN_HOLE) { + err = ntfs_prepare_mft_folio_units(ni, folio_byte, file_limit, + record_selected, units, &nr_units, + max_units, &defer); + if (err) + goto out_noio; + if (!nr_units) { + if (defer) + goto out_noio; + folio_mark_uptodate(folio); + kunmap_local(kaddr); folio_start_writeback(folio); folio_unlock(folio); folio_end_writeback(folio); - return -EIO; + ntfs_release_mft_write_refs(locked_nis, nr_locked_nis, ref_inos, + nr_ref_inos); + return 0; } - /* Map folio so we can access its contents. */ - kaddr = kmap_local_folio(folio, 0); - /* Clear the page uptodate flag whilst the mst fixups are applied. */ - folio_clear_uptodate(folio); + parent = ntfs_alloc_mft_parent_bio(vol, folio, NULL); + if (!parent) { + err = -ENOMEM; + goto out_noio; + } + ctx = container_of(parent, struct ntfs_mft_write_ctx, bio); - for (mft_ofs = 0; mft_ofs < PAGE_SIZE && vcn < end_vcn; - mft_ofs += vol->mft_record_size) { - /* Get the mft record number. */ - mft_no = (((s64)folio->index << PAGE_SHIFT) + mft_ofs) >> - vol->mft_record_size_bits; - vcn = ntfs_mft_no_to_cluster(vol, mft_no); - /* Check whether to write this mft record. */ - tni = NULL; - if (ntfs_may_write_mft_record(vol, mft_no, - (struct mft_record *)(kaddr + mft_ofs), - &tni, &ref_inos[nr_ref_inos])) { - unsigned int mft_record_off = 0; - s64 vcn_off = vcn; - s64 rl_len = 0; + if (!ntfs_add_mft_io_unit(parent, folio, &units[0])) { + err = -EIO; + goto out_noio; + } - /* - * The record should be written. If a locked ntfs - * inode was returned, add it to the array of locked - * ntfs inodes. - */ - if (tni) - locked_nis[nr_locked_nis++] = tni; - else if (ref_inos[nr_ref_inos]) - nr_ref_inos++; - - if (bio && (mft_ofs != prev_mft_ofs + vol->mft_record_size)) { -flush_bio: - bio->bi_end_io = ntfs_bio_end_io; - submit_bio(bio); - bio = NULL; - } + for (unit_idx = 0; unit_idx < nr_units; unit_idx++) { + struct ntfs_mft_io_unit *unit = &units[unit_idx]; + u64 file_ofs = folio_byte + unit->folio_ofs; - if (vol->cluster_size < folio_size(folio)) { - struct runlist_element *rl; + if (unit_idx) { + struct bio *target = child ? child : parent; - down_write(&ni->runlist.lock); - rl = ntfs_attr_vcn_to_rl(ni, vcn_off, &lcn); - if (!IS_ERR(rl)) - rl_len = rl->length - (vcn_off - rl->vcn); - up_write(&ni->runlist.lock); - if (IS_ERR(rl) || lcn < 0) { - err = -EIO; - goto unm_done; + if (!ntfs_add_mft_io_unit(target, folio, unit)) { + if (child) { + ntfs_start_mft_writeback(ctx); + bio_chain(child, parent); + submit_bio(child); + child = NULL; } - - if (bio && - (bio_end_sector(bio) >> (vol->cluster_size_bits - 9)) != - lcn) { - bio->bi_end_io = ntfs_bio_end_io; - submit_bio(bio); - bio = NULL; + child = bio_alloc(vol->sb->s_bdev, 1, + REQ_OP_WRITE, GFP_NOIO); + if (!child) { + defer = true; + break; + } + if (!ntfs_add_mft_io_unit(child, folio, unit)) { + bio_put(child); + child = NULL; + if (!ctx->error) + ctx->error = -EIO; + redirty = true; + err = -EIO; + break; } } + } + if (file_ofs < mirror_size) { + int mirror_err = ntfs_sync_mft_mirror_unit( + vol, folio, file_ofs, unit); - if (!bio) { - unsigned int off; - - off = ((mft_no << vol->mft_record_size_bits) + - mft_record_off) & vol->cluster_size_mask; - - bio = bio_alloc(vol->sb->s_bdev, 1, REQ_OP_WRITE, - GFP_NOIO); - bio->bi_iter.bi_sector = - ntfs_bytes_to_bio_sector( - ntfs_cluster_to_bytes(vol, lcn) + off); - } - - if (vol->cluster_size == NTFS_BLOCK_SIZE && - (mft_record_off || - rl_len == 1 || - mft_ofs + NTFS_BLOCK_SIZE >= PAGE_SIZE)) - folio_sz = NTFS_BLOCK_SIZE; - else - folio_sz = vol->mft_record_size; - if (!bio_add_folio(bio, folio, folio_sz, - mft_ofs + mft_record_off)) { - err = -EIO; - bio_put(bio); - goto unm_done; - } - mft_record_off += folio_sz; - - if (mft_record_off != vol->mft_record_size) { - vcn_off++; - goto flush_bio; - } - prev_mft_ofs = mft_ofs; - - if (mft_no < vol->mftmirr_size) { - int sub_err = ntfs_sync_mft_mirror(vol, mft_no, - (struct mft_record *)(kaddr + mft_ofs)); - - if (unlikely(sub_err) && !err) - err = sub_err; - } - } else if (ref_inos[nr_ref_inos]) - nr_ref_inos++; + if (mirror_err && !ctx->error) + ctx->error = mirror_err; + } } - if (bio) { - bio->bi_end_io = ntfs_bio_end_io; - submit_bio(bio); + if (child) { + ntfs_start_mft_writeback(ctx); + bio_chain(child, parent); + submit_bio(child); } -unm_done: + ntfs_start_mft_writeback(ctx); folio_mark_uptodate(folio); + if (defer || redirty) + folio_redirty_for_writepage(wbc, folio); kunmap_local(kaddr); - - folio_start_writeback(folio); + kaddr = NULL; folio_unlock(folio); - folio_end_writeback(folio); + submit_bio(parent); + ntfs_release_mft_write_refs(locked_nis, nr_locked_nis, ref_inos, + nr_ref_inos); - /* Unlock any locked inodes. */ - while (nr_locked_nis-- > 0) { - struct ntfs_inode *base_tni; - - tni = locked_nis[nr_locked_nis]; - mutex_unlock(&tni->mrec_lock); - - /* Get the base inode. */ - mutex_lock(&tni->extent_lock); - if (tni->nr_extents >= 0) - base_tni = tni; - else - base_tni = tni->ext.base_ntfs_ino; - mutex_unlock(&tni->extent_lock); - ntfs_debug("Unlocking %s inode 0x%llx.", - tni == base_tni ? "base" : "extent", - tni->mft_no); - atomic_dec(&tni->count); - iput(VFS_I(base_tni)); - } + return 0; - /* Dropping deferred references */ - while (nr_ref_inos-- > 0) { - if (ref_inos[nr_ref_inos]) - iput(ref_inos[nr_ref_inos]); +out_noio: + if (kaddr) { + folio_mark_uptodate(folio); + kunmap_local(kaddr); } - - if (unlikely(err && err != -ENOMEM)) - NVolSetErrors(vol); - if (likely(!err)) - ntfs_debug("Done."); + if (err && err != -ENOMEM) + ntfs_mft_write_error(vol, mapping, err); + if (parent) + bio_put(parent); + if (redirty || defer || err == -ENOMEM) + folio_redirty_for_writepage(wbc, folio); + folio_unlock(folio); + ntfs_release_mft_write_refs(locked_nis, nr_locked_nis, ref_inos, + nr_ref_inos); + if (err == -ENOMEM || (defer && !err)) + return 0; return err; } diff --git a/fs/ntfs/mft.h b/fs/ntfs/mft.h index ed5c1d595c0d..d2a31205e08c 100644 --- a/fs/ntfs/mft.h +++ b/fs/ntfs/mft.h @@ -42,8 +42,8 @@ static inline void mark_mft_record_dirty(struct ntfs_inode *ni) __mark_mft_record_dirty(ni); } -int ntfs_sync_mft_mirror(struct ntfs_volume *vol, const u64 mft_no, - struct mft_record *m); +int ntfs_mft_bioset_init(void); +void ntfs_mft_bioset_exit(void); int write_mft_record_nolock(struct ntfs_inode *ni, struct mft_record *m, int sync); /* @@ -53,16 +53,19 @@ int write_mft_record_nolock(struct ntfs_inode *ni, struct mft_record *m, int syn * @sync: if true, wait for i/o completion * * This is just a wrapper for write_mft_record_nolock() (see mft.c), which - * locks the page for the duration of the write. This ensures that there are - * no race conditions between writing the mft record via the dirty inode code - * paths and via the page cache write back code paths or between writing - * neighbouring mft records residing in the same page. + * locks the folio while preparing the write. write_mft_record_nolock() waits + * for prior folio writeback before modifying the folio and keeps PG_writeback + * set until the submitted I/O completes. Together these serialize dirty + * inode writes, page cache writeback, and neighbouring mft record writes in + * the same folio. * * Locking the page also serializes us against ->read_folio() if the page is not * uptodate. * - * On success, clean the mft record and return 0. On error, leave the mft - * record dirty and return -errno. + * On success, clean the mft record and return 0. On allocation failure, + * redirty the record for retry. Asynchronous callers return 0 after + * redirtying, while synchronous callers receive -ENOMEM. On other errors, + * return -errno and mark the volume with errors. */ static inline int write_mft_record(struct ntfs_inode *ni, struct mft_record *m, int sync) { diff --git a/fs/ntfs/mst.c b/fs/ntfs/mst.c index 7f9faad924ad..cc6363413449 100644 --- a/fs/ntfs/mst.c +++ b/fs/ntfs/mst.c @@ -19,7 +19,7 @@ * magic of the ntfs record header being processed with "BAAD" (in memory only!) * and abort processing. * - * Return 0 on success and -EINVAL on error ("BAAD" magic will be present). + * Return 0 on success and -EIO on error ("BAAD" magic will be present). * * NOTE: We consider the absence / invalidity of an update sequence array to * mean that the structure is not protected at all and hence doesn't need to @@ -71,7 +71,7 @@ int post_read_mst_fixup(struct ntfs_record *b, const u32 size) * Note that magic_BAAD is already converted to le32. */ b->magic = magic_BAAD; - return -EINVAL; + return -EIO; } data_pos += NTFS_BLOCK_SIZE / sizeof(u16); } diff --git a/fs/ntfs/namei.c b/fs/ntfs/namei.c index fdf52fac4329..028efad7d9c0 100644 --- a/fs/ntfs/namei.c +++ b/fs/ntfs/namei.c @@ -391,7 +391,7 @@ static int ntfs_sd_add_everyone(struct ntfs_inode *ni) return ret; } -static struct ntfs_inode *__ntfs_create(struct mnt_idmap *idmap, struct inode *dir, +static struct ntfs_inode *__ntfs_create(const struct mnt_idmap *idmap, struct inode *dir, __le16 *name, u8 name_len, mode_t mode, dev_t dev, const char *target, int target_len) { @@ -433,9 +433,8 @@ static struct ntfs_inode *__ntfs_create(struct mnt_idmap *idmap, struct inode *d ni->itype.index.vcn_size_bits = vol->cluster_size_bits; } else { - ni->itype.index.vcn_size = vol->sector_size; - ni->itype.index.vcn_size_bits = - vol->sector_size_bits; + ni->itype.index.vcn_size = NTFS_BLOCK_SIZE; + ni->itype.index.vcn_size_bits = NTFS_BLOCK_SIZE_BITS; } } @@ -569,7 +568,7 @@ static struct ntfs_inode *__ntfs_create(struct mnt_idmap *idmap, struct inode *d NTFS_B_TO_CLU(vol, ni->vol->index_record_size); else ir->clusters_per_index_block = - ni->vol->index_record_size >> ni->vol->sector_size_bits; + ni->vol->index_record_size >> NTFS_BLOCK_SIZE_BITS; ir->index.entries_offset = cpu_to_le32(sizeof(struct index_header)); ir->index.index_length = cpu_to_le32(index_len); ir->index.allocated_size = cpu_to_le32(index_len); @@ -732,7 +731,7 @@ err_out: return ERR_PTR(err); } -static int ntfs_create(struct mnt_idmap *idmap, struct inode *dir, +static int ntfs_create(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *dentry, umode_t mode) { struct ntfs_volume *vol = NTFS_SB(dir->i_sb); @@ -757,8 +756,7 @@ static int ntfs_create(struct mnt_idmap *idmap, struct inode *dir, return err; } - if (!(vol->vol_flags & VOLUME_IS_DIRTY)) - ntfs_set_volume_flags(vol, VOLUME_IS_DIRTY); + ntfs_set_volume_flags(vol, VOLUME_IS_DIRTY); ni = __ntfs_create(idmap, dir, uname, uname_len, S_IFREG | mode, 0, NULL, 0); kmem_cache_free(ntfs_name_cache, uname); @@ -1032,8 +1030,7 @@ static int ntfs_unlink(struct inode *dir, struct dentry *dentry) return err; } - if (!(vol->vol_flags & VOLUME_IS_DIRTY)) - ntfs_set_volume_flags(vol, VOLUME_IS_DIRTY); + ntfs_set_volume_flags(vol, VOLUME_IS_DIRTY); err = ntfs_delete(ni, NTFS_I(dir), uname, uname_len, true); if (err) @@ -1049,7 +1046,7 @@ out: return err; } -static struct dentry *ntfs_mkdir(struct mnt_idmap *idmap, struct inode *dir, +static struct dentry *ntfs_mkdir(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *dentry, umode_t mode) { struct super_block *sb = dir->i_sb; @@ -1076,8 +1073,7 @@ static struct dentry *ntfs_mkdir(struct mnt_idmap *idmap, struct inode *dir, return ERR_PTR(err); } - if (!(vol->vol_flags & VOLUME_IS_DIRTY)) - ntfs_set_volume_flags(vol, VOLUME_IS_DIRTY); + ntfs_set_volume_flags(vol, VOLUME_IS_DIRTY); ni = __ntfs_create(idmap, dir, uname, uname_len, mode, 0, NULL, 0); kmem_cache_free(ntfs_name_cache, uname); @@ -1118,8 +1114,7 @@ static int ntfs_rmdir(struct inode *dir, struct dentry *dentry) return err; } - if (!(vol->vol_flags & VOLUME_IS_DIRTY)) - ntfs_set_volume_flags(vol, VOLUME_IS_DIRTY); + ntfs_set_volume_flags(vol, VOLUME_IS_DIRTY); err = ntfs_delete(ni, NTFS_I(dir), uname, uname_len, true); if (err) @@ -1248,7 +1243,7 @@ err_out: return err; } -static int ntfs_rename(struct mnt_idmap *idmap, struct inode *old_dir, +static int ntfs_rename(const struct mnt_idmap *idmap, struct inode *old_dir, struct dentry *old_dentry, struct inode *new_dir, struct dentry *new_dentry, unsigned int flags) { @@ -1305,8 +1300,7 @@ static int ntfs_rename(struct mnt_idmap *idmap, struct inode *old_dir, new_dir_first = is_subdir(new_dentry->d_parent, old_dentry->d_parent); - if (!(vol->vol_flags & VOLUME_IS_DIRTY)) - ntfs_set_volume_flags(vol, VOLUME_IS_DIRTY); + ntfs_set_volume_flags(vol, VOLUME_IS_DIRTY); mutex_lock_nested(&old_ni->mrec_lock, NTFS_INODE_MUTEX_NORMAL); if (new_ni) @@ -1399,7 +1393,7 @@ err_out: return err; } -static int ntfs_symlink(struct mnt_idmap *idmap, struct inode *dir, +static int ntfs_symlink(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *dentry, const char *symname) { struct super_block *sb = dir->i_sb; @@ -1429,8 +1423,7 @@ static int ntfs_symlink(struct mnt_idmap *idmap, struct inode *dir, goto out; } - if (!(vol->vol_flags & VOLUME_IS_DIRTY)) - ntfs_set_volume_flags(vol, VOLUME_IS_DIRTY); + ntfs_set_volume_flags(vol, VOLUME_IS_DIRTY); ni = __ntfs_create(idmap, dir, usrc, usrc_len, S_IFLNK | 0777, 0, symname, symlen); @@ -1447,7 +1440,7 @@ out: return err; } -static int ntfs_mknod(struct mnt_idmap *idmap, struct inode *dir, +static int ntfs_mknod(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *dentry, umode_t mode, dev_t rdev) { struct super_block *sb = dir->i_sb; @@ -1474,8 +1467,7 @@ static int ntfs_mknod(struct mnt_idmap *idmap, struct inode *dir, return err; } - if (!(vol->vol_flags & VOLUME_IS_DIRTY)) - ntfs_set_volume_flags(vol, VOLUME_IS_DIRTY); + ntfs_set_volume_flags(vol, VOLUME_IS_DIRTY); switch (mode & S_IFMT) { case S_IFCHR: @@ -1521,8 +1513,7 @@ static int ntfs_link(struct dentry *old_dentry, struct inode *dir, return -ENOMEM; } - if (!(vol->vol_flags & VOLUME_IS_DIRTY)) - ntfs_set_volume_flags(vol, VOLUME_IS_DIRTY); + ntfs_set_volume_flags(vol, VOLUME_IS_DIRTY); ihold(vi); mutex_lock_nested(&ni->mrec_lock, NTFS_INODE_MUTEX_NORMAL); diff --git a/fs/ntfs/ntfs.h b/fs/ntfs/ntfs.h index 45f77848a9cf..9557a5c36b81 100644 --- a/fs/ntfs/ntfs.h +++ b/fs/ntfs/ntfs.h @@ -80,6 +80,8 @@ enum { NTFS_MAX_LABEL_LEN = 128, }; +#define NTFS_MAX_ATTR_LIST_SIZE (256 * 1024) + enum { CASE_SENSITIVE = 0, IGNORE_CASE = 1, @@ -219,7 +221,6 @@ struct option_t { }; extern const struct option_t on_errors_arr[]; int ntfs_set_volume_flags(struct ntfs_volume *vol, __le16 flags); -int ntfs_clear_volume_flags(struct ntfs_volume *vol, __le16 flags); int ntfs_write_volume_label(struct ntfs_volume *vol, char *label); /* From fs/ntfs/mst.c */ @@ -234,14 +235,14 @@ bool ntfs_are_names_equal(const __le16 *s1, size_t s1_len, const __le16 *upcase, const u32 upcase_size); int ntfs_collate_names(const __le16 *name1, const u32 name1_len, const __le16 *name2, const u32 name2_len, - const int err_val, const u32 ic, + const bool check_invalid, const u32 ic, const __le16 *upcase, const u32 upcase_len); int ntfs_ucsncmp(const __le16 *s1, const __le16 *s2, size_t n); int ntfs_ucsncasecmp(const __le16 *s1, const __le16 *s2, size_t n, const __le16 *upcase, const u32 upcase_size); int ntfs_file_compare_values(const struct file_name_attr *file_name_attr1, const struct file_name_attr *file_name_attr2, - const int err_val, const u32 ic, + const bool check_invalid, const u32 ic, const __le16 *upcase, const u32 upcase_len); int ntfs_nlstoucs(const struct ntfs_volume *vol, const char *ins, const int ins_len, __le16 **outs, int max_name_len); diff --git a/fs/ntfs/reparse.c b/fs/ntfs/reparse.c index 1a6073e22677..2a6480211fe3 100644 --- a/fs/ntfs/reparse.c +++ b/fs/ntfs/reparse.c @@ -676,7 +676,7 @@ static int update_reparse_data(struct ntfs_inode *ni, struct ntfs_index_context rp_inode = ntfs_attr_iget(VFS_I(ni), AT_REPARSE_POINT, AT_UNNAMED, 0); if (IS_ERR(rp_inode)) - return -EINVAL; + return PTR_ERR(rp_inode); rp_ni = NTFS_I(rp_inode); /* remove the existing reparse data */ diff --git a/fs/ntfs/runlist.c b/fs/ntfs/runlist.c index 3a61f19bcbee..f13d28c15887 100644 --- a/fs/ntfs/runlist.c +++ b/fs/ntfs/runlist.c @@ -680,9 +680,8 @@ struct runlist_element *ntfs_runlists_merge(struct runlist *d_runlist, /* Add an unmapped runlist element. */ if (!slots) { drl = ntfs_rl_realloc_nofail(drl, ds, - ds + 2); + ds + 3); slots = 2; - *new_rl_count += 2; } ds++; /* Need to set vcn if it isn't set already. */ @@ -698,11 +697,11 @@ struct runlist_element *ntfs_runlists_merge(struct runlist *d_runlist, ds++; if (!slots) { drl = ntfs_rl_realloc_nofail(drl, ds, ds + 1); - *new_rl_count += 1; } drl[ds].vcn = marker_vcn; drl[ds].lcn = LCN_ENOENT; drl[ds].length = (s64)0; + *new_rl_count = ds + 1; } } } diff --git a/fs/ntfs/super.c b/fs/ntfs/super.c index 4066bacabe37..716ba775f1f3 100644 --- a/fs/ntfs/super.c +++ b/fs/ntfs/super.c @@ -18,6 +18,7 @@ #include "sysctl.h" #include "logfile.h" #include "index.h" +#include "mft.h" #include "ntfs.h" #include "ea.h" #include "volume.h" @@ -262,14 +263,23 @@ static int ntfs_parse_param(struct fs_context *fc, struct fs_parameter *param) return 0; } +static int ntfs_sync_volume_dirty_state(struct ntfs_volume *vol); + static int ntfs_reconfigure(struct fs_context *fc) { struct super_block *sb = fc->root->d_sb; struct ntfs_volume *vol = NTFS_SB(sb); + int err; ntfs_debug("Entering with remount"); - sync_filesystem(sb); + err = sync_filesystem(sb); + if (err) { + ntfs_warning(sb, "Failed to sync the filesystem."); + /* A forced remount must still turn the superblock read-only. */ + if (!(fc->sb_flags & SB_FORCE)) + return err; + } /* * For the read-write compiled driver, if we are remounting read-write, @@ -286,6 +296,12 @@ static int ntfs_reconfigure(struct fs_context *fc) static const char *es = ". Cannot remount read-write."; /* Remounting read-write. */ + if (!vol->mft_write_supported) { + ntfs_error(sb, + "MFT writeback is unsupported for this geometry%s", + es); + return -EROFS; + } if (NVolErrors(vol)) { ntfs_error(sb, "Volume has errors and is read-only%s", es); @@ -312,10 +328,58 @@ static int ntfs_reconfigure(struct fs_context *fc) } } else if (!sb_rdonly(sb) && (fc->sb_flags & SB_RDONLY)) { /* Remounting read-only. */ - if (!NVolErrors(vol)) { - if (ntfs_clear_volume_flags(vol, VOLUME_IS_DIRTY)) + /* + * With errors recorded the dirty bit is set rather than + * cleared, and it is committed right away: the VFS does + * not sync the filesystem during a remount, and once the + * remount succeeds no further persistence point exists - + * ntfs_sync_fs() is only ever invoked for read-write + * superblocks (all its VFS callers skip read-only ones) + * and ntfs_put_super() skips them, so the only remaining + * write would be the evict-time commit at unmount, which + * a crash never reaches. An error recorded only after + * the remount is still never persisted; a failed commit + * or flush fails the remount, leaving the superblock + * read-write so ntfs_put_super() retries at unmount. + */ + /* + * A forced remount does not drain writers in progress, + * so one may still be modifying metadata when the flags + * are committed: never clear the dirty bit then; if + * errors have been recorded, the update preserves or + * sets it; otherwise, skip the update entirely. + */ + if (!(fc->sb_flags & SB_FORCE) || NVolErrors(vol)) { + err = ntfs_sync_volume_dirty_state(vol); + if (err) { + ntfs_warning(sb, + "Failed to update dirty bit in volume information flags. Run chkdsk."); + return err; + } + } + if (NInoDirty(NTFS_I(vol->vol_ino))) { + /* ntfs_commit_inode() would discard the error. */ + err = __ntfs_write_inode(vol->vol_ino, 1); + if (err) { + ntfs_warning(sb, + "Failed to commit volume information flags. Run chkdsk."); + return err; + } + /* + * write_mft_record() redirties the record on + * -ENOMEM and still reports success. + */ + if (NInoDirty(NTFS_I(vol->vol_ino))) { ntfs_warning(sb, - "Failed to clear dirty bit in volume information flags. Run chkdsk."); + "Volume information flags remain dirty after commit. Run chkdsk."); + return -EIO; + } + err = blkdev_issue_flush(sb->s_bdev); + if (err) { + ntfs_warning(sb, + "Failed to flush volume information flags. Run chkdsk."); + return err; + } } } @@ -353,31 +417,52 @@ void ntfs_handle_error(struct super_block *sb) } /* - * ntfs_write_volume_flags - write new flags to the volume information flags + * ntfs_write_volume_flags - apply flag changes to the volume information flags * @vol: ntfs volume on which to modify the flags - * @flags: new flags value for the volume information flags + * @set_bits: bits to set in the volume information flags + * @clear_bits: bits to clear in the volume information flags + * @dirty_if_errors: force VOLUME_IS_DIRTY on when NVolErrors() is set * * Internal function. You probably want to use ntfs_{set,clear}_volume_flags() - * instead (see below). + * or ntfs_sync_volume_dirty_state() instead (see below). * - * Replace the volume information flags on the volume @vol with the value - * supplied in @flags. Note, this overwrites the volume information flags, so - * make sure to combine the flags you want to modify with the old flags and use - * the result when calling ntfs_write_volume_flags(). + * Combine @set_bits and @clear_bits with the current in-memory flag state and + * write the result back. The set/clear helpers pass only the bits to modify, + * not the complete flag state. The read-modify-write happens under + * ni->mrec_lock so that concurrent set/clear operations cannot lose updates. + * All bit manipulation is done on CPU-endian values, and the result is + * converted back to little-endian before storing it. + * + * When @dirty_if_errors is true and errors have been recorded on @vol, + * VOLUME_IS_DIRTY is forced on after the requested changes. NVolErrors() is + * evaluated under the same mrec_lock, which orders this against other + * locked flag updates; the runtime error paths themselves record the flag + * lock-free, so see ntfs_sync_volume_dirty_state() for the guarantee this + * provides against them. * * Return 0 on success and -errno on error. */ -static int ntfs_write_volume_flags(struct ntfs_volume *vol, const __le16 flags) +static int ntfs_write_volume_flags(struct ntfs_volume *vol, + const __le16 set_bits, const __le16 clear_bits, + const bool dirty_if_errors) { struct ntfs_inode *ni = NTFS_I(vol->vol_ino); struct volume_information *vi; struct ntfs_attr_search_ctx *ctx; + u16 flags; int err; - ntfs_debug("Entering, old flags = 0x%x, new flags = 0x%x.", - le16_to_cpu(vol->vol_flags), le16_to_cpu(flags)); mutex_lock(&ni->mrec_lock); - if (vol->vol_flags == flags) + + flags = le16_to_cpu(vol->vol_flags); + flags |= le16_to_cpu(set_bits) & le16_to_cpu(VOLUME_FLAGS_MASK); + flags &= ~(le16_to_cpu(clear_bits) & le16_to_cpu(VOLUME_FLAGS_MASK)); + if (dirty_if_errors && NVolErrors(vol)) + flags |= le16_to_cpu(VOLUME_IS_DIRTY); + ntfs_debug("Entering, old flags = 0x%x, new flags = 0x%x.", + le16_to_cpu(vol->vol_flags), flags); + + if (le16_to_cpu(vol->vol_flags) == flags) goto done; ctx = ntfs_attr_get_search_ctx(ni, NULL); @@ -393,7 +478,7 @@ static int ntfs_write_volume_flags(struct ntfs_volume *vol, const __le16 flags) vi = (struct volume_information *)((u8 *)ctx->attr + le16_to_cpu(ctx->attr->data.resident.value_offset)); - vol->vol_flags = vi->flags = flags; + vol->vol_flags = vi->flags = cpu_to_le16(flags); mark_mft_record_dirty(ctx->ntfs_ino); ntfs_attr_put_search_ctx(ctx); done: @@ -414,29 +499,56 @@ put_unm_err_out: * @flags: flags to set on the volume * * Set the bits in @flags in the volume information flags on the volume @vol. + * The bits are combined with the current flag state under the lock in + * ntfs_write_volume_flags(), so concurrent updates are not lost. * * Return 0 on success and -errno on error. */ int ntfs_set_volume_flags(struct ntfs_volume *vol, __le16 flags) { - flags &= VOLUME_FLAGS_MASK; - return ntfs_write_volume_flags(vol, vol->vol_flags | flags); + return ntfs_write_volume_flags(vol, flags, 0, false); } /* - * ntfs_clear_volume_flags - clear bits in the volume information flags - * @vol: ntfs volume on which to modify the flags - * @flags: flags to clear on the volume + * ntfs_sync_volume_dirty_state - persist the dirty bit per the error state + * @vol: ntfs volume whose dirty bit to persist * - * Clear the bits in @flags in the volume information flags on the volume @vol. + * Set VOLUME_IS_DIRTY if errors have been recorded on @vol and clear it + * otherwise, under the $Volume mrec_lock. + * + * The guarantee this provides is eventual, not instantaneous: the runtime + * error paths record NVolErrors() with a lock-free set_bit(), so a + * persistence point that evaluates the flag just before an error is + * recorded can still leave the on-disk bit clean. This is sound because + * NVolErrors() is sticky (nothing clears it for the lifetime of the mount) + * and every persistence point re-derives the on-disk bit from it; the + * last one, ntfs_put_super(), runs after evict_inodes() on a quiesced + * filesystem, so a volume that is read-write at unmount time cannot + * unmount clean. A volume that is already read-only when the error is + * recorded (errors=remount-ro flips the superblock on the first error, + * as does an earlier remount-ro) has no persistence point left and + * keeps whatever on-disk bit it had; that behaviour is unchanged. The + * residual window is a crash between the error and the next + * persistence point. + * + * This is the single point that persists the in-memory error state to disk. + * The runtime error paths only record NVolErrors() because they run under a + * variety of ntfs locks the dirty-bit write cannot be taken under (runlist + * locks, vol->lcnbmp_lock, vol->mftbmp_lock, mrec_locks); a sync of a + * volume with recorded errors, a remount to read-only, or the unmount, + * then persists the flag here. + * + * A hibernated volume is not written from these persistence paths: + * resuming Windows from a modified image corrupts it, so the dirty bit + * is left as it is on disk and only the in-memory error state is kept. * * Return 0 on success and -errno on error. */ -int ntfs_clear_volume_flags(struct ntfs_volume *vol, __le16 flags) +static int ntfs_sync_volume_dirty_state(struct ntfs_volume *vol) { - flags &= VOLUME_FLAGS_MASK; - flags = vol->vol_flags & cpu_to_le16(~le16_to_cpu(flags)); - return ntfs_write_volume_flags(vol, flags); + if (NVolHibernated(vol)) + return 0; + return ntfs_write_volume_flags(vol, 0, VOLUME_IS_DIRTY, true); } int ntfs_write_volume_label(struct ntfs_volume *vol, char *label) @@ -530,6 +642,8 @@ out: static bool is_boot_sector_ntfs(const struct super_block *sb, const struct ntfs_boot_sector *b, const bool silent) { + u16 sector_size = le16_to_cpu(b->bpb.bytes_per_sector); + /* * Check that checksum == sum of u32 values from b to the checksum * field. If checksum is zero, no checking is done. We will work when @@ -550,8 +664,8 @@ static bool is_boot_sector_ntfs(const struct super_block *sb, if (b->oem_id != magicNTFS) goto not_ntfs; /* Check bytes per sector value is between 256 and 4096. */ - if (le16_to_cpu(b->bpb.bytes_per_sector) < 0x100 || - le16_to_cpu(b->bpb.bytes_per_sector) > 0x1000) + if (sector_size < 0x100 || sector_size > 0x1000 || + !is_power_of_2(sector_size)) goto not_ntfs; /* * Check sectors per cluster value is valid and the cluster size @@ -632,6 +746,60 @@ static char *read_ntfs_boot_sector(struct super_block *sb, return boot_sector; } +static bool ntfs_validate_mft_io_geometry(struct ntfs_volume *vol, + const unsigned int logical_block_size) +{ + struct super_block *sb = vol->sb; + u32 mft_io_unit_size = 0; + bool mft_write_supported = true; + + if (!is_power_of_2(logical_block_size) || + logical_block_size > PAGE_SIZE || + !is_power_of_2(vol->sector_size) || + vol->sector_size < logical_block_size || + vol->sector_size % logical_block_size || + !is_power_of_2(vol->cluster_size) || + !is_power_of_2(vol->mft_record_size) || + vol->mft_record_size > PAGE_SIZE) + goto err; + + mft_io_unit_size = max(logical_block_size, vol->mft_record_size); + if (mft_io_unit_size > PAGE_SIZE || + mft_io_unit_size % logical_block_size || + mft_io_unit_size % vol->mft_record_size) + goto err; + + /* + * The current direct-write path stores at most two MFT runlist + * segments. Keep valid but larger records read-only until that path + * can map an arbitrary number of segments. + */ + if (vol->mft_record_size > 2 * (u64)vol->cluster_size) + mft_write_supported = false; + + /* + * A containing device block is currently mapped through one MFT + * runlist element. Keep valid geometries that require crossing a + * cluster read-only until the mapping is generalized. + */ + if (mft_io_unit_size > vol->mft_record_size && + (vol->cluster_size < mft_io_unit_size || + vol->cluster_size % mft_io_unit_size)) + mft_write_supported = false; + + vol->mft_io_unit_size = mft_io_unit_size; + vol->mft_write_supported = mft_write_supported; + return true; + +err: + ntfs_error(sb, + "Unsupported MFT I/O geometry (logical %u, block %lu, sector %u, cluster %u, MFT record %u, I/O unit %u).", + logical_block_size, sb->s_blocksize, + (unsigned int)vol->sector_size, vol->cluster_size, + vol->mft_record_size, mft_io_unit_size); + return false; +} + /* * parse_ntfs_boot_sector - parse the boot sector and store the data in @vol * @vol: volume structure to initialise with data from boot sector @@ -644,9 +812,11 @@ static bool parse_ntfs_boot_sector(struct ntfs_volume *vol, const struct ntfs_boot_sector *b) { unsigned int sectors_per_cluster, sectors_per_cluster_bits, nr_hidden_sects; + unsigned int logical_block_size; int clusters_per_mft_record, clusters_per_index_record; u64 ll; + logical_block_size = bdev_logical_block_size(vol->sb->s_bdev); vol->sector_size = le16_to_cpu(b->bpb.bytes_per_sector); vol->sector_size_bits = ffs(vol->sector_size) - 1; ntfs_debug("vol->sector_size = %i (0x%x)", vol->sector_size, @@ -719,6 +889,9 @@ static bool parse_ntfs_boot_sector(struct ntfs_volume *vol, ntfs_warning(vol->sb, "Mft record size (%i) is smaller than the sector size (%i).", vol->mft_record_size, vol->sector_size); } + if (!ntfs_validate_mft_io_geometry(vol, logical_block_size)) + return false; + clusters_per_index_record = b->clusters_per_index_record; ntfs_debug("clusters_per_index_record = %i (0x%x)", clusters_per_index_record, clusters_per_index_record); @@ -1235,11 +1408,8 @@ static bool load_and_init_attrdef(struct ntfs_volume *vol) ntfs_debug("Entering."); /* Read attrdef table and setup vol->attrdef and vol->attrdef_size. */ ino = ntfs_iget(sb, FILE_AttrDef); - if (IS_ERR(ino)) { - if (!IS_ERR(ino)) - iput(ino); + if (IS_ERR(ino)) goto failed; - } NInoSetSparseDisabled(NTFS_I(ino)); /* FILE_AttrDef must hold at least one entry and fit inside 31 bits. */ i_size = i_size_read(ino); @@ -1301,11 +1471,8 @@ static bool load_and_init_upcase(struct ntfs_volume *vol) ntfs_debug("Entering."); /* Read upcase table and setup vol->upcase and vol->upcase_len. */ ino = ntfs_iget(sb, FILE_UpCase); - if (IS_ERR(ino)) { - if (!IS_ERR(ino)) - iput(ino); + if (IS_ERR(ino)) goto upcase_failed; - } /* * The upcase size must not be above 64k Unicode characters, must not * be zero and must be a multiple of sizeof(__le16). @@ -1404,6 +1571,7 @@ static bool load_system_files(struct ntfs_volume *vol) struct ntfs_attr_search_ctx *ctx; struct restart_page_header *rp; int err; + u8 saved_on_errors; ntfs_debug("Entering."); /* Get mft mirror inode compare the contents of $MFT and $MFTMirr. */ @@ -1468,8 +1636,7 @@ bitmap_failed: */ vol->vol_ino = ntfs_iget(sb, FILE_Volume); if (IS_ERR(vol->vol_ino)) { - if (!IS_ERR(vol->vol_ino)) - iput(vol->vol_ino); + vol->vol_ino = NULL; volume_failed: ntfs_error(sb, "Failed to load $Volume."); goto iput_lcnbmp_err_out; @@ -1478,6 +1645,7 @@ volume_failed: if (IS_ERR(m)) { iput_volume_failed: iput(vol->vol_ino); + vol->vol_ino = NULL; goto volume_failed; } @@ -1578,8 +1746,19 @@ get_ctx_vol_failed: * NVolErrors() without setting the dirty volume flag and mount * read-only. This will prevent read-write remounting and it will also * prevent all writes. + * + * Nested lookup and inode-loading errors must not panic before the + * read-only fallback has run. Temporarily use errors=remount-ro + * instead of errors=panic, preserving any read-only transition even + * if an error is not propagated to the check's return value or + * recorded in NVolErrors(). This is safe during initial mount, + * before the super block is published. */ + saved_on_errors = vol->on_errors; + if (saved_on_errors == ON_ERRORS_PANIC) + vol->on_errors = ON_ERRORS_REMOUNT_RO; err = check_windows_hibernation_status(vol); + vol->on_errors = saved_on_errors; if (unlikely(err)) { static const char *es1a = "Failed to determine if Windows is hibernated"; static const char *es1b = "Windows is hibernated"; @@ -1587,12 +1766,21 @@ get_ctx_vol_failed: const char *es1; es1 = err < 0 ? es1a : es1b; - /* If a read-write mount, convert it to a read-only mount. */ - if (!sb_rdonly(sb) && vol->on_errors == ON_ERRORS_REMOUNT_RO) { - sb->s_flags |= SB_RDONLY; - ntfs_error(sb, "%s. Mounting read-only%s", es1, es2); - } + /* Hibernation safety takes precedence over the errors= policy. */ + sb->s_flags |= SB_RDONLY; NVolSetErrors(vol); + ntfs_error(sb, "%s. Mounting read-only%s", es1, es2); + + /* + * Remember it for the lifetime of the mount: see + * ntfs_sync_volume_dirty_state(). + */ + NVolSetHibernated(vol); + } else if (!sb_rdonly(sb) && NVolErrors(vol)) { + /* Match the read-write remount restriction for recorded errors. */ + sb->s_flags |= SB_RDONLY; + ntfs_error(sb, + "Errors were recorded during mount. Mounting read-only. Run chkdsk."); } /* If (still) a read-write mount, empty the logfile. */ @@ -1638,6 +1826,8 @@ iput_logfile_err_out: if (vol->logfile_ino) iput(vol->logfile_ino); iput(vol->vol_ino); + /* Do not leave a stale pointer behind for the rest of the teardown. */ + vol->vol_ino = NULL; iput_lcnbmp_err_out: iput(vol->lcnbmp_ino); iput_attrdef_err_out: @@ -1748,28 +1938,21 @@ static void ntfs_put_super(struct super_block *sb) ntfs_commit_inode(vol->mft_ino); /* - * If a read-write mount and no volume errors have occurred, mark the - * volume clean. Also, re-commit all affected inodes. + * If a read-write mount, re-commit all affected inodes once more. + * The dirty state itself is persisted at the end of ntfs_put_super(), + * after the last commits and the final write_inode_now(): those can + * still record errors via __ntfs_write_inode(), and the sync must + * evaluate NVolErrors() with the last setter already run. */ if (!sb_rdonly(sb)) { if (!NVolErrors(vol)) { - if (ntfs_clear_volume_flags(vol, VOLUME_IS_DIRTY)) - ntfs_warning(sb, - "Failed to clear dirty bit in volume information flags. Run chkdsk."); - ntfs_commit_inode(vol->vol_ino); ntfs_commit_inode(vol->root_ino); if (vol->mftmirr_ino) ntfs_commit_inode(vol->mftmirr_ino); ntfs_commit_inode(vol->mft_ino); - } else { - ntfs_warning(sb, - "Volume has errors. Leaving volume marked dirty. Run chkdsk."); } } - iput(vol->vol_ino); - vol->vol_ino = NULL; - /* NTFS 3.0+ specific clean up. */ if (vol->major_ver >= 3) { if (vol->extend_ino) { @@ -1799,8 +1982,6 @@ static void ntfs_put_super(struct super_block *sb) /* Re-commit the mft mirror and mft just in case. */ ntfs_commit_inode(vol->mftmirr_ino); ntfs_commit_inode(vol->mft_ino); - iput(vol->mftmirr_ino); - vol->mftmirr_ino = NULL; } /* * We should have no dirty inodes left, due to @@ -1810,6 +1991,59 @@ static void ntfs_put_super(struct super_block *sb) ntfs_commit_inode(vol->mft_ino); write_inode_now(vol->mft_ino, 1); + /* + * If a read-write mount, persist the error state in the volume flags: + * mark the volume clean if no volume errors have occurred, and make + * sure VOLUME_IS_DIRTY is on disk if any have, so chkdsk runs on the + * next mount. + */ + if (!sb_rdonly(sb)) { + if (ntfs_sync_volume_dirty_state(vol)) { + ntfs_warning(sb, + "Failed to sync dirty bit in volume information flags. Run chkdsk."); + } else { + /* + * __ntfs_write_inode(), not the void + * ntfs_commit_inode() wrapper: the error can only + * be warned about here. The mirror inode is only + * released below: writing the $Volume record (mft + * record number 3, below vol->mftmirr_size) mirrors + * it through ntfs_sync_mft_mirror(), which fails + * with -EIO once vol->mftmirr_ino is gone. + */ + if (__ntfs_write_inode(vol->vol_ino, 1)) { + ntfs_warning(sb, + "Failed to commit volume information flags. Run chkdsk."); + } else if (NInoDirty(NTFS_I(vol->vol_ino))) { + ntfs_warning(sb, + "Volume information flags remain dirty after commit. Run chkdsk."); + } else if (NVolErrors(vol)) { + /* + * Only warn once the commit has succeeded, + * or this could contradict a failure + * reported above. + */ + ntfs_warning(sb, + "Volume has errors. Leaving volume marked dirty. Run chkdsk."); + } + } + } + + /* + * Release $Volume while the mft inode is still available: if the + * commit above failed before it could clear the dirty flag, + * ntfs_evict_big_inode() commits the inode again on its way out, + * and __ntfs_write_inode() needs vol->mft_ino to look up the + * runlist of the record to write. + */ + iput(vol->vol_ino); + vol->vol_ino = NULL; + + if (vol->mftmirr_ino) { + iput(vol->mftmirr_ino); + vol->mftmirr_ino = NULL; + } + iput(vol->mft_ino); vol->mft_ino = NULL; blkdev_issue_flush(sb->s_bdev); @@ -1853,7 +2087,7 @@ static void ntfs_shutdown(struct super_block *sb) static int ntfs_sync_fs(struct super_block *sb, int wait) { struct ntfs_volume *vol = NTFS_SB(sb); - int err = 0; + int ret, err = 0; if (NVolShutdown(vol)) return -EIO; @@ -1861,15 +2095,30 @@ static int ntfs_sync_fs(struct super_block *sb, int wait) if (!wait) return 0; - /* If there are some dirty buffers in the bdev inode */ - if (!NVolErrors(vol) && - ntfs_clear_volume_flags(vol, VOLUME_IS_DIRTY)) { - ntfs_warning(sb, "Failed to clear dirty bit in volume information flags. Run chkdsk."); + /* + * The volume dirty bit is deliberately not cleared here: a sync + * running concurrently with an in-flight modification could clear + * and persist a bit that was just set, leaving the modification + * on a volume that is clean on disk. The bit is only cleared at + * the quiescent state transitions, remounting read-only and clean + * unmount. A recorded error state, however, is persisted right + * away so that it is not lost to a crash on a volume that has + * seen no modification; with NVolErrors() set this can only set + * the bit, never clear it. + */ + if (NVolErrors(vol) && + ntfs_sync_volume_dirty_state(vol)) { + ntfs_warning(sb, "Failed to sync dirty bit in volume information flags. Run chkdsk."); err = -EIO; } sync_inodes_sb(sb); - sync_blockdev(sb->s_bdev); - blkdev_issue_flush(sb->s_bdev); + ret = sync_blockdev(sb->s_bdev); + if (ret && !err) + err = ret; + + ret = blkdev_issue_flush(sb->s_bdev); + if (ret && !err) + err = ret; return err; } @@ -2290,6 +2539,11 @@ static int ntfs_fill_super(struct super_block *sb, struct fs_context *fc) ntfs_error(sb, "Unsupported NTFS filesystem."); goto err_out_now; } + if (!vol->mft_write_supported && !sb_rdonly(sb)) { + sb->s_flags |= SB_RDONLY; + ntfs_warning(sb, + "MFT writeback is unsupported for this geometry. Mounting read-only."); + } if (vol->sector_size > blocksize) { blocksize = sb_set_blocksize(sb, vol->sector_size); @@ -2613,6 +2867,13 @@ static int __init init_ntfs_fs(void) return err; } + err = ntfs_mft_bioset_init(); + if (err) { + pr_crit("Failed to initialize NTFS MFT bioset!\n"); + ntfs_workqueue_destroy(); + return err; + } + ntfs_index_ctx_cache = kmem_cache_create(ntfs_index_ctx_cache_name, sizeof(struct ntfs_index_context), 0 /* offset */, SLAB_HWCACHE_ALIGN, NULL /* ctor */); @@ -2680,6 +2941,7 @@ name_err_out: actx_err_out: kmem_cache_destroy(ntfs_index_ctx_cache); ictx_err_out: + ntfs_mft_bioset_exit(); if (!err) { pr_crit("Aborting NTFS filesystem driver registration...\n"); err = -ENOMEM; @@ -2698,6 +2960,7 @@ static void __exit exit_ntfs_fs(void) * destroy cache. */ rcu_barrier(); + ntfs_mft_bioset_exit(); #ifdef CONFIG_NTFS_FS_WOF_COMPRESSION ntfs_wof_free_workspaces(); #endif diff --git a/fs/ntfs/unistr.c b/fs/ntfs/unistr.c index 7f11a2825527..733bd6fe8599 100644 --- a/fs/ntfs/unistr.c +++ b/fs/ntfs/unistr.c @@ -64,7 +64,8 @@ bool ntfs_are_names_equal(const __le16 *s1, size_t s1_len, * @name1_len: first Unicode name length * @name2: second Unicode name to compare * @name2_len: second Unicode name length - * @err_val: if @name1 contains an invalid character return this value + * @check_invalid: if true and @name1 contains an invalid character, + * return -EINVAL * @ic: either CASE_SENSITIVE or IGNORE_CASE * @upcase: upcase table (ignored if @ic is CASE_SENSITIVE) * @upcase_len: upcase table size (ignored if @ic is CASE_SENSITIVE) @@ -74,13 +75,14 @@ bool ntfs_are_names_equal(const __le16 *s1, size_t s1_len, * -1 if the first name collates before the second one, * 0 if the names match, * 1 if the second name collates before the first one, or - * @err_val if an invalid character is found in @name1 during the comparison. + * -EINVAL if @check_invalid is true and an invalid character is found in + * @name1 during the comparison. * * The following characters are considered invalid: '"', '*', '<', '>' and '?'. */ int ntfs_collate_names(const __le16 *name1, const u32 name1_len, const __le16 *name2, const u32 name2_len, - const int err_val, const u32 ic, + const bool check_invalid, const u32 ic, const __le16 *upcase, const u32 upcase_len) { u32 cnt, min_len; @@ -98,8 +100,8 @@ int ntfs_collate_names(const __le16 *name1, const u32 name1_len, if (c2 < upcase_len) c2 = le16_to_cpu(upcase[c2]); } - if (c1 < 64 && legal_ansi_char_array[c1] & 8) - return err_val; + if (check_invalid && c1 < 64 && legal_ansi_char_array[c1] & 8) + return -EINVAL; if (c1 < c2) return -1; if (c1 > c2) @@ -111,8 +113,8 @@ int ntfs_collate_names(const __le16 *name1, const u32 name1_len, return 0; /* name1_len > name2_len */ c1 = le16_to_cpu(*name1); - if (c1 < 64 && legal_ansi_char_array[c1] & 8) - return err_val; + if (check_invalid && c1 < 64 && legal_ansi_char_array[c1] & 8) + return -EINVAL; return 1; } @@ -191,14 +193,23 @@ int ntfs_ucsncasecmp(const __le16 *s1, const __le16 *s2, size_t n, int ntfs_file_compare_values(const struct file_name_attr *file_name_attr1, const struct file_name_attr *file_name_attr2, - const int err_val, const u32 ic, + const bool check_invalid, const u32 ic, const __le16 *upcase, const u32 upcase_len) { + bool compare_check = check_invalid; + + /* + * POSIX file names may contain characters that are invalid in the + * Windows namespace, so compare them without treating them as errors. + */ + if (file_name_attr1->file_name_type == FILE_NAME_POSIX) + compare_check = false; + return ntfs_collate_names((__le16 *)&file_name_attr1->file_name, file_name_attr1->file_name_length, (__le16 *)&file_name_attr2->file_name, file_name_attr2->file_name_length, - err_val, ic, upcase, upcase_len); + compare_check, ic, upcase, upcase_len); } /* diff --git a/fs/ntfs/volume.h b/fs/ntfs/volume.h index bc85a9592245..0473b602084c 100644 --- a/fs/ntfs/volume.h +++ b/fs/ntfs/volume.h @@ -43,6 +43,7 @@ * @mft_record_size: in bytes * @mft_record_size_mask: mft_record_size - 1 * @mft_record_size_bits: log2(mft_record_size) + * @mft_write_supported: Whether the MFT write paths support the geometry. * @index_record_size: in bytes * @index_record_size_mask: index_record_size - 1 * @index_record_size_bits: log2(index_record_size) @@ -111,6 +112,13 @@ struct ntfs_volume { u32 mft_record_size; u32 mft_record_size_mask; u8 mft_record_size_bits; + bool mft_write_supported; + /* + * Unit size used for MFT I/O. This is the MFT record size when + * it is at least as large as the device logical block, or the + * containing device logical block when the record is smaller. + */ + u32 mft_io_unit_size; u32 index_record_size; u32 index_record_size_mask; u8 index_record_size_bits; @@ -181,9 +189,12 @@ struct ntfs_volume { * Windows-reserved names (CON, AUX, NUL, COM1, * LPT1, etc.) or invalid characters. * + * NV_Hibernated Windows is hibernated on the volume; the sync + * paths must not write the volume flags. * NV_Discard Issue discard/TRIM commands for freed clusters. * NV_DisableSparse Disable creation of sparse regions. * NV_NativeSymlinkRel Translate absolute Windows reparse targets (native_symlink=rel). + * NV_MftBootstrap Mount is still assembling $MFT's own runlist. */ enum { NV_Errors, @@ -199,10 +210,12 @@ enum { NV_ShowHiddenFiles, NV_HideDotFiles, NV_CheckWindowsNames, + NV_Hibernated, NV_Discard, NV_DisableSparse, NV_NativeSymlinkRel, NV_SymlinkNative, + NV_MftBootstrap, }; /* @@ -237,10 +250,12 @@ DEFINE_NVOL_BIT_OPS(SysImmutable) DEFINE_NVOL_BIT_OPS(ShowHiddenFiles) DEFINE_NVOL_BIT_OPS(HideDotFiles) DEFINE_NVOL_BIT_OPS(CheckWindowsNames) +DEFINE_NVOL_BIT_OPS(Hibernated) DEFINE_NVOL_BIT_OPS(Discard) DEFINE_NVOL_BIT_OPS(DisableSparse) DEFINE_NVOL_BIT_OPS(NativeSymlinkRel) DEFINE_NVOL_BIT_OPS(SymlinkNative) +DEFINE_NVOL_BIT_OPS(MftBootstrap) static inline void ntfs_inc_free_clusters(struct ntfs_volume *vol, s64 nr) { diff --git a/fs/ntfs3/attrib.c b/fs/ntfs3/attrib.c index b1c315206ffa..3b4ef15162da 100644 --- a/fs/ntfs3/attrib.c +++ b/fs/ntfs3/attrib.c @@ -61,11 +61,16 @@ static int attr_load_runs(struct ATTRIB *attr, struct ntfs_inode *ni, struct runs_tree *run, const CLST *vcn) { int err; - CLST svcn = le64_to_cpu(attr->nres.svcn); - CLST evcn = le64_to_cpu(attr->nres.evcn); + CLST svcn, evcn; u32 asize; u16 run_off; + if (!attr->non_res) + return -EIO; + + svcn = le64_to_cpu(attr->nres.svcn); + evcn = le64_to_cpu(attr->nres.evcn); + if (svcn >= evcn + 1 || run_is_mapped_full(run, svcn, evcn)) return 0; @@ -394,13 +399,15 @@ static int attr_set_size_res(struct ntfs_inode *ni, struct ATTRIB *attr, char *next = Add2Ptr(attr, asize); s64 dsize = ALIGN(new_size, 8) - ALIGN(rsize, 8); + if (new_size > sbi->record_size || + (dsize > 0 && used + dsize > sbi->max_bytes_per_attr)) { + return attr_make_nonresident(ni, attr, le, mi, new_size, run, + ins_attr, NULL); + } + if (dsize < 0) { memmove(next + dsize, next, tail); } else if (dsize > 0) { - if (used + dsize > sbi->max_bytes_per_attr) - return attr_make_nonresident(ni, attr, le, mi, new_size, - run, ins_attr, NULL); - memmove(next + dsize, next, tail); memset(next, 0, dsize); } @@ -556,6 +563,10 @@ again_1: } next_le_1: + if (!attr->non_res) { + err = -EIO; + goto out; + } svcn = le64_to_cpu(attr->nres.svcn); evcn = le64_to_cpu(attr->nres.evcn); } @@ -1034,6 +1045,21 @@ again: if (!attr_b->non_res) { u32 data_size = le32_to_cpu(attr_b->res.data_size); + + /* + * A resident attribute is copied into a single page below + * (alloc_page() + memcpy()) and mapped as a one-page + * IOMAP_INLINE extent by ntfs_iomap_begin(). mi_enum_attr() + * only bounds the resident value length against the MFT record + * size, so a corrupted volume whose records are larger than a + * page can report data_size > PAGE_SIZE; copying that many bytes + * would overflow the single destination page. Reject it. + */ + if (data_size > PAGE_SIZE) { + err = -EINVAL; + goto out; + } + *lcn = RESIDENT_LCN; *len = data_size; if (res) { @@ -1456,6 +1482,9 @@ int attr_load_runs_vcn(struct ntfs_inode *ni, enum ATTR_TYPE type, return -ENOENT; } + if (!attr->non_res) + return -EIO; + svcn = le64_to_cpu(attr->nres.svcn); evcn = le64_to_cpu(attr->nres.evcn); @@ -1995,7 +2024,7 @@ out: valid_size = le64_to_cpu(attr_b->nres.valid_size); if (new_valid != valid_size) { - attr_b->nres.valid_size = cpu_to_le64(valid_size); + attr_b->nres.valid_size = cpu_to_le64(new_valid); mi_b->dirty = true; } } @@ -2610,8 +2639,8 @@ int attr_insert_range(struct ntfs_inode *ni, u64 vbo, u64 bytes) char *data = Add2Ptr(attr_b, le16_to_cpu(attr_b->res.data_off)); - memmove(data + bytes, data, bytes); - memset(data, 0, bytes); + memmove(data + vbo + bytes, data + vbo, data_size - vbo); + memset(data + vbo, 0, bytes); goto done; } diff --git a/fs/ntfs3/dir.c b/fs/ntfs3/dir.c index eb9152e9fa22..328c9a29df89 100644 --- a/fs/ntfs3/dir.c +++ b/fs/ntfs3/dir.c @@ -34,6 +34,9 @@ int ntfs_utf16_to_nls(struct ntfs_sb_info *sbi, const __le16 *name, u32 len, /* UTF-16 -> UTF-8 */ ret = utf16s_to_utf8s((wchar_t *)name, len, UTF16_LITTLE_ENDIAN, buf, buf_len); + if (ret >= buf_len) { + ret = buf_len-1; + } buf[ret] = '\0'; return ret; } diff --git a/fs/ntfs3/file.c b/fs/ntfs3/file.c index 2abf334bfa0c..4cbdd9e4222b 100644 --- a/fs/ntfs3/file.c +++ b/fs/ntfs3/file.c @@ -129,7 +129,7 @@ int ntfs_fileattr_get(struct dentry *dentry, struct file_kattr *fa) /* * ntfs_fileattr_set - inode_operations::fileattr_set */ -int ntfs_fileattr_set(struct mnt_idmap *idmap, struct dentry *dentry, +int ntfs_fileattr_set(const struct mnt_idmap *idmap, struct dentry *dentry, struct file_kattr *fa) { struct inode *inode = d_inode(dentry); @@ -183,6 +183,9 @@ static int ntfs_ioctl_set_volume_label(struct ntfs_sb_info *sbi, u8 __user *buf) if (!capable(CAP_SYS_ADMIN)) return -EPERM; + if (sb_rdonly(sbi->sb)) + return -EROFS; + if (copy_from_user(user, buf, FSLABEL_MAX)) return -EFAULT; @@ -261,7 +264,7 @@ long ntfs_compat_ioctl(struct file *filp, u32 cmd, unsigned long arg) /* * ntfs_getattr - inode_operations::getattr */ -int ntfs_getattr(struct mnt_idmap *idmap, const struct path *path, +int ntfs_getattr(const struct mnt_idmap *idmap, const struct path *path, struct kstat *stat, u32 request_mask, u32 flags) { struct inode *inode = d_inode(path->dentry); @@ -703,7 +706,7 @@ out: /* * ntfs_setattr - inode_operations::setattr */ -int ntfs_setattr(struct mnt_idmap *idmap, struct dentry *dentry, +int ntfs_setattr(const struct mnt_idmap *idmap, struct dentry *dentry, struct iattr *attr) { struct inode *inode = d_inode(dentry); diff --git a/fs/ntfs3/frecord.c b/fs/ntfs3/frecord.c index bead01a953f3..4651dd8f6464 100644 --- a/fs/ntfs3/frecord.c +++ b/fs/ntfs3/frecord.c @@ -2190,10 +2190,13 @@ remove_wof: /* Clear cached flag. */ ni->ni_flags &= ~NI_FLAG_COMPRESSED_MASK; + /* offs_folio is accessed under run_lock in attr_wof_frame_info(). */ + down_write(&ni->file.run_lock); if (ni->file.offs_folio) { folio_put(ni->file.offs_folio); ni->file.offs_folio = NULL; } + up_write(&ni->file.run_lock); mapping->a_ops = &ntfs_aops; out: @@ -2521,6 +2524,8 @@ out1: out: for (i = 0; i < pages_per_frame; i++) { pg = pages[i]; + if (err) + clear_highpage(pg); SetPageUptodate(pg); } diff --git a/fs/ntfs3/fslog.c b/fs/ntfs3/fslog.c index ed50c1d0c23e..9a44b322ad60 100644 --- a/fs/ntfs3/fslog.c +++ b/fs/ntfs3/fslog.c @@ -7,6 +7,7 @@ #include <linux/blkdev.h> #include <linux/fs.h> +#include <linux/overflow.h> #include <linux/random.h> #include <linux/slab.h> @@ -2457,7 +2458,8 @@ static int find_log_rec(struct ntfs_log *log, u64 lsn, struct lcb *lcb) * Check that the length field isn't greater than the total * available space the log file. */ - rec_len = len + log->record_header_len; + if (check_add_overflow(len, log->record_header_len, &rec_len)) + return -EINVAL; if (rec_len >= log->total_avail) return -EINVAL; diff --git a/fs/ntfs3/fsntfs.c b/fs/ntfs3/fsntfs.c index 97c04ab2763a..b55311ccd229 100644 --- a/fs/ntfs3/fsntfs.c +++ b/fs/ntfs3/fsntfs.c @@ -2153,6 +2153,9 @@ int ntfs_insert_security(struct ntfs_sb_info *sbi, goto out; while (e) { + if (le16_to_cpu(e->de.size) < SIZEOF_SDH_DIRENTRY) + break; + if (le32_to_cpu(e->sec_hdr.size) == new_sec_size) { err = ntfs_read_run_nb(sbi, &ni->file.run, le64_to_cpu(e->sec_hdr.off), diff --git a/fs/ntfs3/index.c b/fs/ntfs3/index.c index 689712d3463d..37e3e16b1f84 100644 --- a/fs/ntfs3/index.c +++ b/fs/ntfs3/index.c @@ -594,6 +594,8 @@ static const struct NTFS_DE *hdr_insert_head(struct INDEX_HDR *hdr, if (!e) return NULL; + if (size_add(used, ins_bytes) > le32_to_cpu(hdr->total)) + return NULL; /* Now we just make room for the inserted entries and jam it in. */ to_move = used - le32_to_cpu(hdr->de_off); @@ -1801,7 +1803,10 @@ static int indx_insert_into_root(struct ntfs_index *indx, struct ntfs_inode *ni, } /* Copy root entries into new buffer. */ - hdr_insert_head(hdr, re, to_move); + if (!hdr_insert_head(hdr, re, to_move)) { + err = -EINVAL; + goto out_put_n; + } /* Update bitmap attribute. */ indx_mark_used(indx, ni, new_vbn >> indx->idx2vbn_bits); @@ -1955,7 +1960,11 @@ static int indx_insert_into_buffer(struct ntfs_index *indx, /* Copy all the entries <= sp into the new buffer. */ de_t = hdr_first_de(hdr1); to_copy = PtrOffset(de_t, sp); - hdr_insert_head(hdr2, de_t, to_copy); + if (!hdr_insert_head(hdr2, de_t, to_copy)) { + err = -EINVAL; + put_indx_node(n2); + goto out; + } /* Remove all entries (sp including) from hdr1. */ used = used1 - to_copy - sp_size; diff --git a/fs/ntfs3/inode.c b/fs/ntfs3/inode.c index 4ac26c80bd34..366af3038248 100644 --- a/fs/ntfs3/inode.c +++ b/fs/ntfs3/inode.c @@ -213,6 +213,11 @@ next_attr: names += 1; fname = Add2Ptr(attr, roff); + + /* Make sure the full name fits in the resident data. */ + if (rsize < fname_full_size(fname)) + goto out; + if (fname->type == FILE_NAME_DOS) goto next_attr; @@ -280,7 +285,9 @@ next_attr: break; case ATTR_ROOT: - if (attr->non_res) + if (attr->non_res || + asize < sizeof(struct INDEX_ROOT) + roff || + rsize < sizeof(struct INDEX_ROOT)) goto out; root = Add2Ptr(attr, roff); @@ -981,6 +988,8 @@ static int ntfs_iomap_begin(struct inode *inode, loff_t offset, loff_t length, if (err) { return err; } + if (!clen) + return -EINVAL; if (lcn == EOF_LCN) { /* request out of file. */ @@ -1016,11 +1025,6 @@ static int ntfs_iomap_begin(struct inode *inode, loff_t offset, loff_t length, return 0; } - if (!clen) { - /* broken file? */ - return -EINVAL; - } - iomap->bdev = inode->i_sb->s_bdev; iomap->offset = offset; iomap->length = ((loff_t)clen << cluster_bits) - off; @@ -1381,7 +1385,7 @@ out: * * NOTE: if fnd != NULL (ntfs_atomic_open) then @dir is locked */ -int ntfs_create_inode(struct mnt_idmap *idmap, struct inode *dir, +int ntfs_create_inode(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *dentry, const struct cpu_str *uni, umode_t mode, dev_t dev, const char *symname, u32 size, struct ntfs_fnd *fnd) @@ -2244,6 +2248,9 @@ static noinline int ntfs_readlink_hlp(const struct dentry *link_de, if (err < 0) goto out; + if (err >= buflen) + err = buflen - 1; + /* Translate Windows '\' into Linux '/'. */ for (i = 0; i < err; i++) { if (buffer[i] == '\\') diff --git a/fs/ntfs3/namei.c b/fs/ntfs3/namei.c index ec59bbabd3c5..ae88cb66fb7e 100644 --- a/fs/ntfs3/namei.c +++ b/fs/ntfs3/namei.c @@ -111,7 +111,7 @@ static struct dentry *ntfs_lookup(struct inode *dir, struct dentry *dentry, /* * ntfs_create - inode_operations::create */ -static int ntfs_create(struct mnt_idmap *idmap, struct inode *dir, +static int ntfs_create(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *dentry, umode_t mode) { return ntfs_create_inode(idmap, dir, dentry, NULL, S_IFREG | mode, 0, @@ -121,7 +121,7 @@ static int ntfs_create(struct mnt_idmap *idmap, struct inode *dir, /* * ntfs_mknod - inode_operations::mknod */ -static int ntfs_mknod(struct mnt_idmap *idmap, struct inode *dir, +static int ntfs_mknod(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *dentry, umode_t mode, dev_t rdev) { return ntfs_create_inode(idmap, dir, dentry, NULL, mode, rdev, NULL, 0, @@ -209,7 +209,7 @@ static int ntfs_unlink(struct inode *dir, struct dentry *dentry) /* * ntfs_symlink - inode_operations::symlink */ -static int ntfs_symlink(struct mnt_idmap *idmap, struct inode *dir, +static int ntfs_symlink(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *dentry, const char *symname) { u32 size = strlen(symname); @@ -228,7 +228,7 @@ static int ntfs_symlink(struct mnt_idmap *idmap, struct inode *dir, /* * ntfs_mkdir - inode_operations::mkdir */ -static struct dentry *ntfs_mkdir(struct mnt_idmap *idmap, struct inode *dir, +static struct dentry *ntfs_mkdir(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *dentry, umode_t mode) { return ERR_PTR(ntfs_create_inode(idmap, dir, dentry, NULL, @@ -262,7 +262,7 @@ static int ntfs_rmdir(struct inode *dir, struct dentry *dentry) /* * ntfs_rename - inode_operations::rename */ -static int ntfs_rename(struct mnt_idmap *idmap, struct inode *dir, +static int ntfs_rename(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *dentry, struct inode *new_dir, struct dentry *new_dentry, u32 flags) { diff --git a/fs/ntfs3/ntfs_fs.h b/fs/ntfs3/ntfs_fs.h index 5811d89d67b3..ea24126c70db 100644 --- a/fs/ntfs3/ntfs_fs.h +++ b/fs/ntfs3/ntfs_fs.h @@ -551,11 +551,11 @@ extern const struct file_operations ntfs_dir_operations; /* Globals from file.c */ int ntfs_fileattr_get(struct dentry *dentry, struct file_kattr *fa); -int ntfs_fileattr_set(struct mnt_idmap *idmap, struct dentry *dentry, +int ntfs_fileattr_set(const struct mnt_idmap *idmap, struct dentry *dentry, struct file_kattr *fa); -int ntfs_getattr(struct mnt_idmap *idmap, const struct path *path, +int ntfs_getattr(const struct mnt_idmap *idmap, const struct path *path, struct kstat *stat, u32 request_mask, u32 flags); -int ntfs_setattr(struct mnt_idmap *idmap, struct dentry *dentry, +int ntfs_setattr(const struct mnt_idmap *idmap, struct dentry *dentry, struct iattr *attr); int ntfs_file_open(struct inode *inode, struct file *file); int ntfs_fiemap(struct inode *inode, struct fiemap_extent_info *fieinfo, @@ -803,7 +803,7 @@ int ntfs_set_size(struct inode *inode, u64 new_size); int ntfs3_write_inode(struct inode *inode, struct writeback_control *wbc); int ntfs_sync_inode(struct inode *inode); int inode_read_data(struct inode *inode, void *data, size_t bytes); -int ntfs_create_inode(struct mnt_idmap *idmap, struct inode *dir, +int ntfs_create_inode(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *dentry, const struct cpu_str *uni, umode_t mode, dev_t dev, const char *symname, u32 size, struct ntfs_fnd *fnd); @@ -963,18 +963,18 @@ unsigned long ntfs_names_hash(const u16 *name, size_t len, const u16 *upcase, /* globals from xattr.c */ #ifdef CONFIG_NTFS3_FS_POSIX_ACL -struct posix_acl *ntfs_get_acl(struct mnt_idmap *idmap, struct dentry *dentry, +struct posix_acl *ntfs_get_acl(const struct mnt_idmap *idmap, struct dentry *dentry, int type); -int ntfs_set_acl(struct mnt_idmap *idmap, struct dentry *dentry, +int ntfs_set_acl(const struct mnt_idmap *idmap, struct dentry *dentry, struct posix_acl *acl, int type); -int ntfs_init_acl(struct mnt_idmap *idmap, struct inode *inode, +int ntfs_init_acl(const struct mnt_idmap *idmap, struct inode *inode, struct inode *dir); #else #define ntfs_get_acl NULL #define ntfs_set_acl NULL #endif -int ntfs_acl_chmod(struct mnt_idmap *idmap, struct dentry *dentry); +int ntfs_acl_chmod(const struct mnt_idmap *idmap, struct dentry *dentry); ssize_t ntfs_listxattr(struct dentry *dentry, char *buffer, size_t size); extern const struct xattr_handler *const ntfs_xattr_handlers[]; diff --git a/fs/ntfs3/record.c b/fs/ntfs3/record.c index 4f12ce15b03b..d7785f9e8e09 100644 --- a/fs/ntfs3/record.c +++ b/fs/ntfs3/record.c @@ -665,18 +665,32 @@ int mi_pack_runs(struct mft_inode *mi, struct ATTRIB *attr, u32 run_size = asize - run_off; u32 tail = used - aoff - asize; u32 dsize = sbi->record_size - used; + u32 avail; /* Make a maximum gap in current record. */ memmove(next + dsize, next, tail); + /* Leave room for the 8-byte ALIGN() to avoid OOB write */ + avail = (run_size + dsize) & ~7u; + + if (!avail) { + memmove(next, next + dsize, tail); + return -ENOSPC; + } + /* Pack as much as possible. */ - err = run_pack(run, svcn, len, Add2Ptr(attr, run_off), run_size + dsize, + err = run_pack(run, svcn, len, Add2Ptr(attr, run_off), avail, &plen); if (err < 0) { memmove(next, next + dsize, tail); return err; } + if (!plen) { + memmove(next, next + dsize, tail); + return -ENOSPC; + } + new_run_size = ALIGN(err, 8); memmove(next + new_run_size - run_size, next + dsize, tail); diff --git a/fs/ntfs3/super.c b/fs/ntfs3/super.c index f4a42a0c73a4..d3093f72ea38 100644 --- a/fs/ntfs3/super.c +++ b/fs/ntfs3/super.c @@ -1210,7 +1210,7 @@ read_boot: #ifdef CONFIG_NTFS3_64BIT_CLUSTER if (clusters >= (1ull << (64 - cluster_bits))) - sbi->maxbytes = -1; + sbi->maxbytes = MAX_LFS_FILESIZE; sbi->maxbytes_sparse = MAX_LFS_FILESIZE; sb->s_maxbytes = MAX_LFS_FILESIZE; #else diff --git a/fs/ntfs3/xattr.c b/fs/ntfs3/xattr.c index 594ef6860b93..e9824bd80322 100644 --- a/fs/ntfs3/xattr.c +++ b/fs/ntfs3/xattr.c @@ -165,6 +165,8 @@ static int ntfs_read_ea(struct ntfs_inode *ni, struct EA_FULL **ea, /* ef->size must fit the list and cover the record. */ if (ea_size > bytes || ea_size < need) goto out1; + if (bytes < offsetof(struct EA_FULL, name)) + goto out1; continue; } @@ -543,7 +545,7 @@ out: /* * ntfs_get_acl - inode_operations::get_acl */ -struct posix_acl *ntfs_get_acl(struct mnt_idmap *idmap, struct dentry *dentry, +struct posix_acl *ntfs_get_acl(const struct mnt_idmap *idmap, struct dentry *dentry, int type) { struct inode *inode = d_inode(dentry); @@ -595,7 +597,7 @@ struct posix_acl *ntfs_get_acl(struct mnt_idmap *idmap, struct dentry *dentry, return acl; } -static noinline int ntfs_set_acl_ex(struct mnt_idmap *idmap, +static noinline int ntfs_set_acl_ex(const struct mnt_idmap *idmap, struct inode *inode, struct posix_acl *acl, int type, bool init_acl) { @@ -677,7 +679,7 @@ out: /* * ntfs_set_acl - inode_operations::set_acl */ -int ntfs_set_acl(struct mnt_idmap *idmap, struct dentry *dentry, +int ntfs_set_acl(const struct mnt_idmap *idmap, struct dentry *dentry, struct posix_acl *acl, int type) { return ntfs_set_acl_ex(idmap, d_inode(dentry), acl, type, false); @@ -688,7 +690,7 @@ int ntfs_set_acl(struct mnt_idmap *idmap, struct dentry *dentry, * * Called from ntfs_create_inode(). */ -int ntfs_init_acl(struct mnt_idmap *idmap, struct inode *inode, +int ntfs_init_acl(const struct mnt_idmap *idmap, struct inode *inode, struct inode *dir) { struct posix_acl *default_acl, *acl; @@ -722,7 +724,7 @@ int ntfs_init_acl(struct mnt_idmap *idmap, struct inode *inode, /* * ntfs_acl_chmod - Helper for ntfs_setattr(). */ -int ntfs_acl_chmod(struct mnt_idmap *idmap, struct dentry *dentry) +int ntfs_acl_chmod(const struct mnt_idmap *idmap, struct dentry *dentry) { struct inode *inode = d_inode(dentry); struct super_block *sb = inode->i_sb; @@ -783,7 +785,7 @@ static int ntfs_getxattr(const struct xattr_handler *handler, struct dentry *de, if (!buffer) { err = sizeof(u8); } else if (size < sizeof(u8)) { - err = -ENODATA; + err = -ERANGE; } else { err = sizeof(u8); *(u8 *)buffer = le32_to_cpu(ni->std_fa); @@ -797,7 +799,7 @@ static int ntfs_getxattr(const struct xattr_handler *handler, struct dentry *de, if (!buffer) { err = sizeof(u32); } else if (size < sizeof(u32)) { - err = -ENODATA; + err = -ERANGE; } else { err = sizeof(u32); *(u32 *)buffer = le32_to_cpu(ni->std_fa); @@ -837,7 +839,7 @@ static int ntfs_getxattr(const struct xattr_handler *handler, struct dentry *de, if (!buffer) { err = sd_size; } else if (size < sd_size) { - err = -ENODATA; + err = -ERANGE; } else { err = sd_size; memcpy(buffer, sd, sd_size); @@ -863,7 +865,7 @@ static bool ntfs_is_reserved_lxattr(const char *name) * ntfs_setxattr - inode_operations::setxattr */ static noinline int ntfs_setxattr(const struct xattr_handler *handler, - struct mnt_idmap *idmap, struct dentry *de, + const struct mnt_idmap *idmap, struct dentry *de, struct inode *inode, const char *name, const void *value, size_t size, int flags) { @@ -976,8 +978,10 @@ set_new_fa: NULL); out: - inode_set_ctime_current(inode); - mark_inode_dirty(inode); + if (!err) { + inode_set_ctime_current(inode); + mark_inode_dirty(inode); + } return err; } diff --git a/fs/ocfs2/acl.c b/fs/ocfs2/acl.c index 090ec60fb576..801a2f56ad08 100644 --- a/fs/ocfs2/acl.c +++ b/fs/ocfs2/acl.c @@ -260,7 +260,7 @@ static int ocfs2_set_acl(handle_t *handle, return ret; } -int ocfs2_iop_set_acl(struct mnt_idmap *idmap, struct dentry *dentry, +int ocfs2_iop_set_acl(const struct mnt_idmap *idmap, struct dentry *dentry, struct posix_acl *acl, int type) { struct buffer_head *bh = NULL; diff --git a/fs/ocfs2/acl.h b/fs/ocfs2/acl.h index a91f9ce278d6..1ed05899cce1 100644 --- a/fs/ocfs2/acl.h +++ b/fs/ocfs2/acl.h @@ -17,7 +17,7 @@ struct ocfs2_acl_entry { }; struct posix_acl *ocfs2_iop_get_acl(struct inode *inode, int type, bool rcu); -int ocfs2_iop_set_acl(struct mnt_idmap *idmap, struct dentry *dentry, +int ocfs2_iop_set_acl(const struct mnt_idmap *idmap, struct dentry *dentry, struct posix_acl *acl, int type); extern int ocfs2_acl_chmod(struct inode *, struct buffer_head *); struct ocfs2_acl_state { diff --git a/fs/ocfs2/buffer_head_io.c b/fs/ocfs2/buffer_head_io.c index 7bfe377af2df..733ceda79ca1 100644 --- a/fs/ocfs2/buffer_head_io.c +++ b/fs/ocfs2/buffer_head_io.c @@ -66,12 +66,14 @@ int ocfs2_write_block(struct ocfs2_super *osb, struct buffer_head *bh, wait_on_buffer(bh); - if (buffer_uptodate(bh)) { + if (!buffer_write_io_error(bh)) { ocfs2_set_buffer_uptodate(ci, bh); } else { - /* We don't need to remove the clustered uptodate - * information for this bh as it's not marked locally - * uptodate. */ + /* + * The buffer still holds what we tried to write, but it did + * not reach the disk, so don't advertise it to the cluster + * as up to date. + */ ret = -EIO; mlog_errno(ret); } @@ -446,7 +448,7 @@ int ocfs2_write_super_or_backup(struct ocfs2_super *osb, wait_on_buffer(bh); - if (!buffer_uptodate(bh)) { + if (buffer_write_io_error(bh)) { ret = -EIO; mlog_errno(ret); } diff --git a/fs/ocfs2/dlmfs/dlmfs.c b/fs/ocfs2/dlmfs/dlmfs.c index 53df5dd10ad0..d3bfcada3e3b 100644 --- a/fs/ocfs2/dlmfs/dlmfs.c +++ b/fs/ocfs2/dlmfs/dlmfs.c @@ -188,7 +188,7 @@ static int dlmfs_file_release(struct inode *inode, * We do ->setattr() just to override size changes. Our size is the size * of the LVB and nothing else. */ -static int dlmfs_file_setattr(struct mnt_idmap *idmap, +static int dlmfs_file_setattr(const struct mnt_idmap *idmap, struct dentry *dentry, struct iattr *attr) { int error; @@ -402,7 +402,7 @@ static struct inode *dlmfs_get_inode(struct inode *parent, * File creation. Allocate an inode, and we're done.. */ /* SMP-safe */ -static struct dentry *dlmfs_mkdir(struct mnt_idmap * idmap, +static struct dentry *dlmfs_mkdir(const struct mnt_idmap * idmap, struct inode * dir, struct dentry * dentry, umode_t mode) @@ -450,7 +450,7 @@ bail: return ERR_PTR(status); } -static int dlmfs_create(struct mnt_idmap *idmap, +static int dlmfs_create(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *dentry, umode_t mode) diff --git a/fs/ocfs2/file.c b/fs/ocfs2/file.c index d6e977ba6565..62f45a1b5ca1 100644 --- a/fs/ocfs2/file.c +++ b/fs/ocfs2/file.c @@ -1117,7 +1117,7 @@ out: return ret; } -int ocfs2_setattr(struct mnt_idmap *idmap, struct dentry *dentry, +int ocfs2_setattr(const struct mnt_idmap *idmap, struct dentry *dentry, struct iattr *attr) { int status = 0, size_change; @@ -1317,7 +1317,7 @@ bail: return status; } -int ocfs2_getattr(struct mnt_idmap *idmap, const struct path *path, +int ocfs2_getattr(const struct mnt_idmap *idmap, const struct path *path, struct kstat *stat, u32 request_mask, unsigned int flags) { struct inode *inode = d_inode(path->dentry); @@ -1349,7 +1349,7 @@ bail: return err; } -int ocfs2_permission(struct mnt_idmap *idmap, struct inode *inode, +int ocfs2_permission(const struct mnt_idmap *idmap, struct inode *inode, int mask) { int ret, had_lock; diff --git a/fs/ocfs2/file.h b/fs/ocfs2/file.h index 41e65e45a9f3..97492ee5789e 100644 --- a/fs/ocfs2/file.h +++ b/fs/ocfs2/file.h @@ -50,11 +50,11 @@ int ocfs2_extend_no_holes(struct inode *inode, struct buffer_head *di_bh, u64 new_i_size, u64 zero_to); int ocfs2_zero_extend(struct inode *inode, struct buffer_head *di_bh, loff_t zero_to); -int ocfs2_setattr(struct mnt_idmap *idmap, struct dentry *dentry, +int ocfs2_setattr(const struct mnt_idmap *idmap, struct dentry *dentry, struct iattr *attr); -int ocfs2_getattr(struct mnt_idmap *idmap, const struct path *path, +int ocfs2_getattr(const struct mnt_idmap *idmap, const struct path *path, struct kstat *stat, u32 request_mask, unsigned int flags); -int ocfs2_permission(struct mnt_idmap *idmap, +int ocfs2_permission(const struct mnt_idmap *idmap, struct inode *inode, int mask); diff --git a/fs/ocfs2/ioctl.c b/fs/ocfs2/ioctl.c index cbe59d231666..36c7c9ac8b5d 100644 --- a/fs/ocfs2/ioctl.c +++ b/fs/ocfs2/ioctl.c @@ -82,7 +82,7 @@ int ocfs2_fileattr_get(struct dentry *dentry, struct file_kattr *fa) return status; } -int ocfs2_fileattr_set(struct mnt_idmap *idmap, +int ocfs2_fileattr_set(const struct mnt_idmap *idmap, struct dentry *dentry, struct file_kattr *fa) { struct inode *inode = d_inode(dentry); diff --git a/fs/ocfs2/ioctl.h b/fs/ocfs2/ioctl.h index 4a1c2313b429..b1cb529fc5f9 100644 --- a/fs/ocfs2/ioctl.h +++ b/fs/ocfs2/ioctl.h @@ -12,7 +12,7 @@ #define OCFS2_IOCTL_PROTO_H int ocfs2_fileattr_get(struct dentry *dentry, struct file_kattr *fa); -int ocfs2_fileattr_set(struct mnt_idmap *idmap, +int ocfs2_fileattr_set(const struct mnt_idmap *idmap, struct dentry *dentry, struct file_kattr *fa); long ocfs2_ioctl(struct file *filp, unsigned int cmd, unsigned long arg); long ocfs2_compat_ioctl(struct file *file, unsigned cmd, unsigned long arg); diff --git a/fs/ocfs2/journal.c b/fs/ocfs2/journal.c index a3938a03e93b..cc60db6b9706 100644 --- a/fs/ocfs2/journal.c +++ b/fs/ocfs2/journal.c @@ -677,19 +677,20 @@ static int __ocfs2_journal_access(handle_t *handle, mlog(ML_ERROR, "giving me a buffer that's not uptodate!\n"); mlog(ML_ERROR, "b_blocknr=%llu, b_state=0x%lx\n", (unsigned long long)bh->b_blocknr, bh->b_state); - + } + /* + * A previous transaction with a couple of buffer heads fail + * to checkpoint, so all the bhs are marked as BH_Write_EIO. + * For current transaction, the bh is just among those error + * bhs which previous transaction handle. We can't just clear + * its BH_Write_EIO and reuse directly, since other bhs are + * not written to disk yet and that will cause metadata + * inconsistency. So we should set fs read-only to avoid + * further damage. + */ + if (buffer_write_io_error(bh)) { lock_buffer(bh); - /* - * A previous transaction with a couple of buffer heads fail - * to checkpoint, so all the bhs are marked as BH_Write_EIO. - * For current transaction, the bh is just among those error - * bhs which previous transaction handle. We can't just clear - * its BH_Write_EIO and reuse directly, since other bhs are - * not written to disk yet and that will cause metadata - * inconsistency. So we should set fs read-only to avoid - * further damage. - */ - if (buffer_write_io_error(bh) && !buffer_uptodate(bh)) { + if (buffer_write_io_error(bh)) { unlock_buffer(bh); return ocfs2_error(osb->sb, "A previous attempt to " "write this buffer head failed\n"); diff --git a/fs/ocfs2/namei.c b/fs/ocfs2/namei.c index 58c6061ed983..fce9a31a3671 100644 --- a/fs/ocfs2/namei.c +++ b/fs/ocfs2/namei.c @@ -227,7 +227,7 @@ static void ocfs2_cleanup_add_entry_failure(struct ocfs2_super *osb, iput(inode); } -static int ocfs2_mknod(struct mnt_idmap *idmap, +static int ocfs2_mknod(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *dentry, umode_t mode, @@ -650,7 +650,7 @@ static int ocfs2_mknod_locked(struct ocfs2_super *osb, suballoc_loc, suballoc_bit); } -static struct dentry *ocfs2_mkdir(struct mnt_idmap *idmap, +static struct dentry *ocfs2_mkdir(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *dentry, umode_t mode) @@ -666,7 +666,7 @@ static struct dentry *ocfs2_mkdir(struct mnt_idmap *idmap, return ERR_PTR(ret); } -static int ocfs2_create(struct mnt_idmap *idmap, +static int ocfs2_create(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *dentry, umode_t mode) @@ -1206,7 +1206,7 @@ static void ocfs2_double_unlock(struct inode *inode1, struct inode *inode2) ocfs2_inode_unlock(inode2, 1); } -static int ocfs2_rename(struct mnt_idmap *idmap, +static int ocfs2_rename(const struct mnt_idmap *idmap, struct inode *old_dir, struct dentry *old_dentry, struct inode *new_dir, @@ -1810,7 +1810,7 @@ bail: return status; } -static int ocfs2_symlink(struct mnt_idmap *idmap, +static int ocfs2_symlink(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *dentry, const char *symname) diff --git a/fs/ocfs2/quota_local.c b/fs/ocfs2/quota_local.c index 76e7dd5aecc8..816be594db41 100644 --- a/fs/ocfs2/quota_local.c +++ b/fs/ocfs2/quota_local.c @@ -171,6 +171,10 @@ static int ocfs2_local_check_quota_file(struct super_block *sb, int type) struct ocfs2_disk_dqheader *dqhead; int status, ret = 0; + /* OCFS2 quota format is supported only for OCFS2 filesystems */ + if (sb->s_magic != OCFS2_SUPER_MAGIC) + goto out_err; + /* First check whether we understand local quota file */ status = ocfs2_read_quota_block(linode, 0, &bh); if (status) { diff --git a/fs/ocfs2/xattr.c b/fs/ocfs2/xattr.c index 9f620f6c6005..f7c486ca6764 100644 --- a/fs/ocfs2/xattr.c +++ b/fs/ocfs2/xattr.c @@ -7622,7 +7622,7 @@ static int ocfs2_xattr_security_get(const struct xattr_handler *handler, } static int ocfs2_xattr_security_set(const struct xattr_handler *handler, - struct mnt_idmap *idmap, + const struct mnt_idmap *idmap, struct dentry *unused, struct inode *inode, const char *name, const void *value, size_t size, int flags) @@ -7717,7 +7717,7 @@ static int ocfs2_xattr_trusted_get(const struct xattr_handler *handler, } static int ocfs2_xattr_trusted_set(const struct xattr_handler *handler, - struct mnt_idmap *idmap, + const struct mnt_idmap *idmap, struct dentry *unused, struct inode *inode, const char *name, const void *value, size_t size, int flags) @@ -7748,7 +7748,7 @@ static int ocfs2_xattr_user_get(const struct xattr_handler *handler, } static int ocfs2_xattr_user_set(const struct xattr_handler *handler, - struct mnt_idmap *idmap, + const struct mnt_idmap *idmap, struct dentry *unused, struct inode *inode, const char *name, const void *value, size_t size, int flags) diff --git a/fs/omfs/dir.c b/fs/omfs/dir.c index 692297cf84e7..18f4b4543cc8 100644 --- a/fs/omfs/dir.c +++ b/fs/omfs/dir.c @@ -279,13 +279,13 @@ out_free_inode: return err; } -static struct dentry *omfs_mkdir(struct mnt_idmap *idmap, struct inode *dir, +static struct dentry *omfs_mkdir(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *dentry, umode_t mode) { return ERR_PTR(omfs_add_node(dir, dentry, mode)); } -static int omfs_create(struct mnt_idmap *idmap, struct inode *dir, +static int omfs_create(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *dentry, umode_t mode) { return omfs_add_node(dir, dentry, mode | S_IFREG); @@ -370,7 +370,7 @@ static bool omfs_fill_chain(struct inode *dir, struct dir_context *ctx, return true; } -static int omfs_rename(struct mnt_idmap *idmap, struct inode *old_dir, +static int omfs_rename(const struct mnt_idmap *idmap, struct inode *old_dir, struct dentry *old_dentry, struct inode *new_dir, struct dentry *new_dentry, unsigned int flags) { diff --git a/fs/omfs/file.c b/fs/omfs/file.c index 28f3b113340e..79a413f1dc0d 100644 --- a/fs/omfs/file.c +++ b/fs/omfs/file.c @@ -338,7 +338,7 @@ const struct file_operations omfs_file_operations = { .splice_read = filemap_splice_read, }; -static int omfs_setattr(struct mnt_idmap *idmap, +static int omfs_setattr(const struct mnt_idmap *idmap, struct dentry *dentry, struct iattr *attr) { struct inode *inode = d_inode(dentry); diff --git a/fs/omfs/inode.c b/fs/omfs/inode.c index 1d915ef72119..bc37029a4afb 100644 --- a/fs/omfs/inode.c +++ b/fs/omfs/inode.c @@ -145,7 +145,7 @@ static int __omfs_write_inode(struct inode *inode, int wait) mark_buffer_dirty(bh); if (wait) { sync_dirty_buffer(bh); - if (buffer_req(bh) && !buffer_uptodate(bh)) + if (buffer_write_io_error(bh)) sync_failed = 1; } @@ -159,7 +159,7 @@ static int __omfs_write_inode(struct inode *inode, int wait) mark_buffer_dirty(bh2); if (wait) { sync_dirty_buffer(bh2); - if (buffer_req(bh2) && !buffer_uptodate(bh2)) + if (buffer_write_io_error(bh2)) sync_failed = 1; } brelse(bh2); diff --git a/fs/open.c b/fs/open.c index 6b1c14e684a9..e43f02ff64ac 100644 --- a/fs/open.c +++ b/fs/open.c @@ -36,7 +36,7 @@ #include "internal.h" -int do_truncate(struct mnt_idmap *idmap, struct dentry *dentry, +int do_truncate(const struct mnt_idmap *idmap, struct dentry *dentry, loff_t length, unsigned int time_attrs, struct file *filp) { int ret; @@ -72,7 +72,7 @@ int do_truncate(struct mnt_idmap *idmap, struct dentry *dentry, int vfs_truncate(const struct path *path, loff_t length) { - struct mnt_idmap *idmap; + const struct mnt_idmap *idmap; struct inode *inode; int error; @@ -787,7 +787,7 @@ static inline bool setattr_vfsgid(struct iattr *attr, kgid_t kgid) int chown_common(const struct path *path, uid_t user, gid_t group) { - struct mnt_idmap *idmap; + const struct mnt_idmap *idmap; struct user_namespace *fs_userns; struct inode *inode = path->dentry->d_inode; struct delegated_inode delegated_inode = { }; @@ -931,6 +931,11 @@ cleanup_inode: return error; } +/* + * Populate struct file + * + * NOTE: it assumes f_path is populated and consumes the caller's reference. + */ static int do_dentry_open(struct file *f, int (*open)(struct inode *, struct file *)) { @@ -938,7 +943,6 @@ static int do_dentry_open(struct file *f, struct inode *inode = f->f_path.dentry->d_inode; int error; - path_get(&f->f_path); f->f_inode = inode; f->f_mapping = inode->i_mapping; f->f_wb_err = filemap_sample_wb_err(f->f_mapping); @@ -1055,6 +1059,7 @@ int finish_open(struct file *file, struct dentry *dentry, BUG_ON(file->f_mode & FMODE_OPENED); /* once it's opened, it's opened */ file->__f_path.dentry = dentry; + path_get(&file->f_path); return do_dentry_open(file, open); } EXPORT_SYMBOL(finish_open); @@ -1098,6 +1103,7 @@ int vfs_open(const struct path *path, struct file *file) int ret; file->__f_path = *path; + path_get(&file->f_path); ret = do_dentry_open(file, NULL); if (!ret) { /* @@ -1110,6 +1116,25 @@ int vfs_open(const struct path *path, struct file *file) return ret; } +/** + * vfs_open_consume - open the file at the given path and consume the reference + * @path: path to open + * @file: newly allocated file with f_flag initialized + */ +int vfs_open_consume(struct path *path, struct file *file) +{ + int ret; + + file->__f_path = *path; + path->mnt = NULL; + path->dentry = NULL; + ret = do_dentry_open(file, NULL); + if (!ret) { + fsnotify_open(file); + } + return ret; +} + struct file *dentry_open(const struct path *path, int flags, const struct cred *cred) { @@ -1537,6 +1562,19 @@ int filp_close(struct file *filp, fl_owner_t id) } EXPORT_SYMBOL(filp_close); +/* Like filp_close() but the last reference is put right here. */ +int filp_close_sync(struct file *filp, fl_owner_t id) +{ + int retval; + + /* Kernel threads must never put their final reference here. */ + VFS_WARN_ON_ONCE(current->flags & PF_KTHREAD); + retval = filp_flush(filp, id); + fput_close_sync(filp); + + return retval; +} + /* * Careful here! We test whether the file pointer is NULL before * releasing the fd. This ensures that one clone task can't release @@ -1551,13 +1589,11 @@ SYSCALL_DEFINE1(close, unsigned int, fd) if (!file) return -EBADF; - retval = filp_flush(file, current->files); - /* * We're returning to user space. Don't bother * with any delayed fput() cases. */ - fput_close_sync(file); + retval = filp_close_sync(file, current->files); if (likely(retval == 0)) return 0; diff --git a/fs/orangefs/acl.c b/fs/orangefs/acl.c index a01ef0c1b1bf..f31196e4bfaa 100644 --- a/fs/orangefs/acl.c +++ b/fs/orangefs/acl.c @@ -112,7 +112,7 @@ int __orangefs_set_acl(struct inode *inode, struct posix_acl *acl, int type) return error; } -int orangefs_set_acl(struct mnt_idmap *idmap, struct dentry *dentry, +int orangefs_set_acl(const struct mnt_idmap *idmap, struct dentry *dentry, struct posix_acl *acl, int type) { int error; diff --git a/fs/orangefs/inode.c b/fs/orangefs/inode.c index cd3273c88e03..b1fed1c81a4d 100644 --- a/fs/orangefs/inode.c +++ b/fs/orangefs/inode.c @@ -181,8 +181,27 @@ static int orangefs_writepages(struct address_space *mapping, { struct orangefs_writepages *ow; struct blk_plug plug; - int error; + int error = 0; struct folio *folio = NULL; + int maxpages; + + maxpages = orangefs_bufmap_size_query() / PAGE_SIZE; + if (maxpages < 1) { + /* + * Probably the client is dead and there's no bufmap. + * Walk writeback_iter anyway so each dirty folio is unlocked + * and writeback is ended. wait_for_direct_io will fail; the + * data is not written. + */ + gossip_err("%s: maxpages < 1. \n", __func__); + while ((folio = writeback_iter(mapping, wbc, folio, &error))) { + error = orangefs_writepage_locked(folio, wbc); + mapping_set_error(mapping, error); + folio_unlock(folio); + folio_end_writeback(folio); + } + return error; + } ow = kzalloc_obj(struct orangefs_writepages); if (!ow) @@ -828,7 +847,7 @@ int __orangefs_setattr_mode(struct dentry *dentry, struct iattr *iattr) /* * Change attributes of an object referenced by dentry. */ -int orangefs_setattr(struct mnt_idmap *idmap, struct dentry *dentry, +int orangefs_setattr(const struct mnt_idmap *idmap, struct dentry *dentry, struct iattr *iattr) { int ret; @@ -848,7 +867,7 @@ out: /* * Obtain attributes of an object given a dentry */ -int orangefs_getattr(struct mnt_idmap *idmap, const struct path *path, +int orangefs_getattr(const struct mnt_idmap *idmap, const struct path *path, struct kstat *stat, u32 request_mask, unsigned int flags) { int ret; @@ -872,7 +891,7 @@ int orangefs_getattr(struct mnt_idmap *idmap, const struct path *path, return ret; } -int orangefs_permission(struct mnt_idmap *idmap, +int orangefs_permission(const struct mnt_idmap *idmap, struct inode *inode, int mask) { int ret; @@ -934,7 +953,7 @@ static int orangefs_fileattr_get(struct dentry *dentry, struct file_kattr *fa) return 0; } -static int orangefs_fileattr_set(struct mnt_idmap *idmap, +static int orangefs_fileattr_set(const struct mnt_idmap *idmap, struct dentry *dentry, struct file_kattr *fa) { u64 val = 0; diff --git a/fs/orangefs/namei.c b/fs/orangefs/namei.c index 8ebc34e112d5..32b7769ea49c 100644 --- a/fs/orangefs/namei.c +++ b/fs/orangefs/namei.c @@ -15,7 +15,7 @@ /* * Get a newly allocated inode to go with a negative dentry. */ -static int orangefs_create(struct mnt_idmap *idmap, +static int orangefs_create(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *dentry, umode_t mode) @@ -211,7 +211,7 @@ static int orangefs_unlink(struct inode *dir, struct dentry *dentry) return ret; } -static int orangefs_symlink(struct mnt_idmap *idmap, +static int orangefs_symlink(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *dentry, const char *symname) @@ -296,7 +296,7 @@ out: return ret; } -static struct dentry *orangefs_mkdir(struct mnt_idmap *idmap, struct inode *dir, +static struct dentry *orangefs_mkdir(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *dentry, umode_t mode) { struct orangefs_inode_s *parent = ORANGEFS_I(dir); @@ -364,7 +364,7 @@ out: return ret ? ERR_PTR(ret) : NULL; } -static int orangefs_rename(struct mnt_idmap *idmap, +static int orangefs_rename(const struct mnt_idmap *idmap, struct inode *old_dir, struct dentry *old_dentry, struct inode *new_dir, diff --git a/fs/orangefs/orangefs-debugfs.c b/fs/orangefs/orangefs-debugfs.c index 9f94919a6bc6..fec7c1e864d2 100644 --- a/fs/orangefs/orangefs-debugfs.c +++ b/fs/orangefs/orangefs-debugfs.c @@ -238,7 +238,7 @@ void orangefs_debugfs_init(int debug_mask) static void orangefs_kernel_debug_init(void) { static char k_buffer[ORANGEFS_MAX_DEBUG_STRING_LEN] = { }; - size_t len = + ssize_t len = strscpy(k_buffer, kernel_debug_string, sizeof(k_buffer) - 1); if (len > 0) { @@ -246,7 +246,8 @@ static void orangefs_kernel_debug_init(void) k_buffer[len + 1] = '\0'; } else { strscpy(k_buffer, "none\n"); - pr_info("%s: overflow 1!\n", __func__); + if (len <0) + pr_info("%s: overflow 1!\n", __func__); } debugfs_create_file_aux_num(ORANGEFS_KMOD_DEBUG_FILE, 0444, debug_dir, k_buffer, @@ -337,7 +338,7 @@ static int help_show(struct seq_file *m, void *v) static void orangefs_client_debug_init(void) { static char c_buffer[ORANGEFS_MAX_DEBUG_STRING_LEN] = { }; - size_t len = + ssize_t len = strscpy(c_buffer, client_debug_string, sizeof(c_buffer) - 1); if (len > 0) { @@ -345,7 +346,8 @@ static void orangefs_client_debug_init(void) c_buffer[len + 1] = '\0'; } else { strscpy(c_buffer, "none\n"); - pr_info("%s: overflow! 2\n", __func__); + if (len <0) + pr_info("%s: overflow! 2\n", __func__); } client_debug_dentry = debugfs_create_file_aux_num( diff --git a/fs/orangefs/orangefs-kernel.h b/fs/orangefs/orangefs-kernel.h index 1451fc2c1917..348fe340c5d5 100644 --- a/fs/orangefs/orangefs-kernel.h +++ b/fs/orangefs/orangefs-kernel.h @@ -98,7 +98,7 @@ enum orangefs_vfs_op_states { extern const struct xattr_handler * const orangefs_xattr_handlers[]; extern struct posix_acl *orangefs_get_acl(struct inode *inode, int type, bool rcu); -extern int orangefs_set_acl(struct mnt_idmap *idmap, +extern int orangefs_set_acl(const struct mnt_idmap *idmap, struct dentry *dentry, struct posix_acl *acl, int type); int __orangefs_set_acl(struct inode *inode, struct posix_acl *acl, int type); @@ -352,12 +352,12 @@ struct inode *orangefs_new_inode(struct super_block *sb, int __orangefs_setattr(struct inode *, struct iattr *); int __orangefs_setattr_mode(struct dentry *dentry, struct iattr *iattr); -int orangefs_setattr(struct mnt_idmap *, struct dentry *, struct iattr *); +int orangefs_setattr(const struct mnt_idmap *, struct dentry *, struct iattr *); -int orangefs_getattr(struct mnt_idmap *idmap, const struct path *path, +int orangefs_getattr(const struct mnt_idmap *idmap, const struct path *path, struct kstat *stat, u32 request_mask, unsigned int flags); -int orangefs_permission(struct mnt_idmap *idmap, +int orangefs_permission(const struct mnt_idmap *idmap, struct inode *inode, int mask); int orangefs_update_time(struct inode *inode, enum fs_update_time type, diff --git a/fs/orangefs/super.c b/fs/orangefs/super.c index 4ec7329b41f6..e6eedb2faef7 100644 --- a/fs/orangefs/super.c +++ b/fs/orangefs/super.c @@ -383,6 +383,24 @@ static int orangefs_unmount(int id, __s32 fs_id, const char *devname) { struct orangefs_kernel_op_s *op; int r; + + /* + * If someone reboots linux without unmounting orangefs first, + * don't bother firing off an unmount service operation that + * will never complete, or the whole system shutdown + * stalls for ORANGEFS_DEFAULT_OP_TIMEOUT_SECS. + */ + + if (!__is_daemon_in_service() || + system_state == SYSTEM_HALT || + system_state == SYSTEM_POWER_OFF || + system_state == SYSTEM_RESTART) { + gossip_debug(GOSSIP_SUPER_DEBUG, + "%s: dirty unmount.\n", + __func__); + return 0; + } + op = op_alloc(ORANGEFS_VFS_OP_FS_UMOUNT); if (!op) return -ENOMEM; diff --git a/fs/orangefs/xattr.c b/fs/orangefs/xattr.c index 885fd3bd5a3d..a49e64566e1e 100644 --- a/fs/orangefs/xattr.c +++ b/fs/orangefs/xattr.c @@ -527,7 +527,7 @@ out_unlock: } static int orangefs_xattr_set_default(const struct xattr_handler *handler, - struct mnt_idmap *idmap, + const struct mnt_idmap *idmap, struct dentry *unused, struct inode *inode, const char *name, diff --git a/fs/overlayfs/dir.c b/fs/overlayfs/dir.c index 7beb0af26498..1194ccf981c3 100644 --- a/fs/overlayfs/dir.c +++ b/fs/overlayfs/dir.c @@ -688,7 +688,7 @@ static int ovl_create_or_link(struct dentry *dentry, struct inode *inode, return err; } -static int ovl_create_object(struct mnt_idmap *idmap, struct dentry *dentry, +static int ovl_create_object(const struct mnt_idmap *idmap, struct dentry *dentry, int mode, dev_t rdev, const char *link) { int err; @@ -730,19 +730,19 @@ out: return err; } -static int ovl_create(struct mnt_idmap *idmap, struct inode *dir, +static int ovl_create(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *dentry, umode_t mode) { return ovl_create_object(idmap, dentry, (mode & 07777) | S_IFREG, 0, NULL); } -static struct dentry *ovl_mkdir(struct mnt_idmap *idmap, struct inode *dir, +static struct dentry *ovl_mkdir(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *dentry, umode_t mode) { return ERR_PTR(ovl_create_object(idmap, dentry, (mode & 07777) | S_IFDIR, 0, NULL)); } -static int ovl_mknod(struct mnt_idmap *idmap, struct inode *dir, +static int ovl_mknod(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *dentry, umode_t mode, dev_t rdev) { /* Don't allow creation of "whiteout" on overlay */ @@ -752,7 +752,7 @@ static int ovl_mknod(struct mnt_idmap *idmap, struct inode *dir, return ovl_create_object(idmap, dentry, mode, rdev, NULL); } -static int ovl_symlink(struct mnt_idmap *idmap, struct inode *dir, +static int ovl_symlink(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *dentry, const char *link) { return ovl_create_object(idmap, dentry, S_IFLNK, 0, link); @@ -1344,7 +1344,7 @@ static void ovl_rename_end(struct ovl_renamedata *ovlrd) ovl_drop_write(ovlrd->old_dentry); } -static int ovl_rename(struct mnt_idmap *idmap, struct inode *olddir, +static int ovl_rename(const struct mnt_idmap *idmap, struct inode *olddir, struct dentry *old, struct inode *newdir, struct dentry *new, unsigned int flags) { @@ -1420,7 +1420,7 @@ static int ovl_dummy_open(struct inode *inode, struct file *file) return 0; } -static int ovl_tmpfile(struct mnt_idmap *idmap, struct inode *dir, +static int ovl_tmpfile(const struct mnt_idmap *idmap, struct inode *dir, struct file *file, umode_t mode) { int err; diff --git a/fs/overlayfs/file.c b/fs/overlayfs/file.c index f3d97eb146e8..7433220d4ad6 100644 --- a/fs/overlayfs/file.c +++ b/fs/overlayfs/file.c @@ -30,7 +30,7 @@ static struct file *ovl_open_realfile(const struct file *file, { struct inode *realinode = d_inode(realpath->dentry); struct inode *inode = file_inode(file); - struct mnt_idmap *real_idmap; + const struct mnt_idmap *real_idmap; struct file *realfile; int flags = file->f_flags | OVL_OPEN_FLAGS; int acc_mode = ACC_MODE(flags); diff --git a/fs/overlayfs/inode.c b/fs/overlayfs/inode.c index 401cb8c75520..70183d516e5e 100644 --- a/fs/overlayfs/inode.c +++ b/fs/overlayfs/inode.c @@ -18,7 +18,7 @@ #include "overlayfs.h" -int ovl_setattr(struct mnt_idmap *idmap, struct dentry *dentry, +int ovl_setattr(const struct mnt_idmap *idmap, struct dentry *dentry, struct iattr *attr) { int err; @@ -168,7 +168,7 @@ static inline int ovl_real_getattr_nosec(struct super_block *sb, return vfs_getattr_nosec(path, stat, request_mask, flags); } -int ovl_getattr(struct mnt_idmap *idmap, const struct path *path, +int ovl_getattr(const struct mnt_idmap *idmap, const struct path *path, struct kstat *stat, u32 request_mask, unsigned int flags) { struct dentry *dentry = path->dentry; @@ -303,7 +303,7 @@ int ovl_getattr(struct mnt_idmap *idmap, const struct path *path, return err; } -int ovl_permission(struct mnt_idmap *idmap, +int ovl_permission(const struct mnt_idmap *idmap, struct inode *inode, int mask) { struct inode *upperinode = ovl_inode_upper(inode); @@ -355,7 +355,7 @@ static const char *ovl_get_link(struct dentry *dentry, * alter the POSIX ACLs for the underlying filesystem. */ static void ovl_idmap_posix_acl(const struct inode *realinode, - struct mnt_idmap *idmap, + const struct mnt_idmap *idmap, struct posix_acl *acl) { struct user_namespace *fs_userns = i_user_ns(realinode); @@ -406,7 +406,7 @@ struct posix_acl *ovl_get_acl_path(const struct path *path, const char *acl_name, bool noperm) { struct posix_acl *real_acl, *clone; - struct mnt_idmap *idmap; + const struct mnt_idmap *idmap; struct inode *realinode = d_inode(path->dentry); idmap = mnt_idmap(path->mnt); @@ -447,7 +447,7 @@ struct posix_acl *ovl_get_acl_path(const struct path *path, * * This is obviously only relevant when idmapped layers are used. */ -struct posix_acl *do_ovl_get_acl(struct mnt_idmap *idmap, +struct posix_acl *do_ovl_get_acl(const struct mnt_idmap *idmap, struct inode *inode, int type, bool rcu, bool noperm) { @@ -536,7 +536,7 @@ out: return err; } -int ovl_set_acl(struct mnt_idmap *idmap, struct dentry *dentry, +int ovl_set_acl(const struct mnt_idmap *idmap, struct dentry *dentry, struct posix_acl *acl, int type) { int err; @@ -650,7 +650,7 @@ int ovl_real_fileattr_set(const struct path *realpath, struct file_kattr *fa) return vfs_fileattr_set(mnt_idmap(realpath->mnt), realpath->dentry, fa); } -int ovl_fileattr_set(struct mnt_idmap *idmap, +int ovl_fileattr_set(const struct mnt_idmap *idmap, struct dentry *dentry, struct file_kattr *fa) { struct inode *inode = d_inode(dentry); diff --git a/fs/overlayfs/overlayfs.h b/fs/overlayfs/overlayfs.h index 7f3558372c59..53fbbe15c31d 100644 --- a/fs/overlayfs/overlayfs.h +++ b/fs/overlayfs/overlayfs.h @@ -804,11 +804,11 @@ int ovl_set_nlink_lower(struct dentry *dentry); unsigned int ovl_get_nlink(struct ovl_fs *ofs, struct dentry *lowerdentry, struct dentry *upperdentry, unsigned int fallback); -int ovl_permission(struct mnt_idmap *idmap, struct inode *inode, +int ovl_permission(const struct mnt_idmap *idmap, struct inode *inode, int mask); #ifdef CONFIG_FS_POSIX_ACL -struct posix_acl *do_ovl_get_acl(struct mnt_idmap *idmap, +struct posix_acl *do_ovl_get_acl(const struct mnt_idmap *idmap, struct inode *inode, int type, bool rcu, bool noperm); static inline struct posix_acl *ovl_get_inode_acl(struct inode *inode, int type, @@ -816,12 +816,12 @@ static inline struct posix_acl *ovl_get_inode_acl(struct inode *inode, int type, { return do_ovl_get_acl(&nop_mnt_idmap, inode, type, rcu, true); } -static inline struct posix_acl *ovl_get_acl(struct mnt_idmap *idmap, +static inline struct posix_acl *ovl_get_acl(const struct mnt_idmap *idmap, struct dentry *dentry, int type) { return do_ovl_get_acl(idmap, d_inode(dentry), type, false, false); } -int ovl_set_acl(struct mnt_idmap *idmap, struct dentry *dentry, +int ovl_set_acl(const struct mnt_idmap *idmap, struct dentry *dentry, struct posix_acl *acl, int type); struct posix_acl *ovl_get_acl_path(const struct path *path, const char *acl_name, bool noperm); @@ -916,7 +916,7 @@ extern const struct file_operations ovl_file_operations; int ovl_real_fileattr_get(const struct path *realpath, struct file_kattr *fa); int ovl_real_fileattr_set(const struct path *realpath, struct file_kattr *fa); int ovl_fileattr_get(struct dentry *dentry, struct file_kattr *fa); -int ovl_fileattr_set(struct mnt_idmap *idmap, +int ovl_fileattr_set(const struct mnt_idmap *idmap, struct dentry *dentry, struct file_kattr *fa); struct ovl_file; struct ovl_file *ovl_file_alloc(struct file *realfile); @@ -950,8 +950,8 @@ static inline bool ovl_force_readonly(struct ovl_fs *ofs) /* xattr.c */ const struct xattr_handler * const *ovl_xattr_handlers(struct ovl_fs *ofs); -int ovl_setattr(struct mnt_idmap *idmap, struct dentry *dentry, +int ovl_setattr(const struct mnt_idmap *idmap, struct dentry *dentry, struct iattr *attr); -int ovl_getattr(struct mnt_idmap *idmap, const struct path *path, +int ovl_getattr(const struct mnt_idmap *idmap, const struct path *path, struct kstat *stat, u32 request_mask, unsigned int flags); ssize_t ovl_listxattr(struct dentry *dentry, char *list, size_t size); diff --git a/fs/overlayfs/ovl_entry.h b/fs/overlayfs/ovl_entry.h index 80cad4ea96a3..ac07e8769f9b 100644 --- a/fs/overlayfs/ovl_entry.h +++ b/fs/overlayfs/ovl_entry.h @@ -105,7 +105,7 @@ static inline struct vfsmount *ovl_upper_mnt(struct ovl_fs *ofs) return ofs->layers[0].mnt; } -static inline struct mnt_idmap *ovl_upper_mnt_idmap(struct ovl_fs *ofs) +static inline const struct mnt_idmap *ovl_upper_mnt_idmap(struct ovl_fs *ofs) { return mnt_idmap(ovl_upper_mnt(ofs)); } diff --git a/fs/overlayfs/util.c b/fs/overlayfs/util.c index b41f4788e4f0..521717209b2e 100644 --- a/fs/overlayfs/util.c +++ b/fs/overlayfs/util.c @@ -657,7 +657,7 @@ bool ovl_path_is_whiteout(struct ovl_fs *ofs, const struct path *path) struct file *ovl_path_open(const struct path *path, int flags) { struct inode *inode = d_inode(path->dentry); - struct mnt_idmap *real_idmap = mnt_idmap(path->mnt); + const struct mnt_idmap *real_idmap = mnt_idmap(path->mnt); int err, acc_mode; if (flags & ~(O_ACCMODE | O_LARGEFILE)) @@ -1496,7 +1496,7 @@ void ovl_copyattr(struct inode *inode) { struct path realpath; struct inode *realinode; - struct mnt_idmap *real_idmap; + const struct mnt_idmap *real_idmap; vfsuid_t vfsuid; vfsgid_t vfsgid; diff --git a/fs/overlayfs/xattrs.c b/fs/overlayfs/xattrs.c index 5ae44b9c8790..acc54f6138ef 100644 --- a/fs/overlayfs/xattrs.c +++ b/fs/overlayfs/xattrs.c @@ -190,7 +190,7 @@ static int ovl_own_xattr_get(const struct xattr_handler *handler, } static int ovl_own_xattr_set(const struct xattr_handler *handler, - struct mnt_idmap *idmap, + const struct mnt_idmap *idmap, struct dentry *dentry, struct inode *inode, const char *name, const void *value, size_t size, int flags) @@ -217,7 +217,7 @@ static int ovl_other_xattr_get(const struct xattr_handler *handler, } static int ovl_other_xattr_set(const struct xattr_handler *handler, - struct mnt_idmap *idmap, + const struct mnt_idmap *idmap, struct dentry *dentry, struct inode *inode, const char *name, const void *value, size_t size, int flags) diff --git a/fs/pidfs.c b/fs/pidfs.c index a6a643f15d08..c37c17bcdbdc 100644 --- a/fs/pidfs.c +++ b/fs/pidfs.c @@ -823,13 +823,13 @@ static struct vfsmount *pidfs_mnt __ro_after_init; * implemented. Let's reject it completely until we have a clean * permission concept for pidfds. */ -static int pidfs_setattr(struct mnt_idmap *idmap, struct dentry *dentry, +static int pidfs_setattr(const struct mnt_idmap *idmap, struct dentry *dentry, struct iattr *attr) { return anon_inode_setattr(idmap, dentry, attr); } -static int pidfs_getattr(struct mnt_idmap *idmap, const struct path *path, +static int pidfs_getattr(const struct mnt_idmap *idmap, const struct path *path, struct kstat *stat, u32 request_mask, unsigned int query_flags) { @@ -1102,7 +1102,7 @@ static int pidfs_xattr_get(const struct xattr_handler *handler, } static int pidfs_xattr_set(const struct xattr_handler *handler, - struct mnt_idmap *idmap, struct dentry *unused, + const struct mnt_idmap *idmap, struct dentry *unused, struct inode *inode, const char *suffix, const void *value, size_t size, int flags) { diff --git a/fs/pipe.c b/fs/pipe.c index 292425a834dd..5791db016e24 100644 --- a/fs/pipe.c +++ b/fs/pipe.c @@ -1433,7 +1433,7 @@ int pipe_resize_ring(struct pipe_inode_info *pipe, unsigned int nr_slots) spin_unlock_irq(&pipe->rd_wait.lock); /* This might have made more room for writers */ - wake_up_interruptible(&pipe->wr_wait); + wake_up_interruptible_poll(&pipe->wr_wait, EPOLLOUT | EPOLLWRNORM); return 0; } diff --git a/fs/pnode.c b/fs/pnode.c index 5d91c3e58d2a..2cd667958efe 100644 --- a/fs/pnode.c +++ b/fs/pnode.c @@ -410,19 +410,99 @@ bool propagation_would_overmount(const struct mount *from, return false; } +/* Does @m receive propagation from @parent? */ +static bool receives_from(struct mount *m, struct mount *parent) +{ + if (m == parent) + return false; + for (; m; m = m->mnt_master) + if (m == parent || peers(m, parent)) + return true; + return false; +} + +/* + * Does @m receive propagation from the victim's parent as well? If so, then + * the mount at the victim's mountpoint inside of @m is a umount candidate as + * well. So it's the next candidate in the chain. Otherwise the chain ends at + * @m. + */ +static struct mount *next_candidate(struct mount *m, struct mount *victim) +{ + if (!receives_from(m, victim->mnt_parent)) + return NULL; + return __lookup_mnt(&m->mnt, victim->mnt_mountpoint); +} + +/* + * Would propagate_umount() pull out a mount of the chain of candidates that + * starts at @c, and does that mount have references beyond its own? + * + * This mirrors how trim_one(), trim_ancestors() and handle_locked() handle a + * synchronous umount: + * + * - single victim + * - without children + * - with MNT_LOCKED already cleared on every candidate by propagate_mount_unlock() + * + * A copy of the victim gets unmounted when each of its children is + * the next candidate in the chain or its overmount, unless the next + * unmount candidate is not its overmount and some unmount candidate further + * down has a child outside the chain. Keep this in sync with + * Documentation/filesystems/propagate_umount.txt. + */ +static bool chain_busy(struct mount *c, struct mount *victim) +{ + struct mount *m, *n, *next, *deepest = NULL; + bool above; + + /* the deepest candidate with a child outside the chain */ + for (m = c; m; m = next) { + next = next_candidate(m, victim); + list_for_each_entry(n, &m->mnt_mounts, mnt_child) { + if (n != next && n != victim) { + deepest = m; + break; + } + } + } + + above = deepest != NULL; /* @deepest is at or below @m */ + for (m = c; m; m = next) { + bool goes = true; + + next = next_candidate(m, victim); + list_for_each_entry(n, &m->mnt_mounts, mnt_child) { + if (n != next && n != m->overmount && n != victim) { + goes = false; + break; + } + } + if (goes && next && next != m->overmount && above && m != deepest) + goes = false; + if (m == deepest) + above = false; + if (goes && do_refcount_check(m, 1)) + return true; + } + return false; +} + /* * check if the mount 'mnt' can be unmounted successfully. * @mnt: the mount to be checked for unmount * NOTE: unmounting 'mnt' would naturally propagate to all * other mounts its parent propagates to. - * Check if any of these mounts that **do not have submounts** - * have more references than 'refcnt'. If so return busy. + * Check if any of the mounts that propagate_umount() would pull out + * along with it have more references than their own. If so return busy. * * vfsmount lock must be held for write */ int propagate_mount_busy(struct mount *mnt, int refcnt) { struct mount *parent = mnt->mnt_parent; + struct dentry *mp = mnt->mnt_mountpoint; + struct mount *m; /* * quickly check if the current mount can be unmounted. @@ -435,24 +515,16 @@ int propagate_mount_busy(struct mount *mnt, int refcnt) if (mnt == parent) return 0; - for (struct mount *m = propagation_next(parent, parent); m; - m = propagation_next(m, parent)) { - struct list_head *head; - struct mount *child = __lookup_mnt(&m->mnt, mnt->mnt_mountpoint); + /* the candidates are the mounts at @mp below the receivers */ + for (m = propagation_next(parent, parent); m; + m = propagation_next(m, parent)) { + struct mount *c = __lookup_mnt(&m->mnt, mp); - if (!child) + /* each chain once, from its top: skip receivers that are candidates */ + if (!c || (mnt_has_parent(m) && m->mnt_mountpoint == mp && + receives_from(m->mnt_parent, parent))) continue; - - head = &child->mnt_mounts; - if (!list_empty(head)) { - /* - * a mount that covers child completely wouldn't prevent - * it being pulled out; any other would. - */ - if (!list_is_singular(head) || !child->overmount) - continue; - } - if (do_refcount_check(child, 1)) + if (chain_busy(c, mnt)) return 1; } return 0; diff --git a/fs/posix_acl.c b/fs/posix_acl.c index 18b302f94174..fe77934ea8f2 100644 --- a/fs/posix_acl.c +++ b/fs/posix_acl.c @@ -118,7 +118,7 @@ void forget_all_cached_acls(struct inode *inode) } EXPORT_SYMBOL(forget_all_cached_acls); -static struct posix_acl *__get_acl(struct mnt_idmap *idmap, +static struct posix_acl *__get_acl(const struct mnt_idmap *idmap, struct dentry *dentry, struct inode *inode, int type) { @@ -378,7 +378,7 @@ EXPORT_SYMBOL(posix_acl_from_mode); * by the acl. Returns -E... otherwise. */ int -posix_acl_permission(struct mnt_idmap *idmap, struct inode *inode, +posix_acl_permission(const struct mnt_idmap *idmap, struct inode *inode, const struct posix_acl *acl, int want) { const struct posix_acl_entry *pa, *pe, *mask_obj; @@ -608,7 +608,7 @@ EXPORT_SYMBOL(__posix_acl_chmod); * performed on the raw inode simply pass @nop_mnt_idmap. */ int - posix_acl_chmod(struct mnt_idmap *idmap, struct dentry *dentry, + posix_acl_chmod(const struct mnt_idmap *idmap, struct dentry *dentry, umode_t mode) { struct inode *inode = d_inode(dentry); @@ -709,7 +709,7 @@ EXPORT_SYMBOL_GPL(posix_acl_create); * * Called from set_acl inode operations. */ -int posix_acl_update_mode(struct mnt_idmap *idmap, +int posix_acl_update_mode(const struct mnt_idmap *idmap, struct inode *inode, umode_t *mode_p, struct posix_acl **acl) { @@ -889,7 +889,7 @@ EXPORT_SYMBOL (posix_acl_to_xattr); * Return: On success, the size of the stored uapi posix acls, on error a * negative errno. */ -static ssize_t vfs_posix_acl_to_xattr(struct mnt_idmap *idmap, +static ssize_t vfs_posix_acl_to_xattr(const struct mnt_idmap *idmap, struct inode *inode, const struct posix_acl *acl, void *buffer, size_t size) @@ -937,7 +937,7 @@ static ssize_t vfs_posix_acl_to_xattr(struct mnt_idmap *idmap, } int -set_posix_acl(struct mnt_idmap *idmap, struct dentry *dentry, +set_posix_acl(const struct mnt_idmap *idmap, struct dentry *dentry, int type, struct posix_acl *acl) { struct inode *inode = d_inode(dentry); @@ -1018,7 +1018,7 @@ const struct xattr_handler nop_posix_acl_default = { }; EXPORT_SYMBOL_GPL(nop_posix_acl_default); -int simple_set_acl(struct mnt_idmap *idmap, struct dentry *dentry, +int simple_set_acl(const struct mnt_idmap *idmap, struct dentry *dentry, struct posix_acl *acl, int type) { int error; @@ -1057,7 +1057,7 @@ int simple_acl_create(struct inode *dir, struct inode *inode) return 0; } -static int vfs_set_acl_idmapped_mnt(struct mnt_idmap *idmap, +static int vfs_set_acl_idmapped_mnt(const struct mnt_idmap *idmap, struct user_namespace *fs_userns, struct posix_acl *acl) { @@ -1091,7 +1091,7 @@ static int vfs_set_acl_idmapped_mnt(struct mnt_idmap *idmap, * * Return: On success 0, on error negative errno. */ -int vfs_set_acl(struct mnt_idmap *idmap, struct dentry *dentry, +int vfs_set_acl(const struct mnt_idmap *idmap, struct dentry *dentry, const char *acl_name, struct posix_acl *kacl) { int acl_type; @@ -1168,7 +1168,7 @@ EXPORT_SYMBOL_GPL(vfs_set_acl); * * Return: On success POSIX ACLs in VFS format, on error negative errno. */ -struct posix_acl *vfs_get_acl(struct mnt_idmap *idmap, +struct posix_acl *vfs_get_acl(const struct mnt_idmap *idmap, struct dentry *dentry, const char *acl_name) { struct inode *inode = d_inode(dentry); @@ -1212,7 +1212,7 @@ EXPORT_SYMBOL_GPL(vfs_get_acl); * * Return: On success 0, on error negative errno. */ -int vfs_remove_acl(struct mnt_idmap *idmap, struct dentry *dentry, +int vfs_remove_acl(const struct mnt_idmap *idmap, struct dentry *dentry, const char *acl_name) { int acl_type; @@ -1265,7 +1265,7 @@ out_inode_unlock: } EXPORT_SYMBOL_GPL(vfs_remove_acl); -int do_set_acl(struct mnt_idmap *idmap, struct dentry *dentry, +int do_set_acl(const struct mnt_idmap *idmap, struct dentry *dentry, const char *acl_name, const void *kvalue, size_t size) { int error; @@ -1286,7 +1286,7 @@ int do_set_acl(struct mnt_idmap *idmap, struct dentry *dentry, return error; } -ssize_t do_get_acl(struct mnt_idmap *idmap, struct dentry *dentry, +ssize_t do_get_acl(const struct mnt_idmap *idmap, struct dentry *dentry, const char *acl_name, void *kvalue, size_t size) { ssize_t error; diff --git a/fs/proc/base.c b/fs/proc/base.c index 0f9efd25bb05..8c04fcd436d2 100644 --- a/fs/proc/base.c +++ b/fs/proc/base.c @@ -703,7 +703,7 @@ static int proc_pid_syscall(struct seq_file *m, struct pid_namespace *ns, /* Here the fs part begins */ /************************************************************************/ -int proc_nochmod_setattr(struct mnt_idmap *idmap, struct dentry *dentry, +int proc_nochmod_setattr(const struct mnt_idmap *idmap, struct dentry *dentry, struct iattr *attr) { int error; @@ -744,7 +744,7 @@ static bool has_pid_permissions(struct proc_fs_info *fs_info, } -static int proc_pid_permission(struct mnt_idmap *idmap, +static int proc_pid_permission(const struct mnt_idmap *idmap, struct inode *inode, int mask) { struct proc_fs_info *fs_info = proc_sb_info(inode->i_sb); @@ -849,15 +849,35 @@ static int __mem_open(struct inode *inode, struct file *file, unsigned int mode) return 0; } +/* private_data for proc_mem_operations */ +struct mem_private { + struct mm_struct *mm; + /* + * Was the ptrace access check on open bypassed because the opener used + * the same MM (introspection)? + */ + bool opened_by_owner; +}; + static int mem_open(struct inode *inode, struct file *file) { + struct mem_private *priv __free(kfree) = kmalloc_obj(struct mem_private); + + if (!priv) + return -ENOMEM; if (WARN_ON_ONCE(!(file->f_op->fop_flags & FOP_UNSIGNED_OFFSET))) return -EINVAL; - return __mem_open(inode, file, PTRACE_MODE_ATTACH); + priv->mm = proc_mem_open(inode, PTRACE_MODE_ATTACH); + if (IS_ERR_OR_NULL(priv->mm)) + return priv->mm ? PTR_ERR(priv->mm) : -ESRCH; + priv->opened_by_owner = priv->mm == current->mm; + file->private_data = no_free_ptr(priv); + return 0; } static bool proc_mem_foll_force(struct file *file, struct mm_struct *mm) { + struct mem_private *priv = file->private_data; struct task_struct *task; bool ptrace_active = false; @@ -872,16 +892,20 @@ static bool proc_mem_foll_force(struct file *file, struct mm_struct *mm) READ_ONCE(task->parent) == current; put_task_struct(task); } - return ptrace_active; + if (!ptrace_active) + return false; + break; default: - return true; + break; } + return security_mem_foll_force(file->f_cred, priv->opened_by_owner) == 0; } static ssize_t mem_rw(struct file *file, char __user *buf, size_t count, loff_t *ppos, int write) { - struct mm_struct *mm = file->private_data; + struct mem_private *priv = file->private_data; + struct mm_struct *mm = priv->mm; unsigned long addr = *ppos; ssize_t copied; char *page; @@ -971,12 +995,21 @@ static int mem_release(struct inode *inode, struct file *file) return 0; } +static int mem_release_with_private(struct inode *inode, struct file *file) +{ + struct mem_private *priv = file->private_data; + + mmdrop(priv->mm); + kfree(priv); + return 0; +} + static const struct file_operations proc_mem_operations = { .llseek = mem_lseek, .read = mem_read, .write = mem_write, .open = mem_open, - .release = mem_release, + .release = mem_release_with_private, .fop_flags = FOP_UNSIGNED_OFFSET, }; @@ -1995,7 +2028,7 @@ static struct inode *proc_pid_make_base_inode(struct super_block *sb, return inode; } -int pid_getattr(struct mnt_idmap *idmap, const struct path *path, +int pid_getattr(const struct mnt_idmap *idmap, const struct path *path, struct kstat *stat, u32 request_mask, unsigned int query_flags) { struct inode *inode = d_inode(path->dentry); @@ -3614,7 +3647,7 @@ int proc_pid_readdir(struct file *file, struct dir_context *ctx) * This function makes sure that the node is always accessible for members of * same thread group. */ -static int proc_tid_comm_permission(struct mnt_idmap *idmap, +static int proc_tid_comm_permission(const struct mnt_idmap *idmap, struct inode *inode, int mask) { bool is_same_tgroup; @@ -3943,7 +3976,7 @@ static int proc_task_readdir(struct file *file, struct dir_context *ctx) return 0; } -static int proc_task_getattr(struct mnt_idmap *idmap, +static int proc_task_getattr(const struct mnt_idmap *idmap, const struct path *path, struct kstat *stat, u32 request_mask, unsigned int query_flags) { diff --git a/fs/proc/fd.c b/fs/proc/fd.c index 0f9a1556f2a3..7214a7495380 100644 --- a/fs/proc/fd.c +++ b/fs/proc/fd.c @@ -82,7 +82,7 @@ static int seq_fdinfo_open(struct inode *inode, struct file *file) * that the current task has PTRACE_MODE_READ in addition to the normal * POSIX-like checks. */ -static int proc_fdinfo_permission(struct mnt_idmap *idmap, struct inode *inode, +static int proc_fdinfo_permission(const struct mnt_idmap *idmap, struct inode *inode, int mask) { bool allowed = false; @@ -323,7 +323,7 @@ static struct dentry *proc_lookupfd(struct inode *dir, struct dentry *dentry, * /proc/pid/fd needs a special permission handler so that a process can still * access /proc/self/fd after it has executed a setuid(). */ -int proc_fd_permission(struct mnt_idmap *idmap, +int proc_fd_permission(const struct mnt_idmap *idmap, struct inode *inode, int mask) { struct task_struct *p; @@ -342,7 +342,7 @@ int proc_fd_permission(struct mnt_idmap *idmap, return rv; } -static int proc_fd_getattr(struct mnt_idmap *idmap, +static int proc_fd_getattr(const struct mnt_idmap *idmap, const struct path *path, struct kstat *stat, u32 request_mask, unsigned int query_flags) { diff --git a/fs/proc/fd.h b/fs/proc/fd.h index 7e7265f7e06f..77f2e4f38592 100644 --- a/fs/proc/fd.h +++ b/fs/proc/fd.h @@ -10,7 +10,7 @@ extern const struct inode_operations proc_fd_inode_operations; extern const struct file_operations proc_fdinfo_operations; extern const struct inode_operations proc_fdinfo_inode_operations; -extern int proc_fd_permission(struct mnt_idmap *idmap, +extern int proc_fd_permission(const struct mnt_idmap *idmap, struct inode *inode, int mask); static inline unsigned int proc_fd(struct inode *inode) diff --git a/fs/proc/generic.c b/fs/proc/generic.c index 26086a283672..2b1971da4a85 100644 --- a/fs/proc/generic.c +++ b/fs/proc/generic.c @@ -117,7 +117,7 @@ static bool pde_subdir_insert(struct proc_dir_entry *dir, return true; } -static int proc_setattr(struct mnt_idmap *idmap, struct dentry *dentry, +static int proc_setattr(const struct mnt_idmap *idmap, struct dentry *dentry, struct iattr *iattr) { struct inode *inode = d_inode(dentry); @@ -135,7 +135,7 @@ static int proc_setattr(struct mnt_idmap *idmap, struct dentry *dentry, return 0; } -static int proc_getattr(struct mnt_idmap *idmap, +static int proc_getattr(const struct mnt_idmap *idmap, const struct path *path, struct kstat *stat, u32 request_mask, unsigned int query_flags) { diff --git a/fs/proc/internal.h b/fs/proc/internal.h index 623bb43ede55..df873df9afc5 100644 --- a/fs/proc/internal.h +++ b/fs/proc/internal.h @@ -258,9 +258,9 @@ extern int proc_pid_statm(struct seq_file *, struct pid_namespace *, * base.c */ extern const struct dentry_operations pid_dentry_operations; -extern int pid_getattr(struct mnt_idmap *, const struct path *, +extern int pid_getattr(const struct mnt_idmap *, const struct path *, struct kstat *, u32, unsigned int); -int proc_nochmod_setattr(struct mnt_idmap *idmap, struct dentry *dentry, +int proc_nochmod_setattr(const struct mnt_idmap *idmap, struct dentry *dentry, struct iattr *attr); extern void proc_pid_evict_inode(struct proc_inode *); extern struct inode *proc_pid_make_inode(struct super_block *, struct task_struct *, umode_t); diff --git a/fs/proc/proc_net.c b/fs/proc/proc_net.c index 00cc385bce21..b1f5eafb069a 100644 --- a/fs/proc/proc_net.c +++ b/fs/proc/proc_net.c @@ -308,7 +308,7 @@ static struct dentry *proc_tgid_net_lookup(struct inode *dir, return de; } -static int proc_tgid_net_getattr(struct mnt_idmap *idmap, +static int proc_tgid_net_getattr(const struct mnt_idmap *idmap, const struct path *path, struct kstat *stat, u32 request_mask, unsigned int query_flags) { diff --git a/fs/proc/proc_sysctl.c b/fs/proc/proc_sysctl.c index 04a382178c65..d1cfd2941359 100644 --- a/fs/proc/proc_sysctl.c +++ b/fs/proc/proc_sysctl.c @@ -788,7 +788,7 @@ out: return 0; } -static int proc_sys_permission(struct mnt_idmap *idmap, +static int proc_sys_permission(const struct mnt_idmap *idmap, struct inode *inode, int mask) { /* @@ -817,7 +817,7 @@ static int proc_sys_permission(struct mnt_idmap *idmap, return error; } -static int proc_sys_setattr(struct mnt_idmap *idmap, +static int proc_sys_setattr(const struct mnt_idmap *idmap, struct dentry *dentry, struct iattr *attr) { struct inode *inode = d_inode(dentry); @@ -834,7 +834,7 @@ static int proc_sys_setattr(struct mnt_idmap *idmap, return 0; } -static int proc_sys_getattr(struct mnt_idmap *idmap, +static int proc_sys_getattr(const struct mnt_idmap *idmap, const struct path *path, struct kstat *stat, u32 request_mask, unsigned int query_flags) { diff --git a/fs/proc/root.c b/fs/proc/root.c index 99adddfeb4a4..7fbbe92bf73a 100644 --- a/fs/proc/root.c +++ b/fs/proc/root.c @@ -402,7 +402,7 @@ void __init proc_root_init(void) register_filesystem(&proc_fs_type); } -static int proc_root_getattr(struct mnt_idmap *idmap, +static int proc_root_getattr(const struct mnt_idmap *idmap, const struct path *path, struct kstat *stat, u32 request_mask, unsigned int query_flags) { diff --git a/fs/proc/vmcore.c b/fs/proc/vmcore.c index 2dc3803aa11d..1fe70c4ecc54 100644 --- a/fs/proc/vmcore.c +++ b/fs/proc/vmcore.c @@ -1710,6 +1710,24 @@ static void vmcore_free_device_dumps(void) #endif /* CONFIG_PROC_VMCORE_DEVICE_DUMP */ } +#define VMCOREINFO_OSRELEASE_KEY "OSRELEASE=" + +static void __init vmcore_report_crashed_release(void) +{ + const char *ver, *eol; + + ver = strnstr(elfnotes_buf, VMCOREINFO_OSRELEASE_KEY, elfnotes_sz); + if (!ver) + return; + + ver += sizeof(VMCOREINFO_OSRELEASE_KEY) - 1; + eol = memchr(ver, '\n', elfnotes_buf + elfnotes_sz - ver); + if (!eol) + return; + + pr_notice("dump is from kernel %.*s\n", (int)(eol - ver), ver); +} + /* Init function for vmcore module. */ static int __init vmcore_init(void) { @@ -1734,6 +1752,8 @@ static int __init vmcore_init(void) elfcorehdr_free(elfcorehdr_addr); elfcorehdr_addr = ELFCORE_ADDR_ERR; + vmcore_report_crashed_release(); + proc_vmcore = proc_create("vmcore", S_IRUSR, NULL, &vmcore_proc_ops); if (proc_vmcore) proc_vmcore->size = vmcore_size; diff --git a/fs/quota/dquot.c b/fs/quota/dquot.c index 1c78c695d0dd..3ad19d7e472f 100644 --- a/fs/quota/dquot.c +++ b/fs/quota/dquot.c @@ -809,6 +809,7 @@ static unsigned long dqcache_shrink_scan(struct shrinker *shrink, struct shrink_control *sc) { struct dquot *dquot; + unsigned long orig_nr_to_scan = sc->nr_to_scan; unsigned long freed = 0; spin_lock(&dq_list_lock); @@ -822,6 +823,21 @@ dqcache_shrink_scan(struct shrinker *shrink, struct shrink_control *sc) freed++; } spin_unlock(&dq_list_lock); + + sc->nr_scanned = orig_nr_to_scan - sc->nr_to_scan; + + /* + * DQST_FREE_DQUOTS is a percpu counter, so count_objects() can + * report a stale/approximate value that is still positive even + * though free_dquots is actually empty by the time we get here. + * When that happens sc->nr_scanned comes back 0 and we have + * nothing further to offer this reclaim pass, so tell + * do_shrink_slab() to stop calling us instead of letting it burn + * through its one-shot scan budget against an empty list. + */ + if (sc->nr_scanned == 0) + return SHRINK_STOP; + return freed; } @@ -2079,7 +2095,7 @@ EXPORT_SYMBOL(__dquot_transfer); /* Wrapper for transferring ownership of an inode for uid/gid only * Called from FSXXX_setattr() */ -int dquot_transfer(struct mnt_idmap *idmap, struct inode *inode, +int dquot_transfer(const struct mnt_idmap *idmap, struct inode *inode, struct iattr *iattr) { struct dquot *transfer_to[MAXQUOTAS] = {}; diff --git a/fs/ramfs/file-nommu.c b/fs/ramfs/file-nommu.c index fb471bf88ab7..7ed6fba134c6 100644 --- a/fs/ramfs/file-nommu.c +++ b/fs/ramfs/file-nommu.c @@ -22,7 +22,7 @@ #include <linux/uaccess.h> #include "internal.h" -static int ramfs_nommu_setattr(struct mnt_idmap *, struct dentry *, struct iattr *); +static int ramfs_nommu_setattr(const struct mnt_idmap *, struct dentry *, struct iattr *); static unsigned long ramfs_nommu_get_unmapped_area(struct file *file, unsigned long addr, unsigned long len, @@ -161,7 +161,7 @@ static int ramfs_nommu_resize(struct inode *inode, loff_t newsize, loff_t size) * handle a change of attributes * - we're specifically interested in a change of size */ -static int ramfs_nommu_setattr(struct mnt_idmap *idmap, +static int ramfs_nommu_setattr(const struct mnt_idmap *idmap, struct dentry *dentry, struct iattr *ia) { struct inode *inode = d_inode(dentry); diff --git a/fs/ramfs/inode.c b/fs/ramfs/inode.c index 0a88ede48e0a..ef91b933d4e6 100644 --- a/fs/ramfs/inode.c +++ b/fs/ramfs/inode.c @@ -95,7 +95,7 @@ struct inode *ramfs_get_inode(struct super_block *sb, */ /* SMP-safe */ static int -ramfs_mknod(struct mnt_idmap *idmap, struct inode *dir, +ramfs_mknod(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *dentry, umode_t mode, dev_t dev) { struct inode * inode = ramfs_get_inode(dir->i_sb, dir, mode, dev); @@ -118,7 +118,7 @@ out: return error; } -static struct dentry *ramfs_mkdir(struct mnt_idmap *idmap, struct inode *dir, +static struct dentry *ramfs_mkdir(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *dentry, umode_t mode) { int retval = ramfs_mknod(&nop_mnt_idmap, dir, dentry, mode, 0); @@ -127,13 +127,13 @@ static struct dentry *ramfs_mkdir(struct mnt_idmap *idmap, struct inode *dir, return ERR_PTR(retval); } -static int ramfs_create(struct mnt_idmap *idmap, struct inode *dir, +static int ramfs_create(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *dentry, umode_t mode) { return ramfs_mknod(&nop_mnt_idmap, dir, dentry, mode | S_IFREG, 0); } -static int ramfs_symlink(struct mnt_idmap *idmap, struct inode *dir, +static int ramfs_symlink(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *dentry, const char *symname) { struct inode *inode; @@ -163,7 +163,7 @@ out: return error; } -static int ramfs_tmpfile(struct mnt_idmap *idmap, +static int ramfs_tmpfile(const struct mnt_idmap *idmap, struct inode *dir, struct file *file, umode_t mode) { struct inode *inode; diff --git a/fs/read_write.c b/fs/read_write.c index e8c14e2760b2..4da846a2bf17 100644 --- a/fs/read_write.c +++ b/fs/read_write.c @@ -274,7 +274,7 @@ loff_t fixed_size_llseek(struct file *file, loff_t offset, int whence, loff_t si EXPORT_SYMBOL(fixed_size_llseek); /** - * no_seek_end_llseek - llseek implementation for fixed-sized devices + * no_seek_end_llseek - llseek implementation for files without SEEK_END * @file: file structure to seek on * @offset: file offset to seek to * @whence: type of seek @@ -293,7 +293,7 @@ loff_t no_seek_end_llseek(struct file *file, loff_t offset, int whence) EXPORT_SYMBOL(no_seek_end_llseek); /** - * no_seek_end_llseek_size - llseek implementation for fixed-sized devices + * no_seek_end_llseek_size - llseek implementation for files without SEEK_END * @file: file structure to seek on * @offset: file offset to seek to * @whence: type of seek @@ -1761,8 +1761,8 @@ EXPORT_SYMBOL(generic_write_checks_count); * Performs necessary checks before doing a write * * Can adjust writing position or amount of bytes to write. - * Returns appropriate error code that caller should return or - * zero in case that write should be allowed. + * Returns the number of bytes that may be written on success (which + * may be less than requested if truncated), or a negative error code. */ ssize_t generic_write_checks(struct kiocb *iocb, struct iov_iter *from) { diff --git a/fs/remap_range.c b/fs/remap_range.c index 26afbbbfb10c..6eb7d845de5d 100644 --- a/fs/remap_range.c +++ b/fs/remap_range.c @@ -415,7 +415,7 @@ EXPORT_SYMBOL(vfs_clone_file_range); /* Check whether we are allowed to dedupe the destination file */ static bool may_dedupe_file(struct file *file) { - struct mnt_idmap *idmap = file_mnt_idmap(file); + const struct mnt_idmap *idmap = file_mnt_idmap(file); struct inode *inode = file_inode(file); if (capable(CAP_SYS_ADMIN)) diff --git a/fs/smb/client/cifsacl.c b/fs/smb/client/cifsacl.c index c5e47a835f99..d1a92bb4d9a5 100644 --- a/fs/smb/client/cifsacl.c +++ b/fs/smb/client/cifsacl.c @@ -1904,7 +1904,7 @@ id_mode_to_cifs_acl_exit: return rc; } -struct posix_acl *cifs_get_acl(struct mnt_idmap *idmap, +struct posix_acl *cifs_get_acl(const struct mnt_idmap *idmap, struct dentry *dentry, int type) { #if defined(CONFIG_CIFS_ALLOW_INSECURE_LEGACY) && defined(CONFIG_CIFS_POSIX) @@ -1968,7 +1968,7 @@ out: #endif } -int cifs_set_acl(struct mnt_idmap *idmap, struct dentry *dentry, +int cifs_set_acl(const struct mnt_idmap *idmap, struct dentry *dentry, struct posix_acl *acl, int type) { #if defined(CONFIG_CIFS_ALLOW_INSECURE_LEGACY) && defined(CONFIG_CIFS_POSIX) diff --git a/fs/smb/client/cifsfs.c b/fs/smb/client/cifsfs.c index 7ecd70efdfea..b1ecbcfb154e 100644 --- a/fs/smb/client/cifsfs.c +++ b/fs/smb/client/cifsfs.c @@ -402,7 +402,7 @@ out_unlock: return rc; } -static int cifs_permission(struct mnt_idmap *idmap, +static int cifs_permission(const struct mnt_idmap *idmap, struct inode *inode, int mask) { unsigned int sbflags = cifs_sb_flags(CIFS_SB(inode)); diff --git a/fs/smb/client/cifsfs.h b/fs/smb/client/cifsfs.h index 0c85daa8386e..51692e14c4dd 100644 --- a/fs/smb/client/cifsfs.h +++ b/fs/smb/client/cifsfs.h @@ -53,23 +53,23 @@ void cifs_sb_deactive(struct super_block *sb); /* Functions related to inodes */ extern const struct inode_operations cifs_dir_inode_ops; struct inode *cifs_root_iget(struct super_block *sb); -int cifs_create(struct mnt_idmap *idmap, struct inode *dir, +int cifs_create(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *direntry, umode_t mode); int cifs_atomic_open(struct inode *dir, struct dentry *direntry, struct file *file, unsigned int oflags, umode_t mode); -int cifs_tmpfile(struct mnt_idmap *idmap, struct inode *dir, +int cifs_tmpfile(const struct mnt_idmap *idmap, struct inode *dir, struct file *file, umode_t mode); struct dentry *cifs_lookup(struct inode *parent_dir_inode, struct dentry *direntry, unsigned int flags); int cifs_unlink(struct inode *dir, struct dentry *dentry); int cifs_hardlink(struct dentry *old_file, struct inode *inode, struct dentry *direntry); -int cifs_mknod(struct mnt_idmap *idmap, struct inode *inode, +int cifs_mknod(const struct mnt_idmap *idmap, struct inode *inode, struct dentry *direntry, umode_t mode, dev_t device_number); -struct dentry *cifs_mkdir(struct mnt_idmap *idmap, struct inode *inode, +struct dentry *cifs_mkdir(const struct mnt_idmap *idmap, struct inode *inode, struct dentry *direntry, umode_t mode); int cifs_rmdir(struct inode *inode, struct dentry *direntry); -int cifs_rename2(struct mnt_idmap *idmap, struct inode *source_dir, +int cifs_rename2(const struct mnt_idmap *idmap, struct inode *source_dir, struct dentry *source_dentry, struct inode *target_dir, struct dentry *target_dentry, unsigned int flags); int cifs_revalidate_file_attr(struct file *filp); @@ -78,9 +78,9 @@ int cifs_revalidate_file(struct file *filp); int cifs_revalidate_dentry(struct dentry *dentry); int cifs_revalidate_mapping(struct inode *inode); int cifs_zap_mapping(struct inode *inode); -int cifs_getattr(struct mnt_idmap *idmap, const struct path *path, +int cifs_getattr(const struct mnt_idmap *idmap, const struct path *path, struct kstat *stat, u32 request_mask, unsigned int flags); -int cifs_setattr(struct mnt_idmap *idmap, struct dentry *direntry, +int cifs_setattr(const struct mnt_idmap *idmap, struct dentry *direntry, struct iattr *attrs); int cifs_fiemap(struct inode *inode, struct fiemap_extent_info *fei, u64 start, u64 len); @@ -129,7 +129,7 @@ struct vfsmount *cifs_d_automount(struct path *path); /* Functions related to symlinks */ const char *cifs_get_link(struct dentry *dentry, struct inode *inode, struct delayed_call *done); -int cifs_symlink(struct mnt_idmap *idmap, struct inode *inode, +int cifs_symlink(const struct mnt_idmap *idmap, struct inode *inode, struct dentry *direntry, const char *symname); #ifdef CONFIG_CIFS_XATTR diff --git a/fs/smb/client/cifsproto.h b/fs/smb/client/cifsproto.h index 00168839c123..e6beff8aafe0 100644 --- a/fs/smb/client/cifsproto.h +++ b/fs/smb/client/cifsproto.h @@ -212,9 +212,9 @@ struct smb_ntsd *get_cifs_acl(struct cifs_sb_info *cifs_sb, struct smb_ntsd *get_cifs_acl_by_fid(struct cifs_sb_info *cifs_sb, const struct cifs_fid *cifsfid, u32 *pacllen, u32 info); -struct posix_acl *cifs_get_acl(struct mnt_idmap *idmap, struct dentry *dentry, +struct posix_acl *cifs_get_acl(const struct mnt_idmap *idmap, struct dentry *dentry, int type); -int cifs_set_acl(struct mnt_idmap *idmap, struct dentry *dentry, +int cifs_set_acl(const struct mnt_idmap *idmap, struct dentry *dentry, struct posix_acl *acl, int type); int set_cifs_acl(struct smb_ntsd *pnntsd, __u32 acllen, struct inode *inode, const char *path, int aclflag); diff --git a/fs/smb/client/dir.c b/fs/smb/client/dir.c index 1a56fa4d0e89..a2cf3d35ca5d 100644 --- a/fs/smb/client/dir.c +++ b/fs/smb/client/dir.c @@ -664,7 +664,7 @@ out_free_xid: * The initial dentry state is hashed-negative. On success, dentry will become * hashed-positive by calling d_instantiate(). */ -int cifs_create(struct mnt_idmap *idmap, struct inode *dir, +int cifs_create(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *direntry, umode_t mode) { struct cifs_sb_info *cifs_sb = CIFS_SB(dir); @@ -721,7 +721,7 @@ out_free_xid: return rc; } -int cifs_mknod(struct mnt_idmap *idmap, struct inode *inode, +int cifs_mknod(const struct mnt_idmap *idmap, struct inode *inode, struct dentry *direntry, umode_t mode, dev_t device_number) { int rc = -EPERM; @@ -1084,7 +1084,7 @@ static int set_tmpfile_attr(const unsigned int xid, unsigned int oflags, * The initial dentry state is unhashed-negative. On success, dentry will * become unhashed-positive by calling d_instantiate(). */ -int cifs_tmpfile(struct mnt_idmap *idmap, struct inode *dir, +int cifs_tmpfile(const struct mnt_idmap *idmap, struct inode *dir, struct file *file, umode_t mode) { struct dentry *dentry = file->f_path.dentry; diff --git a/fs/smb/client/inode.c b/fs/smb/client/inode.c index 1fe0ef0a95db..4d7a87c7b73a 100644 --- a/fs/smb/client/inode.c +++ b/fs/smb/client/inode.c @@ -2281,7 +2281,7 @@ posix_mkdir_get_info: } #endif /* CONFIG_CIFS_ALLOW_INSECURE_LEGACY */ -struct dentry *cifs_mkdir(struct mnt_idmap *idmap, struct inode *inode, +struct dentry *cifs_mkdir(const struct mnt_idmap *idmap, struct inode *inode, struct dentry *direntry, umode_t mode) { int rc = 0; @@ -2531,7 +2531,7 @@ do_rename_exit: } int -cifs_rename2(struct mnt_idmap *idmap, struct inode *source_dir, +cifs_rename2(const struct mnt_idmap *idmap, struct inode *source_dir, struct dentry *source_dentry, struct inode *target_dir, struct dentry *target_dentry, unsigned int flags) { @@ -2937,7 +2937,7 @@ int cifs_revalidate_dentry(struct dentry *dentry) return cifs_revalidate_mapping(inode); } -int cifs_getattr(struct mnt_idmap *idmap, const struct path *path, +int cifs_getattr(const struct mnt_idmap *idmap, const struct path *path, struct kstat *stat, u32 request_mask, unsigned int flags) { struct cifs_sb_info *cifs_sb = CIFS_SB(path->dentry); @@ -3554,7 +3554,7 @@ cifs_setattr_exit: } int -cifs_setattr(struct mnt_idmap *idmap, struct dentry *direntry, +cifs_setattr(const struct mnt_idmap *idmap, struct dentry *direntry, struct iattr *attrs) { struct cifs_sb_info *cifs_sb = CIFS_SB(direntry->d_sb); diff --git a/fs/smb/client/link.c b/fs/smb/client/link.c index 8d5d6aca742a..76df31abeaea 100644 --- a/fs/smb/client/link.c +++ b/fs/smb/client/link.c @@ -533,7 +533,7 @@ cifs_hl_exit: } int -cifs_symlink(struct mnt_idmap *idmap, struct inode *inode, +cifs_symlink(const struct mnt_idmap *idmap, struct inode *inode, struct dentry *direntry, const char *symname) { struct cifs_sb_info *cifs_sb = CIFS_SB(inode); diff --git a/fs/smb/client/transport.c b/fs/smb/client/transport.c index 6e21b5f8754a..93ff6a4dbb35 100644 --- a/fs/smb/client/transport.c +++ b/fs/smb/client/transport.c @@ -22,7 +22,6 @@ #include <linux/mempool.h> #include <linux/sched/signal.h> #include <linux/task_io_accounting_ops.h> -#include <linux/task_work.h> #include "cifsglob.h" #include "cifsproto.h" #include "cifs_debug.h" @@ -171,15 +170,11 @@ smb_send_kvec(struct TCP_Server_Info *server, struct msghdr *smb_msg, * after the retries we will kill the socket and * reconnect which may clear the network problem. * - * Even if regular signals are masked, EINTR might be - * propagated from sk_stream_wait_memory() to here when - * TIF_NOTIFY_SIGNAL is used for task work. For example, - * certain io_uring completions will use that. Treat - * having EINTR with pending task work the same as EAGAIN - * to avoid unnecessary reconnects. + * Task work must not abort the send, see signal_pending(). */ - rc = sock_sendmsg(ssocket, smb_msg); - if (rc == -EAGAIN || unlikely(rc == -EINTR && task_work_pending(current))) { + scoped_guard(no_notify_signal) + rc = sock_sendmsg(ssocket, smb_msg); + if (rc == -EAGAIN) { retries++; if (retries >= 14 || (!server->noblocksnd && (retries > 2))) { diff --git a/fs/smb/client/xattr.c b/fs/smb/client/xattr.c index 5091f6c0d7fe..f6c9343016f7 100644 --- a/fs/smb/client/xattr.c +++ b/fs/smb/client/xattr.c @@ -91,7 +91,7 @@ static int cifs_creation_time_set(unsigned int xid, struct cifs_tcon *pTcon, } static int cifs_xattr_set(const struct xattr_handler *handler, - struct mnt_idmap *idmap, + const struct mnt_idmap *idmap, struct dentry *dentry, struct inode *inode, const char *name, const void *value, size_t size, int flags) diff --git a/fs/smb/common/compress/compress.c b/fs/smb/common/compress/compress.c index a4123c8f1c0a..4e5164b6dc33 100644 --- a/fs/smb/common/compress/compress.c +++ b/fs/smb/common/compress/compress.c @@ -130,12 +130,18 @@ static int smb_decompress_chained(__le16 alg, bool allow_chained, len = le32_to_cpu(payload->Length); /* - * CHAINED marks only the first payload. Requiring NONE on every - * later payload rejects ambiguous or independently chained data. + * Conforming chains must set CHAINED on the first payload. + * Windows 11 clients leave uninitialized residual bits in + * Flags on subsequent payload headers. Only reject trailing + * payloads that attempt to initiate an invalid nested chain. */ - if ((first && flags != cpu_to_le16(SMB2_COMPRESSION_FLAG_CHAINED)) || - (!first && flags != cpu_to_le16(SMB2_COMPRESSION_FLAG_NONE))) - return -EINVAL; + if (first) { + if (flags != cpu_to_le16(SMB2_COMPRESSION_FLAG_CHAINED)) + return -EINVAL; + } else { + if (flags == cpu_to_le16(SMB2_COMPRESSION_FLAG_CHAINED)) + return -EINVAL; + } src += SMB2_COMPRESSION_PAYLOAD_BASE_LEN; remaining -= SMB2_COMPRESSION_PAYLOAD_BASE_LEN; diff --git a/fs/smb/common/fscc.h b/fs/smb/common/fscc.h index e46d3379b779..00f09e1a3afe 100644 --- a/fs/smb/common/fscc.h +++ b/fs/smb/common/fscc.h @@ -399,7 +399,7 @@ static_assert(offsetof(struct smb2_file_rename_info, FileName) == sizeof(struct #define FS_OBJECT_ID_INFORMATION 8 /* Query, Set */ #define FS_DRIVER_PATH_INFORMATION 9 /* Query */ #define FS_SECTOR_SIZE_INFORMATION 11 /* SMB3 or later. Query */ -/* See POSIX Extensions to MS-FSCC 2.3.1.1 */ +/* See POSIX-FSCC 2.3 */ #define FS_POSIX_INFORMATION 100 /* SMB3.1.1 POSIX. Query */ /* See MS-FSCC 2.5.1 */ @@ -575,8 +575,7 @@ struct file_notify_information { } __packed; /* - * See POSIX Extensions to MS-FSCC 2.3.2.1 - * Link: https://gitlab.com/samba-team/smb3-posix-spec/-/blob/master/fscc_posix_extensions.md + * See POSIX-FSCC 2.3.1 */ typedef struct { /* For undefined recommended transfer size return -1 in that field */ diff --git a/fs/smb/common/smbglob.h b/fs/smb/common/smbglob.h index d9c7e6e7af29..fdf840888062 100644 --- a/fs/smb/common/smbglob.h +++ b/fs/smb/common/smbglob.h @@ -40,6 +40,7 @@ struct smb_version_values { size_t create_disk_id_size; size_t create_posix_size; size_t create_aapl_size; + size_t create_rsp_size; }; static inline unsigned int get_rfc1002_len(void *buf) diff --git a/fs/smb/server/Kconfig b/fs/smb/server/Kconfig index 221ec9717a83..b7665e0e4942 100644 --- a/fs/smb/server/Kconfig +++ b/fs/smb/server/Kconfig @@ -71,3 +71,5 @@ config SMB_SERVER_KERBEROS5 bool "Support for Kerberos 5" depends on SMB_SERVER default y + +source "fs/smb/server/tests/Kconfig" diff --git a/fs/smb/server/Makefile b/fs/smb/server/Makefile index a3e9306055e8..9bc87695a53c 100644 --- a/fs/smb/server/Makefile +++ b/fs/smb/server/Makefile @@ -19,3 +19,4 @@ $(obj)/ksmbd_spnego_negtokentarg.asn1.o: $(obj)/ksmbd_spnego_negtokentarg.asn1.c ksmbd-$(CONFIG_SMB_SERVER_SMBDIRECT) += transport_rdma.o ksmbd-$(CONFIG_PROC_FS) += proc.o +obj-$(CONFIG_SMB_SERVER_KUNIT_TESTS) += tests/ diff --git a/fs/smb/server/compress.c b/fs/smb/server/compress.c index 5162fb84c755..3fcb3e44302e 100644 --- a/fs/smb/server/compress.c +++ b/fs/smb/server/compress.c @@ -65,6 +65,7 @@ static int __ksmbd_decompress_request(struct ksmbd_conn *conn, return -ENOMEM; *(__be32 *)out = cpu_to_be32(out_size); + out[out_size + 4] = 0; rc = smb_compression_decompress(conn->compress_algorithm, conn->compress_chained, conn->compress_pattern, diff --git a/fs/smb/server/connection.c b/fs/smb/server/connection.c index d211861ff86f..d1cc725eeb06 100644 --- a/fs/smb/server/connection.c +++ b/fs/smb/server/connection.c @@ -114,7 +114,7 @@ static int proc_show_clients(struct seq_file *m, void *v) sessions++; rcu_read_unlock(); #if IS_ENABLED(CONFIG_IPV6) - if (!conn->inet_addr) + if (conn->is_ipv6) seq_printf(m, "client:\t%pI6c\n", &conn->inet6_addr); else #endif @@ -752,6 +752,7 @@ recheck: if (!conn->request_buf) break; + conn->request_buf[pdu_size + 4] = 0; memcpy(conn->request_buf, hdr_buf, sizeof(hdr_buf)); /* diff --git a/fs/smb/server/connection.h b/fs/smb/server/connection.h index 371f17b4f02a..2026a2afa5e4 100644 --- a/fs/smb/server/connection.h +++ b/fs/smb/server/connection.h @@ -68,6 +68,9 @@ struct ksmbd_conn { u8 inet6_addr[16]; #endif }; +#if IS_ENABLED(CONFIG_IPV6) + bool is_ipv6; +#endif unsigned int inet_hash; char *request_buf; struct ksmbd_transport *transport; diff --git a/fs/smb/server/ksmbd_work.c b/fs/smb/server/ksmbd_work.c index d307aefe0aec..52ea1a674a67 100644 --- a/fs/smb/server/ksmbd_work.c +++ b/fs/smb/server/ksmbd_work.c @@ -57,7 +57,6 @@ struct ksmbd_work *ksmbd_alloc_work_struct(void) INIT_LIST_HEAD(&work->request_entry); INIT_LIST_HEAD(&work->async_request_entry); INIT_LIST_HEAD(&work->fp_entry); - INIT_LIST_HEAD(&work->notify_entry); INIT_LIST_HEAD(&work->aux_read_list); work->iov_alloc_cnt = ARRAY_SIZE(work->iov_inline); work->iov = work->iov_inline; @@ -87,8 +86,6 @@ void ksmbd_free_work_struct(struct ksmbd_work *work) if (work->async_id) ksmbd_release_id(&work->conn->async_ida, work->async_id); - if (work->owns_conn_ref) - ksmbd_conn_put(work->conn); ksmbd_fd_put(work, work->request_open); kmem_cache_free(work_cache, work); } diff --git a/fs/smb/server/ksmbd_work.h b/fs/smb/server/ksmbd_work.h index 0844aa929f55..3a14e4d69aa1 100644 --- a/fs/smb/server/ksmbd_work.h +++ b/fs/smb/server/ksmbd_work.h @@ -91,8 +91,6 @@ struct ksmbd_work { bool compress_response:1; /* Is this SYNC or ASYNC ksmbd_work */ bool asynchronous:1; - /* Work owns a reference to @conn. */ - bool owns_conn_ref:1; bool need_invalidate_rkey:1; bool request_open_chseq_tracked:1; bool session_setup_reauth:1; @@ -115,9 +113,8 @@ struct ksmbd_work { struct list_head request_entry; /* List head at conn->async_requests */ struct list_head async_request_entry; + /* List head at ksmbd_file->blocked_works */ struct list_head fp_entry; - /* List head at ksmbd_file->notify_pendings */ - struct list_head notify_entry; }; /** diff --git a/fs/smb/server/mgmt/user_session.c b/fs/smb/server/mgmt/user_session.c index 44dc3f800cd4..1321cd6c84e7 100644 --- a/fs/smb/server/mgmt/user_session.c +++ b/fs/smb/server/mgmt/user_session.c @@ -130,7 +130,7 @@ static int show_proc_session(struct seq_file *m, void *v) const char *name; #if IS_ENABLED(CONFIG_IPV6) - if (chan->conn->inet_addr) + if (!chan->conn->is_ipv6) seq_printf(m, "client:\t%pI4\n", &chan->conn->inet_addr); else @@ -231,7 +231,7 @@ static int show_proc_sessions(struct seq_file *m, void *v) ksmbd_user_session_get(session); #if IS_ENABLED(CONFIG_IPV6) - if (!chan->conn->inet_addr) + if (chan->conn->is_ipv6) seq_printf(m, "client:\t%pI6c\n", &chan->conn->inet6_addr); else #endif diff --git a/fs/smb/server/ndr.c b/fs/smb/server/ndr.c index 58d71560f626..7e546c22e284 100644 --- a/fs/smb/server/ndr.c +++ b/fs/smb/server/ndr.c @@ -338,7 +338,7 @@ static int ndr_encode_posix_acl_entry(struct ndr *n, struct xattr_smb_acl *acl) } int ndr_encode_posix_acl(struct ndr *n, - struct mnt_idmap *idmap, + const struct mnt_idmap *idmap, struct inode *inode, struct xattr_smb_acl *acl, struct xattr_smb_acl *def_acl) diff --git a/fs/smb/server/ndr.h b/fs/smb/server/ndr.h index f3c108c8cf4d..646568c42e4d 100644 --- a/fs/smb/server/ndr.h +++ b/fs/smb/server/ndr.h @@ -14,7 +14,7 @@ struct ndr { int ndr_encode_dos_attr(struct ndr *n, struct xattr_dos_attrib *da); int ndr_decode_dos_attr(struct ndr *n, struct xattr_dos_attrib *da); -int ndr_encode_posix_acl(struct ndr *n, struct mnt_idmap *idmap, +int ndr_encode_posix_acl(struct ndr *n, const struct mnt_idmap *idmap, struct inode *inode, struct xattr_smb_acl *acl, struct xattr_smb_acl *def_acl); int ndr_encode_v4_ntacl(struct ndr *n, struct xattr_ntacl *acl); diff --git a/fs/smb/server/oplock.c b/fs/smb/server/oplock.c index 1b8c3482d1e4..d0f18ebf471b 100644 --- a/fs/smb/server/oplock.c +++ b/fs/smb/server/oplock.c @@ -2270,7 +2270,7 @@ void create_posix_rsp_buf(char *cc, struct ksmbd_file *fp) { struct create_posix_rsp *buf; struct inode *inode = file_inode(fp->filp); - struct mnt_idmap *idmap = file_mnt_idmap(fp->filp); + const struct mnt_idmap *idmap = file_mnt_idmap(fp->filp); vfsuid_t vfsuid = i_uid_into_vfsuid(idmap, inode); vfsgid_t vfsgid = i_gid_into_vfsgid(idmap, inode); diff --git a/fs/smb/server/smb2ops.c b/fs/smb/server/smb2ops.c index 4578291fb172..747276b171bd 100644 --- a/fs/smb/server/smb2ops.c +++ b/fs/smb/server/smb2ops.c @@ -13,6 +13,35 @@ #include "server.h" #include "stats.h" +/* + * work->response_sz includes the RFC1002 length field while smb_get_msg() + * skips over it. Durable v1 and v2 response contexts are mutually exclusive, + * and POSIX CREATE contexts are only negotiated for SMB3.1.1. + */ +#define SMB2_CREATE_RSP_SIZE(lease_size, durable_size) \ + (sizeof(__be32) + offsetof(struct smb2_create_rsp, Buffer) + \ + (lease_size) + (durable_size) + \ + sizeof(struct create_mxac_rsp) + \ + sizeof(struct create_disk_id_rsp) + \ + AAPL_RSP_MAX_SIZE) + +#define SMB21_CREATE_RSP_SIZE \ + SMB2_CREATE_RSP_SIZE(sizeof(struct create_lease), \ + sizeof(struct create_durable_rsp)) + +#define SMB3_CREATE_DURABLE_RSP_SIZE \ + ((sizeof(struct create_durable_rsp) > \ + sizeof(struct create_durable_rsp_v2)) ? \ + sizeof(struct create_durable_rsp) : \ + sizeof(struct create_durable_rsp_v2)) + +#define SMB3_CREATE_RSP_SIZE \ + SMB2_CREATE_RSP_SIZE(sizeof(struct create_lease_v2), \ + SMB3_CREATE_DURABLE_RSP_SIZE) + +#define SMB311_CREATE_RSP_SIZE \ + (SMB3_CREATE_RSP_SIZE + sizeof(struct create_posix_rsp)) + static struct smb_version_values smb21_server_values = { .version_string = SMB21_VERSION_STRING, .protocol_id = SMB21_PROT_ID, @@ -38,6 +67,7 @@ static struct smb_version_values smb21_server_values = { .create_disk_id_size = sizeof(struct create_disk_id_rsp), .create_posix_size = sizeof(struct create_posix_rsp), .create_aapl_size = AAPL_RSP_MAX_SIZE, + .create_rsp_size = SMB21_CREATE_RSP_SIZE, }; static struct smb_version_values smb30_server_values = { @@ -66,6 +96,7 @@ static struct smb_version_values smb30_server_values = { .create_disk_id_size = sizeof(struct create_disk_id_rsp), .create_posix_size = sizeof(struct create_posix_rsp), .create_aapl_size = AAPL_RSP_MAX_SIZE, + .create_rsp_size = SMB3_CREATE_RSP_SIZE, }; static struct smb_version_values smb302_server_values = { @@ -94,6 +125,7 @@ static struct smb_version_values smb302_server_values = { .create_disk_id_size = sizeof(struct create_disk_id_rsp), .create_posix_size = sizeof(struct create_posix_rsp), .create_aapl_size = AAPL_RSP_MAX_SIZE, + .create_rsp_size = SMB3_CREATE_RSP_SIZE, }; static struct smb_version_values smb311_server_values = { @@ -122,6 +154,7 @@ static struct smb_version_values smb311_server_values = { .create_disk_id_size = sizeof(struct create_disk_id_rsp), .create_posix_size = sizeof(struct create_posix_rsp), .create_aapl_size = AAPL_RSP_MAX_SIZE, + .create_rsp_size = SMB311_CREATE_RSP_SIZE, }; static struct smb_version_ops smb2_0_server_ops = { diff --git a/fs/smb/server/smb2pdu.c b/fs/smb/server/smb2pdu.c index 6b8809f67b92..45dce9c30b6b 100644 --- a/fs/smb/server/smb2pdu.c +++ b/fs/smb/server/smb2pdu.c @@ -16,6 +16,7 @@ #include <linux/mount.h> #include <linux/filelock.h> #include <linux/fileattr.h> +#include <linux/math.h> #include <linux/timekeeping.h> #include <linux/unaligned.h> @@ -58,10 +59,6 @@ static void __wbuf(struct ksmbd_work *work, void **req, void **rsp) } } -static struct ksmbd_work *smb2_notify_cancel_claim(void **argv); -static void smb2_notify_cancel_fn(void **argv); -static void smb2_complete_notify_cancel(struct ksmbd_work *in_work); - #define WORK_BUFFERS(w, rq, rs) __wbuf((w), (void **)&(rq), (void **)&(rs)) #define SMB2_CREATE_FILE_ATTRIBUTE_MASK \ @@ -858,14 +855,18 @@ static void smb2_update_lock_sequence(struct ksmbd_work *work, int smb2_allocate_rsp_buf(struct ksmbd_work *work) { struct smb2_hdr *hdr = smb_get_msg(work->request_buf); + struct smb_version_values *vals = work->conn->vals; size_t small_sz = MAX_CIFS_SMALL_BUFFER_SIZE; - size_t large_sz = small_sz + work->conn->vals->max_trans_size; + size_t large_sz = small_sz + vals->max_trans_size; size_t sz = small_sz; int cmd = le16_to_cpu(hdr->Command); if (cmd == SMB2_IOCTL_HE || cmd == SMB2_QUERY_DIRECTORY_HE) sz = large_sz; + if (cmd == SMB2_CREATE_HE) + sz = max_t(size_t, sz, vals->create_rsp_size); + if (cmd == SMB2_QUERY_INFO_HE) { struct smb2_query_info_req *req; @@ -1246,6 +1247,7 @@ void smb2_send_interim_resp(struct ksmbd_work *work, __le32 status) { struct smb2_hdr *rsp_hdr; struct ksmbd_work *in_work = ksmbd_alloc_work_struct(); + u16 command; if (!in_work) return; @@ -1269,6 +1271,23 @@ void smb2_send_interim_resp(struct ksmbd_work *work, __le32 status) smb2_set_err_rsp(in_work); rsp_hdr->Status = status; + /* + * Async interim responses are unsigned, but final responses must + * follow the normal signing rules. The synthetic work has no + * request buffer, so use the original work for request signing + * checks and the response header for SMB3 command selection. + */ + command = work->conn->ops->get_cmd_val(work); + if (status != STATUS_PENDING && !work->encrypted && work->sess && + work->conn->ops->set_sign_rsp && + (work->sess->sign || + (work->conn->ops->is_sign_req && + work->conn->ops->is_sign_req(work, command)))) { + in_work->sess = work->sess; + work->conn->ops->set_sign_rsp(in_work); + in_work->sess = NULL; + } + if (smb2_send_interim_work(in_work, work, true)) ksmbd_debug(SMB, "failed to send interim response\n"); ksmbd_free_work_struct(in_work); @@ -3278,7 +3297,7 @@ static bool smb2_is_private_ea(const char *name, size_t name_len) static int smb2_set_ea(struct smb2_ea_info *eabuf, unsigned int buf_len, const struct path *path, bool get_write) { - struct mnt_idmap *idmap = mnt_idmap(path->mnt); + const struct mnt_idmap *idmap = mnt_idmap(path->mnt); char *attr_name = NULL, *value; int rc = 0; unsigned int next = 0; @@ -3379,7 +3398,7 @@ static noinline int smb2_set_stream_name_xattr(const struct path *path, struct ksmbd_file *fp, char *stream_name, int s_type) { - struct mnt_idmap *idmap = mnt_idmap(path->mnt); + const struct mnt_idmap *idmap = mnt_idmap(path->mnt); size_t xattr_stream_size; char *xattr_stream_name; int rc; @@ -3419,8 +3438,9 @@ static noinline int smb2_set_stream_name_xattr(const struct path *path, * AAPL there too. */ static const u8 afpinfo_empty[60] = { - 0x00, 0x05, 0x16, 0x07, /* magic 0x00051607 BE */ - 0x00, 0x02, 0x00, 0x00, /* version 0x00020000 BE */ + 'A', 'F', 'P', 0x00, /* signature */ + 0x00, 0x00, 0x01, 0x00, /* version */ + [15] = 0x80, /* backup time */ }; rc = ksmbd_vfs_setxattr(idmap, path, xattr_stream_name, (void *)afpinfo_empty, @@ -3455,7 +3475,7 @@ static loff_t ksmbd_stream_eof(struct ksmbd_file *fp) static int smb2_remove_smb_xattrs(const struct path *path) { - struct mnt_idmap *idmap = mnt_idmap(path->mnt); + const struct mnt_idmap *idmap = mnt_idmap(path->mnt); char *name, *xattr_list = NULL; ssize_t xattr_list_len; int err = 0; @@ -3647,10 +3667,11 @@ static int smb2_create_sd_buffer(struct ksmbd_work *work, le32_to_cpu(sd_buf->ccontext.DataLength), true, false); } -static void ksmbd_acls_fattr(struct smb_fattr *fattr, - struct mnt_idmap *idmap, - struct inode *inode) +static int ksmbd_acls_fattr(struct smb_fattr *fattr, + const struct mnt_idmap *idmap, + struct inode *inode) { + struct posix_acl *acl; vfsuid_t vfsuid = i_uid_into_vfsuid(idmap, inode); vfsgid_t vfsgid = i_gid_into_vfsgid(idmap, inode); @@ -3661,10 +3682,28 @@ static void ksmbd_acls_fattr(struct smb_fattr *fattr, fattr->cf_dacls = NULL; if (IS_ENABLED(CONFIG_FS_POSIX_ACL)) { - fattr->cf_acls = get_inode_acl(inode, ACL_TYPE_ACCESS); - if (S_ISDIR(inode->i_mode)) - fattr->cf_dacls = get_inode_acl(inode, ACL_TYPE_DEFAULT); + acl = get_inode_acl(inode, ACL_TYPE_ACCESS); + if (IS_ERR(acl)) { + if (acl != ERR_PTR(-EOPNOTSUPP)) + return PTR_ERR(acl); + acl = NULL; + } + fattr->cf_acls = acl; + + if (S_ISDIR(inode->i_mode)) { + acl = get_inode_acl(inode, ACL_TYPE_DEFAULT); + if (IS_ERR(acl)) { + if (acl != ERR_PTR(-EOPNOTSUPP)) { + posix_acl_release(fattr->cf_acls); + return PTR_ERR(acl); + } + acl = NULL; + } + fattr->cf_dacls = acl; + } } + + return 0; } enum { @@ -4132,7 +4171,7 @@ int smb2_open(struct ksmbd_work *work) struct ksmbd_share_config *share = tcon->share_conf; struct ksmbd_file *fp = NULL; struct file *filp = NULL; - struct mnt_idmap *idmap = NULL; + const struct mnt_idmap *idmap = NULL; struct kstat stat; struct create_context *context; struct lease_ctx_info *lc = NULL; @@ -4791,7 +4830,10 @@ int smb2_open(struct ksmbd_work *work) int pntsd_size; size_t scratch_len; - ksmbd_acls_fattr(&fattr, idmap, inode); + rc = ksmbd_acls_fattr(&fattr, idmap, inode); + if (rc) + goto err_out; + scratch_len = smb_acl_sec_desc_scratch_len(&fattr, NULL, 0, OWNER_SECINFO | GROUP_SECINFO | @@ -5797,7 +5839,7 @@ struct smb2_query_dir_private { static int process_query_dir_entries(struct smb2_query_dir_private *priv) { - struct mnt_idmap *idmap = file_mnt_idmap(priv->dir_fp->filp); + const struct mnt_idmap *idmap = file_mnt_idmap(priv->dir_fp->filp); struct kstat kstat; struct ksmbd_kstat ksmbd_kstat; int rc; @@ -6390,7 +6432,7 @@ static int smb2_get_ea(struct ksmbd_work *work, struct ksmbd_file *fp, ssize_t buf_free_len, alignment_bytes, next_offset, rsp_data_cnt = 0; struct smb2_ea_info_req *ea_req = NULL; const struct path *path; - struct mnt_idmap *idmap = file_mnt_idmap(fp->filp); + const struct mnt_idmap *idmap = file_mnt_idmap(fp->filp); if (!(fp->daccess & FILE_READ_EA_LE)) { pr_err("Not permitted to read ext attr : 0x%x\n", @@ -6741,12 +6783,12 @@ static void get_file_alternate_info(struct ksmbd_work *work, void *rsp_org) { struct ksmbd_conn *conn = work->conn; - struct smb2_file_alt_name_info *file_info; + struct smb2_file_name_info *file_info; struct dentry *dentry = fp->filp->f_path.dentry; int conv_len; spin_lock(&dentry->d_lock); - file_info = (struct smb2_file_alt_name_info *)rsp->Buffer; + file_info = (struct smb2_file_name_info *)rsp->Buffer; conv_len = ksmbd_extract_shortname(conn, dentry->d_name.name, file_info->FileName); @@ -6793,7 +6835,7 @@ static int get_file_normalized_name_info(struct ksmbd_work *work, struct smb2_query_info_rsp *rsp, struct ksmbd_file *fp) { - struct smb2_file_alt_name_info *file_info; + struct smb2_file_name_info *file_info; char *filename, *normalized, *stream_name; int buf_free_len, conv_len, filename_len; @@ -6827,7 +6869,7 @@ static int get_file_normalized_name_info(struct ksmbd_work *work, return -EINVAL; } - file_info = (struct smb2_file_alt_name_info *)rsp->Buffer; + file_info = (struct smb2_file_name_info *)rsp->Buffer; conv_len = smbConvertToUTF16((__le16 *)file_info->FileName, normalized, filename_len, work->conn->local_nls, 0); @@ -6912,17 +6954,8 @@ static int get_file_stream_info(struct ksmbd_work *work, streamlen *= 2; kfree(stream_buf); file_info->StreamNameLength = cpu_to_le32(streamlen); - /* - * stream_name_len is the byte length of the xattr's *name*, - * not its value -- same class of bug ksmbd_stream_eof() - * (smb2pdu.c) already fixes for EndOfFile/AllocationSize on - * a stream handle; this enumeration path needs the same - * real xattr value length, not the name length reused as a - * size. - */ - slen = ksmbd_vfs_casexattr_len(file_mnt_idmap(fp->filp), - path->dentry, stream_name, - strlen(stream_name) + 1); + slen = ksmbd_vfs_xattr_len(file_mnt_idmap(fp->filp), + path->dentry, stream_name); ssize = slen < 0 ? 0 : (loff_t)slen; file_info->StreamSize = cpu_to_le64(ssize); file_info->StreamAllocationSize = cpu_to_le64(ssize); @@ -7108,7 +7141,7 @@ static int find_file_posix_info(struct smb2_query_info_rsp *rsp, { struct smb311_posix_qinfo *file_info; struct inode *inode = file_inode(fp->filp); - struct mnt_idmap *idmap = file_mnt_idmap(fp->filp); + const struct mnt_idmap *idmap = file_mnt_idmap(fp->filp); vfsuid_t vfsuid = i_uid_into_vfsuid(idmap, inode); vfsgid_t vfsgid = i_gid_into_vfsgid(idmap, inode); struct kstat stat; @@ -7190,7 +7223,7 @@ static int smb2_get_info_file(struct ksmbd_work *work, struct ksmbd_file *fp; int fileinfoclass = 0; int rc = 0; - unsigned int fixed_len; + unsigned int fixed_len, req_output_len; unsigned int id = KSMBD_NO_FID, pid = KSMBD_NO_FID; if (test_share_config_flag(work->tcon->share_conf, @@ -7297,24 +7330,24 @@ static int smb2_get_info_file(struct ksmbd_work *work, rc = -EOPNOTSUPP; } if (!rc) { + req_output_len = le32_to_cpu(req->OutputBufferLength); fixed_len = le32_to_cpu(rsp->OutputBufferLength); switch (fileinfoclass) { case FILE_ALL_INFORMATION: fixed_len = FILE_ALL_INFORMATION_SIZE; break; case FILE_ALTERNATE_NAME_INFORMATION: - fixed_len = FILE_ALTERNATE_NAME_INFORMATION_SIZE; + fixed_len = FILE_NAME_INFORMATION_SIZE; break; case FILE_NORMALIZED_NAME_INFORMATION: - fixed_len = FILE_NORMALIZED_NAME_INFORMATION_SIZE; + fixed_len = FILE_NAME_INFORMATION_SIZE; + req_output_len = round_down(req_output_len, 2); break; case FILE_STREAM_INFORMATION: fixed_len = FILE_STREAM_INFORMATION_SIZE; break; } - rc = buffer_check_err(le32_to_cpu(req->OutputBufferLength), - fixed_len, - rsp); + rc = buffer_check_err(req_output_len, fixed_len, rsp); } ksmbd_fd_put(work, fp); @@ -7600,7 +7633,7 @@ static int smb2_get_info_sec(struct ksmbd_work *work, struct smb2_query_info_rsp *rsp) { struct ksmbd_file *fp; - struct mnt_idmap *idmap; + const struct mnt_idmap *idmap; struct smb_ntsd *pntsd = NULL, *ppntsd = NULL; struct smb_fattr fattr = {{0}}; struct inode *inode; @@ -7652,7 +7685,11 @@ static int smb2_get_info_sec(struct ksmbd_work *work, idmap = file_mnt_idmap(fp->filp); inode = file_inode(fp->filp); - ksmbd_acls_fattr(&fattr, idmap, inode); + rc = ksmbd_acls_fattr(&fattr, idmap, inode); + if (rc) { + ksmbd_fd_put(work, fp); + return rc; + } if (test_share_config_flag(work->tcon->share_conf, KSMBD_SHARE_FLAG_ACL_XATTR)) @@ -7982,9 +8019,11 @@ static int smb2_rename(struct ksmbd_work *work, return PTR_ERR(new_name); if (fp->is_posix_ctxt == false && strchr(new_name, ':')) { - int s_type; + int s_type = 0; char *xattr_stream_name, *stream_name = NULL; + char *stream_buf = NULL; size_t xattr_stream_size; + ssize_t stream_buf_len = 0; int len; rc = parse_stream_name(new_name, &stream_name, &s_type); @@ -7998,6 +8037,10 @@ static int smb2_rename(struct ksmbd_work *work, goto out; } + /* An empty stream name is the base file's default stream. */ + if (!stream_name || !stream_name[0]) + goto out; + rc = ksmbd_vfs_xattr_stream_name(stream_name, &xattr_stream_name, &xattr_stream_size, @@ -8005,15 +8048,34 @@ static int smb2_rename(struct ksmbd_work *work, if (rc) goto out; + /* A handle opened without a stream has no source to copy. */ + if (ksmbd_stream_fd(fp)) { + if (!strcasecmp(xattr_stream_name, fp->stream.name)) { + kfree(xattr_stream_name); + goto out; + } + + stream_buf_len = ksmbd_vfs_getcasexattr(file_mnt_idmap(fp->filp), + fp->filp->f_path.dentry, + fp->stream.name, + fp->stream.size, + &stream_buf); + if (stream_buf_len < 0) { + rc = stream_buf_len; + kfree(xattr_stream_name); + goto out; + } + } + rc = ksmbd_vfs_setxattr(file_mnt_idmap(fp->filp), &fp->filp->f_path, xattr_stream_name, - NULL, 0, 0, true); - if (rc < 0) { + stream_buf, stream_buf_len, 0, true); + kfree(stream_buf); + if (rc < 0) pr_err("failed to store stream name in xattr: %d\n", rc); - rc = -EINVAL; - } + kfree(xattr_stream_name); goto out; } @@ -8113,7 +8175,7 @@ static int set_file_basic_info(struct ksmbd_file *fp, struct iattr attrs; struct file *filp; struct inode *inode; - struct mnt_idmap *idmap; + const struct mnt_idmap *idmap; __le32 attrs_mask = FILE_ATTRIBUTE_DIRECTORY_LE | FILE_ATTRIBUTE_COMPRESSED_LE; int rc = 0; @@ -9276,23 +9338,25 @@ static noinline int smb2_write_pipe(struct ksmbd_work *work) int err = 0, ret = 0; char *data_buf; size_t length; + unsigned int data_offset, req_len; WORK_BUFFERS(work, req, rsp); length = le32_to_cpu(req->Length); id = req->VolatileFileId; + data_offset = le16_to_cpu(req->DataOffset); + req_len = smb2_current_req_len(work, &req->hdr); - if ((u64)le16_to_cpu(req->DataOffset) + length > - get_rfc1002_len(work->request_buf)) { - pr_err("invalid write data offset %u, smb_len %u\n", - le16_to_cpu(req->DataOffset), - get_rfc1002_len(work->request_buf)); + if (data_offset < offsetof(struct smb2_write_req, Buffer) || + data_offset > req_len || length > req_len - data_offset) { + pr_err("invalid write data offset %u, length %zu, req_len %u\n", + data_offset, length, req_len); err = -EINVAL; goto out; } data_buf = (char *)(((char *)&req->hdr.ProtocolId) + - le16_to_cpu(req->DataOffset)); + data_offset); rpc_resp = ksmbd_rpc_write(work->sess, id, data_buf, length); if (rpc_resp) { @@ -9579,14 +9643,21 @@ int smb2_write(struct ksmbd_work *work) writethrough = true; if (is_rdma_channel == false) { - if (le16_to_cpu(req->DataOffset) < - offsetof(struct smb2_write_req, Buffer)) { + unsigned int data_offset = le16_to_cpu(req->DataOffset); + unsigned int req_len = smb2_current_req_len(work, &req->hdr); + + if (data_offset < offsetof(struct smb2_write_req, Buffer) || + data_offset > req_len || + length > req_len - data_offset) { + ksmbd_debug(SMB, + "invalid write data offset %u, length %zu, req_len %u\n", + data_offset, length, req_len); err = -EINVAL; goto out; } data_buf = (char *)(((char *)&req->hdr.ProtocolId) + - le16_to_cpu(req->DataOffset)); + data_offset); ksmbd_debug(SMB, "filename %pD, offset %lld, len %zu\n", fp->filp, offset, length); @@ -9708,7 +9779,6 @@ int smb2_cancel(struct ksmbd_work *work) struct smb2_hdr *hdr = smb_get_msg(work->request_buf); struct smb2_hdr *chdr; struct ksmbd_work *iter; - struct ksmbd_work *cancelled_notify = NULL; struct list_head *command_list; if (work->next_smb2_rcv_hdr_off) @@ -9746,23 +9816,11 @@ int smb2_cancel(struct ksmbd_work *work) "smb2 with AsyncId %llu cancelled command = 0x%x\n", le64_to_cpu(hdr->Id.AsyncId), le16_to_cpu(chdr->Command)); - if (iter->cancel_fn == smb2_notify_cancel_fn) - cancelled_notify = - smb2_notify_cancel_claim(iter->cancel_argv); - else if (iter->cancel_fn) + if (iter->cancel_fn) iter->cancel_fn(iter->cancel_argv); break; } spin_unlock(&conn->request_lock); - - /* - * Complete a cancelled notify before this CANCEL handler returns. - * Deferring it to the system workqueue lets a following request and - * its response overtake STATUS_CANCELLED, leaving clients waiting - * for the original notify even though the cancellation was accepted. - */ - if (cancelled_notify) - smb2_complete_notify_cancel(cancelled_notify); } else { command_list = &conn->requests; @@ -10466,6 +10524,27 @@ static __be32 idev_ipv4_address(struct in_device *idev) return addr; } +static struct network_interface_info_ioctl_rsp * +ksmbd_iface_entry_init(struct smb2_ioctl_rsp *rsp, int nbytes, + struct net_device *netdev, unsigned long long speed) +{ + struct network_interface_info_ioctl_rsp *nii_rsp; + + nii_rsp = (struct network_interface_info_ioctl_rsp *)&rsp->Buffer[nbytes]; + nii_rsp->IfIndex = cpu_to_le32(netdev->ifindex); + nii_rsp->Capability = 0; + if (netdev->real_num_tx_queues > 1) + nii_rsp->Capability |= RSS_CAPABLE; + if (ksmbd_rdma_capable_netdev(netdev)) + nii_rsp->Capability |= RDMA_CAPABLE; + nii_rsp->Next = cpu_to_le32(152); + nii_rsp->Reserved = 0; + nii_rsp->LinkSpeed = cpu_to_le64(speed); + memset(nii_rsp->SockAddr_Storage, 0, 128); + + return nii_rsp; +} + static int fsctl_query_iface_info_ioctl(struct ksmbd_conn *conn, struct smb2_ioctl_rsp *rsp, unsigned int out_buf_len) @@ -10476,10 +10555,16 @@ static int fsctl_query_iface_info_ioctl(struct ksmbd_conn *conn, struct sockaddr_storage_rsp *sockaddr_storage; unsigned int flags; unsigned long long speed; + struct ethtool_link_ksettings cmd; rtnl_lock(); for_each_netdev(&init_net, netdev) { - bool ipv4_set = false; + struct inet6_ifaddr *ifa; + struct inet6_dev *idev6; + struct in_device *idev; + struct in6_addr ip6 = { }; + bool have_ip6 = false; + __be32 ip4 = 0; if (netdev->type == ARPHRD_LOOPBACK) continue; @@ -10490,87 +10575,80 @@ static int fsctl_query_iface_info_ioctl(struct ksmbd_conn *conn, flags = netif_get_flags(netdev); if (!(flags & IFF_RUNNING)) continue; -ipv6_retry: - if (out_buf_len < - nbytes + sizeof(struct network_interface_info_ioctl_rsp)) { - rtnl_unlock(); - return -ENOSPC; - } - - nii_rsp = (struct network_interface_info_ioctl_rsp *) - &rsp->Buffer[nbytes]; - nii_rsp->IfIndex = cpu_to_le32(netdev->ifindex); - - nii_rsp->Capability = 0; - if (netdev->real_num_tx_queues > 1) - nii_rsp->Capability |= RSS_CAPABLE; - if (ksmbd_rdma_capable_netdev(netdev)) - nii_rsp->Capability |= RDMA_CAPABLE; - nii_rsp->Next = cpu_to_le32(152); - nii_rsp->Reserved = 0; - - if (netdev->ethtool_ops->get_link_ksettings) { - struct ethtool_link_ksettings cmd; - - netdev->ethtool_ops->get_link_ksettings(netdev, &cmd); + if (!__ethtool_get_link_ksettings(netdev, &cmd) && + cmd.base.speed && cmd.base.speed != SPEED_UNKNOWN) { speed = cmd.base.speed; } else { ksmbd_debug(SMB, "%s %s\n", netdev->name, "speed is unknown, defaulting to 1Gb/sec"); speed = SPEED_1000; } - speed *= 1000000; - nii_rsp->LinkSpeed = cpu_to_le64(speed); - sockaddr_storage = (struct sockaddr_storage_rsp *) - nii_rsp->SockAddr_Storage; - memset(sockaddr_storage, 0, 128); + /* + * Query IPv4 and IPv6 independently; emit an entry only when a + * usable address exists, so an interface missing one family is + * still reported for the other and 0.0.0.0 / :: placeholders are + * never advertised. + */ + idev = __in_dev_get_rtnl(netdev); + if (idev) + ip4 = idev_ipv4_address(idev); + + idev6 = __in6_dev_get(netdev); + if (idev6) { + rcu_read_lock(); + list_for_each_entry_rcu(ifa, &idev6->addr_list, if_list) { + if (ifa->flags & (IFA_F_TENTATIVE | IFA_F_DEPRECATED)) + continue; + memcpy(&ip6, ifa->addr.s6_addr, sizeof(ip6)); + have_ip6 = true; + break; + } + rcu_read_unlock(); + } - if (!ipv4_set) { - struct in_device *idev; + if (ip4) { + if (out_buf_len < + nbytes + sizeof(struct network_interface_info_ioctl_rsp)) { + rtnl_unlock(); + return -ENOSPC; + } + nii_rsp = ksmbd_iface_entry_init(rsp, nbytes, netdev, speed); + sockaddr_storage = (struct sockaddr_storage_rsp *)nii_rsp->SockAddr_Storage; sockaddr_storage->Family = INTERNETWORK; sockaddr_storage->addr4.Port = 0; - - idev = __in_dev_get_rtnl(netdev); - if (!idev) - continue; - sockaddr_storage->addr4.IPv4Address = - idev_ipv4_address(idev); + sockaddr_storage->addr4.IPv4Address = ip4; nbytes += sizeof(struct network_interface_info_ioctl_rsp); - ipv4_set = true; - goto ipv6_retry; - } else { - struct inet6_dev *idev6; - struct inet6_ifaddr *ifa; - __u8 *ipv6_addr = sockaddr_storage->addr6.IPv6Address; + } + if (have_ip6) { + if (out_buf_len < + nbytes + sizeof(struct network_interface_info_ioctl_rsp)) { + rtnl_unlock(); + return -ENOSPC; + } + + nii_rsp = ksmbd_iface_entry_init(rsp, nbytes, netdev, speed); + sockaddr_storage = (struct sockaddr_storage_rsp *)nii_rsp->SockAddr_Storage; sockaddr_storage->Family = INTERNETWORKV6; sockaddr_storage->addr6.Port = 0; sockaddr_storage->addr6.FlowInfo = 0; - - idev6 = __in6_dev_get(netdev); - if (!idev6) - continue; - - list_for_each_entry(ifa, &idev6->addr_list, if_list) { - if (ifa->flags & (IFA_F_TENTATIVE | - IFA_F_DEPRECATED)) - continue; - memcpy(ipv6_addr, ifa->addr.s6_addr, 16); - break; - } + memcpy(sockaddr_storage->addr6.IPv6Address, ip6.s6_addr, 16); sockaddr_storage->addr6.ScopeId = 0; nbytes += sizeof(struct network_interface_info_ioctl_rsp); } } rtnl_unlock(); - /* zero if this is last one */ - if (nii_rsp) + /* Clear Next of the last committed entry to terminate the list. */ + if (nbytes > 0) { + nii_rsp = (struct network_interface_info_ioctl_rsp *) + &rsp->Buffer[nbytes - sizeof(*nii_rsp)]; nii_rsp->Next = 0; + } rsp->PersistentFileId = SMB2_NO_FID; rsp->VolatileFileId = SMB2_NO_FID; @@ -10715,7 +10793,7 @@ static inline int fsctl_set_sparse(struct ksmbd_work *work, u64 id, struct file_sparse *sparse) { struct ksmbd_file *fp; - struct mnt_idmap *idmap; + const struct mnt_idmap *idmap; int ret = 0; __le32 old_fattr; @@ -10857,9 +10935,28 @@ int smb2_ioctl(struct ksmbd_work *work) case FSCTL_QUERY_NETWORK_INTERFACE_INFO: case FSCTL_VALIDATE_NEGOTIATE_INFO: case FSCTL_PIPE_WAIT: - case FSCTL_PIPE_TRANSCEIVE: no_fileid_ioctl = true; break; + case FSCTL_PIPE_TRANSCEIVE: + if (!test_share_config_flag(work->tcon->share_conf, + KSMBD_SHARE_FLAG_PIPE)) { + ret = -EOPNOTSUPP; + goto out; + } + + /* RPC pipe handles are not in the regular file table. */ + if (has_file_id(id) && !pid) { + down_read(&work->sess->rpc_lock); + no_fileid_ioctl = + ksmbd_session_rpc_method(work->sess, id) != 0; + up_read(&work->sess->rpc_lock); + } + if (!no_fileid_ioctl) { + ret = -EBADF; + rsp->hdr.Status = STATUS_FILE_CLOSED; + goto out2; + } + break; default: break; } @@ -11018,6 +11115,8 @@ int smb2_ioctl(struct ksmbd_work *work) break; } case FSCTL_PIPE_TRANSCEIVE: + rsp->PersistentFileId = pid; + rsp->VolatileFileId = id; out_buf_len = min_t(u32, KSMBD_IPC_MAX_PAYLOAD, out_buf_len); nbytes = fsctl_pipe_transceive(work, id, out_buf_len, req, rsp); break; @@ -11719,137 +11818,26 @@ int smb2_oplock_break(struct ksmbd_work *work) return 0; } -/* - * Cancel handler for a deferred CHANGE_NOTIFY. Races against - * __ksmbd_close_fd()'s notify_pendings drain (vfs_cache.c), which can run - * concurrently on a different connection closing the same handle -- only - * one of the two may claim and free in_work, so both sides check - * list_empty() under fp->f_lock before touching it (list_del_init() - * leaves a node empty, so whichever side removes it first is the owner; - * the loser must not touch in_work again, since the winner may already be - * freeing it). - * - * smb2_cancel() holds conn->request_lock (a spinlock) for the entire - * time it walks conn->async_requests and calls this function -- so this - * runs with preemption disabled and must not sleep or re-acquire that - * same lock. release_async_work() does both (it takes conn->request_lock - * itself, and frees things that can involve sleeping paths), so calling - * it from here would self-deadlock the very thread processing the - * client's CANCEL command. ksmbd_conn_write() can also sleep (it takes - * conn's write mutex). So: do only the non-sleeping, no-relock cleanup - * inline here. smb2_cancel() sends and frees the claimed notify after it - * drops request_lock, preserving response order for a client CANCEL. The - * connection teardown caller has no such post-unlock path, so its wrapper - * defers the send and free to a workqueue. - */ -struct notify_cancel_ctx { - struct work_struct work; - struct ksmbd_work *in_work; +struct ksmbd_notify_req { + wait_queue_head_t wait; }; -static void smb2_send_notify_cancelled(struct ksmbd_work *work) -{ - struct smb2_hdr *hdr = smb_get_msg(work->response_buf); - struct ksmbd_conn *conn = work->conn; - struct ksmbd_session *sess; - - sess = ksmbd_session_lookup(conn, le64_to_cpu(hdr->SessionId)); - if (sess) { - work->sess = sess; - if (work->encrypted && sess->enc && conn->ops->encrypt_resp) { - conn->ops->encrypt_resp(work); - } else if (conn->ops->is_sign_req && conn->ops->set_sign_rsp && - conn->ops->is_sign_req(work, - conn->ops->get_cmd_val(work))) { - conn->ops->set_sign_rsp(work); - } - } - - ksmbd_conn_write(work); - if (sess) { - ksmbd_user_session_put(sess); - work->sess = NULL; - } -} - -static void smb2_notify_cancel_deferred(struct work_struct *w) -{ - struct notify_cancel_ctx *ctx = - container_of(w, struct notify_cancel_ctx, work); - struct ksmbd_conn *conn = ctx->in_work->conn; - - smb2_complete_notify_cancel(ctx->in_work); - kfree(ctx); - /* - * The connection teardown waits for r_count before destroying - * connection sessions and their proc entries. - */ - ksmbd_conn_r_count_dec(conn); -} - -static struct ksmbd_work *smb2_notify_cancel_claim(void **argv) -{ - struct ksmbd_work *in_work = (struct ksmbd_work *)argv[0]; - struct ksmbd_file *fp = (struct ksmbd_file *)argv[1]; - bool claimed; - - spin_lock(&fp->f_lock); - claimed = !list_empty(&in_work->notify_entry); - if (claimed) - list_del_init(&in_work->notify_entry); - spin_unlock(&fp->f_lock); - - if (!claimed) - return NULL; - - /* conn->request_lock is held by smb2_cancel() or connection teardown. */ - in_work->cancel_fn = NULL; - kfree(in_work->cancel_argv); - in_work->cancel_argv = NULL; - return in_work; -} - -static void smb2_complete_notify_cancel(struct ksmbd_work *in_work) -{ - struct smb2_hdr *in_hdr = smb_get_msg(in_work->response_buf); - - in_hdr->Status = STATUS_CANCELLED; - smb2_send_notify_cancelled(in_work); - release_async_work(in_work); - ksmbd_free_work_struct(in_work); -} - -static void smb2_notify_cancel_fn(void **argv) +/* + * Cancel handler for a pending CHANGE_NOTIFY. Called either by + * smb2_cancel() (conn->request_lock held, work->state already set to + * KSMBD_WORK_CANCELLED by the caller) or by + * set_close_state_blocked_works() (vfs_cache.c, fp->f_lock held, + * work->state already set to KSMBD_WORK_CLOSED by the caller) -- both + * callers hold a spinlock across this call, so it must not sleep. + * wake_up() only wakes the waiter in smb2_notify(); it does not touch + * fp->blocked_works itself, matching smb2_remove_blocked_lock()'s same + * non-mutating style for the equivalent byte-range-lock wait. + */ +static void smb2_notify_cancel(void **argv) { - struct ksmbd_work *in_work = smb2_notify_cancel_claim(argv); - struct ksmbd_conn *conn; - struct notify_cancel_ctx *ctx; - - if (!in_work) - return; - conn = in_work->conn; + struct ksmbd_notify_req *notify_req = argv[0]; - ctx = kmalloc_obj(*ctx, GFP_ATOMIC); - if (!ctx) { - /* Can't defer the response -- free without sending one. */ - list_del_init(&in_work->async_request_entry); - in_work->asynchronous = false; - if (in_work->async_id) { - ksmbd_release_id(&conn->async_ida, in_work->async_id); - in_work->async_id = 0; - } - ksmbd_free_work_struct(in_work); - return; - } - ctx->in_work = in_work; - INIT_WORK(&ctx->work, smb2_notify_cancel_deferred); - /* - * This deferred work can outlive the connection handler's receive loop. - * Keep teardown from destroying the connection's sessions until the - * deferred response has finished using them. - */ - ksmbd_conn_r_count_inc(conn); - schedule_work(&ctx->work); + wake_up(¬ify_req->wait); } /** @@ -11862,9 +11850,11 @@ int smb2_notify(struct ksmbd_work *work) { struct smb2_change_notify_req *req; struct smb2_change_notify_rsp *rsp; - struct ksmbd_work *in_work; - struct smb2_hdr *in_hdr; - struct ksmbd_file *fp; + struct ksmbd_notify_req notify_req; + struct ksmbd_file *fp = NULL; + void **argv = NULL; + bool async_work = false; + int err = 0; ksmbd_debug(SMB, "Received smb2 notify\n"); @@ -11875,164 +11865,83 @@ int smb2_notify(struct ksmbd_work *work) if (work->next_smb2_rcv_hdr_off && req->hdr.NextCommand) { rsp->hdr.Status = STATUS_INTERNAL_ERROR; - smb2_set_err_rsp(work); - return -EIO; - } - - /* - * macOS backupd sends CHANGE_NOTIFY with FileId=FFFF...FFFF (share-root - * sentinel) to watch for changes on the share root without holding an - * open handle. Respond STATUS_PENDING + STATUS_NOTIFY_CLEANUP immediately; - * without this, backupd aborts Time Machine setup on STATUS_FILE_CLOSED. - */ - if (req->VolatileFileId == SMB2_NO_FID && - req->PersistentFileId == SMB2_NO_FID) { - in_work = ksmbd_alloc_work_struct(); - if (!in_work || allocate_interim_rsp_buf(in_work)) { - if (in_work) - ksmbd_free_work_struct(in_work); - rsp->hdr.Status = STATUS_INSUFFICIENT_RESOURCES; - smb2_set_err_rsp(work); - return 0; - } - if (setup_async_work(work, NULL, NULL)) { - ksmbd_free_work_struct(in_work); - rsp->hdr.Status = STATUS_INSUFFICIENT_RESOURCES; - smb2_set_err_rsp(work); - return 0; - } - smb2_send_interim_resp(work, STATUS_PENDING); - in_work->conn = work->conn; - in_hdr = smb_get_msg(in_work->response_buf); - memcpy(in_hdr, ksmbd_resp_buf_next(work), - __SMB2_HEADER_STRUCTURE_SIZE); - in_hdr->Flags |= SMB2_FLAGS_ASYNC_COMMAND; - in_hdr->Id.AsyncId = cpu_to_le64(work->async_id); - smb2_set_err_rsp(in_work); - in_hdr->Status = STATUS_NOTIFY_CLEANUP; - in_work->async_id = work->async_id; - work->async_id = 0; - release_async_work(work); - if (smb2_send_interim_work(in_work, work, false)) - ksmbd_debug(SMB, "failed to send notify cleanup\n"); - ksmbd_free_work_struct(in_work); - work->send_no_response = 1; - return 0; + err = -EIO; + goto out; } - /* - * KSMBD does not implement a real change-notification backend. - * Genuine SMB2 servers (and macOS smbfs) never complete a - * CHANGE_NOTIFY spontaneously: it is satisfied only by a real - * directory change, or with STATUS_NOTIFY_CLEANUP when the watched - * handle is closed. Completing it early (e.g. on a timer) makes - * Finder treat the cleanup as "directory changed" and re-enumerate - * the directory forever, leaving items unopenable. Returning - * STATUS_NOT_IMPLEMENTED here (like stock ksmbd) makes macOS smbfs - * hard-freeze on unmount, so this must stay deferred. - */ fp = ksmbd_lookup_fd_slow(work, req->VolatileFileId, req->PersistentFileId); if (!fp) { rsp->hdr.Status = STATUS_FILE_CLOSED; - smb2_set_err_rsp(work); - return 0; + err = -ENOENT; + goto out; } - in_work = ksmbd_alloc_work_struct(); - if (!in_work || allocate_interim_rsp_buf(in_work)) { - if (in_work) - ksmbd_free_work_struct(in_work); - ksmbd_fd_put(work, fp); - rsp->hdr.Status = STATUS_INSUFFICIENT_RESOURCES; - smb2_set_err_rsp(work); - return 0; - } - /* - * in_work is synthetic (not from the normal request-receiving - * pipeline), so it has no request_buf of its own. It gets registered - * into conn->async_requests below, and smb2_cancel() unconditionally - * computes smb_get_msg(iter->request_buf) for every entry in that - * list while searching for a match -- give it its own small buffer - * (not an alias of response_buf: ksmbd_free_work_struct() kvfree()s - * both separately, so aliasing them would double-free) so that stays - * a harmless read instead of a near-NULL dereference. - */ - in_work->request_buf = kzalloc(MAX_CIFS_SMALL_BUFFER_SIZE, KSMBD_DEFAULT_GFP); - if (!in_work->request_buf) { - ksmbd_free_work_struct(in_work); - ksmbd_fd_put(work, fp); + argv = kmalloc_obj(*argv, KSMBD_DEFAULT_GFP); + if (!argv) { rsp->hdr.Status = STATUS_INSUFFICIENT_RESOURCES; - smb2_set_err_rsp(work); - return 0; + err = -ENOMEM; + goto out; } - memcpy(smb_get_msg(in_work->request_buf), req, - __SMB2_HEADER_STRUCTURE_SIZE); + init_waitqueue_head(¬ify_req.wait); + argv[0] = ¬ify_req; - if (setup_async_work(work, NULL, NULL)) { - ksmbd_free_work_struct(in_work); - ksmbd_fd_put(work, fp); + err = setup_async_work(work, smb2_notify_cancel, argv); + if (err) { rsp->hdr.Status = STATUS_INSUFFICIENT_RESOURCES; - smb2_set_err_rsp(work); - return 0; + goto out; } - - smb2_send_interim_resp(work, STATUS_PENDING); - - /* Keep the async IDA alive until the deferred work is released. */ - in_work->conn = ksmbd_conn_get(work->conn); - in_work->owns_conn_ref = true; - in_work->encrypted = work->encrypted; - in_hdr = smb_get_msg(in_work->response_buf); - memcpy(in_hdr, ksmbd_resp_buf_next(work), __SMB2_HEADER_STRUCTURE_SIZE); - in_hdr->Flags |= SMB2_FLAGS_ASYNC_COMMAND; - in_hdr->Id.AsyncId = cpu_to_le64(work->async_id); - smb2_set_err_rsp(in_work); - in_hdr->Status = STATUS_NOTIFY_CLEANUP; + async_work = true; /* - * Transfer ownership of the async id to in_work; it stays reserved - * until in_work is freed after the deferred response is sent on - * close, so it can't be reused for an unrelated async response. + * Handle close holds the file-table write lock while it marks the + * handle closed and walks blocked_works. Hold the matching read lock + * across the state check and registration so close cannot finish its + * walk between the lookup above and this list insertion. */ - in_work->async_id = work->async_id; - work->async_id = 0; - release_async_work(work); + read_lock(&work->sess->file_table.lock); + if (fp->f_state != FP_INITED) { + read_unlock(&work->sess->file_table.lock); + rsp->hdr.Status = STATUS_NOTIFY_CLEANUP; + err = -ENOENT; + goto out; + } + spin_lock(&fp->f_lock); + list_add_tail(&work->fp_entry, &fp->blocked_works); + spin_unlock(&fp->f_lock); + read_unlock(&work->sess->file_table.lock); - /* - * work itself is about to be recycled by the normal request-processing - * pipeline, so it can't stay the target of a future CANCEL -- register - * in_work instead, reusing the same async_id, so a client-sent CANCEL - * for this notify actually finds something to cancel instead of - * silently doing nothing until the handle eventually closes. - */ - in_work->asynchronous = true; - in_work->cancel_argv = kmalloc_array(2, sizeof(void *), KSMBD_DEFAULT_GFP); - if (in_work->cancel_argv) { - in_work->cancel_argv[0] = in_work; - in_work->cancel_argv[1] = fp; - in_work->cancel_fn = smb2_notify_cancel_fn; - } - - if (!ksmbd_conn_link_async_request(work->conn, in_work)) { - kfree(in_work->cancel_argv); - in_work->cancel_argv = NULL; - in_work->cancel_fn = NULL; - in_work->asynchronous = false; - ksmbd_fd_put(work, fp); - if (smb2_send_interim_work(in_work, work, false)) - ksmbd_debug(SMB, "failed to send notify cleanup\n"); - ksmbd_free_work_struct(in_work); - work->send_no_response = 1; - return 0; + smb2_send_interim_resp(work, STATUS_PENDING); + + err = wait_event_interruptible(notify_req.wait, + READ_ONCE(work->state) != KSMBD_WORK_ACTIVE); + if (err && READ_ONCE(work->state) == KSMBD_WORK_ACTIVE) { + /* + * Woken by a signal, not a real cancel/close. There is no + * notification backend yet to report anything else against, + * so treat this the same as a client-side cancel. + */ + WRITE_ONCE(work->state, KSMBD_WORK_CANCELLED); } spin_lock(&fp->f_lock); - list_add_tail(&in_work->notify_entry, &fp->notify_pendings); + list_del_init(&work->fp_entry); spin_unlock(&fp->f_lock); - ksmbd_fd_put(work, fp); + rsp->hdr.Status = work->state == KSMBD_WORK_CLOSED ? + STATUS_NOTIFY_CLEANUP : STATUS_CANCELLED; + smb2_send_interim_resp(work, rsp->hdr.Status); work->send_no_response = 1; - return 0; + +out: + if (rsp->hdr.Status != STATUS_SUCCESS && !work->send_no_response) + smb2_set_err_rsp(work); + if (async_work) + release_async_work(work); + else + kfree(argv); + if (fp) + ksmbd_fd_put(work, fp); + return err; } /** @@ -12227,11 +12136,12 @@ void smb3_set_sign_rsp(struct ksmbd_work *work) struct channel *chann; char signature[SMB2_CMACAES_SIZE]; struct kvec *iov; - u16 command = conn->ops->get_cmd_val(work); + u16 command; int n_vec; char *signing_key; hdr = ksmbd_resp_buf_curr(work); + command = le16_to_cpu(hdr->Command); if (command == SMB2_SESSION_SETUP_HE && (!conn->binding || hdr->Status != STATUS_SUCCESS)) { diff --git a/fs/smb/server/smb2pdu.h b/fs/smb/server/smb2pdu.h index ca8e27f7b712..40d61745ef41 100644 --- a/fs/smb/server/smb2pdu.h +++ b/fs/smb/server/smb2pdu.h @@ -156,8 +156,7 @@ struct create_durable_rsp { } __packed; /* - * See POSIX-SMB2 2.2.14.2.16 - * Link: https://gitlab.com/samba-team/smb3-posix-spec/-/blob/master/smb3_posix_extensions.md + * See POSIX-SMB2 2.1.3.2.1 */ struct create_posix_rsp { struct create_context_hdr ccontext; @@ -199,7 +198,7 @@ struct file_sparse { #define FILE_INTERNAL_INFORMATION_SIZE 8 #define FILE_EA_INFORMATION_SIZE 4 #define FILE_ACCESS_INFORMATION_SIZE 4 -#define FILE_NAME_INFORMATION_SIZE 9 +#define FILE_NAME_INFORMATION_SIZE 8 #define FILE_RENAME_INFORMATION_SIZE 10 #define FILE_LINK_INFORMATION_SIZE 11 #define FILE_NAMES_INFORMATION_SIZE 12 @@ -211,8 +210,6 @@ struct file_sparse { #define FILE_ALL_INFORMATION_SIZE 104 #define FILE_ALLOCATION_INFORMATION_SIZE 19 #define FILE_END_OF_FILE_INFORMATION_SIZE 20 -#define FILE_ALTERNATE_NAME_INFORMATION_SIZE 8 -#define FILE_NORMALIZED_NAME_INFORMATION_SIZE 8 #define FILE_STREAM_INFORMATION_SIZE 32 #define FILE_PIPE_INFORMATION_SIZE 23 #define FILE_PIPE_LOCAL_INFORMATION_SIZE 24 @@ -259,7 +256,7 @@ struct smb2_file_alignment_info { __le32 AlignmentRequirement; } __packed; -struct smb2_file_alt_name_info { +struct smb2_file_name_info { __le32 FileNameLength; char FileName[]; } __packed; @@ -351,6 +348,7 @@ struct create_sd_buf_req { struct smb_ntsd ntsd; } __packed; +/* See POSIX-FSCC 2.2.1 */ struct smb2_posix_info { __le32 NextEntryOffset; __u32 Ignored; @@ -364,20 +362,18 @@ struct smb2_posix_info { __le64 Inode; __le32 DeviceId; __le32 Zero; - /* beginning of POSIX Create Context Response */ + /* + * Beginning of POSIX Create Context Response + * See POSIX-SMB2 2.1.3.2.1 + */ __le32 HardLinks; __le32 ReparseTag; __le32 Mode; /* SidBuffer contain two sids (UNIX user sid(16), UNIX group sid(16)) */ u8 SidBuffer[32]; + /* End of POSIX Create Context Response */ __le32 name_len; u8 name[]; - /* - * var sized owner SID - * var sized group SID - * le32 filenamelength - * u8 filename[] - */ } __packed; /* functions */ diff --git a/fs/smb/server/smb_common.c b/fs/smb/server/smb_common.c index 4c2da65510bc..7dbcfa658edd 100644 --- a/fs/smb/server/smb_common.c +++ b/fs/smb/server/smb_common.c @@ -467,7 +467,7 @@ int ksmbd_populate_dot_dotdot_entries(struct ksmbd_work *work, int info_level, { int i, rc = 0; struct ksmbd_conn *conn = work->conn; - struct mnt_idmap *idmap = file_mnt_idmap(dir->filp); + const struct mnt_idmap *idmap = file_mnt_idmap(dir->filp); for (i = 0; i < 2; i++) { struct kstat kstat; @@ -523,8 +523,6 @@ int ksmbd_populate_dot_dotdot_entries(struct ksmbd_work *work, int info_level, * @shortname: destination short filename * * Return: shortname length or 0 when source long name is '.' or '..' - * TODO: Though this function conforms the restriction of 8.3 Filename spec, - * but the result is different with Windows 7's one. need to check. */ int ksmbd_extract_shortname(struct ksmbd_conn *conn, const char *longname, char *shortname) @@ -588,7 +586,7 @@ int ksmbd_extract_shortname(struct ksmbd_conn *conn, const char *longname, if (dot_present) memcpy(out + baselen + 4, extension, 4); else - out[baselen + 4] = '\0'; + out[baselen + 3] = '\0'; smbConvertToUTF16((__le16 *)shortname, out, PATH_MAX, conn->local_nls, 0); len = strlen(out) * 2; diff --git a/fs/smb/server/smbacl.c b/fs/smb/server/smbacl.c index 1fad6ccf3a72..96428df33b43 100644 --- a/fs/smb/server/smbacl.c +++ b/fs/smb/server/smbacl.c @@ -7,6 +7,7 @@ */ #include <linux/fs.h> +#include <kunit/visibility.h> #include <linux/slab.h> #include <linux/string.h> #include <linux/mnt_idmapping.h> @@ -257,7 +258,7 @@ void id_to_sid(unsigned int cid, uint sidtype, struct smb_sid *ssid) ssid->num_subauth++; } -static int sid_to_id(struct mnt_idmap *idmap, +static int sid_to_id(const struct mnt_idmap *idmap, struct smb_sid *psid, uint sidtype, struct smb_fattr *fattr) { @@ -383,7 +384,7 @@ void free_acl_state(struct posix_acl_state *state) kfree(state->groups); } -static int parse_dacl(struct mnt_idmap *idmap, +static int parse_dacl(const struct mnt_idmap *idmap, struct smb_acl *pdacl, char *end_of_acl, struct smb_sid *pownersid, struct smb_sid *pgrpsid, struct smb_fattr *fattr) @@ -619,7 +620,7 @@ out: return ret; } -static void set_posix_acl_entries_dacl(struct mnt_idmap *idmap, +static void set_posix_acl_entries_dacl(const struct mnt_idmap *idmap, struct smb_ace *pndace, struct smb_fattr *fattr, u16 *num_aces, u16 *size, u16 existing_nt_aces, @@ -750,7 +751,7 @@ posix_default_acl: } } -static void set_ntacl_dacl(struct mnt_idmap *idmap, +static void set_ntacl_dacl(const struct mnt_idmap *idmap, struct smb_acl *pndacl, struct smb_acl *nt_dacl, unsigned int aces_size, @@ -809,7 +810,7 @@ next_ace: pndacl->size = cpu_to_le16(le16_to_cpu(pndacl->size) + size); } -static void set_mode_dacl(struct mnt_idmap *idmap, +static void set_mode_dacl(const struct mnt_idmap *idmap, struct smb_acl *pndacl, struct smb_fattr *fattr) { struct smb_ace *pace, *pndace; @@ -895,7 +896,7 @@ static int parse_sid(struct smb_sid *psid, char *end_of_acl) } /* Convert CIFS ACL to POSIX form */ -int parse_sec_desc(struct mnt_idmap *idmap, struct smb_ntsd *pntsd, +int parse_sec_desc(const struct mnt_idmap *idmap, struct smb_ntsd *pntsd, int acl_len, struct smb_fattr *fattr) { int rc = 0; @@ -1030,7 +1031,7 @@ size_t smb_acl_sec_desc_scratch_len(struct smb_fattr *fattr, } /* Convert permission bits from mode to equivalent CIFS ACL */ -int build_sec_desc(struct mnt_idmap *idmap, +int build_sec_desc(const struct mnt_idmap *idmap, struct smb_ntsd *pntsd, struct smb_ntsd *ppntsd, int ppntsd_size, int addition_info, __u32 *secdesclen, struct smb_fattr *fattr) @@ -1096,9 +1097,12 @@ int build_sec_desc(struct mnt_idmap *idmap, struct smb_acl *ppdacl_ptr; unsigned int dacl_offset = le32_to_cpu(ppntsd->dacloffset); int ppdacl_size, ntacl_size = ppntsd_size - dacl_offset; + size_t dacl_struct_end; if (!dacl_offset || - (dacl_offset + sizeof(struct smb_acl) > ppntsd_size)) + check_add_overflow(dacl_offset, sizeof(struct smb_acl), + &dacl_struct_end) || + dacl_struct_end > (size_t)ppntsd_size) goto out; ppdacl_ptr = (struct smb_acl *)((char *)ppntsd + dacl_offset); @@ -1196,7 +1200,7 @@ int smb_inherit_dacl(struct ksmbd_conn *conn, struct smb_ntsd *parent_pntsd = NULL; struct smb_sid owner_sid, group_sid; struct dentry *parent = path->dentry->d_parent; - struct mnt_idmap *idmap = mnt_idmap(path->mnt); + const struct mnt_idmap *idmap = mnt_idmap(path->mnt); int inherited_flags = 0, flags = 0, i, nt_size = 0, pdacl_size; int rc = 0, pntsd_type, ppntsd_size, acl_len, aces_size; unsigned int dacloffset; @@ -1451,7 +1455,7 @@ int smb_check_perm_dacl(struct ksmbd_conn *conn, const struct path *path, __le32 *pdaccess, __le32 raw_daccess, int uid, bool strict) { - struct mnt_idmap *idmap = mnt_idmap(path->mnt); + const struct mnt_idmap *idmap = mnt_idmap(path->mnt); struct smb_ntsd *pntsd = NULL; struct smb_acl *pdacl; struct posix_acl *posix_acls; @@ -1665,6 +1669,7 @@ err_out: kfree(pntsd); return rc; } +EXPORT_SYMBOL_IF_KUNIT(smb_check_perm_dacl); int set_info_sec(struct ksmbd_conn *conn, struct ksmbd_tree_connect *tcon, const struct path *path, struct smb_ntsd *pntsd, int ntsd_len, @@ -1673,7 +1678,7 @@ int set_info_sec(struct ksmbd_conn *conn, struct ksmbd_tree_connect *tcon, int rc; struct smb_fattr fattr = {{0}}; struct inode *inode = d_inode(path->dentry); - struct mnt_idmap *idmap = mnt_idmap(path->mnt); + const struct mnt_idmap *idmap = mnt_idmap(path->mnt); struct iattr newattrs; fattr.cf_uid = INVALID_UID; diff --git a/fs/smb/server/smbacl.h b/fs/smb/server/smbacl.h index 01810c16cc04..28d215807faa 100644 --- a/fs/smb/server/smbacl.h +++ b/fs/smb/server/smbacl.h @@ -81,9 +81,9 @@ struct posix_acl_state { struct posix_ace_state_array *groups; }; -int parse_sec_desc(struct mnt_idmap *idmap, struct smb_ntsd *pntsd, +int parse_sec_desc(const struct mnt_idmap *idmap, struct smb_ntsd *pntsd, int acl_len, struct smb_fattr *fattr); -int build_sec_desc(struct mnt_idmap *idmap, struct smb_ntsd *pntsd, +int build_sec_desc(const struct mnt_idmap *idmap, struct smb_ntsd *pntsd, struct smb_ntsd *ppntsd, int ppntsd_size, int addition_info, __u32 *secdesclen, struct smb_fattr *fattr); int init_acl_state(struct posix_acl_state *state, u16 cnt); @@ -105,7 +105,7 @@ void ksmbd_init_domain(u32 *sub_auth); size_t smb_acl_sec_desc_scratch_len(struct smb_fattr *fattr, struct smb_ntsd *ppntsd, int ppntsd_size, int addition_info); -static inline uid_t posix_acl_uid_translate(struct mnt_idmap *idmap, +static inline uid_t posix_acl_uid_translate(const struct mnt_idmap *idmap, struct posix_acl_entry *pace) { vfsuid_t vfsuid; @@ -117,7 +117,7 @@ static inline uid_t posix_acl_uid_translate(struct mnt_idmap *idmap, return from_kuid(&init_user_ns, vfsuid_into_kuid(vfsuid)); } -static inline gid_t posix_acl_gid_translate(struct mnt_idmap *idmap, +static inline gid_t posix_acl_gid_translate(const struct mnt_idmap *idmap, struct posix_acl_entry *pace) { vfsgid_t vfsgid; diff --git a/fs/smb/server/tests/Kconfig b/fs/smb/server/tests/Kconfig new file mode 100644 index 000000000000..ad7a4e94ceaa --- /dev/null +++ b/fs/smb/server/tests/Kconfig @@ -0,0 +1,15 @@ +# SPDX-License-Identifier: GPL-2.0-or-later +# Copyright (C) 2026 Hang Nan <nanx95726@gmail.com> + +config SMB_SERVER_KUNIT_TESTS + tristate "KUnit tests for SMB3 server helpers" if !KUNIT_ALL_TESTS + depends on SMB_SERVER && SMB_KUNIT_TESTS && TMPFS_XATTR + default SMB_KUNIT_TESTS + help + This builds the KUnit tests for ksmbd server helpers. The tests + exercise internal server functionality and help detect regressions + in server-side behavior. They are intended for kernel developers + and are not suitable for production systems. + + For more information on KUnit and unit tests in the kernel, + please read Documentation/dev-tools/kunit/index.rst. diff --git a/fs/smb/server/tests/Makefile b/fs/smb/server/tests/Makefile new file mode 100644 index 000000000000..8738ab0b0667 --- /dev/null +++ b/fs/smb/server/tests/Makefile @@ -0,0 +1,4 @@ +# SPDX-License-Identifier: GPL-2.0-or-later +# Copyright (C) 2026 Hang Nan <nanx95726@gmail.com> + +obj-$(CONFIG_SMB_SERVER_KUNIT_TESTS) += smbacl_kunit.o diff --git a/fs/smb/server/tests/smbacl_kunit.c b/fs/smb/server/tests/smbacl_kunit.c new file mode 100644 index 000000000000..33496b4d31a3 --- /dev/null +++ b/fs/smb/server/tests/smbacl_kunit.c @@ -0,0 +1,301 @@ +// SPDX-License-Identifier: GPL-2.0-or-later +/* + * KUnit tests for ksmbd security descriptor (DACL) handling. + * + * Copyright (C) 2026 Hang Nan <nanx95726@gmail.com> + * + * The tests pin the DACL declared-size boundary in smb_check_perm_dacl(): + * + * - ksmbd_dacl_walk_must_stop_at_declared_size: a pure semantic harness + * that models the ACE walk. Walking to the end of the enclosing + * security descriptor (the pre-fix behaviour) selects an ACE that + * sits beyond struct smb_acl::size; stopping at the declared DACL + * size (the fixed behaviour) rejects it. + * + * - ksmbd_smb_check_perm_dacl_boundary and + * ksmbd_smb_check_perm_dacl_maximal_boundary: drive the real + * smb_check_perm_dacl() with a descriptor stored through ksmbd's own + * NTACL xattr path on a tmpfs file, and assert that a post-boundary + * ACE is not selected for either a regular or maximal access check. + */ + +#include <kunit/test.h> +#include <linux/fs.h> +#include <linux/mm.h> +#include <linux/shmem_fs.h> +#include <linux/slab.h> + +#include "../smbacl.h" +#include "../smb_common.h" +#include "../vfs.h" + +struct ksmbd_acl_walk_result { + bool found; + bool allowed; + const struct smb_ace *selected; +}; + +static const struct smb_sid test_nonmatching_sid = { + 1, 5, {0, 0, 0, 0, 0, 5}, + { cpu_to_le32(21), cpu_to_le32(1), cpu_to_le32(2), + cpu_to_le32(3), cpu_to_le32(9999) } +}; + +/* + * S-1-22-1-0: the SID id_to_sid(0, SIDUNIX_USER) resolves to, i.e. what + * smb_check_perm_dacl() looks for when called with uid == 0. + */ +static const struct smb_sid test_owner_sid = { + 1, 2, {0, 0, 0, 0, 0, 22}, + { cpu_to_le32(1), cpu_to_le32(0) } +}; + +static int test_compare_sids(const struct smb_sid *a, const struct smb_sid *b) +{ + int i; + + if (a->revision != b->revision || a->num_subauth != b->num_subauth) + return 1; + for (i = 0; i < NUM_AUTHS; i++) { + if (a->authority[i] != b->authority[i]) + return 1; + } + for (i = 0; i < a->num_subauth; i++) { + if (a->sub_auth[i] != b->sub_auth[i]) + return 1; + } + return 0; +} + +static u16 test_ace_size(const struct smb_sid *sid) +{ + return offsetof(struct smb_ace, sid) + CIFS_SID_BASE_SIZE + + sid->num_subauth * sizeof(__le32); +} + +static u16 fill_test_ace(struct smb_ace *ace, const struct smb_sid *sid, + u32 access_req) +{ + u16 size = test_ace_size(sid); + + ace->type = ACCESS_ALLOWED_ACE_TYPE; + ace->flags = 0; + ace->size = cpu_to_le16(size); + ace->access_req = cpu_to_le32(access_req); + memcpy(&ace->sid, sid, size - offsetof(struct smb_ace, sid)); + return size; +} + +static struct ksmbd_acl_walk_result test_walk_dacl(struct smb_acl *pdacl, + int walk_boundary, + const struct smb_sid *target, + u32 requested) +{ + struct ksmbd_acl_walk_result result = {}; + struct smb_ace *ace; + int aces_size; + int i; + + ace = (struct smb_ace *)((char *)pdacl + sizeof(struct smb_acl)); + aces_size = walk_boundary - sizeof(struct smb_acl); + for (i = 0; i < le16_to_cpu(pdacl->num_aces); i++) { + u16 ace_size; + + if (aces_size < offsetof(struct smb_ace, sid) + CIFS_SID_BASE_SIZE) + break; + ace_size = le16_to_cpu(ace->size); + if (ace_size > aces_size || + ace_size < offsetof(struct smb_ace, sid) + CIFS_SID_BASE_SIZE) + break; + aces_size -= ace_size; + + if (ace->sid.num_subauth > SID_MAX_SUB_AUTHORITIES || + ace_size < offsetof(struct smb_ace, sid) + CIFS_SID_BASE_SIZE + + sizeof(__le32) * ace->sid.num_subauth) + break; + + if (!test_compare_sids(target, &ace->sid)) { + result.found = true; + result.selected = ace; + result.allowed = !(requested & ~le32_to_cpu(ace->access_req)); + return result; + } + + ace = (struct smb_ace *)((char *)ace + ace_size); + } + + return result; +} + +static void ksmbd_dacl_walk_must_stop_at_declared_size(struct kunit *test) +{ + struct ksmbd_acl_walk_result declared, enclosing; + struct smb_acl *acl; + struct smb_ace *ace1, *fake; + u16 ace1_size, fake_size; + u16 pdacl_size; + u16 acl_size; + + acl = kunit_kzalloc(test, 128, GFP_KERNEL); + KUNIT_ASSERT_NOT_NULL(test, acl); + + acl->revision = cpu_to_le16(2); + acl->num_aces = cpu_to_le16(2); + + ace1 = (struct smb_ace *)((char *)acl + sizeof(*acl)); + ace1_size = fill_test_ace(ace1, &test_nonmatching_sid, 0); + fake = (struct smb_ace *)((char *)ace1 + ace1_size); + fake_size = fill_test_ace(fake, &test_owner_sid, FILE_READ_DATA); + + pdacl_size = sizeof(*acl) + ace1_size; + acl_size = pdacl_size + fake_size; + acl->size = cpu_to_le16(pdacl_size); + + declared = test_walk_dacl(acl, pdacl_size, &test_owner_sid, + FILE_READ_DATA); + enclosing = test_walk_dacl(acl, acl_size, &test_owner_sid, + FILE_READ_DATA); + + KUNIT_EXPECT_FALSE(test, declared.found); + KUNIT_EXPECT_FALSE(test, declared.allowed); + + /* Demonstrates that the buggy acl_size boundary selects fake ACE #2. */ + KUNIT_EXPECT_TRUE(test, enclosing.found); + KUNIT_EXPECT_TRUE(test, enclosing.allowed); +} + +/* + * Build an NTSD whose DACL declares one ACE (pdacl->size) but actually + * contains two: the second ACE sits beyond the declared DACL boundary + * yet inside the enclosing security descriptor. The trailing ACE applies + * to S-1-22-1-0, which smb_check_perm_dacl() looks for when uid is zero. + */ +static struct smb_ntsd *build_boundary_ntsd(struct kunit *test, + const struct smb_sid *first_sid, + u32 first_access, + u32 trailing_access, + int *ntsd_size) +{ + struct smb_ntsd *pntsd; + struct smb_acl *pdacl; + struct smb_ace *ace; + u16 first_size = test_ace_size(first_sid); + u16 trailing_size = test_ace_size(&test_owner_sid); + + *ntsd_size = sizeof(struct smb_ntsd) + sizeof(struct smb_acl) + + first_size + trailing_size; + pntsd = kunit_kzalloc(test, *ntsd_size, GFP_KERNEL); + if (!pntsd) + return NULL; + + pntsd->revision = cpu_to_le16(SD_REVISION); + pntsd->type = cpu_to_le16(DACL_PRESENT); + pntsd->dacloffset = cpu_to_le32(sizeof(struct smb_ntsd)); + + pdacl = (struct smb_acl *)((char *)pntsd + sizeof(struct smb_ntsd)); + pdacl->revision = cpu_to_le16(2); + pdacl->num_aces = cpu_to_le16(2); + pdacl->size = cpu_to_le16(sizeof(struct smb_acl) + first_size); + + ace = (struct smb_ace *)((char *)pdacl + sizeof(struct smb_acl)); + fill_test_ace(ace, first_sid, first_access); + + ace = (struct smb_ace *)((char *)ace + first_size); + fill_test_ace(ace, &test_owner_sid, trailing_access); + + return pntsd; +} + +static void ksmbd_smb_check_perm_dacl_boundary_test(struct kunit *test) +{ + struct file *file; + struct smb_ntsd *pntsd; + __le32 daccess = cpu_to_le32(FILE_READ_DATA); + int ntsd_size, rc; + + pntsd = build_boundary_ntsd(test, &test_nonmatching_sid, 0, + FILE_READ_DATA, &ntsd_size); + KUNIT_ASSERT_NOT_NULL(test, pntsd); + + file = shmem_file_setup("ksmbd-kunit-dacl", 0, + mk_vma_flags(VMA_NORESERVE_BIT)); + KUNIT_ASSERT_NOT_ERR_OR_NULL(test, file); + + rc = ksmbd_vfs_set_sd_xattr(NULL, mnt_idmap(file->f_path.mnt), + &file->f_path, pntsd, ntsd_size, + false); + KUNIT_EXPECT_EQ(test, 0, rc); + if (rc) + goto out; + + rc = smb_check_perm_dacl(NULL, &file->f_path, &daccess, + cpu_to_le32(FILE_READ_DATA), 0, false); + + /* + * The post-boundary ACE (ACE #2, beyond pdacl->size) grants + * FILE_READ_DATA to the caller's SID, but it must not be + * selected: the walk stops at the declared DACL size and access + * is denied. Before the fix the walk used the enclosing + * descriptor length, selected ACE #2 and returned 0. + */ + KUNIT_EXPECT_EQ(test, -EACCES, rc); +out: + fput(file); +} + +static void +ksmbd_smb_check_perm_dacl_maximal_boundary_test(struct kunit *test) +{ + struct file *file; + struct smb_ntsd *pntsd; + __le32 daccess = FILE_MAXIMAL_ACCESS_LE; + int ntsd_size, rc; + + /* + * The in-boundary ACE grants read access. The trailing ACE grants + * write access, which must not be included in the maximal access mask. + */ + pntsd = build_boundary_ntsd(test, &test_owner_sid, FILE_READ_DATA, + FILE_WRITE_DATA, &ntsd_size); + KUNIT_ASSERT_NOT_NULL(test, pntsd); + + file = shmem_file_setup("ksmbd-kunit-dacl-maximal", 0, + mk_vma_flags(VMA_NORESERVE_BIT)); + KUNIT_ASSERT_NOT_ERR_OR_NULL(test, file); + + rc = ksmbd_vfs_set_sd_xattr(NULL, mnt_idmap(file->f_path.mnt), + &file->f_path, pntsd, ntsd_size, + false); + KUNIT_EXPECT_EQ(test, 0, rc); + if (rc) + goto out; + + rc = smb_check_perm_dacl(NULL, &file->f_path, &daccess, + FILE_MAXIMAL_ACCESS_LE, 0, false); + KUNIT_EXPECT_EQ(test, 0, rc); + if (rc) + goto out; + + KUNIT_EXPECT_TRUE(test, le32_to_cpu(daccess) & FILE_READ_DATA); + KUNIT_EXPECT_FALSE(test, le32_to_cpu(daccess) & FILE_WRITE_DATA); +out: + fput(file); +} + +static struct kunit_case ksmbd_smbacl_test_cases[] = { + KUNIT_CASE(ksmbd_dacl_walk_must_stop_at_declared_size), + KUNIT_CASE(ksmbd_smb_check_perm_dacl_boundary_test), + KUNIT_CASE(ksmbd_smb_check_perm_dacl_maximal_boundary_test), + {} +}; + +static struct kunit_suite ksmbd_smbacl_test_suite = { + .name = "ksmbd-smbacl", + .test_cases = ksmbd_smbacl_test_cases, +}; + +kunit_test_suite(ksmbd_smbacl_test_suite); + +MODULE_DESCRIPTION("KUnit tests for ksmbd smbacl helpers"); +MODULE_LICENSE("GPL"); +MODULE_IMPORT_NS("EXPORTED_FOR_KUNIT_TESTING"); diff --git a/fs/smb/server/transport_tcp.c b/fs/smb/server/transport_tcp.c index 4968cfc1a572..488ca113b431 100644 --- a/fs/smb/server/transport_tcp.c +++ b/fs/smb/server/transport_tcp.c @@ -77,6 +77,7 @@ static struct tcp_transport *alloc_transport(struct socket *client_sk) if (client_sk->sk->sk_family == AF_INET6) { memcpy(&conn->inet6_addr, &client_sk->sk->sk_v6_daddr, 16); conn->inet_hash = ipv6_addr_hash(&client_sk->sk->sk_v6_daddr); + conn->is_ipv6 = true; } else { conn->inet_addr = inet_sk(client_sk->sk)->inet_daddr; conn->inet_hash = ipv4_addr_hash(inet_sk(client_sk->sk)->inet_daddr); @@ -656,10 +657,15 @@ static void ksmbd_tcp_stop_listener(struct interface *iface) void ksmbd_tcp_destroy(void) { struct interface *iface, *tmp; + LIST_HEAD(iface_list_to_free); unregister_netdevice_notifier(&ksmbd_netdev_notifier); - list_for_each_entry_safe(iface, tmp, &iface_list, entry) { + rtnl_lock(); + list_splice_init(&iface_list, &iface_list_to_free); + rtnl_unlock(); + + list_for_each_entry_safe(iface, tmp, &iface_list_to_free, entry) { ksmbd_tcp_stop_listener(iface); list_del(&iface->entry); kfree(iface->name); diff --git a/fs/smb/server/vfs.c b/fs/smb/server/vfs.c index c2c9aaa5de1b..2e2c554bc2c1 100644 --- a/fs/smb/server/vfs.c +++ b/fs/smb/server/vfs.c @@ -5,6 +5,7 @@ */ #include <crypto/sha2.h> +#include <kunit/visibility.h> #include <linux/kernel.h> #include <linux/fs.h> #include <linux/fs_struct.h> @@ -115,7 +116,7 @@ static int ksmbd_vfs_path_lookup(struct ksmbd_share_config *share_conf, return 0; } -void ksmbd_vfs_query_maximal_access(struct mnt_idmap *idmap, +void ksmbd_vfs_query_maximal_access(const struct mnt_idmap *idmap, struct dentry *dentry, __le32 *daccess) { *daccess = cpu_to_le32(FILE_READ_ATTRIBUTES | READ_CONTROL); @@ -183,7 +184,7 @@ int ksmbd_vfs_create(struct ksmbd_work *work, const char *name, umode_t mode) */ int ksmbd_vfs_mkdir(struct ksmbd_work *work, const char *name, umode_t mode) { - struct mnt_idmap *idmap; + const struct mnt_idmap *idmap; struct path path; struct dentry *dentry, *d; int err = 0; @@ -216,9 +217,9 @@ int ksmbd_vfs_mkdir(struct ksmbd_work *work, const char *name, umode_t mode) return err; } -static ssize_t ksmbd_vfs_getcasexattr(struct mnt_idmap *idmap, - struct dentry *dentry, char *attr_name, - int attr_name_len, char **attr_value) +ssize_t ksmbd_vfs_getcasexattr(const struct mnt_idmap *idmap, + struct dentry *dentry, char *attr_name, + int attr_name_len, char **attr_value) { char *name, *xattr_list = NULL; ssize_t value_len = -ENOENT, xattr_list_len; @@ -386,7 +387,7 @@ static int ksmbd_vfs_stream_write(struct ksmbd_file *fp, char *buf, loff_t *pos, { const struct cred *saved_cred; char *stream_buf = NULL, *wbuf; - struct mnt_idmap *idmap = file_mnt_idmap(fp->filp); + const struct mnt_idmap *idmap = file_mnt_idmap(fp->filp); size_t size; ssize_t v_len; int err = 0; @@ -577,7 +578,7 @@ int ksmbd_vfs_fsync(struct ksmbd_work *work, u64 fid, u64 p_id) */ int ksmbd_vfs_remove_file(struct ksmbd_work *work, const struct path *path) { - struct mnt_idmap *idmap; + const struct mnt_idmap *idmap; struct dentry *parent = path->dentry->d_parent; int err; @@ -841,8 +842,8 @@ ssize_t ksmbd_vfs_listxattr(struct dentry *dentry, char **list) return size; } -static ssize_t ksmbd_vfs_xattr_len(struct mnt_idmap *idmap, - struct dentry *dentry, char *xattr_name) +ssize_t ksmbd_vfs_xattr_len(const struct mnt_idmap *idmap, + struct dentry *dentry, char *xattr_name) { return vfs_getxattr(idmap, dentry, xattr_name, NULL, 0); } @@ -856,7 +857,7 @@ static ssize_t ksmbd_vfs_xattr_len(struct mnt_idmap *idmap, * * Return: read xattr value length on success, otherwise error */ -ssize_t ksmbd_vfs_getxattr(struct mnt_idmap *idmap, +ssize_t ksmbd_vfs_getxattr(const struct mnt_idmap *idmap, struct dentry *dentry, char *xattr_name, char **xattr_buf) { @@ -893,7 +894,7 @@ ssize_t ksmbd_vfs_getxattr(struct mnt_idmap *idmap, * * Return: 0 on success, otherwise error */ -int ksmbd_vfs_setxattr(struct mnt_idmap *idmap, +int ksmbd_vfs_setxattr(const struct mnt_idmap *idmap, const struct path *path, const char *attr_name, void *attr_value, size_t attr_size, int flags, bool get_write) @@ -1177,7 +1178,7 @@ int ksmbd_vfs_query_allocated_ranges(struct ksmbd_file *fp, loff_t start, return ret; } -int ksmbd_vfs_remove_xattr(struct mnt_idmap *idmap, +int ksmbd_vfs_remove_xattr(const struct mnt_idmap *idmap, const struct path *path, char *attr_name, bool get_write) { @@ -1202,7 +1203,7 @@ int ksmbd_vfs_unlink(struct file *filp) const struct cred *saved_cred; int err = 0; struct dentry *dir, *dentry = filp->f_path.dentry; - struct mnt_idmap *idmap = file_mnt_idmap(filp); + const struct mnt_idmap *idmap = file_mnt_idmap(filp); saved_cred = override_creds(filp->f_cred); err = mnt_want_write(filp->f_path.mnt); @@ -1471,7 +1472,7 @@ struct dentry *ksmbd_vfs_kern_path_create(struct ksmbd_work *work, return dent; } -int ksmbd_vfs_remove_acl_xattrs(struct mnt_idmap *idmap, +int ksmbd_vfs_remove_acl_xattrs(const struct mnt_idmap *idmap, const struct path *path) { char *name, *xattr_list = NULL; @@ -1511,7 +1512,7 @@ out: return err; } -int ksmbd_vfs_remove_sd_xattrs(struct mnt_idmap *idmap, const struct path *path) +int ksmbd_vfs_remove_sd_xattrs(const struct mnt_idmap *idmap, const struct path *path) { char *name, *xattr_list = NULL; ssize_t xattr_list_len; @@ -1540,7 +1541,7 @@ out: return err; } -static struct xattr_smb_acl *ksmbd_vfs_make_xattr_posix_acl(struct mnt_idmap *idmap, +static struct xattr_smb_acl *ksmbd_vfs_make_xattr_posix_acl(const struct mnt_idmap *idmap, struct inode *inode, int acl_type) { @@ -1606,7 +1607,7 @@ out: } int ksmbd_vfs_set_sd_xattr(struct ksmbd_conn *conn, - struct mnt_idmap *idmap, + const struct mnt_idmap *idmap, const struct path *path, struct smb_ntsd *pntsd, int len, bool get_write) @@ -1671,9 +1672,10 @@ out: kfree(def_smb_acl); return rc; } +EXPORT_SYMBOL_IF_KUNIT(ksmbd_vfs_set_sd_xattr); int ksmbd_vfs_get_sd_xattr(struct ksmbd_conn *conn, - struct mnt_idmap *idmap, + const struct mnt_idmap *idmap, struct dentry *dentry, struct smb_ntsd **pntsd) { @@ -1742,7 +1744,7 @@ out_free: return rc; } -int ksmbd_vfs_set_dos_attrib_xattr(struct mnt_idmap *idmap, +int ksmbd_vfs_set_dos_attrib_xattr(const struct mnt_idmap *idmap, const struct path *path, struct xattr_dos_attrib *da, bool get_write) @@ -1764,7 +1766,7 @@ out: return err; } -int ksmbd_vfs_get_dos_attrib_xattr(struct mnt_idmap *idmap, +int ksmbd_vfs_get_dos_attrib_xattr(const struct mnt_idmap *idmap, struct dentry *dentry, struct xattr_dos_attrib *da) { @@ -1820,7 +1822,7 @@ void *ksmbd_vfs_init_kstat(char **p, struct ksmbd_kstat *ksmbd_kstat) } int ksmbd_vfs_fill_dentry_attrs(struct ksmbd_work *work, - struct mnt_idmap *idmap, + const struct mnt_idmap *idmap, struct dentry *dentry, struct ksmbd_kstat *ksmbd_kstat) { @@ -1895,7 +1897,7 @@ int ksmbd_vfs_fill_dentry_attrs(struct ksmbd_work *work, return 0; } -ssize_t ksmbd_vfs_casexattr_len(struct mnt_idmap *idmap, +ssize_t ksmbd_vfs_casexattr_len(const struct mnt_idmap *idmap, struct dentry *dentry, char *attr_name, int attr_name_len) { @@ -2210,7 +2212,7 @@ void ksmbd_vfs_posix_lock_unblock(struct file_lock *flock) locks_delete_block(flock); } -int ksmbd_vfs_set_init_posix_acl(struct mnt_idmap *idmap, +int ksmbd_vfs_set_init_posix_acl(const struct mnt_idmap *idmap, const struct path *path) { struct posix_acl_state acl_state; @@ -2263,7 +2265,7 @@ int ksmbd_vfs_set_init_posix_acl(struct mnt_idmap *idmap, return rc; } -int ksmbd_vfs_inherit_posix_acl(struct mnt_idmap *idmap, +int ksmbd_vfs_inherit_posix_acl(const struct mnt_idmap *idmap, const struct path *path, struct inode *parent_inode) { struct posix_acl *acls; @@ -2326,7 +2328,7 @@ static int __ksmbd_vfs_set_compression(struct ksmbd_work *work, const struct cred *saved_cred = NULL; struct file_kattr fa; struct dentry *dentry = fp->filp->f_path.dentry; - struct mnt_idmap *idmap = file_mnt_idmap(fp->filp); + const struct mnt_idmap *idmap = file_mnt_idmap(fp->filp); u32 flags; __le32 old_fattr; int rc; diff --git a/fs/smb/server/vfs.h b/fs/smb/server/vfs.h index 55d099de71f5..216c76291fc1 100644 --- a/fs/smb/server/vfs.h +++ b/fs/smb/server/vfs.h @@ -74,7 +74,7 @@ struct ksmbd_kstat { }; int ksmbd_vfs_lock_parent(struct dentry *parent, struct dentry *child); -void ksmbd_vfs_query_maximal_access(struct mnt_idmap *idmap, +void ksmbd_vfs_query_maximal_access(const struct mnt_idmap *idmap, struct dentry *dentry, __le32 *daccess); int ksmbd_vfs_create(struct ksmbd_work *work, const char *name, umode_t mode); int ksmbd_vfs_mkdir(struct ksmbd_work *work, const char *name, umode_t mode); @@ -104,20 +104,25 @@ int ksmbd_vfs_copy_file_ranges(struct ksmbd_work *work, unsigned int *chunk_size_written, loff_t *total_size_written); ssize_t ksmbd_vfs_listxattr(struct dentry *dentry, char **list); -ssize_t ksmbd_vfs_getxattr(struct mnt_idmap *idmap, +ssize_t ksmbd_vfs_getxattr(const struct mnt_idmap *idmap, struct dentry *dentry, char *xattr_name, char **xattr_buf); -ssize_t ksmbd_vfs_casexattr_len(struct mnt_idmap *idmap, +ssize_t ksmbd_vfs_xattr_len(const struct mnt_idmap *idmap, + struct dentry *dentry, char *xattr_name); +ssize_t ksmbd_vfs_getcasexattr(const struct mnt_idmap *idmap, + struct dentry *dentry, char *attr_name, + int attr_name_len, char **attr_value); +ssize_t ksmbd_vfs_casexattr_len(const struct mnt_idmap *idmap, struct dentry *dentry, char *attr_name, int attr_name_len); -int ksmbd_vfs_setxattr(struct mnt_idmap *idmap, +int ksmbd_vfs_setxattr(const struct mnt_idmap *idmap, const struct path *path, const char *attr_name, void *attr_value, size_t attr_size, int flags, bool get_write); int ksmbd_vfs_xattr_stream_name(char *stream_name, char **xattr_stream_name, size_t *xattr_stream_name_size, int s_type); -int ksmbd_vfs_remove_xattr(struct mnt_idmap *idmap, +int ksmbd_vfs_remove_xattr(const struct mnt_idmap *idmap, const struct path *path, char *attr_name, bool get_write); int ksmbd_vfs_kern_path(struct ksmbd_work *work, char *name, @@ -147,33 +152,33 @@ int ksmbd_vfs_query_allocated_ranges(struct ksmbd_file *fp, loff_t start, int ksmbd_vfs_unlink(struct file *filp); void *ksmbd_vfs_init_kstat(char **p, struct ksmbd_kstat *ksmbd_kstat); int ksmbd_vfs_fill_dentry_attrs(struct ksmbd_work *work, - struct mnt_idmap *idmap, + const struct mnt_idmap *idmap, struct dentry *dentry, struct ksmbd_kstat *ksmbd_kstat); void ksmbd_vfs_posix_lock_wait(struct file_lock *flock); void ksmbd_vfs_posix_lock_unblock(struct file_lock *flock); -int ksmbd_vfs_remove_acl_xattrs(struct mnt_idmap *idmap, +int ksmbd_vfs_remove_acl_xattrs(const struct mnt_idmap *idmap, const struct path *path); -int ksmbd_vfs_remove_sd_xattrs(struct mnt_idmap *idmap, const struct path *path); +int ksmbd_vfs_remove_sd_xattrs(const struct mnt_idmap *idmap, const struct path *path); int ksmbd_vfs_set_sd_xattr(struct ksmbd_conn *conn, - struct mnt_idmap *idmap, + const struct mnt_idmap *idmap, const struct path *path, struct smb_ntsd *pntsd, int len, bool get_write); int ksmbd_vfs_get_sd_xattr(struct ksmbd_conn *conn, - struct mnt_idmap *idmap, + const struct mnt_idmap *idmap, struct dentry *dentry, struct smb_ntsd **pntsd); -int ksmbd_vfs_set_dos_attrib_xattr(struct mnt_idmap *idmap, +int ksmbd_vfs_set_dos_attrib_xattr(const struct mnt_idmap *idmap, const struct path *path, struct xattr_dos_attrib *da, bool get_write); -int ksmbd_vfs_get_dos_attrib_xattr(struct mnt_idmap *idmap, +int ksmbd_vfs_get_dos_attrib_xattr(const struct mnt_idmap *idmap, struct dentry *dentry, struct xattr_dos_attrib *da); -int ksmbd_vfs_set_init_posix_acl(struct mnt_idmap *idmap, +int ksmbd_vfs_set_init_posix_acl(const struct mnt_idmap *idmap, const struct path *path); -int ksmbd_vfs_inherit_posix_acl(struct mnt_idmap *idmap, +int ksmbd_vfs_inherit_posix_acl(const struct mnt_idmap *idmap, const struct path *path, struct inode *parent_inode); void ksmbd_vfs_update_compressed_fattr(struct dentry *dentry, __le32 *fattr); diff --git a/fs/smb/server/vfs_cache.c b/fs/smb/server/vfs_cache.c index fd2c595f0486..7bc95fae4ac0 100644 --- a/fs/smb/server/vfs_cache.c +++ b/fs/smb/server/vfs_cache.c @@ -617,7 +617,6 @@ static void __ksmbd_close_fd(struct ksmbd_file_table *ft, struct ksmbd_file *fp) { struct file *filp; struct ksmbd_lock *smb_lock, *tmp_lock; - struct ksmbd_work *cn_work; fd_limit_close(); ksmbd_remove_durable_fd(fp); @@ -653,52 +652,6 @@ static void __ksmbd_close_fd(struct ksmbd_file_table *ft, struct ksmbd_file *fp) } /* - * Complete any CHANGE_NOTIFY left pending on this handle now that - * it is closed. KSMBD never completes CHANGE_NOTIFY spontaneously - * (no real change-notification backend), only on close -- matching - * genuine SMB2/macOS smbfs semantics and avoiding the Finder - * "directory changed, re-enumerate everything" loop. - * - * smb2_notify() on another connection can be adding to - * notify_pendings under fp->f_lock at the same time this handle is - * closed, and a client-sent CANCEL can concurrently be racing to - * claim the same entry via smb2_notify_cancel_fn() (smb2pdu.c). - * Pop one entry at a time under the lock via list_del_init() rather - * than a bulk list_splice_init(): list_del_init() leaves the node - * self-linked ("empty"), which is what the cancel path checks under - * the same lock to tell whether it lost the race -- a bulk splice - * would instead relink every entry into a shared local list, so an - * entry claimed here would still read as "not empty" to a racing - * cancel_fn, and both sides could end up freeing the same work. - * ksmbd_conn_write() can sleep (it takes conn's write mutex), so it - * must not be called while fp->f_lock is held -- release the lock - * before processing each popped entry, then reacquire it for the - * next. - */ - for (;;) { - spin_lock(&fp->f_lock); - if (list_empty(&fp->notify_pendings)) { - spin_unlock(&fp->f_lock); - break; - } - cn_work = list_first_entry(&fp->notify_pendings, - struct ksmbd_work, notify_entry); - list_del_init(&cn_work->notify_entry); - spin_unlock(&fp->f_lock); - - ksmbd_conn_write(cn_work); - /* - * release_async_work() removes cn_work from - * conn->async_requests, frees cancel_argv, and releases+zeroes - * async_id -- all needed before ksmbd_free_work_struct(), which - * only releases async_id itself if still nonzero (i.e. if this - * hadn't already been done). - */ - release_async_work(cn_work); - ksmbd_free_work_struct(cn_work); - } - - /* * Drop fp's strong reference on conn (taken in ksmbd_open_fd() / * ksmbd_reopen_durable_fd()). Durable fps that reached the * scavenger have already had fp->conn cleared by session_fd_check(), @@ -1266,7 +1219,6 @@ struct ksmbd_file *ksmbd_open_fd(struct ksmbd_work *work, struct file *filp) INIT_LIST_HEAD(&fp->blocked_works); INIT_LIST_HEAD(&fp->node); INIT_LIST_HEAD(&fp->lock_list); - INIT_LIST_HEAD(&fp->notify_pendings); spin_lock_init(&fp->f_lock); mutex_init(&fp->readdir_lock); atomic_set(&fp->refcount, 1); @@ -1877,9 +1829,21 @@ int ksmbd_validate_name_reconnect(struct ksmbd_share_config *share, return -EACCES; } - if (name && strcmp(&ab_pathname[share->path_sz + 1], name)) { - ksmbd_debug(SMB, "invalid name reconnect %s\n", name); - ret = -EINVAL; + if (name) { + size_t len = strlen(ab_pathname); + + if (len == share->path_sz && !strncmp(ab_pathname, share->path, len)) { + /* the durable fp is the share root itself */ + if (name[0]) + ret = -EINVAL; + } else if (len <= share->path_sz || + strncmp(ab_pathname, share->path, share->path_sz) || + ab_pathname[share->path_sz] != '/' || + strcmp(&ab_pathname[share->path_sz + 1], name)) { + ret = -EINVAL; + } + if (ret) + ksmbd_debug(SMB, "invalid name reconnect %s\n", name); } kfree(pathname); diff --git a/fs/smb/server/vfs_cache.h b/fs/smb/server/vfs_cache.h index 1884f6deb9d0..732ae26dd6a7 100644 --- a/fs/smb/server/vfs_cache.h +++ b/fs/smb/server/vfs_cache.h @@ -162,12 +162,6 @@ struct ksmbd_file { unsigned int outstanding_requests; unsigned int outstanding_pre_requests; struct ksmbd_lock_sequence lock_seq[KSMBD_LOCK_SEQ_ARRAY_SIZE]; - - /* - * Pending CHANGE_NOTIFY completions for this handle, sent with - * STATUS_NOTIFY_CLEANUP when the handle is closed. - */ - struct list_head notify_pendings; }; static inline void set_ctx_actor(struct dir_context *ctx, diff --git a/fs/splice.c b/fs/splice.c index 9d8f63e2fd1a..bc243ab8dc43 100644 --- a/fs/splice.c +++ b/fs/splice.c @@ -177,9 +177,9 @@ static const struct pipe_buf_operations user_page_pipe_buf_ops = { static void wakeup_pipe_readers(struct pipe_inode_info *pipe) { - smp_mb(); - if (waitqueue_active(&pipe->rd_wait)) - wake_up_interruptible(&pipe->rd_wait); + if (wq_has_sleeper(&pipe->rd_wait)) + wake_up_interruptible_poll(&pipe->rd_wait, + EPOLLIN | EPOLLRDNORM); kill_fasync(&pipe->fasync_readers, SIGIO, POLL_IN); } @@ -413,9 +413,9 @@ EXPORT_SYMBOL(nosteal_pipe_buf_ops); static void wakeup_pipe_writers(struct pipe_inode_info *pipe) { - smp_mb(); - if (waitqueue_active(&pipe->wr_wait)) - wake_up_interruptible(&pipe->wr_wait); + if (wq_has_sleeper(&pipe->wr_wait)) + wake_up_interruptible_poll(&pipe->wr_wait, + EPOLLOUT | EPOLLWRNORM); kill_fasync(&pipe->fasync_writers, SIGIO, POLL_OUT); } @@ -1009,21 +1009,14 @@ ssize_t vfs_splice_read(struct file *in, loff_t *ppos, } EXPORT_SYMBOL_GPL(vfs_splice_read); -/** - * splice_direct_to_actor - splices data directly between two non-pipes - * @in: file to splice from - * @sd: actor information on where to splice to - * @actor: handles the data splicing - * - * Description: - * This is a special case helper to splice directly between two - * points, without requiring an explicit pipe. Internally an allocated - * pipe is cached in the process, and reused during the lifetime of - * that process. - * +/* + * This is a special case helper to splice directly between two + * points, without requiring an explicit pipe. Internally an allocated + * pipe is cached in the process, and reused during the lifetime of + * that process. */ -ssize_t splice_direct_to_actor(struct file *in, struct splice_desc *sd, - splice_direct_actor *actor) +static ssize_t splice_direct_to_actor(struct file *in, struct splice_desc *sd, + splice_direct_actor *actor) { struct pipe_inode_info *pipe; ssize_t ret, bytes; @@ -1147,7 +1140,42 @@ out_release: goto done; } -EXPORT_SYMBOL(splice_direct_to_actor); + +/** + * vfs_splice_to_actor - call an actor on data read from a file + * @in: file to read from + * @pos: file offset + * @count: maximum number of bytes to read + * @actor: callback to process a pipe's worth of data + * @private: private data passed to @actor + * + * Read up to @count worth of data from @in at @pos, and call @actor + * when the hidden pipe used to buffer the data is full. Ensures the + * read is allowed using rw_verify_area() and emits fsnotify access + * events. @in must be seekable (FMODE_LSEEK). + * + * Return: The number of bytes spliced, or a negative errno. + */ +ssize_t vfs_splice_to_actor(struct file *in, loff_t pos, size_t count, + splice_direct_actor *actor, void *private) +{ + struct splice_desc sd = { + .total_len = count, + .pos = pos, + .u.data = private, + }; + ssize_t ret; + + ret = rw_verify_area(READ, in, &sd.pos, sd.total_len); + if (ret < 0) + return ret; + + ret = splice_direct_to_actor(in, &sd, actor); + if (ret >= 0) + fsnotify_access(in); + return ret; +} +EXPORT_SYMBOL(vfs_splice_to_actor); static int direct_splice_actor(struct pipe_inode_info *pipe, struct splice_desc *sd) diff --git a/fs/stat.c b/fs/stat.c index c461c3054234..a9b7383d538d 100644 --- a/fs/stat.c +++ b/fs/stat.c @@ -79,7 +79,7 @@ EXPORT_SYMBOL(fill_mg_cmtime); * uid and gid filds. On non-idmapped mounts or if permission checking is to be * performed on the raw inode simply pass @nop_mnt_idmap. */ -void generic_fillattr(struct mnt_idmap *idmap, u32 request_mask, +void generic_fillattr(const struct mnt_idmap *idmap, u32 request_mask, struct inode *inode, struct kstat *stat) { vfsuid_t vfsuid = i_uid_into_vfsuid(idmap, inode); @@ -181,7 +181,7 @@ EXPORT_SYMBOL_GPL(generic_fill_statx_atomic_writes); int vfs_getattr_nosec(const struct path *path, struct kstat *stat, u32 request_mask, unsigned int query_flags) { - struct mnt_idmap *idmap; + const struct mnt_idmap *idmap; struct inode *inode = d_backing_inode(path->dentry); memset(stat, 0, sizeof(*stat)); diff --git a/fs/super.c b/fs/super.c index 1d5ccf540a9b..b1d4add11b77 100644 --- a/fs/super.c +++ b/fs/super.c @@ -1374,7 +1374,17 @@ static int test_single_super(struct super_block *s, struct fs_context *fc) return 1; } -static int vfs_get_super(struct fs_context *fc, +/** + * get_tree_super - Get a superblock, optionally sharing an existing one + * @fc: The filesystem context holding the parameters + * @test: Comparison function to find a matching existing superblock, or NULL + * @fill_super: Helper to initialise a new superblock + * + * If @test is non-NULL and matches an existing superblock, that superblock is + * reused; otherwise a new anonymous superblock is created and initialised with + * @fill_super. Passing NULL for @test always creates a new superblock. + */ +int get_tree_super(struct fs_context *fc, int (*test)(struct super_block *, struct fs_context *), int (*fill_super)(struct super_block *sb, struct fs_context *fc)) @@ -1401,12 +1411,13 @@ error: deactivate_locked_super(sb); return err; } +EXPORT_SYMBOL(get_tree_super); int get_tree_nodev(struct fs_context *fc, int (*fill_super)(struct super_block *sb, struct fs_context *fc)) { - return vfs_get_super(fc, NULL, fill_super); + return get_tree_super(fc, NULL, fill_super); } EXPORT_SYMBOL(get_tree_nodev); @@ -1414,7 +1425,7 @@ int get_tree_single(struct fs_context *fc, int (*fill_super)(struct super_block *sb, struct fs_context *fc)) { - return vfs_get_super(fc, test_single_super, fill_super); + return get_tree_super(fc, test_single_super, fill_super); } EXPORT_SYMBOL(get_tree_single); @@ -1424,7 +1435,7 @@ int get_tree_keyed(struct fs_context *fc, void *key) { fc->s_fs_info = key; - return vfs_get_super(fc, test_keyed_super, fill_super); + return get_tree_super(fc, test_keyed_super, fill_super); } EXPORT_SYMBOL(get_tree_keyed); diff --git a/fs/tests/.kunitconfig b/fs/tests/.kunitconfig new file mode 100644 index 000000000000..de67125a9421 --- /dev/null +++ b/fs/tests/.kunitconfig @@ -0,0 +1,2 @@ +CONFIG_KUNIT=y +CONFIG_FDTABLE_KUNIT_TEST=y diff --git a/fs/tests/fdtable_kunit.c b/fs/tests/fdtable_kunit.c new file mode 100644 index 000000000000..c5b028264557 --- /dev/null +++ b/fs/tests/fdtable_kunit.c @@ -0,0 +1,72 @@ +// SPDX-License-Identifier: GPL-2.0-only +#include <kunit/test.h> +#include <linux/fdtable.h> +#include <linux/file.h> + +static void test_alloc_fdtable(struct kunit *test) +{ + struct fdtable *fdt; + unsigned int slots = 64; + + fdt = alloc_fdtable(slots); + KUNIT_ASSERT_NOT_ERR_OR_NULL(test, fdt); + + /* Check that max_fds is set correctly and is >= slots */ + KUNIT_EXPECT_GE(test, fdt->max_fds, slots); + + /* Check that fd is allocated */ + KUNIT_ASSERT_NOT_ERR_OR_NULL(test, fdt->fd); + + /* + * Check dynamic object size of fdt->fd if compiler supports + * __counted_by_ptr. + */ +#ifdef CONFIG_CC_HAS_COUNTED_BY_PTR + KUNIT_EXPECT_EQ(test, __struct_size(fdt->fd), + fdt->max_fds * sizeof(struct file *)); +#endif + + __free_fdtable(fdt); +} + +static void test_dup_fd(struct kunit *test) +{ + struct files_struct *newf; + struct fdtable *fdt; + + newf = dup_fd(&init_files, NULL); + KUNIT_ASSERT_NOT_ERR_OR_NULL(test, newf); + + fdt = rcu_dereference_raw(newf->fdt); + KUNIT_ASSERT_NOT_ERR_OR_NULL(test, fdt); + + /* Check that max_fds is set correctly and is >= NR_OPEN_DEFAULT */ + KUNIT_EXPECT_GE(test, fdt->max_fds, NR_OPEN_DEFAULT); + + /* Check that fd is allocated */ + KUNIT_ASSERT_NOT_ERR_OR_NULL(test, fdt->fd); + + /* + * Check dynamic object size of fdt->fd if compiler supports + * __counted_by_ptr. + */ +#ifdef CONFIG_CC_HAS_COUNTED_BY_PTR + KUNIT_EXPECT_EQ(test, __struct_size(fdt->fd), + fdt->max_fds * sizeof(struct file *)); +#endif + + put_files_struct(newf); +} + +static struct kunit_case fdtable_test_cases[] = { + KUNIT_CASE(test_alloc_fdtable), + KUNIT_CASE(test_dup_fd), + {} +}; + +static struct kunit_suite fdtable_test_suite = { + .name = "fdtable", + .test_cases = fdtable_test_cases, +}; + +kunit_test_suite(fdtable_test_suite); diff --git a/fs/tracefs/event_inode.c b/fs/tracefs/event_inode.c index 6e3513b13cfa..6f6daac88621 100644 --- a/fs/tracefs/event_inode.c +++ b/fs/tracefs/event_inode.c @@ -182,7 +182,7 @@ static void update_attr(struct eventfs_attr *attr, struct iattr *iattr) } } -static int eventfs_set_attr(struct mnt_idmap *idmap, struct dentry *dentry, +static int eventfs_set_attr(const struct mnt_idmap *idmap, struct dentry *dentry, struct iattr *iattr) { const struct eventfs_entry *entry; diff --git a/fs/tracefs/inode.c b/fs/tracefs/inode.c index f3d6188a3b7b..5020365ca704 100644 --- a/fs/tracefs/inode.c +++ b/fs/tracefs/inode.c @@ -94,7 +94,7 @@ static struct tracefs_dir_ops { int (*rmdir)(const char *name); } tracefs_ops __ro_after_init; -static struct dentry *tracefs_syscall_mkdir(struct mnt_idmap *idmap, +static struct dentry *tracefs_syscall_mkdir(const struct mnt_idmap *idmap, struct inode *inode, struct dentry *dentry, umode_t mode) { @@ -189,14 +189,14 @@ static void set_tracefs_inode_owner(struct inode *inode) inode->i_gid = gid; } -static int tracefs_permission(struct mnt_idmap *idmap, +static int tracefs_permission(const struct mnt_idmap *idmap, struct inode *inode, int mask) { set_tracefs_inode_owner(inode); return generic_permission(idmap, inode, mask); } -static int tracefs_getattr(struct mnt_idmap *idmap, +static int tracefs_getattr(const struct mnt_idmap *idmap, const struct path *path, struct kstat *stat, u32 request_mask, unsigned int flags) { @@ -207,7 +207,7 @@ static int tracefs_getattr(struct mnt_idmap *idmap, return 0; } -static int tracefs_setattr(struct mnt_idmap *idmap, struct dentry *dentry, +static int tracefs_setattr(const struct mnt_idmap *idmap, struct dentry *dentry, struct iattr *attr) { unsigned int ia_valid = attr->ia_valid; diff --git a/fs/ubifs/dir.c b/fs/ubifs/dir.c index 23ec924162d6..c67954f6bee0 100644 --- a/fs/ubifs/dir.c +++ b/fs/ubifs/dir.c @@ -302,7 +302,7 @@ static int ubifs_prepare_create(struct inode *dir, struct dentry *dentry, return fscrypt_setup_filename(dir, &dentry->d_name, 0, nm); } -static int ubifs_create(struct mnt_idmap *idmap, struct inode *dir, +static int ubifs_create(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *dentry, umode_t mode) { struct inode *inode; @@ -440,7 +440,7 @@ static void unlock_2_inodes(struct inode *inode1, struct inode *inode2) mutex_unlock(&ubifs_inode(inode1)->ui_mutex); } -static int ubifs_tmpfile(struct mnt_idmap *idmap, struct inode *dir, +static int ubifs_tmpfile(const struct mnt_idmap *idmap, struct inode *dir, struct file *file, umode_t mode) { struct dentry *dentry = file->f_path.dentry; @@ -1002,7 +1002,7 @@ out_fname: return err; } -static struct dentry *ubifs_mkdir(struct mnt_idmap *idmap, struct inode *dir, +static struct dentry *ubifs_mkdir(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *dentry, umode_t mode) { struct inode *inode; @@ -1077,7 +1077,7 @@ out_budg: return ERR_PTR(err); } -static int ubifs_mknod(struct mnt_idmap *idmap, struct inode *dir, +static int ubifs_mknod(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *dentry, umode_t mode, dev_t rdev) { struct inode *inode; @@ -1170,7 +1170,7 @@ out_budg: return err; } -static int ubifs_symlink(struct mnt_idmap *idmap, struct inode *dir, +static int ubifs_symlink(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *dentry, const char *symname) { struct inode *inode; @@ -1642,7 +1642,7 @@ out: return err; } -static int ubifs_rename(struct mnt_idmap *idmap, +static int ubifs_rename(const struct mnt_idmap *idmap, struct inode *old_dir, struct dentry *old_dentry, struct inode *new_dir, struct dentry *new_dentry, unsigned int flags) @@ -1667,7 +1667,7 @@ static int ubifs_rename(struct mnt_idmap *idmap, return do_rename(old_dir, old_dentry, new_dir, new_dentry, flags); } -int ubifs_getattr(struct mnt_idmap *idmap, const struct path *path, +int ubifs_getattr(const struct mnt_idmap *idmap, const struct path *path, struct kstat *stat, u32 request_mask, unsigned int flags) { loff_t size; diff --git a/fs/ubifs/file.c b/fs/ubifs/file.c index aa0298ce451e..99dd52715eee 100644 --- a/fs/ubifs/file.c +++ b/fs/ubifs/file.c @@ -1251,7 +1251,7 @@ static int do_setattr(struct ubifs_info *c, struct inode *inode, return err; } -int ubifs_setattr(struct mnt_idmap *idmap, struct dentry *dentry, +int ubifs_setattr(const struct mnt_idmap *idmap, struct dentry *dentry, struct iattr *attr) { int err; @@ -1611,7 +1611,7 @@ static const char *ubifs_get_link(struct dentry *dentry, return fscrypt_get_symlink(inode, ui->data, ui->data_len, done); } -static int ubifs_symlink_getattr(struct mnt_idmap *idmap, +static int ubifs_symlink_getattr(const struct mnt_idmap *idmap, const struct path *path, struct kstat *stat, u32 request_mask, unsigned int query_flags) { diff --git a/fs/ubifs/ioctl.c b/fs/ubifs/ioctl.c index 79536b2e3d7a..5c34f895bd4e 100644 --- a/fs/ubifs/ioctl.c +++ b/fs/ubifs/ioctl.c @@ -144,7 +144,7 @@ int ubifs_fileattr_get(struct dentry *dentry, struct file_kattr *fa) return 0; } -int ubifs_fileattr_set(struct mnt_idmap *idmap, +int ubifs_fileattr_set(const struct mnt_idmap *idmap, struct dentry *dentry, struct file_kattr *fa) { struct inode *inode = d_inode(dentry); diff --git a/fs/ubifs/ubifs.h b/fs/ubifs/ubifs.h index 00db0d19a85e..b85b45a0564b 100644 --- a/fs/ubifs/ubifs.h +++ b/fs/ubifs/ubifs.h @@ -2020,7 +2020,7 @@ int ubifs_calc_dark(const struct ubifs_info *c, int spc); /* file.c */ int ubifs_fsync(struct file *file, loff_t start, loff_t end, int datasync); -int ubifs_setattr(struct mnt_idmap *idmap, struct dentry *dentry, +int ubifs_setattr(const struct mnt_idmap *idmap, struct dentry *dentry, struct iattr *attr); int ubifs_update_time(struct inode *inode, enum fs_update_time type, unsigned int flags); @@ -2028,7 +2028,7 @@ int ubifs_update_time(struct inode *inode, enum fs_update_time type, /* dir.c */ struct inode *ubifs_new_inode(struct ubifs_info *c, struct inode *dir, umode_t mode, bool is_xattr); -int ubifs_getattr(struct mnt_idmap *idmap, const struct path *path, +int ubifs_getattr(const struct mnt_idmap *idmap, const struct path *path, struct kstat *stat, u32 request_mask, unsigned int flags); int ubifs_check_dir_empty(struct inode *dir); @@ -2083,7 +2083,7 @@ void ubifs_destroy_size_tree(struct ubifs_info *c); /* ioctl.c */ int ubifs_fileattr_get(struct dentry *dentry, struct file_kattr *fa); -int ubifs_fileattr_set(struct mnt_idmap *idmap, +int ubifs_fileattr_set(const struct mnt_idmap *idmap, struct dentry *dentry, struct file_kattr *fa); long ubifs_ioctl(struct file *file, unsigned int cmd, unsigned long arg); void ubifs_set_inode_flags(struct inode *inode); diff --git a/fs/ubifs/xattr.c b/fs/ubifs/xattr.c index b5a9ab9d8a10..3b0e8a270f59 100644 --- a/fs/ubifs/xattr.c +++ b/fs/ubifs/xattr.c @@ -660,7 +660,7 @@ static int xattr_get(const struct xattr_handler *handler, } static int xattr_set(const struct xattr_handler *handler, - struct mnt_idmap *idmap, + const struct mnt_idmap *idmap, struct dentry *dentry, struct inode *inode, const char *name, const void *value, size_t size, int flags) diff --git a/fs/udf/file.c b/fs/udf/file.c index 57d11606a2a7..02e9314818dc 100644 --- a/fs/udf/file.c +++ b/fs/udf/file.c @@ -212,7 +212,7 @@ const struct file_operations udf_file_operations = { .setlease = generic_setlease, }; -static int udf_setattr(struct mnt_idmap *idmap, struct dentry *dentry, +static int udf_setattr(const struct mnt_idmap *idmap, struct dentry *dentry, struct iattr *attr) { struct inode *inode = d_inode(dentry); diff --git a/fs/udf/inode.c b/fs/udf/inode.c index e45e546a739a..71386e7ac796 100644 --- a/fs/udf/inode.c +++ b/fs/udf/inode.c @@ -442,6 +442,13 @@ int udf_expand_file_adinicb(struct inode *inode) err = udf_map_block(inode, &map); if (err < 0) goto restore; + /* + * The block may have held metadata that is still dirty in the block + * device page cache (e.g. the file entry of a deleted inode). Make + * sure writeback of that buffer cannot overwrite our data. + */ + if (map.oflags & UDF_BLK_NEW) + clean_bdev_aliases(inode->i_sb->s_bdev, map.pblk, 1); folio_mark_dirty(folio); folio_unlock(folio); @@ -1337,6 +1344,26 @@ update_time: } /* + * Verify validity of struct deviceSpec on disk. udf_get_extendedattr() has + * already verified the generic header and made sure attribute fits in the + * inode so we just have to make sure attribute space is large enough for + * deviceSpec struct and required impUse information. + */ +static bool udf_device_spec_valid(struct deviceSpec *dsea) +{ + u32 attr_length, imp_use_length; + + attr_length = le32_to_cpu(dsea->attrLength); + imp_use_length = le32_to_cpu(dsea->impUseLength); + if (attr_length < sizeof(struct deviceSpec) || + imp_use_length < sizeof(struct regid) || + imp_use_length > attr_length - sizeof(struct deviceSpec)) + return false; + + return true; +} + +/* * Maximum length of linked list formed by ICB hierarchy. The chosen number is * arbitrary - just that we hopefully don't limit any real use of rewritten * inode on write-once media but avoid looping for too long on corrupted media. @@ -1654,13 +1681,19 @@ reread: if (S_ISCHR(inode->i_mode) || S_ISBLK(inode->i_mode)) { struct deviceSpec *dsea = (struct deviceSpec *)udf_get_extendedattr(inode, 12, 1); - if (dsea) { - init_special_inode(inode, inode->i_mode, + + if (IS_ERR(dsea)) { + ret = PTR_ERR(dsea); + goto out; + } + /* Device inodes must have a device spec attribute */ + if (!dsea || !udf_device_spec_valid(dsea)) { + ret = -EFSCORRUPTED; + goto out; + } + init_special_inode(inode, inode->i_mode, MKDEV(le32_to_cpu(dsea->majorDeviceIdent), le32_to_cpu(dsea->minorDeviceIdent))); - /* Developer ID ??? */ - } else - goto out; } ret = 0; out: @@ -1757,6 +1790,7 @@ int udf_write_inode(struct inode *inode, struct writeback_control *wbc) struct udf_sb_info *sbi = UDF_SB(inode->i_sb); unsigned char blocksize_bits = inode->i_sb->s_blocksize_bits; struct udf_inode_info *iinfo = UDF_I(inode); + int err; bh = sb_getblk(inode->i_sb, udf_get_lb_pblock(inode->i_sb, &iinfo->i_location, 0)); @@ -1816,11 +1850,21 @@ int udf_write_inode(struct inode *inode, struct writeback_control *wbc) struct regid *eid; struct deviceSpec *dsea = (struct deviceSpec *)udf_get_extendedattr(inode, 12, 1); + + /* Validity of extended attrs was checked on load */ + if (WARN_ON_ONCE(IS_ERR(dsea))) { + err = PTR_ERR(dsea); + goto out_unlock; + } if (!dsea) { dsea = (struct deviceSpec *) udf_add_extendedattr(inode, sizeof(struct deviceSpec) + sizeof(struct regid), 12, 0x3); + if (IS_ERR(dsea)) { + err = PTR_ERR(dsea); + goto out_unlock; + } dsea->attrType = cpu_to_le32(12); dsea->attrSubtype = 1; dsea->attrLength = cpu_to_le32( @@ -1962,6 +2006,11 @@ finish: set_inode_metadata_writeback(inode); return 0; + +out_unlock: + unlock_buffer(bh); + brelse(bh); + return err; } struct inode *__udf_iget(struct super_block *sb, struct kernel_lb_addr *ino, diff --git a/fs/udf/misc.c b/fs/udf/misc.c index 6928e378fbbd..a2084dfbfbd6 100644 --- a/fs/udf/misc.c +++ b/fs/udf/misc.c @@ -58,7 +58,7 @@ struct genericFormat *udf_add_extendedattr(struct inode *inode, uint32_t size, cpu_to_le16(TAG_IDENT_EAHD) || le32_to_cpu(eahd->descTag.tagLocation) != iinfo->i_location.logicalBlockNum) - return NULL; + return ERR_PTR(-EFSCORRUPTED); } else { struct udf_sb_info *sbi = UDF_SB(inode->i_sb); @@ -122,7 +122,7 @@ struct genericFormat *udf_add_extendedattr(struct inode *inode, uint32_t size, return (struct genericFormat *)&ea[offset]; } - return NULL; + return ERR_PTR(-ENOSPC); } struct genericFormat *udf_get_extendedattr(struct inode *inode, uint32_t type, @@ -144,7 +144,7 @@ struct genericFormat *udf_get_extendedattr(struct inode *inode, uint32_t type, cpu_to_le16(TAG_IDENT_EAHD) || le32_to_cpu(eahd->descTag.tagLocation) != iinfo->i_location.logicalBlockNum) - return NULL; + return ERR_PTR(-EFSCORRUPTED); if (type < 2048) offset = sizeof(struct extendedAttrHeaderDesc); @@ -153,16 +153,17 @@ struct genericFormat *udf_get_extendedattr(struct inode *inode, uint32_t type, else offset = le32_to_cpu(eahd->appAttrLocation); - while (offset + sizeof(*gaf) < iinfo->i_lenEAttr) { + while (offset < + iinfo->i_lenEAttr - sizeof(struct genericFormat)) { uint32_t attrLength; gaf = (struct genericFormat *)&ea[offset]; attrLength = le32_to_cpu(gaf->attrLength); /* Detect undersized elements and buffer overflows */ - if ((attrLength < sizeof(*gaf)) || - (attrLength > (iinfo->i_lenEAttr - offset))) - break; + if (attrLength < sizeof(struct genericFormat) || + attrLength > iinfo->i_lenEAttr - offset) + return ERR_PTR(-EFSCORRUPTED); if (le32_to_cpu(gaf->attrType) == type && gaf->attrSubtype == subtype) diff --git a/fs/udf/namei.c b/fs/udf/namei.c index b90841ac0a40..42d000fe9d6b 100644 --- a/fs/udf/namei.c +++ b/fs/udf/namei.c @@ -370,7 +370,7 @@ static int udf_add_nondir(struct dentry *dentry, struct inode *inode) return 0; } -static int udf_create(struct mnt_idmap *idmap, struct inode *dir, +static int udf_create(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *dentry, umode_t mode) { struct inode *inode = udf_new_inode(dir, mode); @@ -386,7 +386,7 @@ static int udf_create(struct mnt_idmap *idmap, struct inode *dir, return udf_add_nondir(dentry, inode); } -static int udf_tmpfile(struct mnt_idmap *idmap, struct inode *dir, +static int udf_tmpfile(const struct mnt_idmap *idmap, struct inode *dir, struct file *file, umode_t mode) { struct inode *inode = udf_new_inode(dir, mode); @@ -403,7 +403,7 @@ static int udf_tmpfile(struct mnt_idmap *idmap, struct inode *dir, return finish_open_simple(file, 0); } -static int udf_mknod(struct mnt_idmap *idmap, struct inode *dir, +static int udf_mknod(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *dentry, umode_t mode, dev_t rdev) { struct inode *inode; @@ -419,7 +419,7 @@ static int udf_mknod(struct mnt_idmap *idmap, struct inode *dir, return udf_add_nondir(dentry, inode); } -static struct dentry *udf_mkdir(struct mnt_idmap *idmap, struct inode *dir, +static struct dentry *udf_mkdir(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *dentry, umode_t mode) { struct inode *inode; @@ -567,7 +567,7 @@ out: return ret; } -static int udf_symlink(struct mnt_idmap *idmap, struct inode *dir, +static int udf_symlink(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *dentry, const char *symname) { struct inode *inode; @@ -762,7 +762,7 @@ static int udf_link(struct dentry *old_dentry, struct inode *dir, /* Anybody can rename anything with this: the permission checks are left to the * higher-level routines. */ -static int udf_rename(struct mnt_idmap *idmap, struct inode *old_dir, +static int udf_rename(const struct mnt_idmap *idmap, struct inode *old_dir, struct dentry *old_dentry, struct inode *new_dir, struct dentry *new_dentry, unsigned int flags) { diff --git a/fs/udf/super.c b/fs/udf/super.c index 2ba5973ef4dd..5351755aca3e 100644 --- a/fs/udf/super.c +++ b/fs/udf/super.c @@ -1617,7 +1617,10 @@ static bool udf_lvid_valid(struct super_block *sb, parts = le32_to_cpu(lvid->numOfPartitions); impuselen = le32_to_cpu(lvid->lengthOfImpUse); - if (parts >= sb->s_blocksize || impuselen >= sb->s_blocksize || + if (sizeof(struct logicalVolIntegrityDescImpUse) > impuselen || + impuselen >= sb->s_blocksize) + return false; + if (parts >= sb->s_blocksize || sizeof(struct logicalVolIntegrityDesc) + impuselen + 2 * parts * sizeof(u32) > sb->s_blocksize) return false; diff --git a/fs/udf/symlink.c b/fs/udf/symlink.c index a05d1888a2ba..df41bf5a05b1 100644 --- a/fs/udf/symlink.c +++ b/fs/udf/symlink.c @@ -133,7 +133,7 @@ out: return err; } -static int udf_symlink_getattr(struct mnt_idmap *idmap, +static int udf_symlink_getattr(const struct mnt_idmap *idmap, const struct path *path, struct kstat *stat, u32 request_mask, unsigned int flags) { diff --git a/fs/ufs/dir.c b/fs/ufs/dir.c index ce43cf20b07c..e96174b738b4 100644 --- a/fs/ufs/dir.c +++ b/fs/ufs/dir.c @@ -213,7 +213,7 @@ fail: static unsigned ufs_last_byte(struct inode *inode, unsigned long page_nr) { - unsigned last_byte = inode->i_size; + u64 last_byte = inode->i_size; last_byte -= page_nr << PAGE_SHIFT; if (last_byte > PAGE_SIZE) diff --git a/fs/ufs/inode.c b/fs/ufs/inode.c index 440d014cc5ed..c9ff8673fa66 100644 --- a/fs/ufs/inode.c +++ b/fs/ufs/inode.c @@ -1195,7 +1195,7 @@ out: return err; } -int ufs_setattr(struct mnt_idmap *idmap, struct dentry *dentry, +int ufs_setattr(const struct mnt_idmap *idmap, struct dentry *dentry, struct iattr *attr) { struct inode *inode = d_inode(dentry); diff --git a/fs/ufs/namei.c b/fs/ufs/namei.c index 6703f3bcf76f..d45347a25741 100644 --- a/fs/ufs/namei.c +++ b/fs/ufs/namei.c @@ -69,7 +69,7 @@ static struct dentry *ufs_lookup(struct inode * dir, struct dentry *dentry, unsi * If the create succeeds, we fill in the inode information * with d_instantiate(). */ -static int ufs_create (struct mnt_idmap * idmap, +static int ufs_create (const struct mnt_idmap * idmap, struct inode * dir, struct dentry * dentry, umode_t mode) { struct inode *inode; @@ -85,7 +85,7 @@ static int ufs_create (struct mnt_idmap * idmap, return ufs_add_nondir(dentry, inode); } -static int ufs_mknod(struct mnt_idmap *idmap, struct inode *dir, +static int ufs_mknod(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *dentry, umode_t mode, dev_t rdev) { struct inode *inode; @@ -105,7 +105,7 @@ static int ufs_mknod(struct mnt_idmap *idmap, struct inode *dir, return err; } -static int ufs_symlink (struct mnt_idmap * idmap, struct inode * dir, +static int ufs_symlink (const struct mnt_idmap * idmap, struct inode * dir, struct dentry * dentry, const char * symname) { struct super_block * sb = dir->i_sb; @@ -165,7 +165,7 @@ static int ufs_link (struct dentry * old_dentry, struct inode * dir, return error; } -static struct dentry *ufs_mkdir(struct mnt_idmap * idmap, struct inode * dir, +static struct dentry *ufs_mkdir(const struct mnt_idmap * idmap, struct inode * dir, struct dentry * dentry, umode_t mode) { struct inode * inode; @@ -240,7 +240,7 @@ static int ufs_rmdir (struct inode * dir, struct dentry *dentry) return err; } -static int ufs_rename(struct mnt_idmap *idmap, struct inode *old_dir, +static int ufs_rename(const struct mnt_idmap *idmap, struct inode *old_dir, struct dentry *old_dentry, struct inode *new_dir, struct dentry *new_dentry, unsigned int flags) { diff --git a/fs/ufs/ufs.h b/fs/ufs/ufs.h index 788e025056b2..541566f5b5fb 100644 --- a/fs/ufs/ufs.h +++ b/fs/ufs/ufs.h @@ -120,7 +120,7 @@ extern struct inode *ufs_iget(struct super_block *, unsigned long); extern int ufs_write_inode (struct inode *, struct writeback_control *); extern int ufs_sync_inode (struct inode *); extern void ufs_evict_inode (struct inode *); -extern int ufs_setattr(struct mnt_idmap *idmap, struct dentry *dentry, +extern int ufs_setattr(const struct mnt_idmap *idmap, struct dentry *dentry, struct iattr *attr); /* namei.c */ diff --git a/fs/vboxsf/dir.c b/fs/vboxsf/dir.c index 0b9eab157432..b9c46c769c4e 100644 --- a/fs/vboxsf/dir.c +++ b/fs/vboxsf/dir.c @@ -296,14 +296,14 @@ out: return err; } -static int vboxsf_dir_mkfile(struct mnt_idmap *idmap, +static int vboxsf_dir_mkfile(const struct mnt_idmap *idmap, struct inode *parent, struct dentry *dentry, umode_t mode) { return vboxsf_dir_create(parent, dentry, mode, false, true, NULL); } -static struct dentry *vboxsf_dir_mkdir(struct mnt_idmap *idmap, +static struct dentry *vboxsf_dir_mkdir(const struct mnt_idmap *idmap, struct inode *parent, struct dentry *dentry, umode_t mode) { @@ -382,7 +382,7 @@ static int vboxsf_dir_unlink(struct inode *parent, struct dentry *dentry) return 0; } -static int vboxsf_dir_rename(struct mnt_idmap *idmap, +static int vboxsf_dir_rename(const struct mnt_idmap *idmap, struct inode *old_parent, struct dentry *old_dentry, struct inode *new_parent, @@ -425,7 +425,7 @@ err_put_old_path: return err; } -static int vboxsf_dir_symlink(struct mnt_idmap *idmap, +static int vboxsf_dir_symlink(const struct mnt_idmap *idmap, struct inode *parent, struct dentry *dentry, const char *symname) { diff --git a/fs/vboxsf/utils.c b/fs/vboxsf/utils.c index 298bfc93255c..8775fbee1ce6 100644 --- a/fs/vboxsf/utils.c +++ b/fs/vboxsf/utils.c @@ -233,7 +233,7 @@ int vboxsf_inode_revalidate(struct dentry *dentry) return 0; } -int vboxsf_getattr(struct mnt_idmap *idmap, const struct path *path, +int vboxsf_getattr(const struct mnt_idmap *idmap, const struct path *path, struct kstat *kstat, u32 request_mask, unsigned int flags) { int err; @@ -258,7 +258,7 @@ int vboxsf_getattr(struct mnt_idmap *idmap, const struct path *path, return 0; } -int vboxsf_setattr(struct mnt_idmap *idmap, struct dentry *dentry, +int vboxsf_setattr(const struct mnt_idmap *idmap, struct dentry *dentry, struct iattr *iattr) { struct vboxsf_inode *sf_i = VBOXSF_I(d_inode(dentry)); diff --git a/fs/vboxsf/vfsmod.h b/fs/vboxsf/vfsmod.h index b61afd0ce842..59a4e44c4005 100644 --- a/fs/vboxsf/vfsmod.h +++ b/fs/vboxsf/vfsmod.h @@ -98,10 +98,10 @@ int vboxsf_stat(struct vboxsf_sbi *sbi, struct shfl_string *path, struct shfl_fsobjinfo *info); int vboxsf_stat_dentry(struct dentry *dentry, struct shfl_fsobjinfo *info); int vboxsf_inode_revalidate(struct dentry *dentry); -int vboxsf_getattr(struct mnt_idmap *idmap, const struct path *path, +int vboxsf_getattr(const struct mnt_idmap *idmap, const struct path *path, struct kstat *kstat, u32 request_mask, unsigned int query_flags); -int vboxsf_setattr(struct mnt_idmap *idmap, struct dentry *dentry, +int vboxsf_setattr(const struct mnt_idmap *idmap, struct dentry *dentry, struct iattr *iattr); struct shfl_string *vboxsf_path_from_dentry(struct vboxsf_sbi *sbi, struct dentry *dentry); diff --git a/fs/xattr.c b/fs/xattr.c index d58979115200..d9f035610f0b 100644 --- a/fs/xattr.c +++ b/fs/xattr.c @@ -100,7 +100,7 @@ xattr_resolve_name(struct inode *inode, const char **name) * * Return: On success zero is returned. On error a negative errno is returned. */ -int may_write_xattr(struct mnt_idmap *idmap, struct inode *inode) +int may_write_xattr(const struct mnt_idmap *idmap, struct inode *inode) { if (IS_IMMUTABLE(inode)) return -EPERM; @@ -123,7 +123,7 @@ static inline int xattr_permission_error(int mask) * because different namespaces have very different rules. */ static int -xattr_permission(struct mnt_idmap *idmap, struct inode *inode, +xattr_permission(const struct mnt_idmap *idmap, struct inode *inode, const char *name, int mask) { if (mask & MAY_WRITE) { @@ -204,7 +204,7 @@ xattr_supports_user_prefix(struct inode *inode) EXPORT_SYMBOL(xattr_supports_user_prefix); int -__vfs_setxattr(struct mnt_idmap *idmap, struct dentry *dentry, +__vfs_setxattr(const struct mnt_idmap *idmap, struct dentry *dentry, struct inode *inode, const char *name, const void *value, size_t size, int flags) { @@ -242,7 +242,7 @@ EXPORT_SYMBOL(__vfs_setxattr); * is executed. It also assumes that the caller will make the appropriate * permission checks. */ -int __vfs_setxattr_noperm(struct mnt_idmap *idmap, +int __vfs_setxattr_noperm(const struct mnt_idmap *idmap, struct dentry *dentry, const char *name, const void *value, size_t size, int flags) { @@ -295,7 +295,7 @@ int __vfs_setxattr_noperm(struct mnt_idmap *idmap, * a delegation was broken on, NULL if none. */ int -__vfs_setxattr_locked(struct mnt_idmap *idmap, struct dentry *dentry, +__vfs_setxattr_locked(const struct mnt_idmap *idmap, struct dentry *dentry, const char *name, const void *value, size_t size, int flags, struct delegated_inode *delegated_inode) { @@ -324,7 +324,7 @@ out: EXPORT_SYMBOL_GPL(__vfs_setxattr_locked); int -vfs_setxattr(struct mnt_idmap *idmap, struct dentry *dentry, +vfs_setxattr(const struct mnt_idmap *idmap, struct dentry *dentry, const char *name, const void *value, size_t size, int flags) { struct inode *inode = dentry->d_inode; @@ -358,7 +358,7 @@ retry_deleg: EXPORT_SYMBOL_GPL(vfs_setxattr); static ssize_t -xattr_getsecurity(struct mnt_idmap *idmap, struct inode *inode, +xattr_getsecurity(const struct mnt_idmap *idmap, struct inode *inode, const char *name, void *value, size_t size) { void *buffer = NULL; @@ -395,7 +395,7 @@ out_noalloc: * Returns the result of alloc, if failed, or the getxattr operation. */ int -vfs_getxattr_alloc(struct mnt_idmap *idmap, struct dentry *dentry, +vfs_getxattr_alloc(const struct mnt_idmap *idmap, struct dentry *dentry, const char *name, char **xattr_value, size_t xattr_size, gfp_t flags) { @@ -448,7 +448,7 @@ __vfs_getxattr(struct dentry *dentry, struct inode *inode, const char *name, EXPORT_SYMBOL(__vfs_getxattr); ssize_t -vfs_getxattr(struct mnt_idmap *idmap, struct dentry *dentry, +vfs_getxattr(const struct mnt_idmap *idmap, struct dentry *dentry, const char *name, void *value, size_t size) { struct inode *inode = dentry->d_inode; @@ -527,7 +527,7 @@ vfs_listxattr(struct dentry *dentry, char *list, size_t size) EXPORT_SYMBOL_GPL(vfs_listxattr); int -__vfs_removexattr(struct mnt_idmap *idmap, struct dentry *dentry, +__vfs_removexattr(const struct mnt_idmap *idmap, struct dentry *dentry, const char *name) { struct inode *inode = d_inode(dentry); @@ -557,7 +557,7 @@ EXPORT_SYMBOL(__vfs_removexattr); * a delegation was broken on, NULL if none. */ int -__vfs_removexattr_locked(struct mnt_idmap *idmap, +__vfs_removexattr_locked(const struct mnt_idmap *idmap, struct dentry *dentry, const char *name, struct delegated_inode *delegated_inode) { @@ -589,7 +589,7 @@ out: EXPORT_SYMBOL_GPL(__vfs_removexattr_locked); int -vfs_removexattr(struct mnt_idmap *idmap, struct dentry *dentry, +vfs_removexattr(const struct mnt_idmap *idmap, struct dentry *dentry, const char *name) { struct inode *inode = dentry->d_inode; @@ -652,7 +652,7 @@ int setxattr_copy(const char __user *name, struct kernel_xattr_ctx *ctx) return error; } -static int do_setxattr(struct mnt_idmap *idmap, struct dentry *dentry, +static int do_setxattr(const struct mnt_idmap *idmap, struct dentry *dentry, struct kernel_xattr_ctx *ctx) { if (is_posix_acl_xattr(ctx->kname->name)) @@ -787,7 +787,7 @@ SYSCALL_DEFINE5(fsetxattr, int, fd, const char __user *, name, * Extended attribute GET operations */ static ssize_t -do_getxattr(struct mnt_idmap *idmap, struct dentry *d, +do_getxattr(const struct mnt_idmap *idmap, struct dentry *d, struct kernel_xattr_ctx *ctx) { ssize_t error; @@ -1029,7 +1029,7 @@ SYSCALL_DEFINE3(flistxattr, int, fd, char __user *, list, size_t, size) * Extended attribute REMOVE operations */ static long -removexattr(struct mnt_idmap *idmap, struct dentry *d, const char *name) +removexattr(const struct mnt_idmap *idmap, struct dentry *d, const char *name) { if (is_posix_acl_xattr(name)) return vfs_remove_acl(idmap, d, name); diff --git a/fs/xfs/libxfs/xfs_attr.c b/fs/xfs/libxfs/xfs_attr.c index b3f7b2c34ad7..c4bb59633ce8 100644 --- a/fs/xfs/libxfs/xfs_attr.c +++ b/fs/xfs/libxfs/xfs_attr.c @@ -59,8 +59,7 @@ STATIC void xfs_attr_restore_rmt_blk(struct xfs_da_args *args); static int xfs_attr_node_try_addname(struct xfs_attr_intent *attr); STATIC int xfs_attr_node_addname_find_attr(struct xfs_attr_intent *attr); STATIC int xfs_attr_node_remove_attr(struct xfs_attr_intent *attr); -STATIC int xfs_attr_node_lookup(struct xfs_da_args *args, - struct xfs_da_state *state); +STATIC int xfs_attr_node_lookup(struct xfs_da_state *state); int xfs_inode_hasattr( @@ -709,7 +708,7 @@ int xfs_attr_node_removename_setup( int error; xfs_attr_item_init_da_state(attr); - error = xfs_attr_node_lookup(args, attr->xattri_da_state); + error = xfs_attr_node_lookup(attr->xattri_da_state); if (error != -EEXIST) goto out; error = 0; @@ -985,7 +984,7 @@ xfs_attr_lookup( } state = xfs_da_state_alloc(args); - error = xfs_attr_node_lookup(args, state); + error = xfs_attr_node_lookup(state); xfs_da_state_free(state); return error; } @@ -1014,7 +1013,7 @@ xfs_attr_add_fork( if (xfs_inode_has_attr_fork(ip)) goto trans_cancel; - error = xfs_bmap_add_attrfork(tp, ip, size, rsvd); + error = xfs_bmap_add_attrfork(tp, ip, size); if (error) goto trans_cancel; @@ -1386,7 +1385,6 @@ xfs_attr_leaf_get( /* Return EEXIST if attr is found, or ENOATTR if not. */ STATIC int xfs_attr_node_lookup( - struct xfs_da_args *args, struct xfs_da_state *state) { int retval, error; @@ -1417,7 +1415,7 @@ xfs_attr_node_addname_find_attr( * to where it should go. */ xfs_attr_item_init_da_state(attr); - error = xfs_attr_node_lookup(args, attr->xattri_da_state); + error = xfs_attr_node_lookup(attr->xattri_da_state); switch (error) { case -ENOATTR: if (args->op_flags & XFS_DA_OP_REPLACE) @@ -1588,7 +1586,7 @@ xfs_attr_node_get( * Search to see if name exists, and get back a pointer to it. */ state = xfs_da_state_alloc(args); - error = xfs_attr_node_lookup(args, state); + error = xfs_attr_node_lookup(state); if (error != -EEXIST) goto out_release; diff --git a/fs/xfs/libxfs/xfs_bmap.c b/fs/xfs/libxfs/xfs_bmap.c index d64defeda645..ae91f63455c5 100644 --- a/fs/xfs/libxfs/xfs_bmap.c +++ b/fs/xfs/libxfs/xfs_bmap.c @@ -1026,8 +1026,7 @@ int /* error code */ xfs_bmap_add_attrfork( struct xfs_trans *tp, struct xfs_inode *ip, /* incore inode pointer */ - int size, /* space new attribute needs */ - int rsvd) /* xact may use reserved blks */ + int size) /* space new attribute needs */ { struct xfs_mount *mp = tp->t_mountp; int logflags; /* logging flags */ @@ -6088,7 +6087,7 @@ xfs_bmap_validate_extent_raw( int whichfork, struct xfs_bmbt_irec *irec) { - if (!xfs_verify_fileext(mp, irec->br_startoff, irec->br_blockcount)) + if (!xfs_verify_fileext(irec->br_startoff, irec->br_blockcount)) return __this_address; if (rtfile && whichfork == XFS_DATA_FORK) { diff --git a/fs/xfs/libxfs/xfs_bmap.h b/fs/xfs/libxfs/xfs_bmap.h index d5f2729305fa..60f6df4ce057 100644 --- a/fs/xfs/libxfs/xfs_bmap.h +++ b/fs/xfs/libxfs/xfs_bmap.h @@ -181,7 +181,7 @@ void xfs_trim_extent(struct xfs_bmbt_irec *irec, xfs_fileoff_t bno, xfs_filblks_t len); unsigned int xfs_bmap_compute_attr_offset(struct xfs_mount *mp); int xfs_bmap_add_attrfork(struct xfs_trans *tp, struct xfs_inode *ip, - int size, int rsvd); + int size); void xfs_bmap_local_to_extents_empty(struct xfs_trans *tp, struct xfs_inode *ip, int whichfork); int xfs_bmap_local_to_extents(struct xfs_trans *tp, struct xfs_inode *ip, diff --git a/fs/xfs/libxfs/xfs_btree_mem.c b/fs/xfs/libxfs/xfs_btree_mem.c index 1d83a4251cee..3d5c7a014d4f 100644 --- a/fs/xfs/libxfs/xfs_btree_mem.c +++ b/fs/xfs/libxfs/xfs_btree_mem.c @@ -73,9 +73,7 @@ xfbtree_destroy( /* Compute the number of bytes available for records. */ static inline unsigned int -xfbtree_rec_bytes( - struct xfs_mount *mp, - const struct xfs_btree_ops *ops) +xfbtree_rec_bytes(void) { return XMBUF_BLOCKSIZE - XFS_BTREE_LBLOCK_CRC_LEN; } @@ -118,7 +116,7 @@ xfbtree_init( const struct xfs_btree_ops *ops) { unsigned long long owner = xfbt->owner; - unsigned int blocklen = xfbtree_rec_bytes(mp, ops); + unsigned int blocklen = xfbtree_rec_bytes(); unsigned int keyptr_len; int error; diff --git a/fs/xfs/libxfs/xfs_dquot_buf.c b/fs/xfs/libxfs/xfs_dquot_buf.c index 77954d1d924c..6c8a9afdfb1c 100644 --- a/fs/xfs/libxfs/xfs_dquot_buf.c +++ b/fs/xfs/libxfs/xfs_dquot_buf.c @@ -454,19 +454,8 @@ xfs_dqinode_metadir_link( .path = xfs_dqinode_path(type), .ip = ip, }; - int error; - error = xfs_metadir_start_link(&upd); - if (error) - return error; - - error = xfs_metadir_link(&upd); - if (error) - return error; - - xfs_trans_log_inode(upd.tp, upd.ip, XFS_ILOG_CORE); - - return xfs_metadir_commit(&upd); + return xfs_metadir_link_file(&upd); } #endif /* __KERNEL__ */ diff --git a/fs/xfs/libxfs/xfs_errortag.h b/fs/xfs/libxfs/xfs_errortag.h index f0c83f1f0b3b..d14aa289699f 100644 --- a/fs/xfs/libxfs/xfs_errortag.h +++ b/fs/xfs/libxfs/xfs_errortag.h @@ -75,7 +75,8 @@ #define XFS_ERRTAG_METAFILE_RESV_CRITICAL 45 #define XFS_ERRTAG_FORCE_ZERO_RANGE 46 #define XFS_ERRTAG_ZONE_RESET 47 -#define XFS_ERRTAG_MAX 48 +#define XFS_ERRTAG_BOUNCE_REREAD 48 +#define XFS_ERRTAG_MAX 49 /* * Random factors for above tags, 1 means always, 2 means 1/2 time, etc. @@ -137,7 +138,8 @@ XFS_ERRTAG(WRITE_DELAY_MS, write_delay_ms, 3000) \ XFS_ERRTAG(EXCHMAPS_FINISH_ONE, exchmaps_finish_one, 1) \ XFS_ERRTAG(METAFILE_RESV_CRITICAL, metafile_resv_crit, 4) \ XFS_ERRTAG(FORCE_ZERO_RANGE, force_zero_range, 4) \ -XFS_ERRTAG(ZONE_RESET, zone_reset, 1) +XFS_ERRTAG(ZONE_RESET, zone_reset, 1) \ +XFS_ERRTAG(BOUNCE_REREAD, bounce_reread, XFS_RANDOM_DEFAULT) #endif /* XFS_ERRTAG */ #endif /* __XFS_ERRORTAG_H_ */ diff --git a/fs/xfs/libxfs/xfs_exchmaps.c b/fs/xfs/libxfs/xfs_exchmaps.c index 6a66b6075e0a..ccb97da3765f 100644 --- a/fs/xfs/libxfs/xfs_exchmaps.c +++ b/fs/xfs/libxfs/xfs_exchmaps.c @@ -135,7 +135,6 @@ xmi_has_postop_work(const struct xfs_exchmaps_intent *xmi) /* Check all mappings to make sure we can actually exchange them. */ int xfs_exchmaps_check_forks( - struct xfs_mount *mp, const struct xfs_exchmaps_req *req) { struct xfs_ifork *ifp1, *ifp2; diff --git a/fs/xfs/libxfs/xfs_exchmaps.h b/fs/xfs/libxfs/xfs_exchmaps.h index fa822dff202a..055b8dbabd5b 100644 --- a/fs/xfs/libxfs/xfs_exchmaps.h +++ b/fs/xfs/libxfs/xfs_exchmaps.h @@ -115,8 +115,7 @@ void xfs_exchmaps_upgrade_extent_counts(struct xfs_trans *tp, int xfs_exchmaps_finish_one(struct xfs_trans *tp, struct xfs_exchmaps_intent *xmi); -int xfs_exchmaps_check_forks(struct xfs_mount *mp, - const struct xfs_exchmaps_req *req); +int xfs_exchmaps_check_forks(const struct xfs_exchmaps_req *req); void xfs_exchange_mappings(struct xfs_trans *tp, const struct xfs_exchmaps_req *req); diff --git a/fs/xfs/libxfs/xfs_ialloc.c b/fs/xfs/libxfs/xfs_ialloc.c index 58dac4d505ba..19b513b11692 100644 --- a/fs/xfs/libxfs/xfs_ialloc.c +++ b/fs/xfs/libxfs/xfs_ialloc.c @@ -615,7 +615,7 @@ xfs_inobt_insert_sprec( trace_xfs_irec_merge_post(pag, nrec); - error = xfs_inobt_rec_check_count(mp, nrec); + error = xfs_inobt_rec_check_count(nrec); if (error) goto error; diff --git a/fs/xfs/libxfs/xfs_ialloc_btree.c b/fs/xfs/libxfs/xfs_ialloc_btree.c index 1376e8630449..1f0bace2f144 100644 --- a/fs/xfs/libxfs/xfs_ialloc_btree.c +++ b/fs/xfs/libxfs/xfs_ialloc_btree.c @@ -687,7 +687,6 @@ xfs_inobt_irec_to_allocmask( */ int xfs_inobt_rec_check_count( - struct xfs_mount *mp, struct xfs_inobt_rec_incore *rec) { int inocount = 0; diff --git a/fs/xfs/libxfs/xfs_ialloc_btree.h b/fs/xfs/libxfs/xfs_ialloc_btree.h index 300edf5bc009..e04c63c66f39 100644 --- a/fs/xfs/libxfs/xfs_ialloc_btree.h +++ b/fs/xfs/libxfs/xfs_ialloc_btree.h @@ -57,10 +57,9 @@ unsigned int xfs_inobt_maxrecs(struct xfs_mount *mp, unsigned int blocklen, uint64_t xfs_inobt_irec_to_allocmask(const struct xfs_inobt_rec_incore *irec); #if defined(DEBUG) || defined(XFS_WARN) -int xfs_inobt_rec_check_count(struct xfs_mount *, - struct xfs_inobt_rec_incore *); +int xfs_inobt_rec_check_count(struct xfs_inobt_rec_incore *); #else -#define xfs_inobt_rec_check_count(mp, rec) 0 +#define xfs_inobt_rec_check_count(rec) 0 #endif /* DEBUG */ int xfs_finobt_calc_reserves(struct xfs_perag *perag, struct xfs_trans *tp, diff --git a/fs/xfs/libxfs/xfs_inode_fork.c b/fs/xfs/libxfs/xfs_inode_fork.c index 606a36526ce2..486fe7ab8b61 100644 --- a/fs/xfs/libxfs/xfs_inode_fork.c +++ b/fs/xfs/libxfs/xfs_inode_fork.c @@ -683,9 +683,7 @@ xfs_ifork_verify_local_data( break; } case S_IFLNK: { - struct xfs_ifork *ifp = xfs_ifork_ptr(ip, XFS_DATA_FORK); - - fa = xfs_symlink_shortform_verify(ifp->if_data, ifp->if_bytes); + fa = xfs_symlink_shortform_verify(ip); break; } default: diff --git a/fs/xfs/libxfs/xfs_inode_util.h b/fs/xfs/libxfs/xfs_inode_util.h index 060242998a23..e9eac35159c3 100644 --- a/fs/xfs/libxfs/xfs_inode_util.h +++ b/fs/xfs/libxfs/xfs_inode_util.h @@ -27,7 +27,7 @@ prid_t xfs_get_initial_prid(struct xfs_inode *dp); * idmap to NULL. To create a tree root, set pip to NULL. */ struct xfs_icreate_args { - struct mnt_idmap *idmap; + const struct mnt_idmap *idmap; struct xfs_inode *pip; /* parent inode or null */ dev_t rdev; umode_t mode; diff --git a/fs/xfs/libxfs/xfs_metadir.c b/fs/xfs/libxfs/xfs_metadir.c index 7c6b086b73db..79b37cb27273 100644 --- a/fs/xfs/libxfs/xfs_metadir.c +++ b/fs/xfs/libxfs/xfs_metadir.c @@ -163,7 +163,7 @@ xfs_metadir_teardown( trace_xfs_metadir_teardown(upd, error); if (upd->ppargs) { - xfs_parent_finish(upd->dp->i_mount, upd->ppargs); + xfs_parent_finish(upd->ppargs); upd->ppargs = NULL; } @@ -317,7 +317,7 @@ xfs_metadir_create( * Begin the process of linking a metadata file by allocating transactions * and locking whatever resources we're going to need. */ -int +static int xfs_metadir_start_link( struct xfs_metadir_update *upd) { @@ -364,7 +364,7 @@ out_teardown: * The path (up to the final component) must already exist, but the final * component must not already exist. */ -int +static int xfs_metadir_link( struct xfs_metadir_update *upd) { @@ -409,7 +409,7 @@ xfs_metadir_link( #endif /* ! __KERNEL__ */ /* Commit a metadir update and unlock/drop all resources. */ -int +static int xfs_metadir_commit( struct xfs_metadir_update *upd) { @@ -499,3 +499,27 @@ xfs_metadir_mkdir( return xfs_metadir_create_file(&upd, S_IFDIR, NULL, NULL, ipp); } + +#ifndef __KERNEL__ +/* Link a metadata file into a metadata directory. */ +int +xfs_metadir_link_file( + struct xfs_metadir_update *upd) +{ + int error; + + error = xfs_metadir_start_link(upd); + if (error) + return error; + + error = xfs_metadir_link(upd); + if (error) { + xfs_metadir_cancel(upd, error); + return error; + } + + xfs_trans_log_inode(upd->tp, upd->ip, XFS_ILOG_CORE); + + return xfs_metadir_commit(upd); +} +#endif /* ! __KERNEL__ */ diff --git a/fs/xfs/libxfs/xfs_metadir.h b/fs/xfs/libxfs/xfs_metadir.h index e434b9d1c932..b64f9fc5ca78 100644 --- a/fs/xfs/libxfs/xfs_metadir.h +++ b/fs/xfs/libxfs/xfs_metadir.h @@ -38,10 +38,7 @@ int xfs_metadir_create_file(struct xfs_metadir_update *upd, umode_t mode, xfs_metadir_createfn create, void *priv, struct xfs_inode **ipp); -int xfs_metadir_start_link(struct xfs_metadir_update *upd); -int xfs_metadir_link(struct xfs_metadir_update *upd); - -int xfs_metadir_commit(struct xfs_metadir_update *upd); +int xfs_metadir_link_file(struct xfs_metadir_update *upd); int xfs_metadir_mkdir(struct xfs_inode *dp, const char *path, struct xfs_inode **ipp); diff --git a/fs/xfs/libxfs/xfs_parent.h b/fs/xfs/libxfs/xfs_parent.h index 8eb4de9c5f1a..1dd3968a78b1 100644 --- a/fs/xfs/libxfs/xfs_parent.h +++ b/fs/xfs/libxfs/xfs_parent.h @@ -72,7 +72,6 @@ xfs_parent_start( /* Finish a parent pointer update by freeing the context object. */ static inline void xfs_parent_finish( - struct xfs_mount *mp, struct xfs_parent_args *ppargs) { if (ppargs) diff --git a/fs/xfs/libxfs/xfs_refcount_btree.c b/fs/xfs/libxfs/xfs_refcount_btree.c index 7e5f92c1ac56..2c6148a97994 100644 --- a/fs/xfs/libxfs/xfs_refcount_btree.c +++ b/fs/xfs/libxfs/xfs_refcount_btree.c @@ -127,7 +127,7 @@ xfs_refcountbt_get_maxrecs( return cur->bc_mp->m_refc_mxr[level != 0]; } -STATIC void +void xfs_refcountbt_init_key_from_rec( union xfs_btree_key *key, const union xfs_btree_rec *rec) @@ -135,7 +135,7 @@ xfs_refcountbt_init_key_from_rec( key->refc.rc_startblock = rec->refc.rc_startblock; } -STATIC void +void xfs_refcountbt_init_high_key_from_rec( union xfs_btree_key *key, const union xfs_btree_rec *rec) @@ -147,7 +147,7 @@ xfs_refcountbt_init_high_key_from_rec( key->refc.rc_startblock = cpu_to_be32(x); } -STATIC void +void xfs_refcountbt_init_rec_from_cur( struct xfs_btree_cur *cur, union xfs_btree_rec *rec) @@ -174,7 +174,7 @@ xfs_refcountbt_init_ptr_from_cur( ptr->s = agf->agf_refcount_root; } -STATIC int +int xfs_refcountbt_cmp_key_with_cur( struct xfs_btree_cur *cur, const union xfs_btree_key *key) @@ -188,7 +188,7 @@ xfs_refcountbt_cmp_key_with_cur( return cmp_int(be32_to_cpu(kp->rc_startblock), start); } -STATIC int +int xfs_refcountbt_cmp_two_keys( struct xfs_btree_cur *cur, const union xfs_btree_key *k1, @@ -283,7 +283,7 @@ const struct xfs_buf_ops xfs_refcountbt_buf_ops = { .verify_struct = xfs_refcountbt_verify, }; -STATIC int +int xfs_refcountbt_keys_inorder( struct xfs_btree_cur *cur, const union xfs_btree_key *k1, @@ -293,7 +293,7 @@ xfs_refcountbt_keys_inorder( be32_to_cpu(k2->refc.rc_startblock); } -STATIC int +int xfs_refcountbt_recs_inorder( struct xfs_btree_cur *cur, const union xfs_btree_rec *r1, @@ -304,7 +304,7 @@ xfs_refcountbt_recs_inorder( be32_to_cpu(r2->refc.rc_startblock); } -STATIC enum xbtree_key_contig +enum xbtree_key_contig xfs_refcountbt_keys_contiguous( struct xfs_btree_cur *cur, const union xfs_btree_key *key1, diff --git a/fs/xfs/libxfs/xfs_refcount_btree.h b/fs/xfs/libxfs/xfs_refcount_btree.h index beb93bef6a81..40eedfae6846 100644 --- a/fs/xfs/libxfs/xfs_refcount_btree.h +++ b/fs/xfs/libxfs/xfs_refcount_btree.h @@ -15,6 +15,8 @@ struct xfs_btree_cur; struct xfs_mount; struct xfs_perag; struct xbtree_afakeroot; +union xfs_btree_key; +union xfs_btree_rec; /* * Btree block header size @@ -69,4 +71,28 @@ unsigned int xfs_refcountbt_maxlevels_ondisk(void); int __init xfs_refcountbt_init_cur_cache(void); void xfs_refcountbt_destroy_cur_cache(void); +/* + * Key and record btree ops. The refcount on-disk key/record format is + * identical for the AG refcount btree and the realtime refcount btree, so + * these are shared by both. + */ +void xfs_refcountbt_init_key_from_rec(union xfs_btree_key *key, + const union xfs_btree_rec *rec); +void xfs_refcountbt_init_high_key_from_rec(union xfs_btree_key *key, + const union xfs_btree_rec *rec); +void xfs_refcountbt_init_rec_from_cur(struct xfs_btree_cur *cur, + union xfs_btree_rec *rec); +int xfs_refcountbt_cmp_key_with_cur(struct xfs_btree_cur *cur, + const union xfs_btree_key *key); +int xfs_refcountbt_cmp_two_keys(struct xfs_btree_cur *cur, + const union xfs_btree_key *k1, const union xfs_btree_key *k2, + const union xfs_btree_key *mask); +int xfs_refcountbt_keys_inorder(struct xfs_btree_cur *cur, + const union xfs_btree_key *k1, const union xfs_btree_key *k2); +int xfs_refcountbt_recs_inorder(struct xfs_btree_cur *cur, + const union xfs_btree_rec *r1, const union xfs_btree_rec *r2); +enum xbtree_key_contig xfs_refcountbt_keys_contiguous(struct xfs_btree_cur *cur, + const union xfs_btree_key *key1, const union xfs_btree_key *key2, + const union xfs_btree_key *mask); + #endif /* __XFS_REFCOUNT_BTREE_H__ */ diff --git a/fs/xfs/libxfs/xfs_rmap.c b/fs/xfs/libxfs/xfs_rmap.c index 34d218de21a9..14aef87837a9 100644 --- a/fs/xfs/libxfs/xfs_rmap.c +++ b/fs/xfs/libxfs/xfs_rmap.c @@ -260,7 +260,7 @@ xfs_rmap_check_irec( /* Check for a valid fork offset, if applicable. */ if (is_inode && !is_bmbt && - !xfs_verify_fileext(mp, irec->rm_offset, irec->rm_blockcount)) + !xfs_verify_fileext(irec->rm_offset, irec->rm_blockcount)) return __this_address; return NULL; @@ -310,7 +310,7 @@ xfs_rtrmap_check_inode_irec( return __this_address; if (!xfs_verify_rgbext(rtg, irec->rm_startblock, irec->rm_blockcount)) return __this_address; - if (!xfs_verify_fileext(mp, irec->rm_offset, irec->rm_blockcount)) + if (!xfs_verify_fileext(irec->rm_offset, irec->rm_blockcount)) return __this_address; return NULL; } @@ -904,7 +904,6 @@ xfs_rmap_hook_enable(void) /* Call downstream hooks for a reverse mapping update. */ static inline void xfs_rmap_update_hook( - struct xfs_trans *tp, struct xfs_group *xg, enum xfs_rmap_intent_type op, xfs_agblock_t startblock, @@ -952,7 +951,7 @@ xfs_rmap_hook_setup( xfs_hook_setup(&hook->rmap_hook, mod_fn); } #else -# define xfs_rmap_update_hook(t, p, o, s, b, u, oi) do { } while (0) +# define xfs_rmap_update_hook(p, o, s, b, u, oi) do { } while (0) #endif /* CONFIG_XFS_LIVE_HOOKS */ /* @@ -975,7 +974,7 @@ xfs_rmap_free( return 0; cur = xfs_rmapbt_init_cursor(mp, tp, agbp, pag); - xfs_rmap_update_hook(tp, pag_group(pag), XFS_RMAP_UNMAP, bno, len, + xfs_rmap_update_hook(pag_group(pag), XFS_RMAP_UNMAP, bno, len, false, oinfo); error = xfs_rmap_unmap(cur, bno, len, false, oinfo); @@ -1220,7 +1219,7 @@ xfs_rmap_alloc( return 0; cur = xfs_rmapbt_init_cursor(mp, tp, agbp, pag); - xfs_rmap_update_hook(tp, pag_group(pag), XFS_RMAP_MAP, bno, len, false, + xfs_rmap_update_hook(pag_group(pag), XFS_RMAP_MAP, bno, len, false, oinfo); error = xfs_rmap_map(cur, bno, len, false, oinfo); @@ -2721,7 +2720,7 @@ xfs_rmap_finish_one( if (error) return error; - xfs_rmap_update_hook(tp, ri->ri_group, ri->ri_type, bno, + xfs_rmap_update_hook(ri->ri_group, ri->ri_type, bno, ri->ri_bmap.br_blockcount, unwritten, &oinfo); return 0; } diff --git a/fs/xfs/libxfs/xfs_rmap_btree.c b/fs/xfs/libxfs/xfs_rmap_btree.c index 10b3272238eb..5b283a5ddd13 100644 --- a/fs/xfs/libxfs/xfs_rmap_btree.c +++ b/fs/xfs/libxfs/xfs_rmap_btree.c @@ -170,7 +170,7 @@ static inline __be64 ondisk_rec_offset_to_key(const union xfs_btree_rec *rec) return rec->rmap.rm_offset & ~cpu_to_be64(XFS_RMAP_OFF_UNWRITTEN); } -STATIC void +void xfs_rmapbt_init_key_from_rec( union xfs_btree_key *key, const union xfs_btree_rec *rec) @@ -187,7 +187,7 @@ xfs_rmapbt_init_key_from_rec( * the startblock for all records, and if the record is for a data/attr * fork mapping, we add blockcount-1 to the offset too. */ -STATIC void +void xfs_rmapbt_init_high_key_from_rec( union xfs_btree_key *key, const union xfs_btree_rec *rec) @@ -209,7 +209,7 @@ xfs_rmapbt_init_high_key_from_rec( key->rmap.rm_offset = cpu_to_be64(off); } -STATIC void +void xfs_rmapbt_init_rec_from_cur( struct xfs_btree_cur *cur, union xfs_btree_rec *rec) @@ -243,7 +243,7 @@ static inline uint64_t offset_keymask(uint64_t offset) return offset & ~XFS_RMAP_OFF_UNWRITTEN; } -STATIC int +int xfs_rmapbt_cmp_key_with_cur( struct xfs_btree_cur *cur, const union xfs_btree_key *key) @@ -257,7 +257,7 @@ xfs_rmapbt_cmp_key_with_cur( offset_keymask(xfs_rmap_irec_offset_pack(rec))); } -STATIC int +int xfs_rmapbt_cmp_two_keys( struct xfs_btree_cur *cur, const union xfs_btree_key *k1, @@ -390,7 +390,7 @@ const struct xfs_buf_ops xfs_rmapbt_buf_ops = { .verify_struct = xfs_rmapbt_verify, }; -STATIC int +int xfs_rmapbt_keys_inorder( struct xfs_btree_cur *cur, const union xfs_btree_key *k1, @@ -420,7 +420,7 @@ xfs_rmapbt_keys_inorder( return 0; } -STATIC int +int xfs_rmapbt_recs_inorder( struct xfs_btree_cur *cur, const union xfs_btree_rec *r1, @@ -450,7 +450,7 @@ xfs_rmapbt_recs_inorder( return 0; } -STATIC enum xbtree_key_contig +enum xbtree_key_contig xfs_rmapbt_keys_contiguous( struct xfs_btree_cur *cur, const union xfs_btree_key *key1, diff --git a/fs/xfs/libxfs/xfs_rmap_btree.h b/fs/xfs/libxfs/xfs_rmap_btree.h index 119b1567cd0e..7071dac745e1 100644 --- a/fs/xfs/libxfs/xfs_rmap_btree.h +++ b/fs/xfs/libxfs/xfs_rmap_btree.h @@ -11,6 +11,8 @@ struct xfs_btree_cur; struct xfs_mount; struct xbtree_afakeroot; struct xfbtree; +union xfs_btree_key; +union xfs_btree_rec; /* rmaps only exist on crc enabled filesystems */ #define XFS_RMAP_BLOCK_LEN XFS_BTREE_SBLOCK_CRC_LEN @@ -69,4 +71,28 @@ struct xfs_btree_cur *xfs_rmapbt_mem_cursor(struct xfs_perag *pag, int xfs_rmapbt_mem_init(struct xfs_mount *mp, struct xfbtree *xfbtree, struct xfs_buftarg *btp, xfs_agnumber_t agno); +/* + * Key and record btree ops. The rmap on-disk key/record format is identical + * for the AG rmap btree and the realtime rmap btree, so these are shared by + * both. + */ +void xfs_rmapbt_init_key_from_rec(union xfs_btree_key *key, + const union xfs_btree_rec *rec); +void xfs_rmapbt_init_high_key_from_rec(union xfs_btree_key *key, + const union xfs_btree_rec *rec); +void xfs_rmapbt_init_rec_from_cur(struct xfs_btree_cur *cur, + union xfs_btree_rec *rec); +int xfs_rmapbt_cmp_key_with_cur(struct xfs_btree_cur *cur, + const union xfs_btree_key *key); +int xfs_rmapbt_cmp_two_keys(struct xfs_btree_cur *cur, + const union xfs_btree_key *k1, const union xfs_btree_key *k2, + const union xfs_btree_key *mask); +int xfs_rmapbt_keys_inorder(struct xfs_btree_cur *cur, + const union xfs_btree_key *k1, const union xfs_btree_key *k2); +int xfs_rmapbt_recs_inorder(struct xfs_btree_cur *cur, + const union xfs_btree_rec *r1, const union xfs_btree_rec *r2); +enum xbtree_key_contig xfs_rmapbt_keys_contiguous(struct xfs_btree_cur *cur, + const union xfs_btree_key *key1, const union xfs_btree_key *key2, + const union xfs_btree_key *mask); + #endif /* __XFS_RMAP_BTREE_H__ */ diff --git a/fs/xfs/libxfs/xfs_rtrefcount_btree.c b/fs/xfs/libxfs/xfs_rtrefcount_btree.c index dcc89b8e149b..697f5e622685 100644 --- a/fs/xfs/libxfs/xfs_rtrefcount_btree.c +++ b/fs/xfs/libxfs/xfs_rtrefcount_btree.c @@ -20,6 +20,7 @@ #include "xfs_btree_staging.h" #include "xfs_rtrefcount_btree.h" #include "xfs_refcount.h" +#include "xfs_refcount_btree.h" #include "xfs_trace.h" #include "xfs_cksum.h" #include "xfs_error.h" @@ -114,41 +115,6 @@ xfs_rtrefcountbt_get_dmaxrecs( } STATIC void -xfs_rtrefcountbt_init_key_from_rec( - union xfs_btree_key *key, - const union xfs_btree_rec *rec) -{ - key->refc.rc_startblock = rec->refc.rc_startblock; -} - -STATIC void -xfs_rtrefcountbt_init_high_key_from_rec( - union xfs_btree_key *key, - const union xfs_btree_rec *rec) -{ - __u32 x; - - x = be32_to_cpu(rec->refc.rc_startblock); - x += be32_to_cpu(rec->refc.rc_blockcount) - 1; - key->refc.rc_startblock = cpu_to_be32(x); -} - -STATIC void -xfs_rtrefcountbt_init_rec_from_cur( - struct xfs_btree_cur *cur, - union xfs_btree_rec *rec) -{ - const struct xfs_refcount_irec *irec = &cur->bc_rec.rc; - uint32_t start; - - start = xfs_refcount_encode_startblock(irec->rc_startblock, - irec->rc_domain); - rec->refc.rc_startblock = cpu_to_be32(start); - rec->refc.rc_blockcount = cpu_to_be32(cur->bc_rec.rc.rc_blockcount); - rec->refc.rc_refcount = cpu_to_be32(cur->bc_rec.rc.rc_refcount); -} - -STATIC void xfs_rtrefcountbt_init_ptr_from_cur( struct xfs_btree_cur *cur, union xfs_btree_ptr *ptr) @@ -156,33 +122,6 @@ xfs_rtrefcountbt_init_ptr_from_cur( ptr->l = 0; } -STATIC int -xfs_rtrefcountbt_cmp_key_with_cur( - struct xfs_btree_cur *cur, - const union xfs_btree_key *key) -{ - const struct xfs_refcount_key *kp = &key->refc; - const struct xfs_refcount_irec *irec = &cur->bc_rec.rc; - uint32_t start; - - start = xfs_refcount_encode_startblock(irec->rc_startblock, - irec->rc_domain); - return cmp_int(be32_to_cpu(kp->rc_startblock), start); -} - -STATIC int -xfs_rtrefcountbt_cmp_two_keys( - struct xfs_btree_cur *cur, - const union xfs_btree_key *k1, - const union xfs_btree_key *k2, - const union xfs_btree_key *mask) -{ - ASSERT(!mask || mask->refc.rc_startblock); - - return cmp_int(be32_to_cpu(k1->refc.rc_startblock), - be32_to_cpu(k2->refc.rc_startblock)); -} - static xfs_failaddr_t xfs_rtrefcountbt_verify( struct xfs_buf *bp) @@ -249,40 +188,6 @@ const struct xfs_buf_ops xfs_rtrefcountbt_buf_ops = { .verify_struct = xfs_rtrefcountbt_verify, }; -STATIC int -xfs_rtrefcountbt_keys_inorder( - struct xfs_btree_cur *cur, - const union xfs_btree_key *k1, - const union xfs_btree_key *k2) -{ - return be32_to_cpu(k1->refc.rc_startblock) < - be32_to_cpu(k2->refc.rc_startblock); -} - -STATIC int -xfs_rtrefcountbt_recs_inorder( - struct xfs_btree_cur *cur, - const union xfs_btree_rec *r1, - const union xfs_btree_rec *r2) -{ - return be32_to_cpu(r1->refc.rc_startblock) + - be32_to_cpu(r1->refc.rc_blockcount) <= - be32_to_cpu(r2->refc.rc_startblock); -} - -STATIC enum xbtree_key_contig -xfs_rtrefcountbt_keys_contiguous( - struct xfs_btree_cur *cur, - const union xfs_btree_key *key1, - const union xfs_btree_key *key2, - const union xfs_btree_key *mask) -{ - ASSERT(!mask || mask->refc.rc_startblock); - - return xbtree_key_contig(be32_to_cpu(key1->refc.rc_startblock), - be32_to_cpu(key2->refc.rc_startblock)); -} - static inline void xfs_rtrefcountbt_move_ptrs( struct xfs_mount *mp, @@ -311,7 +216,7 @@ xfs_rtrefcountbt_broot_realloc( unsigned int old_size = ifp->if_broot_bytes; const unsigned int level = cur->bc_nlevels - 1; - new_size = xfs_rtrefcount_broot_space_calc(mp, level, new_numrecs); + new_size = xfs_rtrefcount_broot_space_calc(level, new_numrecs); /* Handle the nop case quietly. */ if (new_size == old_size) @@ -383,16 +288,16 @@ const struct xfs_btree_ops xfs_rtrefcountbt_ops = { .get_minrecs = xfs_rtrefcountbt_get_minrecs, .get_maxrecs = xfs_rtrefcountbt_get_maxrecs, .get_dmaxrecs = xfs_rtrefcountbt_get_dmaxrecs, - .init_key_from_rec = xfs_rtrefcountbt_init_key_from_rec, - .init_high_key_from_rec = xfs_rtrefcountbt_init_high_key_from_rec, - .init_rec_from_cur = xfs_rtrefcountbt_init_rec_from_cur, + .init_key_from_rec = xfs_refcountbt_init_key_from_rec, + .init_high_key_from_rec = xfs_refcountbt_init_high_key_from_rec, + .init_rec_from_cur = xfs_refcountbt_init_rec_from_cur, .init_ptr_from_cur = xfs_rtrefcountbt_init_ptr_from_cur, - .cmp_key_with_cur = xfs_rtrefcountbt_cmp_key_with_cur, + .cmp_key_with_cur = xfs_refcountbt_cmp_key_with_cur, .buf_ops = &xfs_rtrefcountbt_buf_ops, - .cmp_two_keys = xfs_rtrefcountbt_cmp_two_keys, - .keys_inorder = xfs_rtrefcountbt_keys_inorder, - .recs_inorder = xfs_rtrefcountbt_recs_inorder, - .keys_contiguous = xfs_rtrefcountbt_keys_contiguous, + .cmp_two_keys = xfs_refcountbt_cmp_two_keys, + .keys_inorder = xfs_refcountbt_keys_inorder, + .recs_inorder = xfs_refcountbt_recs_inorder, + .keys_contiguous = xfs_refcountbt_keys_contiguous, .broot_realloc = xfs_rtrefcountbt_broot_realloc, }; @@ -602,7 +507,7 @@ xfs_rtrefcountbt_from_disk( unsigned int maxrecs; unsigned int rblocklen; - rblocklen = xfs_rtrefcount_broot_space(mp, dblock); + rblocklen = xfs_rtrefcount_broot_space(dblock); xfs_btree_init_block(mp, rblock, &xfs_rtrefcountbt_ops, 0, 0, I_INO(ip)); @@ -661,7 +566,7 @@ xfs_iformat_rtrefcount( } broot = xfs_broot_alloc(xfs_ifork_ptr(ip, XFS_DATA_FORK), - xfs_rtrefcount_broot_space_calc(mp, level, numrecs)); + xfs_rtrefcount_broot_space_calc(level, numrecs)); if (broot) xfs_rtrefcountbt_from_disk(ip, dfp, dsize, broot); return 0; @@ -751,7 +656,7 @@ xfs_rtrefcountbt_create( /* Initialize the empty incore btree root. */ broot = xfs_broot_realloc(ifp, - xfs_rtrefcount_broot_space_calc(mp, 0, 0)); + xfs_rtrefcount_broot_space_calc(0, 0)); if (broot) xfs_btree_init_block(mp, broot, &xfs_rtrefcountbt_ops, 0, 0, I_INO(ip)); diff --git a/fs/xfs/libxfs/xfs_rtrefcount_btree.h b/fs/xfs/libxfs/xfs_rtrefcount_btree.h index a99b7a8aec86..9ab6ecf90ba5 100644 --- a/fs/xfs/libxfs/xfs_rtrefcount_btree.h +++ b/fs/xfs/libxfs/xfs_rtrefcount_btree.h @@ -129,7 +129,6 @@ xfs_rtrefcount_broot_ptr_addr( */ static inline size_t xfs_rtrefcount_broot_space_calc( - struct xfs_mount *mp, unsigned int level, unsigned int nrecs) { @@ -146,9 +145,9 @@ xfs_rtrefcount_broot_space_calc( * btree root block. */ static inline size_t -xfs_rtrefcount_broot_space(struct xfs_mount *mp, struct xfs_rtrefcount_root *bb) +xfs_rtrefcount_broot_space(struct xfs_rtrefcount_root *bb) { - return xfs_rtrefcount_broot_space_calc(mp, be16_to_cpu(bb->bb_level), + return xfs_rtrefcount_broot_space_calc(be16_to_cpu(bb->bb_level), be16_to_cpu(bb->bb_numrecs)); } diff --git a/fs/xfs/libxfs/xfs_rtrmap_btree.c b/fs/xfs/libxfs/xfs_rtrmap_btree.c index a15e460a1ec7..2f00d0698ae1 100644 --- a/fs/xfs/libxfs/xfs_rtrmap_btree.c +++ b/fs/xfs/libxfs/xfs_rtrmap_btree.c @@ -20,6 +20,7 @@ #include "xfs_btree_staging.h" #include "xfs_metafile.h" #include "xfs_rmap.h" +#include "xfs_rmap_btree.h" #include "xfs_rtrmap_btree.h" #include "xfs_trace.h" #include "xfs_cksum.h" @@ -113,60 +114,6 @@ xfs_rtrmapbt_get_dmaxrecs( return xfs_rtrmapbt_droot_maxrecs(cur->bc_ino.forksize, level == 0); } -/* - * Convert the ondisk record's offset field into the ondisk key's offset field. - * Fork and bmbt are significant parts of the rmap record key, but written - * status is merely a record attribute. - */ -static inline __be64 ondisk_rec_offset_to_key(const union xfs_btree_rec *rec) -{ - return rec->rmap.rm_offset & ~cpu_to_be64(XFS_RMAP_OFF_UNWRITTEN); -} - -STATIC void -xfs_rtrmapbt_init_key_from_rec( - union xfs_btree_key *key, - const union xfs_btree_rec *rec) -{ - key->rmap.rm_startblock = rec->rmap.rm_startblock; - key->rmap.rm_owner = rec->rmap.rm_owner; - key->rmap.rm_offset = ondisk_rec_offset_to_key(rec); -} - -STATIC void -xfs_rtrmapbt_init_high_key_from_rec( - union xfs_btree_key *key, - const union xfs_btree_rec *rec) -{ - uint64_t off; - int adj; - - adj = be32_to_cpu(rec->rmap.rm_blockcount) - 1; - - key->rmap.rm_startblock = rec->rmap.rm_startblock; - be32_add_cpu(&key->rmap.rm_startblock, adj); - key->rmap.rm_owner = rec->rmap.rm_owner; - key->rmap.rm_offset = ondisk_rec_offset_to_key(rec); - if (XFS_RMAP_NON_INODE_OWNER(be64_to_cpu(rec->rmap.rm_owner)) || - XFS_RMAP_IS_BMBT_BLOCK(be64_to_cpu(rec->rmap.rm_offset))) - return; - off = be64_to_cpu(key->rmap.rm_offset); - off = (XFS_RMAP_OFF(off) + adj) | (off & ~XFS_RMAP_OFF_MASK); - key->rmap.rm_offset = cpu_to_be64(off); -} - -STATIC void -xfs_rtrmapbt_init_rec_from_cur( - struct xfs_btree_cur *cur, - union xfs_btree_rec *rec) -{ - rec->rmap.rm_startblock = cpu_to_be32(cur->bc_rec.r.rm_startblock); - rec->rmap.rm_blockcount = cpu_to_be32(cur->bc_rec.r.rm_blockcount); - rec->rmap.rm_owner = cpu_to_be64(cur->bc_rec.r.rm_owner); - rec->rmap.rm_offset = cpu_to_be64( - xfs_rmap_irec_offset_pack(&cur->bc_rec.r)); -} - STATIC void xfs_rtrmapbt_init_ptr_from_cur( struct xfs_btree_cur *cur, @@ -175,69 +122,6 @@ xfs_rtrmapbt_init_ptr_from_cur( ptr->l = 0; } -/* - * Mask the appropriate parts of the ondisk key field for a key comparison. - * Fork and bmbt are significant parts of the rmap record key, but written - * status is merely a record attribute. - */ -static inline uint64_t offset_keymask(uint64_t offset) -{ - return offset & ~XFS_RMAP_OFF_UNWRITTEN; -} - -STATIC int -xfs_rtrmapbt_cmp_key_with_cur( - struct xfs_btree_cur *cur, - const union xfs_btree_key *key) -{ - struct xfs_rmap_irec *rec = &cur->bc_rec.r; - const struct xfs_rmap_key *kp = &key->rmap; - - return cmp_int(be32_to_cpu(kp->rm_startblock), rec->rm_startblock) ?: - cmp_int(be64_to_cpu(kp->rm_owner), rec->rm_owner) ?: - cmp_int(offset_keymask(be64_to_cpu(kp->rm_offset)), - offset_keymask(xfs_rmap_irec_offset_pack(rec))); -} - -STATIC int -xfs_rtrmapbt_cmp_two_keys( - struct xfs_btree_cur *cur, - const union xfs_btree_key *k1, - const union xfs_btree_key *k2, - const union xfs_btree_key *mask) -{ - const struct xfs_rmap_key *kp1 = &k1->rmap; - const struct xfs_rmap_key *kp2 = &k2->rmap; - int d; - - /* Doesn't make sense to mask off the physical space part */ - ASSERT(!mask || mask->rmap.rm_startblock); - - d = cmp_int(be32_to_cpu(kp1->rm_startblock), - be32_to_cpu(kp2->rm_startblock)); - if (d) - return d; - - if (!mask || mask->rmap.rm_owner) { - d = cmp_int(be64_to_cpu(kp1->rm_owner), - be64_to_cpu(kp2->rm_owner)); - if (d) - return d; - } - - if (!mask || mask->rmap.rm_offset) { - /* Doesn't make sense to allow offset but not owner */ - ASSERT(!mask || mask->rmap.rm_owner); - - d = cmp_int(offset_keymask(be64_to_cpu(kp1->rm_offset)), - offset_keymask(be64_to_cpu(kp2->rm_offset))); - if (d) - return d; - } - - return 0; -} - static xfs_failaddr_t xfs_rtrmapbt_verify( struct xfs_buf *bp) @@ -304,86 +188,6 @@ const struct xfs_buf_ops xfs_rtrmapbt_buf_ops = { .verify_struct = xfs_rtrmapbt_verify, }; -STATIC int -xfs_rtrmapbt_keys_inorder( - struct xfs_btree_cur *cur, - const union xfs_btree_key *k1, - const union xfs_btree_key *k2) -{ - uint32_t x; - uint32_t y; - uint64_t a; - uint64_t b; - - x = be32_to_cpu(k1->rmap.rm_startblock); - y = be32_to_cpu(k2->rmap.rm_startblock); - if (x < y) - return 1; - else if (x > y) - return 0; - a = be64_to_cpu(k1->rmap.rm_owner); - b = be64_to_cpu(k2->rmap.rm_owner); - if (a < b) - return 1; - else if (a > b) - return 0; - a = offset_keymask(be64_to_cpu(k1->rmap.rm_offset)); - b = offset_keymask(be64_to_cpu(k2->rmap.rm_offset)); - if (a <= b) - return 1; - return 0; -} - -STATIC int -xfs_rtrmapbt_recs_inorder( - struct xfs_btree_cur *cur, - const union xfs_btree_rec *r1, - const union xfs_btree_rec *r2) -{ - uint32_t x; - uint32_t y; - uint64_t a; - uint64_t b; - - x = be32_to_cpu(r1->rmap.rm_startblock); - y = be32_to_cpu(r2->rmap.rm_startblock); - if (x < y) - return 1; - else if (x > y) - return 0; - a = be64_to_cpu(r1->rmap.rm_owner); - b = be64_to_cpu(r2->rmap.rm_owner); - if (a < b) - return 1; - else if (a > b) - return 0; - a = offset_keymask(be64_to_cpu(r1->rmap.rm_offset)); - b = offset_keymask(be64_to_cpu(r2->rmap.rm_offset)); - if (a <= b) - return 1; - return 0; -} - -STATIC enum xbtree_key_contig -xfs_rtrmapbt_keys_contiguous( - struct xfs_btree_cur *cur, - const union xfs_btree_key *key1, - const union xfs_btree_key *key2, - const union xfs_btree_key *mask) -{ - ASSERT(!mask || mask->rmap.rm_startblock); - - /* - * We only support checking contiguity of the physical space component. - * If any callers ever need more specificity than that, they'll have to - * implement it here. - */ - ASSERT(!mask || (!mask->rmap.rm_owner && !mask->rmap.rm_offset)); - - return xbtree_key_contig(be32_to_cpu(key1->rmap.rm_startblock), - be32_to_cpu(key2->rmap.rm_startblock)); -} - static inline void xfs_rtrmapbt_move_ptrs( struct xfs_mount *mp, @@ -412,7 +216,7 @@ xfs_rtrmapbt_broot_realloc( unsigned int old_size = ifp->if_broot_bytes; const unsigned int level = cur->bc_nlevels - 1; - new_size = xfs_rtrmap_broot_space_calc(mp, level, new_numrecs); + new_size = xfs_rtrmap_broot_space_calc(level, new_numrecs); /* Handle the nop case quietly. */ if (new_size == old_size) @@ -486,16 +290,16 @@ const struct xfs_btree_ops xfs_rtrmapbt_ops = { .get_minrecs = xfs_rtrmapbt_get_minrecs, .get_maxrecs = xfs_rtrmapbt_get_maxrecs, .get_dmaxrecs = xfs_rtrmapbt_get_dmaxrecs, - .init_key_from_rec = xfs_rtrmapbt_init_key_from_rec, - .init_high_key_from_rec = xfs_rtrmapbt_init_high_key_from_rec, - .init_rec_from_cur = xfs_rtrmapbt_init_rec_from_cur, + .init_key_from_rec = xfs_rmapbt_init_key_from_rec, + .init_high_key_from_rec = xfs_rmapbt_init_high_key_from_rec, + .init_rec_from_cur = xfs_rmapbt_init_rec_from_cur, .init_ptr_from_cur = xfs_rtrmapbt_init_ptr_from_cur, - .cmp_key_with_cur = xfs_rtrmapbt_cmp_key_with_cur, + .cmp_key_with_cur = xfs_rmapbt_cmp_key_with_cur, .buf_ops = &xfs_rtrmapbt_buf_ops, - .cmp_two_keys = xfs_rtrmapbt_cmp_two_keys, - .keys_inorder = xfs_rtrmapbt_keys_inorder, - .recs_inorder = xfs_rtrmapbt_recs_inorder, - .keys_contiguous = xfs_rtrmapbt_keys_contiguous, + .cmp_two_keys = xfs_rmapbt_cmp_two_keys, + .keys_inorder = xfs_rmapbt_keys_inorder, + .recs_inorder = xfs_rmapbt_recs_inorder, + .keys_contiguous = xfs_rmapbt_keys_contiguous, .broot_realloc = xfs_rtrmapbt_broot_realloc, }; @@ -595,16 +399,16 @@ const struct xfs_btree_ops xfs_rtrmapbt_mem_ops = { .free_block = xfbtree_free_block, .get_minrecs = xfbtree_get_minrecs, .get_maxrecs = xfbtree_get_maxrecs, - .init_key_from_rec = xfs_rtrmapbt_init_key_from_rec, - .init_high_key_from_rec = xfs_rtrmapbt_init_high_key_from_rec, - .init_rec_from_cur = xfs_rtrmapbt_init_rec_from_cur, + .init_key_from_rec = xfs_rmapbt_init_key_from_rec, + .init_high_key_from_rec = xfs_rmapbt_init_high_key_from_rec, + .init_rec_from_cur = xfs_rmapbt_init_rec_from_cur, .init_ptr_from_cur = xfbtree_init_ptr_from_cur, - .cmp_key_with_cur = xfs_rtrmapbt_cmp_key_with_cur, + .cmp_key_with_cur = xfs_rmapbt_cmp_key_with_cur, .buf_ops = &xfs_rtrmapbt_mem_buf_ops, - .cmp_two_keys = xfs_rtrmapbt_cmp_two_keys, - .keys_inorder = xfs_rtrmapbt_keys_inorder, - .recs_inorder = xfs_rtrmapbt_recs_inorder, - .keys_contiguous = xfs_rtrmapbt_keys_contiguous, + .cmp_two_keys = xfs_rmapbt_cmp_two_keys, + .keys_inorder = xfs_rmapbt_keys_inorder, + .recs_inorder = xfs_rmapbt_recs_inorder, + .keys_contiguous = xfs_rmapbt_keys_contiguous, }; /* Create a cursor for an in-memory btree. */ @@ -895,7 +699,7 @@ xfs_iformat_rtrmap( } broot = xfs_broot_alloc(xfs_ifork_ptr(ip, XFS_DATA_FORK), - xfs_rtrmap_broot_space_calc(mp, level, numrecs)); + xfs_rtrmap_broot_space_calc(level, numrecs)); if (broot) xfs_rtrmapbt_from_disk(ip, dfp, dsize, broot); return 0; @@ -980,7 +784,7 @@ xfs_rtrmapbt_create( ASSERT(ifp->if_bytes == 0); /* Initialize the empty incore btree root. */ - broot = xfs_broot_realloc(ifp, xfs_rtrmap_broot_space_calc(mp, 0, 0)); + broot = xfs_broot_realloc(ifp, xfs_rtrmap_broot_space_calc(0, 0)); if (broot) xfs_btree_init_block(mp, broot, &xfs_rtrmapbt_ops, 0, 0, I_INO(ip)); diff --git a/fs/xfs/libxfs/xfs_rtrmap_btree.h b/fs/xfs/libxfs/xfs_rtrmap_btree.h index e328fd62a149..c59a144b4bbf 100644 --- a/fs/xfs/libxfs/xfs_rtrmap_btree.h +++ b/fs/xfs/libxfs/xfs_rtrmap_btree.h @@ -140,7 +140,6 @@ xfs_rtrmap_broot_ptr_addr( */ static inline size_t xfs_rtrmap_broot_space_calc( - struct xfs_mount *mp, unsigned int level, unsigned int nrecs) { @@ -159,7 +158,7 @@ xfs_rtrmap_broot_space_calc( static inline size_t xfs_rtrmap_broot_space(struct xfs_mount *mp, struct xfs_rtrmap_root *bb) { - return xfs_rtrmap_broot_space_calc(mp, be16_to_cpu(bb->bb_level), + return xfs_rtrmap_broot_space_calc(be16_to_cpu(bb->bb_level), be16_to_cpu(bb->bb_numrecs)); } diff --git a/fs/xfs/libxfs/xfs_sb.c b/fs/xfs/libxfs/xfs_sb.c index f0341adbb879..d2ef7b20a49f 100644 --- a/fs/xfs/libxfs/xfs_sb.c +++ b/fs/xfs/libxfs/xfs_sb.c @@ -1460,46 +1460,6 @@ xfs_update_secondary_sbs( return saved_error ? saved_error : error; } -/* - * Same behavior as xfs_sync_sb, except that it is always synchronous and it - * also writes the superblock buffer to disk sector 0 immediately. - */ -int -xfs_sync_sb_buf( - struct xfs_mount *mp, - bool update_rtsb) -{ - struct xfs_trans *tp; - int error; - - error = xfs_trans_alloc(mp, &M_RES(mp)->tr_sb, 0, 0, 0, &tp); - if (error) - return error; - - xfs_log_sb(tp); - if (update_rtsb) - xfs_log_rtsb(tp, xfs_trans_getsb(tp)); - xfs_trans_set_sync(tp); - error = xfs_trans_commit(tp); - if (error) - return error; - - /* Re-acquire and write the sb and rtsb to disk. */ - xfs_buf_lock(mp->m_sb_bp); - error = xfs_bwrite(mp->m_sb_bp); - xfs_buf_unlock(mp->m_sb_bp); - if (error) - return error; - - if (update_rtsb && mp->m_rtsb_bp) { - xfs_buf_lock(mp->m_rtsb_bp); - error = xfs_bwrite(mp->m_rtsb_bp); - xfs_buf_unlock(mp->m_rtsb_bp); - } - - return error; -} - void xfs_fs_geometry( struct xfs_mount *mp, diff --git a/fs/xfs/libxfs/xfs_sb.h b/fs/xfs/libxfs/xfs_sb.h index 34d0dd374e9b..77de65922213 100644 --- a/fs/xfs/libxfs/xfs_sb.h +++ b/fs/xfs/libxfs/xfs_sb.h @@ -15,7 +15,6 @@ struct xfs_perag; extern void xfs_log_sb(struct xfs_trans *tp); extern int xfs_sync_sb(struct xfs_mount *mp, bool wait); -extern int xfs_sync_sb_buf(struct xfs_mount *mp, bool update_rtsb); extern void xfs_sb_mount_common(struct xfs_mount *mp, struct xfs_sb *sbp); void xfs_sb_mount_rextsize(struct xfs_mount *mp, struct xfs_sb *sbp); void xfs_mount_sb_set_rextsize(struct xfs_mount *mp, diff --git a/fs/xfs/libxfs/xfs_symlink_remote.c b/fs/xfs/libxfs/xfs_symlink_remote.c index b0dc3888bf1b..0201a3d59b1a 100644 --- a/fs/xfs/libxfs/xfs_symlink_remote.c +++ b/fs/xfs/libxfs/xfs_symlink_remote.c @@ -208,11 +208,15 @@ xfs_symlink_local_to_remote( */ xfs_failaddr_t xfs_symlink_shortform_verify( - void *sfp, - int64_t size) + struct xfs_inode *ip) { + struct xfs_ifork *ifp = xfs_ifork_ptr(ip, XFS_DATA_FORK); + char *sfp = (char *)ifp->if_data; + int size = ifp->if_bytes; char *endp = sfp + size; + ASSERT(ifp->if_format == XFS_DINODE_FMT_LOCAL); + /* * Zero length symlinks should never occur in memory as they are * never allowed to exist on disk. diff --git a/fs/xfs/libxfs/xfs_symlink_remote.h b/fs/xfs/libxfs/xfs_symlink_remote.h index c1672fe1f17b..3f6590602473 100644 --- a/fs/xfs/libxfs/xfs_symlink_remote.h +++ b/fs/xfs/libxfs/xfs_symlink_remote.h @@ -18,7 +18,7 @@ bool xfs_symlink_hdr_ok(xfs_ino_t ino, uint32_t offset, void xfs_symlink_local_to_remote(struct xfs_trans *tp, struct xfs_buf *bp, struct xfs_inode *ip, struct xfs_ifork *ifp, void *priv); -xfs_failaddr_t xfs_symlink_shortform_verify(void *sfp, int64_t size); +xfs_failaddr_t xfs_symlink_shortform_verify(struct xfs_inode *ip); int xfs_symlink_remote_read(struct xfs_inode *ip, char *link); int xfs_symlink_write_target(struct xfs_trans *tp, struct xfs_inode *ip, xfs_ino_t owner, const char *target_path, int pathlen, diff --git a/fs/xfs/libxfs/xfs_trans_resv.c b/fs/xfs/libxfs/xfs_trans_resv.c index 3151e97ca8ff..b09c88aec185 100644 --- a/fs/xfs/libxfs/xfs_trans_resv.c +++ b/fs/xfs/libxfs/xfs_trans_resv.c @@ -606,10 +606,10 @@ static inline unsigned int xfs_calc_pptr_replace_overhead(void) */ STATIC uint xfs_calc_rename_reservation( - struct xfs_mount *mp) + struct xfs_mount *mp, + struct xfs_trans_resv *resp) { unsigned int overhead = XFS_DQUOT_LOGRES; - struct xfs_trans_resv *resp = M_RES(mp); unsigned int t1, t2, t3 = 0; t1 = xfs_calc_inode_res(mp, 5) + @@ -715,10 +715,10 @@ xfs_link_log_count( */ STATIC uint xfs_calc_link_reservation( - struct xfs_mount *mp) + struct xfs_mount *mp, + struct xfs_trans_resv *resp) { unsigned int overhead = XFS_DQUOT_LOGRES; - struct xfs_trans_resv *resp = M_RES(mp); unsigned int t1, t2, t3 = 0; overhead += xfs_calc_iunlink_remove_reservation(mp); @@ -777,10 +777,10 @@ xfs_remove_log_count( */ STATIC uint xfs_calc_remove_reservation( - struct xfs_mount *mp) + struct xfs_mount *mp, + struct xfs_trans_resv *resp) { unsigned int overhead = XFS_DQUOT_LOGRES; - struct xfs_trans_resv *resp = M_RES(mp); unsigned int t1, t2, t3 = 0; overhead += xfs_calc_iunlink_add_reservation(mp); @@ -862,9 +862,9 @@ xfs_icreate_log_count( STATIC uint xfs_calc_icreate_reservation( - struct xfs_mount *mp) + struct xfs_mount *mp, + struct xfs_trans_resv *resp) { - struct xfs_trans_resv *resp = M_RES(mp); unsigned int overhead = XFS_DQUOT_LOGRES; unsigned int t1, t2, t3 = 0; @@ -911,9 +911,10 @@ xfs_mkdir_log_count( */ STATIC uint xfs_calc_mkdir_reservation( - struct xfs_mount *mp) + struct xfs_mount *mp, + struct xfs_trans_resv *resp) { - return xfs_calc_icreate_reservation(mp); + return xfs_calc_icreate_reservation(mp, resp); } static inline unsigned int @@ -940,9 +941,10 @@ xfs_symlink_log_count( */ STATIC uint xfs_calc_symlink_reservation( - struct xfs_mount *mp) + struct xfs_mount *mp, + struct xfs_trans_resv *resp) { - return xfs_calc_icreate_reservation(mp) + + return xfs_calc_icreate_reservation(mp, resp) + xfs_calc_buf_res(1, XFS_SYMLINK_MAXLEN); } @@ -1265,34 +1267,33 @@ xfs_calc_namespace_reservations( { ASSERT(resp->tr_attrsetm.tr_logres > 0); - resp->tr_rename.tr_logres = xfs_calc_rename_reservation(mp); + resp->tr_rename.tr_logres = xfs_calc_rename_reservation(mp, resp); resp->tr_rename.tr_logcount = xfs_rename_log_count(mp, resp); resp->tr_rename.tr_logflags |= XFS_TRANS_PERM_LOG_RES; - resp->tr_link.tr_logres = xfs_calc_link_reservation(mp); + resp->tr_link.tr_logres = xfs_calc_link_reservation(mp, resp); resp->tr_link.tr_logcount = xfs_link_log_count(mp, resp); resp->tr_link.tr_logflags |= XFS_TRANS_PERM_LOG_RES; - resp->tr_remove.tr_logres = xfs_calc_remove_reservation(mp); + resp->tr_remove.tr_logres = xfs_calc_remove_reservation(mp, resp); resp->tr_remove.tr_logcount = xfs_remove_log_count(mp, resp); resp->tr_remove.tr_logflags |= XFS_TRANS_PERM_LOG_RES; - resp->tr_symlink.tr_logres = xfs_calc_symlink_reservation(mp); + resp->tr_symlink.tr_logres = xfs_calc_symlink_reservation(mp, resp); resp->tr_symlink.tr_logcount = xfs_symlink_log_count(mp, resp); resp->tr_symlink.tr_logflags |= XFS_TRANS_PERM_LOG_RES; - resp->tr_create.tr_logres = xfs_calc_icreate_reservation(mp); + resp->tr_create.tr_logres = xfs_calc_icreate_reservation(mp, resp); resp->tr_create.tr_logcount = xfs_icreate_log_count(mp, resp); resp->tr_create.tr_logflags |= XFS_TRANS_PERM_LOG_RES; - resp->tr_mkdir.tr_logres = xfs_calc_mkdir_reservation(mp); + resp->tr_mkdir.tr_logres = xfs_calc_mkdir_reservation(mp, resp); resp->tr_mkdir.tr_logcount = xfs_mkdir_log_count(mp, resp); resp->tr_mkdir.tr_logflags |= XFS_TRANS_PERM_LOG_RES; } STATIC void xfs_calc_default_atomic_ioend_reservation( - struct xfs_mount *mp, struct xfs_trans_resv *resp) { /* Pick a default that will scale reasonably for the log size. */ @@ -1398,7 +1399,7 @@ xfs_trans_resv_calc( * Now that we've finished computing the static reservations, we can * compute the dynamic reservation for atomic writes. */ - xfs_calc_default_atomic_ioend_reservation(mp, resp); + xfs_calc_default_atomic_ioend_reservation(resp); } /* @@ -1508,7 +1509,7 @@ xfs_calc_atomic_write_log_geometry( ASSERT(blockcount > 0); - xfs_calc_default_atomic_ioend_reservation(mp, M_RES(mp)); + xfs_calc_default_atomic_ioend_reservation(M_RES(mp)); per_intent = xfs_calc_atomic_write_ioend_geometry(mp, &step_size); @@ -1545,7 +1546,7 @@ xfs_calc_atomic_write_reservation( * use the defaults. */ if (blockcount == 0) { - xfs_calc_default_atomic_ioend_reservation(mp, M_RES(mp)); + xfs_calc_default_atomic_ioend_reservation(M_RES(mp)); return 0; } diff --git a/fs/xfs/libxfs/xfs_trans_space.c b/fs/xfs/libxfs/xfs_trans_space.c index c4cd547033e5..7edd0d86f0bd 100644 --- a/fs/xfs/libxfs/xfs_trans_space.c +++ b/fs/xfs/libxfs/xfs_trans_space.c @@ -127,10 +127,10 @@ xfs_rename_space_res( if (has_whiteout) ret += xfs_parent_calc_space_res(mp, src_namelen); ret += 2 * xfs_parent_calc_space_res(mp, target_namelen); - } - if (target_exists) - ret += xfs_parent_calc_space_res(mp, target_namelen); + if (target_exists) + ret += xfs_parent_calc_space_res(mp, target_namelen); + } return ret; } diff --git a/fs/xfs/libxfs/xfs_types.c b/fs/xfs/libxfs/xfs_types.c index 67c947a47f14..f195a04dbf66 100644 --- a/fs/xfs/libxfs/xfs_types.c +++ b/fs/xfs/libxfs/xfs_types.c @@ -222,7 +222,6 @@ xfs_verify_icount( /* Sanity-checking of dir/attr block offsets. */ bool xfs_verify_dablk( - struct xfs_mount *mp, xfs_fileoff_t dabno) { xfs_dablk_t max_dablk = -1U; @@ -233,7 +232,6 @@ xfs_verify_dablk( /* Check that a file block offset does not exceed the maximum. */ bool xfs_verify_fileoff( - struct xfs_mount *mp, xfs_fileoff_t off) { return off <= XFS_MAX_FILEOFF; @@ -242,15 +240,14 @@ xfs_verify_fileoff( /* Check that a range of file block offsets do not exceed the maximum. */ bool xfs_verify_fileext( - struct xfs_mount *mp, xfs_fileoff_t off, xfs_fileoff_t len) { if (off + len <= off) return false; - if (!xfs_verify_fileoff(mp, off)) + if (!xfs_verify_fileoff(off)) return false; - return xfs_verify_fileoff(mp, off + len - 1); + return xfs_verify_fileoff(off + len - 1); } diff --git a/fs/xfs/libxfs/xfs_types.h b/fs/xfs/libxfs/xfs_types.h index f6f4f2d4b5db..19dd5e7c8b12 100644 --- a/fs/xfs/libxfs/xfs_types.h +++ b/fs/xfs/libxfs/xfs_types.h @@ -277,11 +277,10 @@ bool xfs_verify_rtbno(struct xfs_mount *mp, xfs_rtblock_t rtbno); bool xfs_verify_rtbext(struct xfs_mount *mp, xfs_rtblock_t rtbno, xfs_filblks_t len); bool xfs_verify_icount(struct xfs_mount *mp, unsigned long long icount); -bool xfs_verify_dablk(struct xfs_mount *mp, xfs_fileoff_t off); +bool xfs_verify_dablk(xfs_fileoff_t off); void xfs_icount_range(struct xfs_mount *mp, unsigned long long *min, unsigned long long *max); -bool xfs_verify_fileoff(struct xfs_mount *mp, xfs_fileoff_t off); -bool xfs_verify_fileext(struct xfs_mount *mp, xfs_fileoff_t off, - xfs_fileoff_t len); +bool xfs_verify_fileoff(xfs_fileoff_t off); +bool xfs_verify_fileext(xfs_fileoff_t off, xfs_fileoff_t len); #endif /* __XFS_TYPES_H__ */ diff --git a/fs/xfs/scrub/alloc_repair.c b/fs/xfs/scrub/alloc_repair.c index 2398e3819597..6e8352405361 100644 --- a/fs/xfs/scrub/alloc_repair.c +++ b/fs/xfs/scrub/alloc_repair.c @@ -108,9 +108,6 @@ struct xrep_abt { struct xfs_scrub *sc; - /* Number of non-null records in @free_records. */ - uint64_t nr_real_records; - /* get_records()'s position in the free space record array. */ xfarray_idx_t array_cur; @@ -403,7 +400,6 @@ xrep_abt_find_freespace( if (error) goto err_agfl; - ra->nr_real_records = xfarray_length(ra->free_records); err_agfl: xfs_trans_brelse(sc->tp, agfl_bp); err: @@ -446,15 +442,17 @@ xrep_abt_reserve_space( uint64_t required; unsigned int desired; unsigned int len; + const uint64_t nr_records = + xfarray_length(ra->free_records); /* Compute how many blocks we'll need. */ error = xfs_btree_bload_compute_geometry(cnt_cur, - &ra->new_cntbt.bload, ra->nr_real_records); + &ra->new_cntbt.bload, nr_records); if (error) break; error = xfs_btree_bload_compute_geometry(bno_cur, - &ra->new_bnobt.bload, ra->nr_real_records); + &ra->new_bnobt.bload, nr_records); if (error) break; @@ -470,7 +468,7 @@ xrep_abt_reserve_space( desired = required - allocated; /* We need space but there's none left; bye! */ - if (ra->nr_real_records == 0) { + if (nr_records == 0) { error = -ENOSPC; break; } @@ -517,10 +515,9 @@ xrep_abt_reserve_space( * records (but doesn't break the sorting order), so we must * go around the loop once more to re-run _bload_init. */ - error = xfarray_unset(ra->free_records, record_nr); + error = xfarray_trim(ra->free_records, 1); if (error) break; - ra->nr_real_records--; record_nr--; } while (1); diff --git a/fs/xfs/scrub/bmap.c b/fs/xfs/scrub/bmap.c index 4f3c7f681bd9..c190590bc562 100644 --- a/fs/xfs/scrub/bmap.c +++ b/fs/xfs/scrub/bmap.c @@ -453,18 +453,17 @@ xchk_bmap_dirattr_extent( struct xchk_bmap_info *info, struct xfs_bmbt_irec *irec) { - struct xfs_mount *mp = ip->i_mount; xfs_fileoff_t off; if (!S_ISDIR(VFS_I(ip)->i_mode) && info->whichfork != XFS_ATTR_FORK) return; - if (!xfs_verify_dablk(mp, irec->br_startoff)) + if (!xfs_verify_dablk(irec->br_startoff)) xchk_fblock_set_corrupt(info->sc, info->whichfork, irec->br_startoff); off = irec->br_startoff + irec->br_blockcount - 1; - if (!xfs_verify_dablk(mp, off)) + if (!xfs_verify_dablk(off)) xchk_fblock_set_corrupt(info->sc, info->whichfork, off); } @@ -486,7 +485,7 @@ xchk_bmap_iextent( xchk_fblock_set_corrupt(info->sc, info->whichfork, irec->br_startoff); - if (!xfs_verify_fileext(mp, irec->br_startoff, irec->br_blockcount)) + if (!xfs_verify_fileext(irec->br_startoff, irec->br_blockcount)) xchk_fblock_set_corrupt(info->sc, info->whichfork, irec->br_startoff); @@ -877,8 +876,6 @@ xchk_bmap_iextent_delalloc( struct xchk_bmap_info *info, struct xfs_bmbt_irec *irec) { - struct xfs_mount *mp = info->sc->mp; - /* * Check for out-of-order extents. This record could have come * from the incore list, for which there is no ordering check. @@ -888,7 +885,7 @@ xchk_bmap_iextent_delalloc( xchk_fblock_set_corrupt(info->sc, info->whichfork, irec->br_startoff); - if (!xfs_verify_fileext(mp, irec->br_startoff, irec->br_blockcount)) + if (!xfs_verify_fileext(irec->br_startoff, irec->br_blockcount)) xchk_fblock_set_corrupt(info->sc, info->whichfork, irec->br_startoff); diff --git a/fs/xfs/scrub/bmap_repair.c b/fs/xfs/scrub/bmap_repair.c index eabffba47776..03af6cb92fcf 100644 --- a/fs/xfs/scrub/bmap_repair.c +++ b/fs/xfs/scrub/bmap_repair.c @@ -211,7 +211,7 @@ xrep_bmap_check_fork_rmap( /* Check the file offset range. */ if (!(rec->rm_flags & XFS_RMAP_BMBT_BLOCK) && - !xfs_verify_fileext(sc->mp, rec->rm_offset, rec->rm_blockcount)) + !xfs_verify_fileext(rec->rm_offset, rec->rm_blockcount)) return -EFSCORRUPTED; /* No contradictory flags. */ @@ -389,7 +389,7 @@ xrep_bmap_check_rtfork_rmap( return -EFSCORRUPTED; /* Check the file offsets and physical extents. */ - if (!xfs_verify_fileext(sc->mp, rec->rm_offset, rec->rm_blockcount)) + if (!xfs_verify_fileext(rec->rm_offset, rec->rm_blockcount)) return -EFSCORRUPTED; /* Check that this is within the rtgroup. */ diff --git a/fs/xfs/scrub/inode_repair.c b/fs/xfs/scrub/inode_repair.c index b87c22146233..c65912d87659 100644 --- a/fs/xfs/scrub/inode_repair.c +++ b/fs/xfs/scrub/inode_repair.c @@ -933,7 +933,7 @@ xrep_dinode_bad_bmbt_fork( fkp = xfs_bmdr_key_addr(dfp, i); fileoff = be64_to_cpu(fkp->br_startoff); - if (!xfs_verify_fileoff(sc->mp, fileoff)) + if (!xfs_verify_fileoff(fileoff)) return true; fpp = xfs_bmdr_ptr_addr(dfp, i, dmxr); @@ -1024,6 +1024,30 @@ xrep_dinode_bad_metabt_fork( return false; } +static xfs_failaddr_t +xrep_symlink_shortform_verify( + void *sfp, + int64_t size) +{ + /* + * Zero length symlinks should never occur in memory as they are + * never allowed to exist on disk. + */ + if (!size) + return __this_address; + + /* No negative sizes or overly long symlink targets. */ + if (size < 0 || size > XFS_SYMLINK_MAXLEN) + return __this_address; + + /* No NULLs in the target either. */ + if (memchr(sfp, 0, size)) + return __this_address; + + /* ondisk symlink target isn't null terminated, unlike incore */ + return NULL; +} + /* * Check the data fork for things that will fail the ifork verifiers or the * ifork formatters. @@ -1099,7 +1123,7 @@ xrep_dinode_check_dfork( return true; /* symlink structure must pass verification. */ if (S_ISLNK(mode) && - xfs_symlink_shortform_verify(dfork_ptr, data_size) != NULL) + xrep_symlink_shortform_verify(dfork_ptr, data_size) != NULL) return true; break; case XFS_DINODE_FMT_EXTENTS: @@ -1405,7 +1429,7 @@ xrep_dinode_ensure_forkoff( break; case XFS_METAFILE_RTREFCOUNT: rcdr = XFS_DFORK_PTR(dip, XFS_DATA_FORK); - dfork_min = xfs_rtrefcount_broot_space(sc->mp, rcdr); + dfork_min = xfs_rtrefcount_broot_space(rcdr); break; default: dfork_min = 0; @@ -1949,7 +1973,7 @@ xrep_inode_pptr( return 0; return xfs_bmap_add_attrfork(sc->tp, ip, - sizeof(struct xfs_attr_sf_hdr), true); + sizeof(struct xfs_attr_sf_hdr)); } /* Fix COW extent size hint problems. */ diff --git a/fs/xfs/scrub/orphanage.c b/fs/xfs/scrub/orphanage.c index 21e31eeaa042..d8e1457f0d5e 100644 --- a/fs/xfs/scrub/orphanage.c +++ b/fs/xfs/scrub/orphanage.c @@ -550,7 +550,7 @@ xrep_adoption_move( if (!xfs_inode_has_attr_fork(sc->ip) && xfs_has_parent(sc->mp)) { int sf_size = xrep_adoption_attr_sizeof(adopt); - error = xfs_bmap_add_attrfork(sc->tp, sc->ip, sf_size, true); + error = xfs_bmap_add_attrfork(sc->tp, sc->ip, sf_size); if (error) return error; } diff --git a/fs/xfs/scrub/quota.c b/fs/xfs/scrub/quota.c index 222812fe202c..8c6ba1240fd5 100644 --- a/fs/xfs/scrub/quota.c +++ b/fs/xfs/scrub/quota.c @@ -89,7 +89,7 @@ xchk_quota_item_bmap( int nmaps = 1; int error; - if (!xfs_verify_fileoff(mp, offset)) { + if (!xfs_verify_fileoff(offset)) { xchk_fblock_set_corrupt(sc, XFS_DATA_FORK, offset); return 0; } diff --git a/fs/xfs/scrub/quota_repair.c b/fs/xfs/scrub/quota_repair.c index 59302e8afc7e..40bb85e8f909 100644 --- a/fs/xfs/scrub/quota_repair.c +++ b/fs/xfs/scrub/quota_repair.c @@ -116,8 +116,8 @@ xrep_quota_item_bmap( int error; /* The computed file offset should always be valid. */ - if (!xfs_verify_fileoff(mp, offset)) { - ASSERT(xfs_verify_fileoff(mp, offset)); + if (!xfs_verify_fileoff(offset)) { + ASSERT(xfs_verify_fileoff(offset)); return -EFSCORRUPTED; } dq->q_fileoffset = offset; @@ -248,10 +248,7 @@ xrep_quota_item( dq->q_flags |= XFS_DQFLAG_DIRTY; xfs_trans_dqjoin(sc->tp, dq); - if (dq->q_id) { - xfs_qm_adjust_dqlimits(dq); - xfs_qm_adjust_dqtimers(dq); - } + xfs_qm_adjust_dqenforcement(dq); xfs_trans_log_dquot(sc->tp, dq); return xfs_trans_roll(&sc->tp); diff --git a/fs/xfs/scrub/quotacheck_repair.c b/fs/xfs/scrub/quotacheck_repair.c index dbb522e1513b..b8334edf380c 100644 --- a/fs/xfs/scrub/quotacheck_repair.c +++ b/fs/xfs/scrub/quotacheck_repair.c @@ -39,6 +39,54 @@ * dquot is locked. */ +static bool +xqcheck_dqres_force_dirty( + const struct xfs_dquot_res *res, + const struct xfs_quota_limits *qlim) +{ + /* zero limits mean that we should set the default limits */ + if (res->softlimit == 0 && qlim->soft != 0) + return true; + if (res->hardlimit == 0 && qlim->hard != 0) + return true; + + /* do we need to adjust the timer setting? */ + if ((res->softlimit && res->count > res->softlimit) || + (res->hardlimit && res->count > res->hardlimit)) { + if (!res->timer) + return true; + } else { + if (res->timer) + return true; + } + + return false; +} + +/* Decide if we need to adjust the dquot limits or timers */ +static bool +xqcheck_dquot_force_dirty( + const struct xfs_dquot *dq) +{ + struct xfs_quotainfo *qi = dq->q_mount->m_quotainfo; + struct xfs_def_quota *defq; + + /* root dquot does not enforce limits */ + if (dq->q_id == 0) + return false; + + defq = xfs_get_defquota(qi, xfs_dquot_type(dq)); + + if (xqcheck_dqres_force_dirty(&dq->q_blk, &defq->blk)) + return true; + if (xqcheck_dqres_force_dirty(&dq->q_ino, &defq->ino)) + return true; + if (xqcheck_dqres_force_dirty(&dq->q_rtb, &defq->rtb)) + return true; + + return false; +} + /* Commit new counters to a dquot. */ static int xqcheck_commit_dquot( @@ -91,6 +139,9 @@ xqcheck_commit_dquot( dirty = true; } + if (!dirty && xqcheck_dquot_force_dirty(dq)) + dirty = true; + xcdq.flags |= (XQCHECK_DQUOT_REPAIR_SCANNED | XQCHECK_DQUOT_WRITTEN); error = xfarray_store(counts, dq->q_id, &xcdq); if (error == -EFBIG) { @@ -110,8 +161,7 @@ xqcheck_commit_dquot( /* Commit the dirty dquot to disk. */ dq->q_flags |= XFS_DQFLAG_DIRTY; - if (dq->q_id) - xfs_qm_adjust_dqtimers(dq); + xfs_qm_adjust_dqenforcement(dq); xfs_trans_log_dquot(xqc->sc->tp, dq); return xrep_trans_commit(xqc->sc); diff --git a/fs/xfs/scrub/rtrefcount_repair.c b/fs/xfs/scrub/rtrefcount_repair.c index 2b939960c7dd..c78a6d2990c5 100644 --- a/fs/xfs/scrub/rtrefcount_repair.c +++ b/fs/xfs/scrub/rtrefcount_repair.c @@ -596,8 +596,7 @@ xrep_rtrefc_iroot_size( unsigned int nr_this_level, void *priv) { - return xfs_rtrefcount_broot_space_calc(cur->bc_mp, level, - nr_this_level); + return xfs_rtrefcount_broot_space_calc(level, nr_this_level); } /* diff --git a/fs/xfs/scrub/rtrmap_repair.c b/fs/xfs/scrub/rtrmap_repair.c index a2b72e61edf5..5cfa4470c57c 100644 --- a/fs/xfs/scrub/rtrmap_repair.c +++ b/fs/xfs/scrub/rtrmap_repair.c @@ -693,7 +693,7 @@ xrep_rtrmap_iroot_size( unsigned int nr_this_level, void *priv) { - return xfs_rtrmap_broot_space_calc(cur->bc_mp, level, nr_this_level); + return xfs_rtrmap_broot_space_calc(level, nr_this_level); } /* diff --git a/fs/xfs/scrub/xfarray.c b/fs/xfs/scrub/xfarray.c index 2ce24bfe4c0f..c94f35560779 100644 --- a/fs/xfs/scrub/xfarray.c +++ b/fs/xfs/scrub/xfarray.c @@ -140,55 +140,23 @@ xfarray_load( xfarray_pos(array, idx)); } -/* Is this array element potentially unset? */ -static inline bool -xfarray_is_unset( - struct xfarray *array, - loff_t pos) -{ - void *temp = xfarray_scratch(array); - int error; - - if (array->unset_slots == 0) - return false; - - error = xfile_load(array->xfile, temp, array->obj_size, pos); - if (!error && xfarray_element_is_null(array, temp)) - return true; - - return false; -} - -/* - * Unset an array element. If @idx is the last element in the array, the - * array will be truncated. Otherwise, the entry will be zeroed. - */ +/* Remove the elements at the end of an array. */ int -xfarray_unset( - struct xfarray *array, - xfarray_idx_t idx) +xfarray_trim( + struct xfarray *array, + unsigned long long nr) { - void *temp = xfarray_scratch(array); - loff_t pos = xfarray_pos(array, idx); - int error; + loff_t new_eof; - if (idx >= array->nr) + if (nr > array->nr) return -ENODATA; - if (idx == array->nr - 1) { - array->nr--; - return 0; - } - - if (xfarray_is_unset(array, pos)) - return 0; - - memset(temp, 0, array->obj_size); - error = xfile_store(array->xfile, temp, array->obj_size, pos); - if (error) - return error; + array->nr -= nr; + if (!array->nr) + array->possibly_sparse = false; - array->unset_slots++; + new_eof = xfarray_pos(array, array->nr); + xfile_discard(array->xfile, new_eof, MAX_LFS_FILESIZE - new_eof); return 0; } @@ -214,6 +182,8 @@ xfarray_store( if (ret) return ret; + if (idx > array->nr) + array->possibly_sparse = true; array->nr = max(array->nr, idx + 1); return 0; } @@ -227,43 +197,6 @@ xfarray_element_is_null( return !memchr_inv(ptr, 0, array->obj_size); } -/* - * Store an element anywhere in the array that is unset. If there are no - * unset slots, append the element to the array. - */ -int -xfarray_store_anywhere( - struct xfarray *array, - const void *ptr) -{ - void *temp = xfarray_scratch(array); - loff_t endpos = xfarray_pos(array, array->nr); - loff_t pos; - int error; - - /* Find an unset slot to put it in. */ - for (pos = 0; - pos < endpos && array->unset_slots > 0; - pos += array->obj_size) { - error = xfile_load(array->xfile, temp, array->obj_size, - pos); - if (error || !xfarray_element_is_null(array, temp)) - continue; - - error = xfile_store(array->xfile, ptr, array->obj_size, - pos); - if (error) - return error; - - array->unset_slots--; - return 0; - } - - /* No unset slots found; attach it on the end. */ - array->unset_slots = 0; - return xfarray_append(array, ptr); -} - /* Return length of array. */ uint64_t xfarray_length( @@ -677,26 +610,10 @@ xfarray_qsort_pivot( /* Load the selected xfarray records into the pivot array. */ for (i = 0; i < XFARRAY_QSORT_PIVOT_NR; i++) { - xfarray_idx_t idx; - recp = xfarray_pivot_array_rec(parray, pivot_rec_sz, i); idxp = xfarray_pivot_array_idx(parray, pivot_rec_sz, i); - /* No unset records; load directly into the array. */ - if (likely(si->array->unset_slots == 0)) { - error = xfarray_sort_load(si, *idxp, recp); - if (error) - return error; - continue; - } - - /* - * Load non-null records into the scratchpad without changing - * the xfarray_idx_t in the pivot array. - */ - idx = *idxp; - xfarray_sort_bump_loads(si); - error = xfarray_load_next(si->array, &idx, recp); + error = xfarray_sort_load(si, *idxp, recp); if (error) return error; } @@ -765,6 +682,18 @@ xfarray_qsort_push( return -EFSCORRUPTED; } + /* + * Avoid the integer underflow below in (lo - 1). This shouldn't + * be possible because the pivot is the median of nine distinct + * filesystem metadata records, so at least four records will be less + * than the pivot, which means the pivot will not be in the low end of + * the range by the time we get here. + */ + if (lo == 0) { + ASSERT(lo != 0); + return -EFSCORRUPTED; + } + si->max_stack_used = max_t(uint8_t, si->max_stack_used, si->stack_depth + 2); @@ -794,6 +723,46 @@ xfarray_sort_scan_done( si->folio = NULL; } +static int +xfarray_sort_load_folio( + struct xfarray_sortinfo *si, + xfarray_idx_t idx, + loff_t idx_pos) +{ + struct folio *folio; + loff_t next_pos; + + folio = xfile_get_folio(si->array->xfile, idx_pos, si->array->obj_size, + XFILE_ALLOC); + if (IS_ERR(folio)) + return PTR_ERR(folio); + si->folio = folio; + + /* No folio? Get the caller to read into the scratchpad. */ + if (!si->folio) + return 0; + + si->first_folio_idx = xfarray_idx(si->array, + folio_pos(si->folio) + si->array->obj_size - 1); + + next_pos = folio_next_pos(si->folio); + si->last_folio_idx = xfarray_idx(si->array, next_pos - 1); + if (xfarray_pos(si->array, si->last_folio_idx + 1) > next_pos) + si->last_folio_idx--; + + /* + * If this folio still doesn't cover the desired element, it must cross + * a folio boundary. Get the caller to read into the scratchpad. + */ + if (idx < si->first_folio_idx || idx > si->last_folio_idx) { + xfarray_sort_scan_done(si); + return 0; + } + + trace_xfarray_sort_scan(si, idx); + return 0; +} + /* * Cache the folio backing the start of the given array element. If the array * element is contained entirely within the folio, return a pointer to the @@ -819,33 +788,18 @@ xfarray_sort_scan( (idx < si->first_folio_idx || idx > si->last_folio_idx)) xfarray_sort_scan_done(si); - /* Grab the first folio that backs this array element. */ + /* Grab the folio that backs this array element. */ if (!si->folio) { - struct folio *folio; - loff_t next_pos; - - folio = xfile_get_folio(si->array->xfile, idx_pos, - si->array->obj_size, XFILE_ALLOC); - if (IS_ERR(folio)) - return PTR_ERR(folio); - si->folio = folio; - - si->first_folio_idx = xfarray_idx(si->array, - folio_pos(si->folio) + si->array->obj_size - 1); - - next_pos = folio_next_pos(si->folio); - si->last_folio_idx = xfarray_idx(si->array, next_pos - 1); - if (xfarray_pos(si->array, si->last_folio_idx + 1) > next_pos) - si->last_folio_idx--; - - trace_xfarray_sort_scan(si, idx); + error = xfarray_sort_load_folio(si, idx, idx_pos); + if (error) + return error; } /* - * If this folio still doesn't cover the desired element, it must cross - * a folio boundary. Read into the scratchpad and we're done. + * If we don't have a folio mapping the entire array element, read into + * the scratchpad and we're done. */ - if (idx < si->first_folio_idx || idx > si->last_folio_idx) { + if (!si->folio) { void *temp = xfarray_scratch(si->array); error = xfile_load(si->array->xfile, temp, si->array->obj_size, @@ -916,6 +870,14 @@ xfarray_sort( return 0; if (array->nr >= QSORT_MAX_RECS) return -E2BIG; + if (array->possibly_sparse) { + /* + * What does it mean to sort an array with holes in it? + * Currently none of the users need this ability. + */ + ASSERT(!array->possibly_sparse); + return -EINVAL; + } error = xfarray_sortinfo_alloc(array, cmp_fn, flags, &si); if (error) @@ -1069,4 +1031,5 @@ xfarray_truncate( { xfile_discard(array->xfile, 0, MAX_LFS_FILESIZE); array->nr = 0; + array->possibly_sparse = false; } diff --git a/fs/xfs/scrub/xfarray.h b/fs/xfs/scrub/xfarray.h index 5eeeeed13ae2..05ff65b09fcf 100644 --- a/fs/xfs/scrub/xfarray.h +++ b/fs/xfs/scrub/xfarray.h @@ -27,23 +27,22 @@ struct xfarray { /* Maximum possible array size. */ xfarray_idx_t max_nr; - /* Number of unset slots in the array below @nr. */ - uint64_t unset_slots; - /* Size of an array element. */ size_t obj_size; /* log2 of array element size, if possible. */ int obj_size_log; + + /* Might there be sparse holes in this array? */ + bool possibly_sparse; }; int xfarray_create(const char *descr, unsigned long long required_capacity, size_t obj_size, struct xfarray **arrayp); void xfarray_destroy(struct xfarray *array); int xfarray_load(struct xfarray *array, xfarray_idx_t idx, void *ptr); -int xfarray_unset(struct xfarray *array, xfarray_idx_t idx); +int xfarray_trim(struct xfarray *array, unsigned long long nr); int xfarray_store(struct xfarray *array, xfarray_idx_t idx, const void *ptr); -int xfarray_store_anywhere(struct xfarray *array, const void *ptr); bool xfarray_element_is_null(struct xfarray *array, const void *ptr); void xfarray_truncate(struct xfarray *array); unsigned long long xfarray_bytes(struct xfarray *array); diff --git a/fs/xfs/xfs_acl.c b/fs/xfs/xfs_acl.c index fdfca6fc75b6..20d87b52c4fc 100644 --- a/fs/xfs/xfs_acl.c +++ b/fs/xfs/xfs_acl.c @@ -243,7 +243,7 @@ xfs_acl_set_mode( } int -xfs_set_acl(struct mnt_idmap *idmap, struct dentry *dentry, +xfs_set_acl(const struct mnt_idmap *idmap, struct dentry *dentry, struct posix_acl *acl, int type) { umode_t mode; diff --git a/fs/xfs/xfs_acl.h b/fs/xfs/xfs_acl.h index bf7f960997d3..183526bec32c 100644 --- a/fs/xfs/xfs_acl.h +++ b/fs/xfs/xfs_acl.h @@ -11,7 +11,7 @@ struct posix_acl; #ifdef CONFIG_XFS_POSIX_ACL extern struct posix_acl *xfs_get_acl(struct inode *inode, int type, bool rcu); -extern int xfs_set_acl(struct mnt_idmap *idmap, struct dentry *dentry, +extern int xfs_set_acl(const struct mnt_idmap *idmap, struct dentry *dentry, struct posix_acl *acl, int type); extern int __xfs_set_acl(struct inode *inode, struct posix_acl *acl, int type); void xfs_forget_acl(struct inode *inode, const char *name); diff --git a/fs/xfs/xfs_aops.c b/fs/xfs/xfs_aops.c index 8b6119776fb3..c30e688cfc9f 100644 --- a/fs/xfs/xfs_aops.c +++ b/fs/xfs/xfs_aops.c @@ -23,7 +23,6 @@ #include "xfs_ioend.h" #include "xfs_zone_alloc.h" #include "xfs_rtgroup.h" -#include <linux/bio-integrity.h> struct xfs_writepage_ctx { struct iomap_writepage_ctx ctx; @@ -498,8 +497,7 @@ xfs_zoned_writeback_submit( bio_endio(&ioend->io_bio); return error; } - if (wpc->iomap.flags & IOMAP_F_INTEGRITY) - fs_bio_integrity_generate(&ioend->io_bio); + xfs_zone_alloc_and_submit(ioend, &XFS_ZWPC(wpc)->open_zone); return 0; } @@ -585,11 +583,10 @@ xfs_bio_submit_read( const struct iomap_iter *iter, struct iomap_read_folio_ctx *ctx) { - struct bio *bio = ctx->read_ctx; - - /* defer read completions to the ioend workqueue */ - iomap_init_ioend(iter->inode, bio, ctx->read_ctx_file_offset, 0); - iomap_bio_submit_read_endio(iter, ctx, xfs_end_bio); + xfs_ioend_submit_read(iter->inode, ctx->read_ctx, + ctx->read_ctx_file_offset, + iomap_ioend_flags(&iter->iomap)); + ctx->read_ctx = NULL; } static const struct iomap_read_ops xfs_iomap_read_ops = { diff --git a/fs/xfs/xfs_bmap_item.c b/fs/xfs/xfs_bmap_item.c index aa5b41629747..62be62342fea 100644 --- a/fs/xfs/xfs_bmap_item.c +++ b/fs/xfs/xfs_bmap_item.c @@ -442,7 +442,7 @@ xfs_bui_validate( if (!xfs_verify_ino(mp, map->me_owner)) return false; - if (!xfs_verify_fileext(mp, map->me_startoff, map->me_len)) + if (!xfs_verify_fileext(map->me_startoff, map->me_len)) return false; if (map->me_flags & XFS_BMAP_EXTENT_REALTIME) diff --git a/fs/xfs/xfs_buf.c b/fs/xfs/xfs_buf.c index 8256c1d13ce2..6c93b4f5629c 100644 --- a/fs/xfs/xfs_buf.c +++ b/fs/xfs/xfs_buf.c @@ -5,6 +5,7 @@ */ #include "xfs_platform.h" #include <linux/backing-dev.h> +#include <linux/blk-integrity.h> #include <linux/dax.h> #include "xfs_shared.h" @@ -1694,6 +1695,7 @@ xfs_configure_buftarg( struct xfs_mount *mp = btp->bt_mount; if (btp->bt_bdev) { + struct blk_integrity *bi = bdev_get_integrity(btp->bt_bdev); int error; error = bdev_validate_blocksize(btp->bt_bdev, sectorsize); @@ -1706,6 +1708,15 @@ xfs_configure_buftarg( if (bdev_can_atomic_write(btp->bt_bdev)) xfs_configure_buftarg_atomic_writes(btp); + + if (!bi) + ; + else if (btp->bt_bdev == btp->bt_mount->m_super->s_bdev) + xfs_info(mp, "using %s integrity profile", + blk_integrity_profile_name(bi)); + else + xfs_info(mp, "using %s integrity profile for %pg", + blk_integrity_profile_name(bi), btp->bt_bdev); } btp->bt_meta_sectorsize = sectorsize; diff --git a/fs/xfs/xfs_dquot.c b/fs/xfs/xfs_dquot.c index e696ee36c2e8..6d22ead562da 100644 --- a/fs/xfs/xfs_dquot.c +++ b/fs/xfs/xfs_dquot.c @@ -115,18 +115,15 @@ xfs_qm_dqdestroy( * We overwrite the dquot limits only if they are zero and this * is not the root dquot. */ -void +static void xfs_qm_adjust_dqlimits( struct xfs_dquot *dq) { struct xfs_mount *mp = dq->q_mount; struct xfs_quotainfo *q = mp->m_quotainfo; - struct xfs_def_quota *defq; + struct xfs_def_quota *defq = xfs_get_defquota(q, xfs_dquot_type(dq)); int prealloc = 0; - ASSERT(dq->q_id); - defq = xfs_get_defquota(q, xfs_dquot_type(dq)); - if (!dq->q_blk.softlimit) { dq->q_blk.softlimit = defq->blk.soft; prealloc = 1; @@ -223,6 +220,19 @@ xfs_qm_adjust_dqtimers( xfs_qm_adjust_res_timer(dq->q_mount, &dq->q_rtb, &defq->rtb); } +/* Adjust enforcement limits and timers after a change in usage. */ +void +xfs_qm_adjust_dqenforcement( + struct xfs_dquot *dq) +{ + if (dq->q_id == 0) + return; + + xfs_qm_adjust_dqlimits(dq); + xfs_qm_adjust_dqtimers(dq); + dq->q_flags |= XFS_DQFLAG_DIRTY; +} + /* * initialize a buffer full of dquots and log the whole thing */ diff --git a/fs/xfs/xfs_dquot.h b/fs/xfs/xfs_dquot.h index bbb824adca82..28c09704acc2 100644 --- a/fs/xfs/xfs_dquot.h +++ b/fs/xfs/xfs_dquot.h @@ -205,7 +205,7 @@ void xfs_qm_dqdestroy(struct xfs_dquot *dqp); int xfs_qm_dqflush(struct xfs_dquot *dqp, struct xfs_buf *bp); void xfs_qm_dqunpin_wait(struct xfs_dquot *dqp); void xfs_qm_adjust_dqtimers(struct xfs_dquot *d); -void xfs_qm_adjust_dqlimits(struct xfs_dquot *d); +void xfs_qm_adjust_dqenforcement(struct xfs_dquot *d); xfs_dqid_t xfs_qm_id_for_quotatype(struct xfs_inode *ip, xfs_dqtype_t type); int xfs_qm_dqget(struct xfs_mount *mp, xfs_dqid_t id, diff --git a/fs/xfs/xfs_exchmaps_item.c b/fs/xfs/xfs_exchmaps_item.c index dd5d92ca1010..dd104bf778ba 100644 --- a/fs/xfs/xfs_exchmaps_item.c +++ b/fs/xfs/xfs_exchmaps_item.c @@ -341,10 +341,10 @@ xfs_xmi_validate( !xfs_verify_ino(mp, xlf->xmi_inode2)) return false; - if (!xfs_verify_fileext(mp, xlf->xmi_startoff1, xlf->xmi_blockcount)) + if (!xfs_verify_fileext(xlf->xmi_startoff1, xlf->xmi_blockcount)) return false; - if (!xfs_verify_fileext(mp, xlf->xmi_startoff2, xlf->xmi_blockcount)) + if (!xfs_verify_fileext(xlf->xmi_startoff2, xlf->xmi_blockcount)) return false; if (xlf->xmi_flags & XFS_EXCHMAPS_SET_SIZES) { diff --git a/fs/xfs/xfs_exchrange.c b/fs/xfs/xfs_exchrange.c index fafb4e3f065c..a1001315d4a7 100644 --- a/fs/xfs/xfs_exchrange.c +++ b/fs/xfs/xfs_exchrange.c @@ -238,7 +238,7 @@ retry: trace_xfs_exchrange_before(ip2, 2); trace_xfs_exchrange_before(ip1, 1); - error = xfs_exchmaps_check_forks(mp, &req); + error = xfs_exchmaps_check_forks(&req); if (error) goto out_trans_cancel; diff --git a/fs/xfs/xfs_file.c b/fs/xfs/xfs_file.c index d8202da15aca..d164de6ff98b 100644 --- a/fs/xfs/xfs_file.c +++ b/fs/xfs/xfs_file.c @@ -37,6 +37,7 @@ #include <linux/fadvise.h> #include <linux/mount.h> #include <linux/filelock.h> +#include <linux/bio-integrity.h> static const struct vm_operations_struct xfs_file_vm_ops; @@ -222,9 +223,8 @@ xfs_dio_read_bounce_submit_io( struct bio *bio, loff_t file_offset) { - iomap_init_ioend(iter->inode, bio, file_offset, IOMAP_IOEND_DIRECT); - bio->bi_end_io = xfs_end_bio; - submit_bio(bio); + xfs_ioend_submit_read(iter->inode, bio, file_offset, + iomap_ioend_flags(&iter->iomap) | IOMAP_IOEND_DIRECT); } static const struct iomap_dio_ops xfs_dio_read_bounce_ops = { @@ -252,8 +252,7 @@ xfs_file_dio_read( return ret; if (mapping_stable_writes(iocb->ki_filp->f_mapping)) { ret = iomap_dio_rw(iocb, to, &xfs_read_iomap_ops, - &xfs_dio_read_bounce_ops, IOMAP_DIO_BOUNCE, - NULL, 0); + &xfs_dio_read_bounce_ops, 0, NULL, 0); } else { ret = iomap_dio_read_simple(iocb, to, xfs_read_iomap_begin); if (ret == -ENOTBLK) @@ -713,7 +712,7 @@ xfs_dio_zoned_submit_io( bio->bi_end_io = xfs_end_bio; ioend = iomap_init_ioend(iter->inode, bio, file_offset, - IOMAP_IOEND_DIRECT); + iomap_ioend_flags(&iter->iomap) | IOMAP_IOEND_DIRECT); xfs_zone_alloc_and_submit(ioend, &ac->open_zone); } @@ -1872,17 +1871,21 @@ xfs_file_release( return 0; /* - * If we can't get the iolock just skip truncating the blocks past EOF - * because we could deadlock with the mmap_lock otherwise. We'll get - * another chance to drop them once the last reference to the inode is - * dropped, so we'll never leak blocks permanently. + * If we can't get the iolock or if the filesystem is frozen, just skip + * truncating the blocks past EOF because we could deadlock with the + * mmap_lock or hang the close() call. We'll get another chance to drop + * them once the last reference to the inode is dropped, so we'll never + * leak blocks permanently. */ if (!xfs_iflags_test(ip, XFS_EOFBLOCKS_RELEASED) && - xfs_ilock_nowait(ip, XFS_IOLOCK_EXCL)) { - if (xfs_can_free_eofblocks(ip) && - !xfs_iflags_test_and_set(ip, XFS_EOFBLOCKS_RELEASED)) - xfs_free_eofblocks(ip); - xfs_iunlock(ip, XFS_IOLOCK_EXCL); + sb_start_write_trylock(mp->m_super)) { + if (xfs_ilock_nowait(ip, XFS_IOLOCK_EXCL)) { + if (xfs_can_free_eofblocks(ip) && + !xfs_iflags_test_and_set(ip, XFS_EOFBLOCKS_RELEASED)) + xfs_free_eofblocks(ip); + xfs_iunlock(ip, XFS_IOLOCK_EXCL); + } + sb_end_write(mp->m_super); } return 0; diff --git a/fs/xfs/xfs_handle.c b/fs/xfs/xfs_handle.c index 0689cade8f74..4924e676ae91 100644 --- a/fs/xfs/xfs_handle.c +++ b/fs/xfs/xfs_handle.c @@ -272,11 +272,11 @@ xfs_open_by_handle( path.mnt = mntget(parfilp->f_path.mnt); FD_PREPARE(fdf, 0, dentry_open(&path, hreq->oflags, cred)); - if (fdf.err) - return fdf.err; + if (fdf->fd < 0) + return fdf->fd; if (S_ISREG(inode->i_mode)) { - struct file *filp = fd_prepare_file(fdf); + struct file *filp = fdf->file; filp->f_flags |= O_NOATIME; filp->f_mode |= FMODE_NOCMTIME; @@ -409,7 +409,7 @@ xfs_ioc_attr_list( void *buffer; int error; - if (bufsize < sizeof(struct xfs_attrlist) || + if (bufsize < struct_size(alist, al_offset, 1) || bufsize > XFS_XATTR_LIST_MAX) return -EINVAL; diff --git a/fs/xfs/xfs_inode.c b/fs/xfs/xfs_inode.c index 621513d7215e..05a14da28031 100644 --- a/fs/xfs/xfs_inode.c +++ b/fs/xfs/xfs_inode.c @@ -755,7 +755,7 @@ xfs_create( *ipp = du.ip; xfs_iunlock(du.ip, XFS_ILOCK_EXCL); xfs_iunlock(dp, XFS_ILOCK_EXCL); - xfs_parent_finish(mp, du.ppargs); + xfs_parent_finish(du.ppargs); return 0; out_trans_cancel: @@ -772,7 +772,7 @@ xfs_create( xfs_irele(du.ip); } out_parent: - xfs_parent_finish(mp, du.ppargs); + xfs_parent_finish(du.ppargs); out_release_dquots: xfs_qm_dqrele(udqp); xfs_qm_dqrele(gdqp); @@ -972,7 +972,7 @@ xfs_link( error = xfs_trans_commit(tp); xfs_iunlock(tdp, XFS_ILOCK_EXCL); xfs_iunlock(sip, XFS_ILOCK_EXCL); - xfs_parent_finish(mp, du.ppargs); + xfs_parent_finish(du.ppargs); return error; error_return: @@ -980,7 +980,7 @@ xfs_link( xfs_iunlock(tdp, XFS_ILOCK_EXCL); xfs_iunlock(sip, XFS_ILOCK_EXCL); out_parent: - xfs_parent_finish(mp, du.ppargs); + xfs_parent_finish(du.ppargs); std_return: if (error == -ENOSPC && nospace_error) error = nospace_error; @@ -1064,7 +1064,7 @@ xfs_itruncate_extents_flags( * the page cache can't scale that far. */ first_unmap_block = XFS_B_TO_FSB(mp, (xfs_ufsize_t)new_size); - if (!xfs_verify_fileoff(mp, first_unmap_block)) { + if (!xfs_verify_fileoff(first_unmap_block)) { WARN_ON_ONCE(first_unmap_block > XFS_MAX_FILEOFF); return 0; } @@ -1985,7 +1985,7 @@ xfs_remove( xfs_iunlock(ip, XFS_ILOCK_EXCL); xfs_iunlock(dp, XFS_ILOCK_EXCL); - xfs_parent_finish(mp, du.ppargs); + xfs_parent_finish(du.ppargs); return 0; out_trans_cancel: @@ -1994,7 +1994,7 @@ xfs_remove( xfs_iunlock(ip, XFS_ILOCK_EXCL); xfs_iunlock(dp, XFS_ILOCK_EXCL); out_parent: - xfs_parent_finish(mp, du.ppargs); + xfs_parent_finish(du.ppargs); std_return: return error; } @@ -2084,7 +2084,7 @@ xfs_sort_inodes( */ static int xfs_rename_alloc_whiteout( - struct mnt_idmap *idmap, + const struct mnt_idmap *idmap, struct xfs_name *src_name, struct xfs_inode *dp, struct xfs_inode **wip) @@ -2130,7 +2130,7 @@ xfs_rename_alloc_whiteout( */ int xfs_rename( - struct mnt_idmap *idmap, + const struct mnt_idmap *idmap, struct xfs_inode *src_dp, struct xfs_name *src_name, struct xfs_inode *src_ip, @@ -2357,11 +2357,11 @@ out_trans_cancel: out_unlock: xfs_iunlock_rename(inodes, num_inodes); out_tgt_ppargs: - xfs_parent_finish(mp, du_tgt.ppargs); + xfs_parent_finish(du_tgt.ppargs); out_wip_ppargs: - xfs_parent_finish(mp, du_wip.ppargs); + xfs_parent_finish(du_wip.ppargs); out_src_ppargs: - xfs_parent_finish(mp, du_src.ppargs); + xfs_parent_finish(du_src.ppargs); out_release_wip: if (du_wip.ip) xfs_irele(du_wip.ip); diff --git a/fs/xfs/xfs_inode.h b/fs/xfs/xfs_inode.h index 1602027cd0aa..ca96ba096359 100644 --- a/fs/xfs/xfs_inode.h +++ b/fs/xfs/xfs_inode.h @@ -568,7 +568,7 @@ int xfs_remove(struct xfs_inode *dp, struct xfs_name *name, struct xfs_inode *ip); int xfs_link(struct xfs_inode *tdp, struct xfs_inode *sip, struct xfs_name *target_name); -int xfs_rename(struct mnt_idmap *idmap, +int xfs_rename(const struct mnt_idmap *idmap, struct xfs_inode *src_dp, struct xfs_name *src_name, struct xfs_inode *src_ip, struct xfs_inode *target_dp, struct xfs_name *target_name, diff --git a/fs/xfs/xfs_ioctl.c b/fs/xfs/xfs_ioctl.c index 96ca3e480cb9..f81b6e52ac40 100644 --- a/fs/xfs/xfs_ioctl.c +++ b/fs/xfs/xfs_ioctl.c @@ -748,7 +748,7 @@ xfs_ioctl_setattr_check_projid( int xfs_fileattr_set( - struct mnt_idmap *idmap, + const struct mnt_idmap *idmap, struct dentry *dentry, struct file_kattr *fa) { @@ -962,10 +962,10 @@ out_free_buf: } int -xfs_ioc_swapext( - xfs_swapext_t *sxp) +xfs_swapext( + struct xfs_swapext *sxp) { - xfs_inode_t *ip, *tip; + struct xfs_inode *ip, *tip; /* Pull information for the target fd */ CLASS(fd, f)((int)sxp->sx_fdtarget); @@ -1035,6 +1035,46 @@ xfs_ioc_getlabel( return 0; } +/* + * Same behavior as xfs_sync_sb, except that it is always synchronous and it + * also writes the superblock buffer to disk sector 0 immediately. + */ +static int +xfs_sync_sb_buf( + struct xfs_mount *mp, + bool update_rtsb) +{ + struct xfs_trans *tp; + int error; + + error = xfs_trans_alloc(mp, &M_RES(mp)->tr_sb, 0, 0, 0, &tp); + if (error) + return error; + + xfs_log_sb(tp); + if (update_rtsb) + xfs_log_rtsb(tp, xfs_trans_getsb(tp)); + xfs_trans_set_sync(tp); + error = xfs_trans_commit(tp); + if (error) + return error; + + /* Re-acquire and write the sb and rtsb to disk. */ + xfs_buf_lock(mp->m_sb_bp); + error = xfs_bwrite(mp->m_sb_bp); + xfs_buf_unlock(mp->m_sb_bp); + if (error) + return error; + + if (update_rtsb && mp->m_rtsb_bp) { + xfs_buf_lock(mp->m_rtsb_bp); + error = xfs_bwrite(mp->m_rtsb_bp); + xfs_buf_unlock(mp->m_rtsb_bp); + } + + return error; +} + static int xfs_ioc_setlabel( struct file *filp, @@ -1144,7 +1184,7 @@ xfs_fs_eofblocks_from_user( } static int -xfs_ioctl_getset_resblocks( +xfs_ioc_getset_resblocks( struct file *filp, unsigned int cmd, void __user *arg) @@ -1183,7 +1223,7 @@ xfs_ioctl_getset_resblocks( } static int -xfs_ioctl_fs_counts( +xfs_ioc_fs_counts( struct xfs_mount *mp, struct xfs_fsop_counts __user *uarg) { @@ -1200,6 +1240,202 @@ xfs_ioctl_fs_counts( return 0; } +static int +xfs_ioc_dioinfo( + struct file *file, + void __user *arg) +{ + struct kstat st; + struct dioattr da; + int error; + + error = vfs_getattr(&file->f_path, &st, STATX_DIOALIGN, 0); + if (error) + return error; + + /* + * Some userspace directly feeds the return value to posix_memalign, + * which fails for values that are smaller than the pointer size. + * Round up the value to not break userspace. + */ + da.d_mem = roundup(st.dio_mem_align, sizeof(void *)); + da.d_miniosz = st.dio_offset_align; + da.d_maxiosz = INT_MAX & ~(da.d_miniosz - 1); + if (copy_to_user(arg, &da, sizeof(da))) + return -EFAULT; + return 0; +} + +static int +xfs_ioc_find_handle( + unsigned int cmd, + void __user *arg) +{ + struct xfs_fsop_handlereq hreq; + + if (copy_from_user(&hreq, arg, sizeof(hreq))) + return -EFAULT; + return xfs_find_handle(cmd, &hreq); +} + +static int +xfs_ioc_open_by_handle( + struct file *file, + void __user *arg) +{ + struct xfs_fsop_handlereq hreq; + + if (copy_from_user(&hreq, arg, sizeof(hreq))) + return -EFAULT; + return xfs_open_by_handle(file, &hreq); +} + +static int +xfs_ioc_readlink_by_handle( + struct file *file, + void __user *arg) +{ + struct xfs_fsop_handlereq hreq; + + if (copy_from_user(&hreq, arg, sizeof(hreq))) + return -EFAULT; + return xfs_readlink_by_handle(file, &hreq); +} + +static int +xfs_ioc_swapext( + struct file *file, + void __user *arg) +{ + struct xfs_swapext sxp; + int error; + + if (copy_from_user(&sxp, arg, sizeof(sxp))) + return -EFAULT; + + error = mnt_want_write_file(file); + if (error) + return error; + error = xfs_swapext(&sxp); + mnt_drop_write_file(file); + return error; +} + +static int +xfs_ioc_growfs_data( + struct file *file, + struct xfs_mount *mp, + void __user *arg) +{ + struct xfs_growfs_data in; + int error; + + if (copy_from_user(&in, arg, sizeof(in))) + return -EFAULT; + + error = mnt_want_write_file(file); + if (error) + return error; + error = xfs_growfs_data(mp, &in); + mnt_drop_write_file(file); + return error; +} + +static int +xfs_ioc_growfs_log( + struct file *file, + struct xfs_mount *mp, + void __user *arg) +{ + struct xfs_growfs_log in; + int error; + + if (copy_from_user(&in, arg, sizeof(in))) + return -EFAULT; + + error = mnt_want_write_file(file); + if (error) + return error; + error = xfs_growfs_log(mp, &in); + mnt_drop_write_file(file); + return error; +} + +static int +xfs_ioc_growfs_rt( + struct file *file, + struct xfs_mount *mp, + void __user *arg) +{ + struct xfs_growfs_rt in; + int error; + + if (copy_from_user(&in, arg, sizeof(in))) + return -EFAULT; + + error = mnt_want_write_file(file); + if (error) + return error; + error = xfs_growfs_rt(mp, &in); + mnt_drop_write_file(file); + return error; +} + +static int +xfs_ioc_goingdown( + struct xfs_mount *mp, + uint32_t __user *arg) +{ + uint32_t in; + + if (!capable(CAP_SYS_ADMIN)) + return -EPERM; + if (get_user(in, arg)) + return -EFAULT; + return xfs_fs_goingdown(mp, in); +} + +static int +xfs_ioc_error_injection( + struct xfs_mount *mp, + struct xfs_error_injection __user *arg) +{ + struct xfs_error_injection in; + + if (!capable(CAP_SYS_ADMIN)) + return -EPERM; + if (copy_from_user(&in, arg, sizeof(in))) + return -EFAULT; + return xfs_errortag_add(mp, in.errtag); +} + +static int +xfs_ioc_free_eofblocks( + struct xfs_mount *mp, + struct xfs_fs_eofblocks __user *arg) +{ + struct xfs_fs_eofblocks eofb; + struct xfs_icwalk icw; + int error; + + if (!capable(CAP_SYS_ADMIN)) + return -EPERM; + if (xfs_is_readonly(mp)) + return -EROFS; + + if (copy_from_user(&eofb, arg, sizeof(eofb))) + return -EFAULT; + + error = xfs_fs_eofblocks_from_user(&eofb, &icw); + if (error) + return error; + + trace_xfs_ioc_free_eofblocks(mp, &icw, _RET_IP_); + + guard(super_write)(mp->m_super); + return xfs_blockgc_free_space(mp, &icw); +} + /* * These long-unused ioctls were removed from the official ioctl API in 5.17, * but retain these definitions so that we can log warnings about them. @@ -1209,12 +1445,6 @@ xfs_ioctl_fs_counts( #define XFS_IOC_ALLOCSP64 _IOW ('X', 36, struct xfs_flock64) #define XFS_IOC_FREESP64 _IOW ('X', 37, struct xfs_flock64) -/* - * Note: some of the ioctl's return positive numbers as a - * byte count indicating success, such as readlink_by_handle. - * So we don't "sign flip" like most other routines. This means - * true errors need to be returned as a negative value. - */ long xfs_file_ioctl( struct file *filp, @@ -1225,7 +1455,6 @@ xfs_file_ioctl( struct xfs_inode *ip = XFS_I(inode); struct xfs_mount *mp = ip->i_mount; void __user *arg = (void __user *)p; - int error; trace_xfs_file_ioctl(ip); @@ -1244,26 +1473,9 @@ xfs_file_ioctl( "%s should use fallocate; XFS_IOC_{ALLOC,FREE}SP ioctl unsupported", current->comm); return -ENOTTY; - case XFS_IOC_DIOINFO: { - struct kstat st; - struct dioattr da; - error = vfs_getattr(&filp->f_path, &st, STATX_DIOALIGN, 0); - if (error) - return error; - - /* - * Some userspace directly feeds the return value to - * posix_memalign, which fails for values that are smaller than - * the pointer size. Round up the value to not break userspace. - */ - da.d_mem = roundup(st.dio_mem_align, sizeof(void *)); - da.d_miniosz = st.dio_offset_align; - da.d_maxiosz = INT_MAX & ~(da.d_miniosz - 1); - if (copy_to_user(arg, &da, sizeof(da))) - return -EFAULT; - return 0; - } + case XFS_IOC_DIOINFO: + return xfs_ioc_dioinfo(filp, arg); case XFS_IOC_FSBULKSTAT_SINGLE: case XFS_IOC_FSBULKSTAT: @@ -1311,148 +1523,45 @@ xfs_file_ioctl( case XFS_IOC_FD_TO_HANDLE: case XFS_IOC_PATH_TO_HANDLE: - case XFS_IOC_PATH_TO_FSHANDLE: { - xfs_fsop_handlereq_t hreq; - - if (copy_from_user(&hreq, arg, sizeof(hreq))) - return -EFAULT; - return xfs_find_handle(cmd, &hreq); - } - case XFS_IOC_OPEN_BY_HANDLE: { - xfs_fsop_handlereq_t hreq; - - if (copy_from_user(&hreq, arg, sizeof(xfs_fsop_handlereq_t))) - return -EFAULT; - return xfs_open_by_handle(filp, &hreq); - } - - case XFS_IOC_READLINK_BY_HANDLE: { - xfs_fsop_handlereq_t hreq; - - if (copy_from_user(&hreq, arg, sizeof(xfs_fsop_handlereq_t))) - return -EFAULT; - return xfs_readlink_by_handle(filp, &hreq); - } + case XFS_IOC_PATH_TO_FSHANDLE: + return xfs_ioc_find_handle(cmd, arg); + case XFS_IOC_OPEN_BY_HANDLE: + return xfs_ioc_open_by_handle(filp, arg); + case XFS_IOC_READLINK_BY_HANDLE: + return xfs_ioc_readlink_by_handle(filp, arg); case XFS_IOC_ATTRLIST_BY_HANDLE: return xfs_attrlist_by_handle(filp, arg); - case XFS_IOC_ATTRMULTI_BY_HANDLE: return xfs_attrmulti_by_handle(filp, arg); - case XFS_IOC_SWAPEXT: { - struct xfs_swapext sxp; - - if (copy_from_user(&sxp, arg, sizeof(xfs_swapext_t))) - return -EFAULT; - error = mnt_want_write_file(filp); - if (error) - return error; - error = xfs_ioc_swapext(&sxp); - mnt_drop_write_file(filp); - return error; - } + case XFS_IOC_SWAPEXT: + return xfs_ioc_swapext(filp, arg); case XFS_IOC_FSCOUNTS: - return xfs_ioctl_fs_counts(mp, arg); + return xfs_ioc_fs_counts(mp, arg); case XFS_IOC_SET_RESBLKS: case XFS_IOC_GET_RESBLKS: - return xfs_ioctl_getset_resblocks(filp, cmd, arg); - - case XFS_IOC_FSGROWFSDATA: { - struct xfs_growfs_data in; - - if (copy_from_user(&in, arg, sizeof(in))) - return -EFAULT; - - error = mnt_want_write_file(filp); - if (error) - return error; - error = xfs_growfs_data(mp, &in); - mnt_drop_write_file(filp); - return error; - } - - case XFS_IOC_FSGROWFSLOG: { - struct xfs_growfs_log in; - - if (copy_from_user(&in, arg, sizeof(in))) - return -EFAULT; - - error = mnt_want_write_file(filp); - if (error) - return error; - error = xfs_growfs_log(mp, &in); - mnt_drop_write_file(filp); - return error; - } - - case XFS_IOC_FSGROWFSRT: { - xfs_growfs_rt_t in; - - if (copy_from_user(&in, arg, sizeof(in))) - return -EFAULT; - - error = mnt_want_write_file(filp); - if (error) - return error; - error = xfs_growfs_rt(mp, &in); - mnt_drop_write_file(filp); - return error; - } - - case XFS_IOC_GOINGDOWN: { - uint32_t in; - - if (!capable(CAP_SYS_ADMIN)) - return -EPERM; - - if (get_user(in, (uint32_t __user *)arg)) - return -EFAULT; - - return xfs_fs_goingdown(mp, in); - } - - case XFS_IOC_ERROR_INJECTION: { - xfs_error_injection_t in; - - if (!capable(CAP_SYS_ADMIN)) - return -EPERM; - - if (copy_from_user(&in, arg, sizeof(in))) - return -EFAULT; - - return xfs_errortag_add(mp, in.errtag); - } - + return xfs_ioc_getset_resblocks(filp, cmd, arg); + + case XFS_IOC_FSGROWFSDATA: + return xfs_ioc_growfs_data(filp, mp, arg); + case XFS_IOC_FSGROWFSLOG: + return xfs_ioc_growfs_log(filp, mp, arg); + case XFS_IOC_FSGROWFSRT: + return xfs_ioc_growfs_rt(filp, mp, arg); + + case XFS_IOC_GOINGDOWN: + return xfs_ioc_goingdown(mp, arg); + case XFS_IOC_ERROR_INJECTION: + return xfs_ioc_error_injection(mp, arg); case XFS_IOC_ERROR_CLEARALL: if (!capable(CAP_SYS_ADMIN)) return -EPERM; - return xfs_errortag_clearall(mp); - case XFS_IOC_FREE_EOFBLOCKS: { - struct xfs_fs_eofblocks eofb; - struct xfs_icwalk icw; - - if (!capable(CAP_SYS_ADMIN)) - return -EPERM; - - if (xfs_is_readonly(mp)) - return -EROFS; - - if (copy_from_user(&eofb, arg, sizeof(eofb))) - return -EFAULT; - - error = xfs_fs_eofblocks_from_user(&eofb, &icw); - if (error) - return error; - - trace_xfs_ioc_free_eofblocks(mp, &icw, _RET_IP_); - - guard(super_write)(mp->m_super); - return xfs_blockgc_free_space(mp, &icw); - } + case XFS_IOC_FREE_EOFBLOCKS: + return xfs_ioc_free_eofblocks(mp, arg); case XFS_IOC_EXCHANGE_RANGE: return xfs_ioc_exchange_range(filp, arg); diff --git a/fs/xfs/xfs_ioctl.h b/fs/xfs/xfs_ioctl.h index f5ed5cf9d3df..6e55cc847654 100644 --- a/fs/xfs/xfs_ioctl.h +++ b/fs/xfs/xfs_ioctl.h @@ -10,9 +10,7 @@ struct xfs_bstat; struct xfs_ibulk; struct xfs_inogrp; -int -xfs_ioc_swapext( - xfs_swapext_t *sxp); +int xfs_swapext(struct xfs_swapext *sxp); extern int xfs_fileattr_get( @@ -21,7 +19,7 @@ xfs_fileattr_get( extern int xfs_fileattr_set( - struct mnt_idmap *idmap, + const struct mnt_idmap *idmap, struct dentry *dentry, struct file_kattr *fa); diff --git a/fs/xfs/xfs_ioctl32.c b/fs/xfs/xfs_ioctl32.c index c66e192448a8..f6875a2cc706 100644 --- a/fs/xfs/xfs_ioctl32.c +++ b/fs/xfs/xfs_ioctl32.c @@ -44,26 +44,46 @@ xfs_compat_ioc_fsgeometry_v1( return 0; } -STATIC int -xfs_compat_growfs_data_copyin( - struct xfs_growfs_data *in, - compat_xfs_growfs_data_t __user *arg32) +static int +xfs_compat_ioc_growfs_data( + struct file *file, + struct xfs_mount *mp, + struct compat_xfs_growfs_data __user *arg32) { - if (get_user(in->newblocks, &arg32->newblocks) || - get_user(in->imaxpct, &arg32->imaxpct)) + struct xfs_growfs_data in = { }; + int error; + + if (get_user(in.newblocks, &arg32->newblocks) || + get_user(in.imaxpct, &arg32->imaxpct)) return -EFAULT; - return 0; + + error = mnt_want_write_file(file); + if (error) + return error; + error = xfs_growfs_data(mp, &in); + mnt_drop_write_file(file); + return error; } -STATIC int -xfs_compat_growfs_rt_copyin( - struct xfs_growfs_rt *in, - compat_xfs_growfs_rt_t __user *arg32) +static int +xfs_compat_ioc_growfs_rt( + struct file *file, + struct xfs_mount *mp, + struct compat_xfs_growfs_rt __user *arg32) { - if (get_user(in->newblocks, &arg32->newblocks) || - get_user(in->extsize, &arg32->extsize)) + struct xfs_growfs_rt in = {}; + int error; + + if (get_user(in.newblocks, &arg32->newblocks) || + get_user(in.extsize, &arg32->extsize)) return -EFAULT; - return 0; + + error = mnt_want_write_file(file); + if (error) + return error; + error = xfs_growfs_rt(mp, &in); + mnt_drop_write_file(file); + return error; } STATIC int @@ -138,6 +158,27 @@ xfs_ioctl32_bstat_copyin( return 0; } +static int +xfs_compat_ioc_swapext( + struct file *file, + struct compat_xfs_swapext __user *sxu) +{ + struct xfs_swapext sxp; + int error; + + /* Bulk copy in up to the sx_stat field, then copy bstat */ + if (copy_from_user(&sxp, sxu, offsetof(struct xfs_swapext, sx_stat)) || + xfs_ioctl32_bstat_copyin(&sxp.sx_stat, &sxu->sx_stat)) + return -EFAULT; + + error = mnt_want_write_file(file); + if (error) + return error; + error = xfs_swapext(&sxp); + mnt_drop_write_file(file); + return error; +} + /* XFS_IOC_FSBULKSTAT and friends */ STATIC int @@ -338,6 +379,43 @@ xfs_compat_handlereq_to_dentry( compat_ptr(hreq->ihandle), hreq->ihandlen); } +static int +xfs_compat_ioc_find_handle( + unsigned int cmd, + void __user *arg) +{ + struct xfs_fsop_handlereq hreq; + + if (xfs_compat_handlereq_copyin(&hreq, arg)) + return -EFAULT; + return xfs_find_handle(_NATIVE_IOC(cmd, struct xfs_fsop_handlereq), + &hreq); +} + +static int +xfs_compat_ioc_open_by_handle( + struct file *file, + void __user *arg) +{ + struct xfs_fsop_handlereq hreq; + + if (xfs_compat_handlereq_copyin(&hreq, arg)) + return -EFAULT; + return xfs_open_by_handle(file, &hreq); +} + +static int +xfs_compat_ioc_readlink_by_handle( + struct file *file, + void __user *arg) +{ + struct xfs_fsop_handlereq hreq; + + if (xfs_compat_handlereq_copyin(&hreq, arg)) + return -EFAULT; + return xfs_readlink_by_handle(file, &hreq); +} + STATIC int xfs_compat_attrlist_by_handle( struct file *parfilp, @@ -426,7 +504,6 @@ xfs_file_compat_ioctl( struct inode *inode = file_inode(filp); struct xfs_inode *ip = XFS_I(inode); void __user *arg = compat_ptr(p); - int error; trace_xfs_file_compat_ioctl(ip); @@ -434,85 +511,34 @@ xfs_file_compat_ioctl( #if defined(BROKEN_X86_ALIGNMENT) case XFS_IOC_FSGEOMETRY_V1_32: return xfs_compat_ioc_fsgeometry_v1(ip->i_mount, arg); - case XFS_IOC_FSGROWFSDATA_32: { - struct xfs_growfs_data in; - - if (xfs_compat_growfs_data_copyin(&in, arg)) - return -EFAULT; - error = mnt_want_write_file(filp); - if (error) - return error; - error = xfs_growfs_data(ip->i_mount, &in); - mnt_drop_write_file(filp); - return error; - } - case XFS_IOC_FSGROWFSRT_32: { - struct xfs_growfs_rt in; - - if (xfs_compat_growfs_rt_copyin(&in, arg)) - return -EFAULT; - error = mnt_want_write_file(filp); - if (error) - return error; - error = xfs_growfs_rt(ip->i_mount, &in); - mnt_drop_write_file(filp); - return error; - } + case XFS_IOC_FSGROWFSDATA_32: + return xfs_compat_ioc_growfs_data(filp, ip->i_mount, arg); + case XFS_IOC_FSGROWFSRT_32: + return xfs_compat_ioc_growfs_rt(filp, ip->i_mount, arg); #endif - /* long changes size, but xfs only copiese out 32 bits */ case XFS_IOC_GETVERSION_32: - cmd = _NATIVE_IOC(cmd, long); - return xfs_file_ioctl(filp, cmd, p); - case XFS_IOC_SWAPEXT_32: { - struct xfs_swapext sxp; - struct compat_xfs_swapext __user *sxu = arg; - - /* Bulk copy in up to the sx_stat field, then copy bstat */ - if (copy_from_user(&sxp, sxu, - offsetof(struct xfs_swapext, sx_stat)) || - xfs_ioctl32_bstat_copyin(&sxp.sx_stat, &sxu->sx_stat)) - return -EFAULT; - error = mnt_want_write_file(filp); - if (error) - return error; - error = xfs_ioc_swapext(&sxp); - mnt_drop_write_file(filp); - return error; - } + /* long changes size, but xfs only copies out 32 bits */ + return xfs_file_ioctl(filp, _NATIVE_IOC(cmd, long), p); + case XFS_IOC_SWAPEXT_32: + return xfs_compat_ioc_swapext(filp, arg); case XFS_IOC_FSBULKSTAT_32: case XFS_IOC_FSBULKSTAT_SINGLE_32: case XFS_IOC_FSINUMBERS_32: return xfs_compat_ioc_fsbulkstat(filp, cmd, arg); case XFS_IOC_FD_TO_HANDLE_32: case XFS_IOC_PATH_TO_HANDLE_32: - case XFS_IOC_PATH_TO_FSHANDLE_32: { - struct xfs_fsop_handlereq hreq; - - if (xfs_compat_handlereq_copyin(&hreq, arg)) - return -EFAULT; - cmd = _NATIVE_IOC(cmd, struct xfs_fsop_handlereq); - return xfs_find_handle(cmd, &hreq); - } - case XFS_IOC_OPEN_BY_HANDLE_32: { - struct xfs_fsop_handlereq hreq; - - if (xfs_compat_handlereq_copyin(&hreq, arg)) - return -EFAULT; - return xfs_open_by_handle(filp, &hreq); - } - case XFS_IOC_READLINK_BY_HANDLE_32: { - struct xfs_fsop_handlereq hreq; - - if (xfs_compat_handlereq_copyin(&hreq, arg)) - return -EFAULT; - return xfs_readlink_by_handle(filp, &hreq); - } + case XFS_IOC_PATH_TO_FSHANDLE_32: + return xfs_compat_ioc_find_handle(cmd, arg); + case XFS_IOC_OPEN_BY_HANDLE_32: + return xfs_compat_ioc_open_by_handle(filp, arg); + case XFS_IOC_READLINK_BY_HANDLE_32: + return xfs_compat_ioc_readlink_by_handle(filp, arg); case XFS_IOC_ATTRLIST_BY_HANDLE_32: return xfs_compat_attrlist_by_handle(filp, arg); case XFS_IOC_ATTRMULTI_BY_HANDLE_32: return xfs_compat_attrmulti_by_handle(filp, arg); default: /* try the native version */ - return xfs_file_ioctl(filp, cmd, (unsigned long)arg); + return xfs_file_ioctl(filp, cmd, p); } } diff --git a/fs/xfs/xfs_ioend.c b/fs/xfs/xfs_ioend.c index 40695d18dac0..e70be5b86f0b 100644 --- a/fs/xfs/xfs_ioend.c +++ b/fs/xfs/xfs_ioend.c @@ -1,6 +1,6 @@ // SPDX-License-Identifier: GPL-2.0 /* - * Copyright (c) 2016-2025 Christoph Hellwig. + * Copyright (c) 2016-2026 Christoph Hellwig. * All Rights Reserved. */ #include "xfs_platform.h" @@ -16,6 +16,135 @@ #include "xfs_reflink.h" #include "xfs_zone_alloc.h" #include "xfs_ioend.h" +#include "xfs_error.h" +#include "xfs_errortag.h" +#include <linux/bio-integrity.h> + +static void +xfs_dio_bounce_end_io( + struct bio *bio) +{ + struct iomap_ioend *ioend = iomap_ioend_from_bio(bio); + int error = blk_status_to_errno(bio->bi_status); + struct bio *orig_bio = bio->bi_private; + + if ((ioend->io_flags & IOMAP_IOEND_INTEGRITY) && !bio->bi_status) + error = iomap_ioend_integrity_verify(ioend); + iomap_bounce_read_end_io(ioend, orig_bio, error); +} + +static void +xfs_bounce_submit_ioend( + struct iomap_ioend *ioend) +{ + if (ioend->io_flags & IOMAP_IOEND_INTEGRITY) + fs_bio_integrity_alloc(&ioend->io_bio); + ioend->io_bio.bi_end_io = xfs_dio_bounce_end_io; + bio_set_flag(&ioend->io_bio, BIO_COMPLETE_IN_TASK); + submit_bio(&ioend->io_bio); +} + +static void +xfs_end_bio_bounced( + struct bio *bio) +{ + /* + * Just complete the original ioends as all verification is done by the + * end_io handlers for the clone bio(s). + */ + iomap_finish_ioends(iomap_ioend_from_bio(bio), + blk_status_to_errno(bio->bi_status)); +} + +static void +xfs_read_bounce_and_resubmit( + struct iomap_ioend *ioend) +{ + struct bio *bio = &ioend->io_bio; + struct xfs_inode *ip = XFS_I(ioend->io_inode); + unsigned int nofs_flag = memalloc_nofs_save(); + + trace_xfs_bounce_reread(ip, ioend->io_offset, ioend->io_size); + + /* + * Free the bio integrity data for the original bio, as we'll allocate + * a new one for each sub-I/O, which could deadlock if we keep the + * integrity data for the original bio around. + */ + if (bio_integrity(bio)) + fs_bio_integrity_free(bio); + + /* + * Resubmit the bio through the iomap bounce machinery. The original + * bio itself is not resubmitted to the block layer, but just used to + * track I/O completion of the cloned bios. + */ + bio_prepare_reissue(bio, xfs_inode_buftarg(ip)->bt_bdev); + bio->bi_iter = (struct bvec_iter) { + .bi_sector = ioend->io_sector, + .bi_size = ioend->io_size, + .bi_offset = ioend->io_bvec_offset, + }; + bio->bi_end_io = xfs_end_bio_bounced; + iomap_bounce_read(ioend, bdev_logical_block_size(bio->bi_bdev), + xfs_bounce_submit_ioend); + memalloc_nofs_restore(nofs_flag); +} + +static void +xfs_end_io_read( + struct bio *bio) +{ + struct iomap_ioend *ioend = iomap_ioend_from_bio(bio); + struct xfs_inode *ip = XFS_I(ioend->io_inode); + struct xfs_mount *mp = ip->i_mount; + int error = blk_status_to_errno(bio->bi_status); + + if (!error && (ioend->io_flags & IOMAP_IOEND_INTEGRITY)) { + error = iomap_ioend_integrity_verify(ioend); + if ((ioend->io_flags & IOMAP_IOEND_DIRECT) && + READ_ONCE(mp->m_read_bounce) == XFS_READ_BOUNCE_LAZY) { + /* + * We only really need to retry for guard tag errors, + * but right now we can't distinguish them from other + * (i.e, reftag) errors. + */ + if (error || + XFS_TEST_ERROR(mp, XFS_ERRTAG_BOUNCE_REREAD)) { + xfs_read_bounce_and_resubmit(ioend); + return; + } + } + } + + iomap_finish_ioends(ioend, error); +} + +void +xfs_ioend_submit_read( + struct inode *inode, + struct bio *bio, + loff_t file_offset, + u16 ioend_flags) +{ + struct xfs_inode *ip = XFS_I(inode); + struct xfs_mount *mp = ip->i_mount; + struct iomap_ioend *ioend; + + ioend = iomap_init_ioend(inode, bio, file_offset, ioend_flags); + if ((ioend_flags & IOMAP_IOEND_DIRECT) && + READ_ONCE(mp->m_read_bounce) == XFS_READ_BOUNCE_ALWAYS) { + iomap_bounce_read(ioend, bdev_logical_block_size(bio->bi_bdev), + xfs_bounce_submit_ioend); + return; + } + + if (ioend_flags & IOMAP_IOEND_INTEGRITY) + fs_bio_integrity_alloc(bio); + bio->bi_end_io = xfs_end_io_read; + bio_set_flag(bio, BIO_COMPLETE_IN_TASK); + submit_bio(bio); +} static void xfs_ioend_put_open_zones( @@ -148,11 +277,7 @@ xfs_end_io( io_list))) { list_del_init(&ioend->io_list); iomap_ioend_try_merge(ioend, &tmp); - if (bio_op(&ioend->io_bio) == REQ_OP_READ) - iomap_finish_ioends(ioend, - blk_status_to_errno(ioend->io_bio.bi_status)); - else - xfs_end_ioend_write(ioend); + xfs_end_ioend_write(ioend); cond_resched(); } } diff --git a/fs/xfs/xfs_ioend.h b/fs/xfs/xfs_ioend.h index 525865767fca..7c2a1ea3e6ed 100644 --- a/fs/xfs/xfs_ioend.h +++ b/fs/xfs/xfs_ioend.h @@ -12,5 +12,7 @@ static inline bool xfs_ioend_is_append(struct iomap_ioend *ioend) } void xfs_end_bio(struct bio *bio); +void xfs_ioend_submit_read(struct inode *inode, struct bio *bio, + loff_t file_offset, u16 ioend_flags); #endif /* __XFS_IOEND_H */ diff --git a/fs/xfs/xfs_iops.c b/fs/xfs/xfs_iops.c index d1306e723899..f67a541ef35a 100644 --- a/fs/xfs/xfs_iops.c +++ b/fs/xfs/xfs_iops.c @@ -169,7 +169,7 @@ xfs_create_need_xattr( STATIC int xfs_generic_create( - struct mnt_idmap *idmap, + const struct mnt_idmap *idmap, struct inode *dir, struct dentry *dentry, umode_t mode, @@ -279,7 +279,7 @@ xfs_generic_create( STATIC int xfs_vn_mknod( - struct mnt_idmap *idmap, + const struct mnt_idmap *idmap, struct inode *dir, struct dentry *dentry, umode_t mode, @@ -290,7 +290,7 @@ xfs_vn_mknod( STATIC int xfs_vn_create( - struct mnt_idmap *idmap, + const struct mnt_idmap *idmap, struct inode *dir, struct dentry *dentry, umode_t mode) @@ -300,7 +300,7 @@ xfs_vn_create( STATIC struct dentry * xfs_vn_mkdir( - struct mnt_idmap *idmap, + const struct mnt_idmap *idmap, struct inode *dir, struct dentry *dentry, umode_t mode) @@ -425,7 +425,7 @@ xfs_vn_unlink( STATIC int xfs_vn_symlink( - struct mnt_idmap *idmap, + const struct mnt_idmap *idmap, struct inode *dir, struct dentry *dentry, const char *symname) @@ -466,7 +466,7 @@ xfs_vn_symlink( STATIC int xfs_vn_rename( - struct mnt_idmap *idmap, + const struct mnt_idmap *idmap, struct inode *odir, struct dentry *odentry, struct inode *ndir, @@ -679,7 +679,7 @@ xfs_report_atomic_write( STATIC int xfs_vn_getattr( - struct mnt_idmap *idmap, + const struct mnt_idmap *idmap, const struct path *path, struct kstat *stat, u32 request_mask, @@ -754,7 +754,7 @@ xfs_vn_getattr( static int xfs_vn_change_ok( - struct mnt_idmap *idmap, + const struct mnt_idmap *idmap, struct dentry *dentry, struct iattr *iattr) { @@ -777,7 +777,7 @@ xfs_vn_change_ok( */ static int xfs_setattr_nonsize( - struct mnt_idmap *idmap, + const struct mnt_idmap *idmap, struct dentry *dentry, struct xfs_inode *ip, struct iattr *iattr) @@ -903,7 +903,7 @@ out_dqrele: */ int xfs_vn_setattr_size( - struct mnt_idmap *idmap, + const struct mnt_idmap *idmap, struct dentry *dentry, struct iattr *iattr) { @@ -1130,7 +1130,7 @@ out_trans_cancel: STATIC int xfs_vn_setattr( - struct mnt_idmap *idmap, + const struct mnt_idmap *idmap, struct dentry *dentry, struct iattr *iattr) { @@ -1250,7 +1250,7 @@ xfs_vn_fiemap( STATIC int xfs_vn_tmpfile( - struct mnt_idmap *idmap, + const struct mnt_idmap *idmap, struct inode *dir, struct file *file, umode_t mode) diff --git a/fs/xfs/xfs_iops.h b/fs/xfs/xfs_iops.h index 0896f6b8b3b8..328305bba19d 100644 --- a/fs/xfs/xfs_iops.h +++ b/fs/xfs/xfs_iops.h @@ -10,7 +10,7 @@ struct xfs_inode; extern ssize_t xfs_vn_listxattr(struct dentry *, char *data, size_t size); -int xfs_vn_setattr_size(struct mnt_idmap *idmap, +int xfs_vn_setattr_size(const struct mnt_idmap *idmap, struct dentry *dentry, struct iattr *vap); int xfs_inode_init_security(struct inode *inode, struct inode *dir, diff --git a/fs/xfs/xfs_itable.c b/fs/xfs/xfs_itable.c index 159295c63e8f..a4cf1effa5e6 100644 --- a/fs/xfs/xfs_itable.c +++ b/fs/xfs/xfs_itable.c @@ -63,7 +63,7 @@ want_metadir_file( STATIC int xfs_bulkstat_one_int( struct xfs_mount *mp, - struct mnt_idmap *idmap, + const struct mnt_idmap *idmap, struct xfs_trans *tp, xfs_ino_t ino, struct xfs_bstat_chunk *bc) diff --git a/fs/xfs/xfs_itable.h b/fs/xfs/xfs_itable.h index 2d0612f14d6e..c0567bfc30fb 100644 --- a/fs/xfs/xfs_itable.h +++ b/fs/xfs/xfs_itable.h @@ -8,7 +8,7 @@ /* In-memory representation of a userspace request for batch inode data. */ struct xfs_ibulk { struct xfs_mount *mp; - struct mnt_idmap *idmap; + const struct mnt_idmap *idmap; void __user *ubuffer; /* user output buffer */ xfs_ino_t startino; /* start with this inode */ unsigned int icount; /* number of elements in ubuffer */ diff --git a/fs/xfs/xfs_mount.h b/fs/xfs/xfs_mount.h index 216a38a354e7..894ff2f4ecbd 100644 --- a/fs/xfs/xfs_mount.h +++ b/fs/xfs/xfs_mount.h @@ -142,6 +142,12 @@ struct xfs_freecounter { uint64_t res_saved; }; +enum xfs_read_bounce { + XFS_READ_BOUNCE_NEVER, + XFS_READ_BOUNCE_ALWAYS, + XFS_READ_BOUNCE_LAZY, +}; + /* * The struct xfsmount layout is optimised to separate read-mostly variables * from variables that are frequently modified. We put the read-mostly variables @@ -177,6 +183,7 @@ typedef struct xfs_mount { struct workqueue_struct *m_sync_workqueue; struct workqueue_struct *m_blockgc_wq; struct workqueue_struct *m_inodegc_wq; + enum xfs_read_bounce m_read_bounce; int m_bsize; /* fs logical block size */ uint8_t m_blkbit_log; /* blocklog + NBBY */ @@ -291,6 +298,7 @@ typedef struct xfs_mount { struct xfs_zone_info *m_zone_info; /* zone allocator information */ struct dentry *m_debugfs; /* debugfs parent */ struct xfs_kobj m_kobj; + struct xfs_kobj m_csum_kobj; struct xfs_kobj m_error_kobj; struct xfs_kobj m_error_meta_kobj; struct xfs_error_cfg m_error_cfg[XFS_ERR_CLASS_MAX][XFS_ERR_ERRNO_MAX]; diff --git a/fs/xfs/xfs_qm.c b/fs/xfs/xfs_qm.c index 54d00d543b51..008fed8624be 100644 --- a/fs/xfs/xfs_qm.c +++ b/fs/xfs/xfs_qm.c @@ -1294,11 +1294,7 @@ xfs_qm_quotacheck_dqadjust( * * There are no timers for the default values set in the root dquot. */ - if (dqp->q_id) { - xfs_qm_adjust_dqlimits(dqp); - xfs_qm_adjust_dqtimers(dqp); - } - + xfs_qm_adjust_dqenforcement(dqp); dqp->q_flags |= XFS_DQFLAG_DIRTY; out_unlock: mutex_unlock(&dqp->q_qlock); diff --git a/fs/xfs/xfs_rmap_item.c b/fs/xfs/xfs_rmap_item.c index 000cff1ce324..be2d5d6fe863 100644 --- a/fs/xfs/xfs_rmap_item.c +++ b/fs/xfs/xfs_rmap_item.c @@ -494,7 +494,7 @@ xfs_rui_validate_map( !xfs_verify_ino(mp, map->me_owner)) return false; - if (!xfs_verify_fileext(mp, map->me_startoff, map->me_len)) + if (!xfs_verify_fileext(map->me_startoff, map->me_len)) return false; if (isrt) diff --git a/fs/xfs/xfs_rtalloc.h b/fs/xfs/xfs_rtalloc.h index 78a690b489ed..1d5108861acd 100644 --- a/fs/xfs/xfs_rtalloc.h +++ b/fs/xfs/xfs_rtalloc.h @@ -46,10 +46,34 @@ int xfs_rtalloc_reinit_frextents(struct xfs_mount *mp); int xfs_growfs_check_rtgeom(const struct xfs_mount *mp, xfs_rfsblock_t dblocks, xfs_rfsblock_t rblocks, xfs_agblock_t rextsize); #else -# define xfs_growfs_rt(mp,in) (-ENOSYS) -# define xfs_rtalloc_reinit_frextents(m) (0) -# define xfs_rtmount_readsb(mp) (0) -# define xfs_rtmount_freesb(mp) ((void)0) +static inline int +xfs_growfs_rt( + struct xfs_mount *mp, + struct xfs_growfs_rt *in) +{ + return -ENOSYS; +} + +static inline int +xfs_rtalloc_reinit_frextents( + struct xfs_mount *mp) +{ + return 0; +} + +static inline int +xfs_rtmount_readsb( + struct xfs_mount *mp) +{ + return 0; +} + +static inline void +xfs_rtmount_freesb( + struct xfs_mount *mp) +{ +} + static inline int /* error */ xfs_rtmount_init( xfs_mount_t *mp) /* file system mount structure */ @@ -60,8 +84,21 @@ xfs_rtmount_init( xfs_warn(mp, "Not built with CONFIG_XFS_RT"); return -ENOSYS; } -# define xfs_rtmount_inodes(m) (((mp)->m_sb.sb_rblocks == 0)? 0 : (-ENOSYS)) -# define xfs_rtunmount_inodes(m) + +static inline int +xfs_rtmount_inodes( + struct xfs_mount *mp) +{ + if (mp->m_sb.sb_rblocks) + return -ENOSYS; + return 0; +} + +static inline void +xfs_rtunmount_inodes( + struct xfs_mount *mp) +{ +} static inline int xfs_growfs_check_rtgeom(const struct xfs_mount *mp, diff --git a/fs/xfs/xfs_super.c b/fs/xfs/xfs_super.c index b24db75eaedc..5a06132aa384 100644 --- a/fs/xfs/xfs_super.c +++ b/fs/xfs/xfs_super.c @@ -1889,7 +1889,7 @@ xfs_fs_fill_super( * Avoid integer overflow by comparing the maximum bmbt offset to the * maximum pagecache offset in units of fs blocks. */ - if (!xfs_verify_fileoff(mp, XFS_B_TO_FSBT(mp, MAX_LFS_FILESIZE))) { + if (!xfs_verify_fileoff(XFS_B_TO_FSBT(mp, MAX_LFS_FILESIZE))) { xfs_warn(mp, "MAX_LFS_FILESIZE block offset (%llu) exceeds extent map maximum (%llu)!", XFS_B_TO_FSBT(mp, MAX_LFS_FILESIZE), @@ -2317,6 +2317,7 @@ xfs_init_fs_context( mp->m_logbufs = -1; mp->m_logbsize = -1; mp->m_allocsize_log = 16; /* 64k */ + mp->m_read_bounce = XFS_READ_BOUNCE_LAZY; xfs_hooks_init(&mp->m_dir_update_hooks); diff --git a/fs/xfs/xfs_symlink.c b/fs/xfs/xfs_symlink.c index 5585ac7f4d16..709cd22248f8 100644 --- a/fs/xfs/xfs_symlink.c +++ b/fs/xfs/xfs_symlink.c @@ -82,7 +82,7 @@ xfs_readlink( int xfs_symlink( - struct mnt_idmap *idmap, + const struct mnt_idmap *idmap, struct xfs_inode *dp, struct xfs_name *link_name, const char *target_path, @@ -219,7 +219,7 @@ xfs_symlink( *ipp = du.ip; xfs_iunlock(du.ip, XFS_ILOCK_EXCL); xfs_iunlock(dp, XFS_ILOCK_EXCL); - xfs_parent_finish(mp, du.ppargs); + xfs_parent_finish(du.ppargs); return 0; out_trans_cancel: @@ -236,7 +236,7 @@ out_release_inode: xfs_irele(du.ip); } out_parent: - xfs_parent_finish(mp, du.ppargs); + xfs_parent_finish(du.ppargs); out_release_dquots: xfs_qm_dqrele(udqp); xfs_qm_dqrele(gdqp); diff --git a/fs/xfs/xfs_symlink.h b/fs/xfs/xfs_symlink.h index 0d29a50e66fd..3c5a969f9fc5 100644 --- a/fs/xfs/xfs_symlink.h +++ b/fs/xfs/xfs_symlink.h @@ -7,7 +7,7 @@ /* Kernel only symlink definitions */ -int xfs_symlink(struct mnt_idmap *idmap, struct xfs_inode *dp, +int xfs_symlink(const struct mnt_idmap *idmap, struct xfs_inode *dp, struct xfs_name *link_name, const char *target_path, umode_t mode, struct xfs_inode **ipp); int xfs_readlink(struct xfs_inode *ip, char *link); diff --git a/fs/xfs/xfs_sysfs.c b/fs/xfs/xfs_sysfs.c index b62712187324..e77917ac179d 100644 --- a/fs/xfs/xfs_sysfs.c +++ b/fs/xfs/xfs_sysfs.c @@ -392,6 +392,71 @@ const struct kobj_type xfs_stats_ktype = { .default_groups = xfs_stats_groups, }; +static inline struct xfs_mount *csum_to_mp(struct kobject *kobj) +{ + return container_of(to_kobj(kobj), struct xfs_mount, m_csum_kobj); +} + +static bool +xfs_has_read_bounce( + struct xfs_mount *mp) +{ + if (bdev_has_integrity_csum(mp->m_ddev_targp->bt_bdev)) + return true; + if (mp->m_rtdev_targp && + bdev_has_integrity_csum(mp->m_rtdev_targp->bt_bdev)) + return true; + return false; +} + +static const char * const bounce_modes[] = { + [XFS_READ_BOUNCE_NEVER] = "never", + [XFS_READ_BOUNCE_ALWAYS] = "always", + [XFS_READ_BOUNCE_LAZY] = "lazy", +}; + +static ssize_t +read_bounce_show( + struct kobject *kobj, + char *buf) +{ + struct xfs_mount *mp = csum_to_mp(kobj); + + return sysfs_emit(buf, "%s\n", + bounce_modes[READ_ONCE(mp->m_read_bounce)]); +} + +static ssize_t +read_bounce_store( + struct kobject *kobj, + const char *buf, + size_t count) +{ + struct xfs_mount *mp = csum_to_mp(kobj); + int ret; + + if (!xfs_has_read_bounce(mp)) + return -EINVAL; + ret = sysfs_match_string(bounce_modes, buf); + if (ret < 0) + return ret; + WRITE_ONCE(mp->m_read_bounce, ret); + return count; +} +XFS_SYSFS_ATTR_RW(read_bounce); + +static struct attribute *xfs_csum_attrs[] = { + ATTR_LIST(read_bounce), + NULL, +}; +ATTRIBUTE_GROUPS(xfs_csum); + +static const struct kobj_type xfs_csum_ktype = { + .release = xfs_sysfs_release, + .sysfs_ops = &xfs_sysfs_ops, + .default_groups = xfs_csum_groups, +}; + /* xlog */ static inline struct xlog * @@ -817,11 +882,17 @@ xfs_mount_sysfs_init( if (error) goto out_remove_fsdir; + /* .../xfs/<dev>/csum/ */ + error = xfs_sysfs_init(&mp->m_csum_kobj, &xfs_csum_ktype, &mp->m_kobj, + "csum"); + if (error) + goto out_remove_stats_dir; + /* .../xfs/<dev>/error/ */ error = xfs_sysfs_init(&mp->m_error_kobj, &xfs_error_ktype, &mp->m_kobj, "error"); if (error) - goto out_remove_stats_dir; + goto out_remove_csum_dir; /* .../xfs/<dev>/error/fail_at_unmount */ error = sysfs_create_file(&mp->m_error_kobj.kobject, @@ -835,12 +906,14 @@ xfs_mount_sysfs_init( "metadata", &mp->m_error_meta_kobj, xfs_error_meta_init); if (error) - goto out_remove_error_dir; + goto out_remove_csum_dir; return 0; out_remove_error_dir: xfs_sysfs_del(&mp->m_error_kobj); +out_remove_csum_dir: + xfs_sysfs_del(&mp->m_csum_kobj); out_remove_stats_dir: xfs_sysfs_del(&mp->m_stats.xs_kobj); out_remove_fsdir: @@ -864,6 +937,7 @@ xfs_mount_sysfs_del( } xfs_sysfs_del(&mp->m_error_meta_kobj); xfs_sysfs_del(&mp->m_error_kobj); + xfs_sysfs_del(&mp->m_csum_kobj); xfs_sysfs_del(&mp->m_stats.xs_kobj); xfs_sysfs_del(&mp->m_kobj); } diff --git a/fs/xfs/xfs_trace.h b/fs/xfs/xfs_trace.h index 6aa379c2cf0c..28d49f158b22 100644 --- a/fs/xfs/xfs_trace.h +++ b/fs/xfs/xfs_trace.h @@ -1896,6 +1896,7 @@ DEFINE_SIMPLE_IO_EVENT(xfs_zero_eof); DEFINE_SIMPLE_IO_EVENT(xfs_end_io_direct_write); DEFINE_SIMPLE_IO_EVENT(xfs_file_splice_read); DEFINE_SIMPLE_IO_EVENT(xfs_zoned_map_blocks); +DEFINE_SIMPLE_IO_EVENT(xfs_bounce_reread); DECLARE_EVENT_CLASS(xfs_itrunc_class, TP_PROTO(struct xfs_inode *ip, xfs_fsize_t new_size), @@ -6030,32 +6031,54 @@ DEFINE_HEALTHMON_EVENT(xfs_healthmon_detach); DEFINE_HEALTHMON_EVENT(xfs_healthmon_report_unmount); #define XFS_HEALTHMON_TYPE_STRINGS \ + { XFS_HEALTHMON_RUNNING, "run" }, \ { XFS_HEALTHMON_LOST, "lost" }, \ { XFS_HEALTHMON_UNMOUNT, "unmount" }, \ + { XFS_HEALTHMON_SHUTDOWN, "shutdown" }, \ { XFS_HEALTHMON_SICK, "sick" }, \ { XFS_HEALTHMON_CORRUPT, "corrupt" }, \ { XFS_HEALTHMON_HEALTHY, "healthy" }, \ - { XFS_HEALTHMON_SHUTDOWN, "shutdown" } + { XFS_HEALTHMON_MEDIA_ERROR, "media" }, \ + { XFS_HEALTHMON_BUFREAD, "bufread" }, \ + { XFS_HEALTHMON_BUFWRITE, "bufwrite" }, \ + { XFS_HEALTHMON_DIOREAD, "dioread" }, \ + { XFS_HEALTHMON_DIOWRITE, "diowrite" }, \ + { XFS_HEALTHMON_DATALOST, "datalost" } #define XFS_HEALTHMON_DOMAIN_STRINGS \ { XFS_HEALTHMON_MOUNT, "mount" }, \ { XFS_HEALTHMON_FS, "fs" }, \ { XFS_HEALTHMON_AG, "ag" }, \ { XFS_HEALTHMON_INODE, "inode" }, \ - { XFS_HEALTHMON_RTGROUP, "rtgroup" } + { XFS_HEALTHMON_RTGROUP, "rtgroup" }, \ + { XFS_HEALTHMON_DATADEV, "datadev" }, \ + { XFS_HEALTHMON_RTDEV, "rtdev" }, \ + { XFS_HEALTHMON_LOGDEV, "logdev" }, \ + { XFS_HEALTHMON_FILERANGE, "filerange" } +TRACE_DEFINE_ENUM(XFS_HEALTHMON_RUNNING); TRACE_DEFINE_ENUM(XFS_HEALTHMON_LOST); -TRACE_DEFINE_ENUM(XFS_HEALTHMON_SHUTDOWN); TRACE_DEFINE_ENUM(XFS_HEALTHMON_UNMOUNT); +TRACE_DEFINE_ENUM(XFS_HEALTHMON_SHUTDOWN); TRACE_DEFINE_ENUM(XFS_HEALTHMON_SICK); TRACE_DEFINE_ENUM(XFS_HEALTHMON_CORRUPT); TRACE_DEFINE_ENUM(XFS_HEALTHMON_HEALTHY); +TRACE_DEFINE_ENUM(XFS_HEALTHMON_MEDIA_ERROR); +TRACE_DEFINE_ENUM(XFS_HEALTHMON_BUFREAD); +TRACE_DEFINE_ENUM(XFS_HEALTHMON_BUFWRITE); +TRACE_DEFINE_ENUM(XFS_HEALTHMON_DIOREAD); +TRACE_DEFINE_ENUM(XFS_HEALTHMON_DIOWRITE); +TRACE_DEFINE_ENUM(XFS_HEALTHMON_DATALOST); TRACE_DEFINE_ENUM(XFS_HEALTHMON_MOUNT); TRACE_DEFINE_ENUM(XFS_HEALTHMON_FS); TRACE_DEFINE_ENUM(XFS_HEALTHMON_AG); TRACE_DEFINE_ENUM(XFS_HEALTHMON_INODE); TRACE_DEFINE_ENUM(XFS_HEALTHMON_RTGROUP); +TRACE_DEFINE_ENUM(XFS_HEALTHMON_DATADEV); +TRACE_DEFINE_ENUM(XFS_HEALTHMON_RTDEV); +TRACE_DEFINE_ENUM(XFS_HEALTHMON_LOGDEV); +TRACE_DEFINE_ENUM(XFS_HEALTHMON_FILERANGE); DECLARE_EVENT_CLASS(xfs_healthmon_event_class, TP_PROTO(const struct xfs_healthmon *hm, diff --git a/fs/xfs/xfs_trans_dquot.c b/fs/xfs/xfs_trans_dquot.c index 1606c614f205..93ec876792cc 100644 --- a/fs/xfs/xfs_trans_dquot.c +++ b/fs/xfs/xfs_trans_dquot.c @@ -566,11 +566,7 @@ xfs_trans_apply_dquot_deltas( * Get any default limits in use. * Start/reset the timer(s) if needed. */ - if (dqp->q_id) { - xfs_qm_adjust_dqlimits(dqp); - xfs_qm_adjust_dqtimers(dqp); - } - + xfs_qm_adjust_dqenforcement(dqp); dqp->q_flags |= XFS_DQFLAG_DIRTY; /* * add this to the list of items to get logged diff --git a/fs/xfs/xfs_xattr.c b/fs/xfs/xfs_xattr.c index 1efe6c8139b2..b059a9714d11 100644 --- a/fs/xfs/xfs_xattr.c +++ b/fs/xfs/xfs_xattr.c @@ -169,7 +169,7 @@ xfs_xattr_flags_to_op( static int xfs_xattr_set(const struct xattr_handler *handler, - struct mnt_idmap *idmap, struct dentry *unused, + const struct mnt_idmap *idmap, struct dentry *unused, struct inode *inode, const char *name, const void *value, size_t size, int flags) { diff --git a/fs/xfs/xfs_zone_alloc.c b/fs/xfs/xfs_zone_alloc.c index b75cf3bfe33c..5f0af0c2c5e5 100644 --- a/fs/xfs/xfs_zone_alloc.c +++ b/fs/xfs/xfs_zone_alloc.c @@ -26,6 +26,7 @@ #include "xfs_zones.h" #include "xfs_trace.h" #include "xfs_mru_cache.h" +#include <linux/bio-integrity.h> static void xfs_open_zone_free_rcu( @@ -911,6 +912,9 @@ xfs_zone_alloc_and_submit( if (xfs_is_shutdown(mp)) goto out_error; + if (ioend->io_flags & IOMAP_IOEND_INTEGRITY) + fs_bio_integrity_generate(&ioend->io_bio); + /* * If we don't have a locally cached zone in this write context, see if * the inode is still associated with a zone and use that if so. diff --git a/fs/zonefs/super.c b/fs/zonefs/super.c index ff43d6d1ea30..b97f1f2b8dda 100644 --- a/fs/zonefs/super.c +++ b/fs/zonefs/super.c @@ -533,7 +533,7 @@ static int zonefs_show_options(struct seq_file *seq, struct dentry *root) return 0; } -static int zonefs_inode_setattr(struct mnt_idmap *idmap, +static int zonefs_inode_setattr(const struct mnt_idmap *idmap, struct dentry *dentry, struct iattr *iattr) { struct inode *inode = d_inode(dentry); diff --git a/include/linux/binfmts.h b/include/linux/binfmts.h index f686a37f7a0a..2e87faf9a8c2 100644 --- a/include/linux/binfmts.h +++ b/include/linux/binfmts.h @@ -128,7 +128,8 @@ struct linux_binfmt { struct module *module; int (*load_binary)(struct linux_binprm *); #ifdef CONFIG_COREDUMP - int (*core_dump)(struct coredump_params *cprm); + /* Returns true if the whole coredump was written. */ + bool (*core_dump)(struct coredump_params *cprm); unsigned long min_coredump; /* minimal dump size */ #endif } __randomize_layout; diff --git a/include/linux/bio-integrity.h b/include/linux/bio-integrity.h index 0ea2a8bf7efb..a954c97be0b3 100644 --- a/include/linux/bio-integrity.h +++ b/include/linux/bio-integrity.h @@ -151,7 +151,6 @@ void bio_integrity_setup_default(struct bio *bio); unsigned int fs_bio_integrity_alloc(struct bio *bio); void fs_bio_integrity_free(struct bio *bio); void fs_bio_integrity_generate(struct bio *bio); -int fs_bio_integrity_verify(struct bio *bio, sector_t sector, - unsigned int size); +int fs_bio_integrity_verify(struct bio *bio, struct bvec_iter *data_iter); #endif /* _LINUX_BIO_INTEGRITY_H */ diff --git a/include/linux/bio.h b/include/linux/bio.h index bb3235497e67..17944e44b584 100644 --- a/include/linux/bio.h +++ b/include/linux/bio.h @@ -479,6 +479,7 @@ static inline void bio_init_inline(struct bio *bio, struct block_device *bdev, extern void bio_uninit(struct bio *); void bio_reset(struct bio *bio, struct block_device *bdev, blk_opf_t opf); void bio_reuse(struct bio *bio, blk_opf_t opf); +void bio_prepare_reissue(struct bio *bio, struct block_device *bdev); void bio_chain(struct bio *, struct bio *); void bio_await(struct bio *bio, void *priv, void (*submit)(struct bio *bio, void *priv)); @@ -516,16 +517,18 @@ int bdev_rw_virt(struct block_device *bdev, sector_t sector, void *data, size_t len, enum req_op op); int bio_iov_iter_get_pages(struct bio *bio, struct iov_iter *iter, - unsigned mem_align_mask, unsigned len_align_mask); + unsigned maxlen, unsigned mem_align_mask, + unsigned len_align_mask); bool bio_iov_iter_set(struct bio *bio, const struct iov_iter *iter); void __bio_release_pages(struct bio *bio, bool mark_dirty); extern void bio_set_pages_dirty(struct bio *bio); extern void bio_check_pages_dirty(struct bio *bio); -int bio_iov_iter_bounce(struct bio *bio, struct iov_iter *iter, size_t maxlen, - size_t minsize); -void bio_iov_iter_unbounce(struct bio *bio, bool is_error, bool mark_dirty); +int bio_alloc_bounce_folios(struct bio *bio, size_t total_len, size_t minsize); +void bio_free_folios(struct bio *bio); +int bio_iov_iter_bounce_write(struct bio *bio, struct iov_iter *iter, + size_t maxlen, size_t minsize); extern void bio_copy_data(struct bio *dst, struct bio *src); extern void bio_free_pages(struct bio *bio); diff --git a/include/linux/blkdev.h b/include/linux/blkdev.h index 4f7905c3412b..098a65f3e48b 100644 --- a/include/linux/blkdev.h +++ b/include/linux/blkdev.h @@ -1816,9 +1816,11 @@ static inline int bio_split_rw_at(struct bio *bio, */ static inline unsigned int max_integrity_io_size(struct queue_limits *lim) { - return min_t(unsigned int, lim->max_segment_size, - (BLK_INTEGRITY_MAX_SIZE / lim->integrity.metadata_size) << - lim->integrity.interval_exp); + u64 max_intervals; + + max_intervals = BLK_INTEGRITY_MAX_SIZE / lim->integrity.metadata_size; + return min_t(u64, lim->max_segment_size, + max_intervals << lim->integrity.interval_exp); } #define DEFINE_IO_COMP_BATCH(name) struct io_comp_batch name = { } diff --git a/include/linux/buffer_head.h b/include/linux/buffer_head.h index 4b0b7188472b..f8782b719026 100644 --- a/include/linux/buffer_head.h +++ b/include/linux/buffer_head.h @@ -59,10 +59,7 @@ struct address_space; struct buffer_head { unsigned long b_state; /* buffer state bitmap (see above) */ struct buffer_head *b_this_page;/* circular list of page's buffers */ - union { - struct page *b_page; /* the page this bh is mapped to */ - struct folio *b_folio; /* the folio this bh is mapped to */ - }; + struct folio *b_folio; /* the folio this bh is mapped to */ sector_t b_blocknr; /* start block number */ size_t b_size; /* size of mapping */ @@ -172,7 +169,36 @@ static __always_inline int buffer_uptodate(const struct buffer_head *bh) static inline unsigned long bh_offset(const struct buffer_head *bh) { - return (unsigned long)(bh)->b_data & (page_size(bh->b_page) - 1); + return (unsigned long)(bh)->b_data & (folio_size(bh->b_folio) - 1); +} + +/** + * kmap_local_bh - Map the data of a buffer. + * @bh: The buffer. + * + * Buffers usually live in the page cache, but a few are built over memory + * which is not. Those carry no folio and b_data is already a kernel address + * which is always mapped, so there is nothing to do for them. Pair with + * kunmap_local_bh(). + * + * Return: A pointer to the buffer's data. + */ +static inline void *kmap_local_bh(const struct buffer_head *bh) +{ + if (!bh->b_folio) + return bh->b_data; + return kmap_local_folio(bh->b_folio, bh_offset(bh)); +} + +/** + * kunmap_local_bh - Unmap the data of a buffer. + * @bh: The buffer. + * @addr: The address returned by kmap_local_bh(). + */ +static inline void kunmap_local_bh(const struct buffer_head *bh, void *addr) +{ + if (bh->b_folio) + kunmap_local(addr); } /* If we *know* folio->private refers to buffer_heads */ @@ -332,20 +358,58 @@ static inline void bforget(struct buffer_head *bh) __bforget(bh); } -static inline struct buffer_head * -sb_bread(struct super_block *sb, sector_t block) +/** + * sb_bread - Read a block. + * @sb: The superblock to read from. + * @block: Block number in units of block size. + * + * Read a specified block, and return the buffer head that refers + * to it. The memory is allocated from the movable area so that it can + * be migrated. The returned buffer head has its refcount increased. + * The caller should call brelse() when it has finished with the buffer. + * + * Context: May sleep waiting for I/O. + * Return: NULL if the block was unreadable. + */ +static inline +struct buffer_head *sb_bread(struct super_block *sb, sector_t block) { return __bread_gfp(sb->s_bdev, block, sb->s_blocksize, __GFP_MOVABLE); } -static inline struct buffer_head * -sb_bread_unmovable(struct super_block *sb, sector_t block) +/** + * sb_bread_unmovable - Read a block. + * @sb: The superblock to read from. + * @block: Block number in units of block size. + * + * Read a specified block, and return the buffer head that refers to it. + * The memory is allocated from the unmovable area so that pointers into + * it remain valid after compaction runs. The returned buffer head has + * its refcount increased. The caller should call brelse() when it has + * finished with the buffer. + * + * Context: May sleep waiting for I/O. + * Return: NULL if the block was unreadable. + */ +static inline +struct buffer_head *sb_bread_unmovable(struct super_block *sb, sector_t block) { return __bread_gfp(sb->s_bdev, block, sb->s_blocksize, 0); } -static inline void -sb_breadahead(struct super_block *sb, sector_t block) +/** + * sb_breadahead - Start readahead. + * @sb: Superblock identifying the block device. + * @block: The block to read. + * + * Read this block. The I/O will be flagged as being readahead rather + * than immediate read, but (unlike the page cache), surrounding blocks + * will not be read. + * + * Context: May sleep in order to allocate memory. + */ +static inline +void sb_breadahead(struct super_block *sb, sector_t block) { __breadahead(sb->s_bdev, block, sb->s_blocksize); } diff --git a/include/linux/capability.h b/include/linux/capability.h index f8532d92fcad..622137f66f09 100644 --- a/include/linux/capability.h +++ b/include/linux/capability.h @@ -186,9 +186,9 @@ static inline bool ns_capable_setid(struct user_namespace *ns, int cap) } #endif /* CONFIG_MULTIUSER */ bool privileged_wrt_inode_uidgid(struct user_namespace *ns, - struct mnt_idmap *idmap, + const struct mnt_idmap *idmap, const struct inode *inode); -bool capable_wrt_inode_uidgid(struct mnt_idmap *idmap, +bool capable_wrt_inode_uidgid(const struct mnt_idmap *idmap, const struct inode *inode, int cap); extern bool file_ns_capable(const struct file *file, struct user_namespace *ns, int cap); extern bool ptracer_capable(struct task_struct *tsk, struct user_namespace *ns); @@ -215,11 +215,11 @@ static inline bool checkpoint_restore_ns_capable_noaudit(struct user_namespace * } /* audit system wants to get cap info from files as well */ -int get_vfs_caps_from_disk(struct mnt_idmap *idmap, +int get_vfs_caps_from_disk(const struct mnt_idmap *idmap, const struct dentry *dentry, struct cpu_vfs_cap_data *cpu_caps); -int cap_convert_nscap(struct mnt_idmap *idmap, struct dentry *dentry, +int cap_convert_nscap(const struct mnt_idmap *idmap, struct dentry *dentry, const void **ivalue, size_t size); #endif /* !_LINUX_CAPABILITY_H */ diff --git a/include/linux/cleanup.h b/include/linux/cleanup.h index b1b5698cbf1b..1fb8058b897d 100644 --- a/include/linux/cleanup.h +++ b/include/linux/cleanup.h @@ -261,10 +261,6 @@ const volatile void * __must_check_fn(const volatile void *val) * CLASS(name, var)(args...): * declare the variable @var as an instance of the named class * - * CLASS_INIT(name, var, init_expr): - * declare the variable @var as an instance of the named class with - * custom initialization expression. - * * Ex. * * DEFINE_CLASS(fdget, struct fd, fdput(_T), fdget(fd), int fd) @@ -302,9 +298,6 @@ static __always_inline class_##_name##_t class_##_name##ext##_constructor(_init_ class_##_name##_t var __cleanup(class_##_name##_destructor) = \ class_##_name##_constructor -#define CLASS_INIT(_name, _var, _init_expr) \ - class_##_name##_t _var __cleanup(class_##_name##_destructor) = (_init_expr) - #define __scoped_class(_name, var, _label, args...) \ for (CLASS(_name, var)(args); ; ({ goto _label; })) \ if (0) { \ diff --git a/include/linux/configfs.h b/include/linux/configfs.h index ef65c75beeaa..5bead9173ec1 100644 --- a/include/linux/configfs.h +++ b/include/linux/configfs.h @@ -66,8 +66,11 @@ struct config_item_type { struct module *ct_owner; const struct configfs_item_operations *ct_item_ops; const struct configfs_group_operations *ct_group_ops; - struct configfs_attribute **ct_attrs; - struct configfs_bin_attribute **ct_bin_attrs; + union { + struct configfs_attribute **ct_attrs; + const struct configfs_attribute *const *ct_attrs_const; + }; + const struct configfs_bin_attribute *const *ct_bin_attrs; }; /** @@ -160,41 +163,41 @@ struct configfs_bin_attribute { ssize_t (*write)(struct config_item *, const void *, size_t); }; -#define CONFIGFS_BIN_ATTR(_pfx, _name, _priv, _maxsz) \ -static struct configfs_bin_attribute _pfx##attr_##_name = { \ - .cb_attr = { \ - .ca_name = __stringify(_name), \ - .ca_mode = S_IRUGO | S_IWUSR, \ - .ca_owner = THIS_MODULE, \ - }, \ - .cb_private = _priv, \ - .cb_max_size = _maxsz, \ - .read = _pfx##_name##_read, \ - .write = _pfx##_name##_write, \ +#define CONFIGFS_BIN_ATTR(_pfx, _name, _priv, _maxsz) \ +static const struct configfs_bin_attribute _pfx##attr_##_name = { \ + .cb_attr = { \ + .ca_name = __stringify(_name), \ + .ca_mode = S_IRUGO | S_IWUSR, \ + .ca_owner = THIS_MODULE, \ + }, \ + .cb_private = _priv, \ + .cb_max_size = _maxsz, \ + .read = _pfx##_name##_read, \ + .write = _pfx##_name##_write, \ } -#define CONFIGFS_BIN_ATTR_RO(_pfx, _name, _priv, _maxsz) \ -static struct configfs_bin_attribute _pfx##attr_##_name = { \ - .cb_attr = { \ - .ca_name = __stringify(_name), \ - .ca_mode = S_IRUGO, \ - .ca_owner = THIS_MODULE, \ - }, \ - .cb_private = _priv, \ - .cb_max_size = _maxsz, \ - .read = _pfx##_name##_read, \ +#define CONFIGFS_BIN_ATTR_RO(_pfx, _name, _priv, _maxsz) \ +static const struct configfs_bin_attribute _pfx##attr_##_name = { \ + .cb_attr = { \ + .ca_name = __stringify(_name), \ + .ca_mode = S_IRUGO, \ + .ca_owner = THIS_MODULE, \ + }, \ + .cb_private = _priv, \ + .cb_max_size = _maxsz, \ + .read = _pfx##_name##_read, \ } -#define CONFIGFS_BIN_ATTR_WO(_pfx, _name, _priv, _maxsz) \ -static struct configfs_bin_attribute _pfx##attr_##_name = { \ - .cb_attr = { \ - .ca_name = __stringify(_name), \ - .ca_mode = S_IWUSR, \ - .ca_owner = THIS_MODULE, \ - }, \ - .cb_private = _priv, \ - .cb_max_size = _maxsz, \ - .write = _pfx##_name##_write, \ +#define CONFIGFS_BIN_ATTR_WO(_pfx, _name, _priv, _maxsz) \ +static const struct configfs_bin_attribute _pfx##attr_##_name = { \ + .cb_attr = { \ + .ca_name = __stringify(_name), \ + .ca_mode = S_IWUSR, \ + .ca_owner = THIS_MODULE, \ + }, \ + .cb_private = _priv, \ + .cb_max_size = _maxsz, \ + .write = _pfx##_name##_write, \ } /* @@ -220,8 +223,8 @@ struct configfs_group_operations { struct config_group *(*make_group)(struct config_group *group, const char *name); void (*disconnect_notify)(struct config_group *group, struct config_item *item); void (*drop_item)(struct config_group *group, struct config_item *item); - bool (*is_visible)(struct config_item *item, struct configfs_attribute *attr, int n); - bool (*is_bin_visible)(struct config_item *item, struct configfs_bin_attribute *attr, + bool (*is_visible)(struct config_item *item, const struct configfs_attribute *attr, int n); + bool (*is_bin_visible)(struct config_item *item, const struct configfs_bin_attribute *attr, int n); }; diff --git a/include/linux/coredump.h b/include/linux/coredump.h index 7b38ee2e7913..74af57b9406b 100644 --- a/include/linux/coredump.h +++ b/include/linux/coredump.h @@ -6,9 +6,20 @@ #include <linux/mm.h> #include <linux/fs.h> #include <linux/sched/coredump.h> +#include <uapi/linux/coredump.h> #include <asm/siginfo.h> #ifdef CONFIG_COREDUMP +/** + * enum coredump_state - what happened while the coredump was written + * @COREDUMP_STATE_STARTED: the dumper committed to writing a coredump + * @COREDUMP_STATE_TRUNCATED: the dumper stopped before it had written all of it + */ +enum coredump_state { + COREDUMP_STATE_STARTED = (1U << 0), + COREDUMP_STATE_TRUNCATED = (1U << 1), +}; + struct core_vma_metadata { unsigned long start, end; vm_flags_t flags; @@ -21,12 +32,20 @@ struct coredump_params { const kernel_siginfo_t *siginfo; struct file *file; unsigned long limit; - /* MMF_DUMP_FILTER_* bits, snapshot of mm->flags at dump start. */ - unsigned long mm_flags; + /* COREDUMP_MEMORY_* types to dump, the task's or the server's. */ + u64 memory_types; /* Snapshot of dumpable at dump start. */ enum task_dumpable dumpable; int cpu; + /* COREDUMP_* options negotiated with the coredump server. */ + u64 mask; + /* COREDUMP_STATE_* raised while the coredump is written. */ + enum coredump_state state; + /* Record header scratch, NULL unless the coredump is a record stream. */ + struct coredump_record_header *record_hdr; + /* Bytes handed to the file, record headers included. */ loff_t written; + /* Offset in the coredump, record headers excluded. */ loff_t pos; loff_t to_skip; int vma_count; @@ -41,13 +60,13 @@ extern unsigned int core_file_note_size_limit; * These are the only things you should do on a core-file: use only these * functions to write out all the necessary info. */ -extern void dump_skip_to(struct coredump_params *cprm, unsigned long to); -extern void dump_skip(struct coredump_params *cprm, size_t nr); -extern int dump_emit(struct coredump_params *cprm, const void *addr, int nr); -extern int dump_align(struct coredump_params *cprm, int align); -int dump_user_range(struct coredump_params *cprm, unsigned long start, - unsigned long len); -extern void vfs_coredump(const kernel_siginfo_t *siginfo); +void dump_skip_to(struct coredump_params *cprm, unsigned long to); +void dump_skip(struct coredump_params *cprm, size_t nr); +bool dump_emit(struct coredump_params *cprm, const void *addr, int nr); +bool dump_align(struct coredump_params *cprm, int align); +bool dump_user_range(struct coredump_params *cprm, unsigned long start, + unsigned long len); +void vfs_coredump(const kernel_siginfo_t *siginfo); /* * Logging for the coredump code, ratelimited. diff --git a/include/linux/dax.h b/include/linux/dax.h index fe6c3ded1b50..f2d47975d905 100644 --- a/include/linux/dax.h +++ b/include/linux/dax.h @@ -155,8 +155,6 @@ int dax_writeback_mapping_range(struct address_space *mapping, struct dax_device *dax_dev, struct writeback_control *wbc); int dax_folio_reset_order(struct folio *folio); -struct page *dax_layout_busy_page(struct address_space *mapping); -struct page *dax_layout_busy_page_range(struct address_space *mapping, loff_t start, loff_t end); dax_entry_t dax_lock_folio(struct folio *folio); void dax_unlock_folio(struct folio *folio, dax_entry_t cookie); dax_entry_t dax_lock_mapping_entry(struct address_space *mapping, @@ -173,16 +171,6 @@ static inline int fs_dax_get(struct dax_device *dax_dev, void *holder, { return -EOPNOTSUPP; } -static inline struct page *dax_layout_busy_page(struct address_space *mapping) -{ - return NULL; -} - -static inline struct page *dax_layout_busy_page_range(struct address_space *mapping, pgoff_t start, pgoff_t nr_pages) -{ - return NULL; -} - static inline int dax_writeback_mapping_range(struct address_space *mapping, struct dax_device *dax_dev, struct writeback_control *wbc) { diff --git a/include/linux/dcache.h b/include/linux/dcache.h index 4b1ff99608e0..adf239f8205f 100644 --- a/include/linux/dcache.h +++ b/include/linux/dcache.h @@ -116,6 +116,8 @@ struct dentry { * possible! */ + /* lockdep tracking of DCACHE_PAR_LOOKUP locks */ + struct lockdep_map lookup_map; struct list_head d_lru; /* LRU list */ struct hlist_node d_sib; /* child of parent list */ struct hlist_head d_children; /* our children */ @@ -236,7 +238,9 @@ enum dentry_flags { DCACHE_PAR_LOOKUP = BIT(24), /* being looked up (with parent locked shared) */ DCACHE_DENTRY_CURSOR = BIT(25), DCACHE_NORCU = BIT(26), /* No RCU delay for freeing */ - DCACHE_PERSISTENT = BIT(27) + DCACHE_PERSISTENT = BIT(27), +/* 28, 29, 30 free */ + DCACHE_PRIVATE = BIT(31) /* fs-specific flag */ }; #define DCACHE_MANAGED_DENTRY \ @@ -257,7 +261,9 @@ extern void d_delete(struct dentry *); extern struct dentry * d_alloc(struct dentry *, const struct qstr *); extern struct dentry * d_alloc_anon(struct super_block *); extern struct dentry * d_alloc_parallel(struct dentry *, const struct qstr *); +extern struct dentry * d_alloc_trylock(struct dentry *, struct qstr *); extern struct dentry * d_splice_alias(struct inode *, struct dentry *); +struct dentry *d_duplicate(struct dentry *dentry); /* weird procfs mess; *NOT* exported */ extern struct dentry * d_splice_alias_ops(struct inode *, struct dentry *, const struct dentry_operations *); @@ -553,6 +559,36 @@ static inline int simple_positive(const struct dentry *dentry) unsigned long vfs_pressure_ratio(unsigned long val); /** + * d_lookup_release - release ownership of DCACHE_PAR_LOOKUP lock + * @dentry: dentry that is locked + * + * If an in-lookup dentry is to be passed to another thread which + * will drop the in-lookup lock, then d_lookup_release() must be called + * to tell lockdep that this thread no lock holds the lock. The + * thread that receives the lock must call d_lookup_acquire() to + * acquire the lock. + */ +static inline void d_lookup_release(struct dentry *dentry) +{ + if (d_in_lookup(dentry)) + lock_map_release(&dentry->lookup_map); +} + +/** + * d_lookup_acquire - acquire ownership of DCACHE_PAR_LOOKUP lock + * @dentry: dentry that is locked + * + * If an in-lookup dentry was passed to this thread, the + * d_lookup_acquire() must be called to tell lockdep that this + * thread now owns the DCACHE_PAR_LOOKUP lock. + */ +static inline void d_lookup_acquire(struct dentry *dentry) +{ + if (d_in_lookup(dentry)) + lock_map_acquire_try(&dentry->lookup_map); +} + +/** * d_inode - Get the actual inode of this dentry * @dentry: The dentry to query * diff --git a/include/linux/f2fs_fs.h b/include/linux/f2fs_fs.h index bb2b6cd5d507..3081702b1ddb 100644 --- a/include/linux/f2fs_fs.h +++ b/include/linux/f2fs_fs.h @@ -14,9 +14,9 @@ #define F2FS_SUPER_OFFSET 1024 /* byte-size offset */ #define F2FS_MIN_LOG_SECTOR_SIZE 9 /* 9 bits for 512 bytes */ #define F2FS_MAX_LOG_SECTOR_SIZE PAGE_SHIFT /* Max is Block Size */ -#define F2FS_LOG_SECTORS_PER_BLOCK (PAGE_SHIFT - 9) /* log number for sector/blk */ -#define F2FS_BLKSIZE PAGE_SIZE /* support only block == page */ -#define F2FS_BLKSIZE_BITS PAGE_SHIFT /* bits for F2FS_BLKSIZE */ +#define F2FS_MIN_LOG_BLOCKSIZE 12 +#define F2FS_MIN_BLKSIZE 4096UL +#define F2FS_MAX_BLKSIZE PAGE_SIZE #define F2FS_MAX_EXTENSION 64 /* # of extension entries */ #define F2FS_EXTENSION_LEN 8 /* max size of extension */ @@ -24,19 +24,25 @@ #define NEW_ADDR ((block_t)-1) /* used as block_t addresses */ #define COMPRESS_ADDR ((block_t)-2) /* used as compressed data flag */ -#define F2FS_BLKSIZE_MASK (F2FS_BLKSIZE - 1) -#define F2FS_BYTES_TO_BLK(bytes) ((unsigned long long)(bytes) >> F2FS_BLKSIZE_BITS) -#define F2FS_BLK_TO_BYTES(blk) ((unsigned long long)(blk) << F2FS_BLKSIZE_BITS) -#define F2FS_BLK_END_BYTES(blk) (F2FS_BLK_TO_BYTES(blk + 1) - 1) -#define F2FS_BLK_ALIGN(x) (F2FS_BYTES_TO_BLK((x) + F2FS_BLKSIZE - 1)) +#define F2FS_BLKSIZE(sbi) ((sbi)->blocksize) +#define F2FS_BLKSIZE_BITS(sbi) ((sbi)->log_blocksize) +#define F2FS_BLKSIZE_MASK(sbi) (F2FS_BLKSIZE(sbi) - 1) +#define F2FS_LOG_SECTORS_PER_BLOCK(sbi) (F2FS_BLKSIZE_BITS(sbi) - 9) +#define F2FS_BLKS_PER_PAGE(sbi) (PAGE_SIZE / F2FS_BLKSIZE(sbi)) +#define F2FS_BYTES_TO_BLK(sbi, bytes) \ + ((unsigned long long)(bytes) >> F2FS_BLKSIZE_BITS(sbi)) +#define F2FS_BLK_TO_BYTES(sbi, blk) \ + ((unsigned long long)(blk) << F2FS_BLKSIZE_BITS(sbi)) +#define F2FS_BLK_END_BYTES(sbi, blk) \ + (F2FS_BLK_TO_BYTES(sbi, (blk) + 1) - 1) +#define F2FS_BLK_ALIGN(sbi, bytes) \ + F2FS_BYTES_TO_BLK(sbi, (unsigned long long)(bytes) + \ + F2FS_BLKSIZE(sbi) - 1) /* 0, 1(node nid), 2(meta nid) are reserved node id */ #define F2FS_RESERVED_NODE_NUM 3 #define F2FS_ROOT_INO(sbi) ((sbi)->root_ino_num) -#define F2FS_NODE_INO(sbi) ((sbi)->node_ino_num) -#define F2FS_META_INO(sbi) ((sbi)->meta_ino_num) -#define F2FS_COMPRESS_INO(sbi) (NM_I(sbi)->max_nid) #define F2FS_MAX_QUOTAS 3 @@ -214,20 +220,27 @@ struct f2fs_checkpoint { unsigned char sit_nat_version_bitmap[]; } __packed; -#define CP_CHKSUM_OFFSET (F2FS_BLKSIZE - sizeof(__le32)) /* default chksum offset in checkpoint */ #define CP_MIN_CHKSUM_OFFSET \ (offsetof(struct f2fs_checkpoint, sit_nat_version_bitmap)) /* * For orphan inode management + * + * The number of inode entries in an orphan block depends on the filesystem + * block size. Its exact on-disk layout is: + * + * 0 blocksize - 16 blocksize + * +--------------------------+--------------------------+ + * | ino[0] ... ino[n - 1] | struct f2fs_orphan_footer | + * +--------------------------+--------------------------+ + * + * n = (blocksize - sizeof(struct f2fs_orphan_footer)) / sizeof(__le32) */ -#define F2FS_ORPHANS_PER_BLOCK ((F2FS_BLKSIZE - 4 * sizeof(__le32)) / sizeof(__le32)) - -#define GET_ORPHAN_BLOCKS(n) (((n) + F2FS_ORPHANS_PER_BLOCK - 1) / \ - F2FS_ORPHANS_PER_BLOCK) - struct f2fs_orphan_block { - __le32 ino[F2FS_ORPHANS_PER_BLOCK]; /* inode numbers */ + DECLARE_FLEX_ARRAY(__le32, ino); +} __packed; + +struct f2fs_orphan_footer { __le32 reserved; /* reserved */ __le16 blk_addr; /* block index in current CP */ __le16 blk_count; /* Number of orphan inode blocks in CP */ @@ -260,26 +273,14 @@ struct node_footer { } __packed; /* Address Pointers in an Inode */ -#define DEF_ADDRS_PER_INODE ((F2FS_BLKSIZE - OFFSET_OF_END_OF_I_EXT \ - - SIZE_OF_I_NID \ - - sizeof(struct node_footer)) / sizeof(__le32)) -#define CUR_ADDRS_PER_INODE(inode) (DEF_ADDRS_PER_INODE - \ - get_extra_isize(inode)) +#define F2FS_DEF_ADDRS_PER_INODE(blocksize) \ + (((blocksize) - OFFSET_OF_END_OF_I_EXT - SIZE_OF_I_NID - \ + sizeof(struct node_footer)) / sizeof(__le32)) #define DEF_NIDS_PER_INODE 5 /* Node IDs in an Inode */ #define ADDRS_PER_INODE(inode) addrs_per_page(inode, true) /* Address Pointers in a Direct Block */ -#define DEF_ADDRS_PER_BLOCK ((F2FS_BLKSIZE - sizeof(struct node_footer)) / sizeof(__le32)) #define ADDRS_PER_BLOCK(inode) addrs_per_page(inode, false) -/* Node IDs in an Indirect Block */ -#define NIDS_PER_BLOCK ((F2FS_BLKSIZE - sizeof(struct node_footer)) / sizeof(__le32)) - -#define ADDRS_PER_PAGE(folio, inode) (addrs_per_page(inode, IS_INODE(folio))) - -#define NODE_DIR1_BLOCK (DEF_ADDRS_PER_INODE + 1) -#define NODE_DIR2_BLOCK (DEF_ADDRS_PER_INODE + 2) -#define NODE_IND1_BLOCK (DEF_ADDRS_PER_INODE + 3) -#define NODE_IND2_BLOCK (DEF_ADDRS_PER_INODE + 4) -#define NODE_DIND_BLOCK (DEF_ADDRS_PER_INODE + 5) +#define ADDRS_PER_PAGE(folio, inode) (addrs_per_page(inode, IS_INODE(F2FS_I_SB(inode), folio))) #define F2FS_INLINE_XATTR 0x01 /* file inline xattr flag */ #define F2FS_INLINE_DATA 0x02 /* file inline data flag */ @@ -339,18 +340,26 @@ struct f2fs_inode { */ __le32 i_extra_end[0]; /* for attribute size calculation */ } __packed; - __le32 i_addr[DEF_ADDRS_PER_INODE]; /* Pointers to data blocks */ + DECLARE_FLEX_ARRAY(__le32, i_addr); /* data block pointers */ }; - __le32 i_nid[DEF_NIDS_PER_INODE]; /* direct(2), indirect(2), - double_indirect(1) node id */ + /* + * __le32 i_nid[DEF_NIDS_PER_INODE]; + * direct(2), indirect(2), double_indirect(1) node IDs + * + * It is stored immediately before the node footer at the end of the + * filesystem block. Its offset depends on the filesystem block size, so + * locate it dynamically with F2FS_INODE_NIDS(). + */ } __packed; struct direct_node { - __le32 addr[DEF_ADDRS_PER_BLOCK]; /* array of data block address */ + /* The address count depends on the filesystem block size. */ + DECLARE_FLEX_ARRAY(__le32, addr); /* array of data block address */ } __packed; struct indirect_node { - __le32 nid[NIDS_PER_BLOCK]; /* array of data block address */ + /* The node ID count depends on the filesystem block size. */ + DECLARE_FLEX_ARRAY(__le32, nid); /* array of data block address */ } __packed; enum { @@ -369,14 +378,18 @@ struct f2fs_node { struct direct_node dn; struct indirect_node in; }; - struct node_footer footer; + /* + * struct node_footer footer; + * + * It is stored at the end of the filesystem block, after the inode or + * direct/indirect node data. Its offset depends on the filesystem block + * size, so locate it dynamically with F2FS_NODE_FOOTER(). + */ } __packed; /* * For NAT entries */ -#define NAT_ENTRY_PER_BLOCK (F2FS_BLKSIZE / sizeof(struct f2fs_nat_entry)) - struct f2fs_nat_entry { __u8 version; /* latest version of cached nat entry */ __le32 ino; /* inode number */ @@ -384,7 +397,8 @@ struct f2fs_nat_entry { } __packed; struct f2fs_nat_block { - struct f2fs_nat_entry entries[NAT_ENTRY_PER_BLOCK]; + /* The entry count depends on the filesystem block size. */ + DECLARE_FLEX_ARRAY(struct f2fs_nat_entry, entries); } __packed; /* @@ -396,8 +410,6 @@ struct f2fs_nat_block { * Not allow to change this. */ #define SIT_VBLOCK_MAP_SIZE 64 -#define SIT_ENTRY_PER_BLOCK (F2FS_BLKSIZE / sizeof(struct f2fs_sit_entry)) - /* * F2FS uses 4 bytes to represent block address. As a result, supported size of * disk is 16 TB for a 4K page size and 64 TB for a 16K page size and it equals @@ -424,8 +436,13 @@ struct f2fs_sit_entry { __le64 mtime; /* segment age for cleaning */ } __packed; +/* + * The on-disk SIT block is a filesystem-block-sized array of SIT entries. + * Its entry count depends on the filesystem block size, so it must be + * calculated by the caller rather than implied by this C structure. + */ struct f2fs_sit_block { - struct f2fs_sit_entry entries[SIT_ENTRY_PER_BLOCK]; + DECLARE_FLEX_ARRAY(struct f2fs_sit_entry, entries); } __packed; /* @@ -595,15 +612,7 @@ typedef __le32 f2fs_hash_t; * dentry, when converting inline dentry we should handle this carefully. */ -/* the number of dentry in a block */ -#define NR_DENTRY_IN_BLOCK ((BITS_PER_BYTE * F2FS_BLKSIZE) / \ - ((SIZE_OF_DIR_ENTRY + F2FS_SLOT_LEN) * BITS_PER_BYTE + 1)) #define SIZE_OF_DIR_ENTRY 11 /* by byte */ -#define SIZE_OF_DENTRY_BITMAP ((NR_DENTRY_IN_BLOCK + BITS_PER_BYTE - 1) / \ - BITS_PER_BYTE) -#define SIZE_OF_RESERVED (F2FS_BLKSIZE - ((SIZE_OF_DIR_ENTRY + \ - F2FS_SLOT_LEN) * \ - NR_DENTRY_IN_BLOCK + SIZE_OF_DENTRY_BITMAP)) #define MIN_INLINE_DENTRY_SIZE 40 /* just include '.' and '..' entries */ /* One directory entry slot representing F2FS_SLOT_LEN-sized file name */ @@ -614,14 +623,21 @@ struct f2fs_dir_entry { __u8 file_type; /* file type */ } __packed; -/* Block-sized directory entry block */ -struct f2fs_dentry_block { - /* validity bitmap for directory entries in each block */ - __u8 dentry_bitmap[SIZE_OF_DENTRY_BITMAP]; - __u8 reserved[SIZE_OF_RESERVED]; - struct f2fs_dir_entry dentry[NR_DENTRY_IN_BLOCK]; - __u8 filename[NR_DENTRY_IN_BLOCK][F2FS_SLOT_LEN]; -} __packed; +/* + * A dentry block is laid out as follows, where the number of entries and all + * offsets are determined by the filesystem block size at runtime: + * + * 0 blocksize + * +--------+----------+-------------------+-----------------------+ + * | bitmap | reserved | dir_entry[entries]| filename[entries][8] | + * +--------+----------+-------------------+-----------------------+ + * + * entries = (BITS_PER_BYTE * blocksize) / + * ((SIZE_OF_DIR_ENTRY + F2FS_SLOT_LEN) * BITS_PER_BYTE + 1) + * bitmap_size = DIV_ROUND_UP(entries, BITS_PER_BYTE) + * reserved_size = blocksize - bitmap_size - + * (SIZE_OF_DIR_ENTRY + F2FS_SLOT_LEN) * entries + */ #define F2FS_DEF_PROJID 0 /* default project ID */ diff --git a/include/linux/fdtable.h b/include/linux/fdtable.h index c45306a9f007..a46781058729 100644 --- a/include/linux/fdtable.h +++ b/include/linux/fdtable.h @@ -25,7 +25,7 @@ struct fdtable { unsigned int max_fds; - struct file __rcu **fd; /* current fd array */ + struct file __rcu **fd __counted_by_ptr(max_fds); /* current fd array */ unsigned long *close_on_exec; unsigned long *open_fds; unsigned long *full_fds_bits; @@ -101,11 +101,22 @@ struct task_struct; void put_files_struct(struct files_struct *fs); int unshare_files(void); +void switch_files_struct(struct task_struct *tsk, struct files_struct *files); +int unshare_fd(unsigned long unshare_flags, struct files_struct **new_fdp); +enum fd_range_flags { + /* Leave behind all descriptors outside of the specified range. */ + FD_RANGE_EXCEPT = (1U << 0), + + /* Only select descriptors that have close-on-exec set. */ + FD_RANGE_CLOEXEC_ONLY = (1U << 1), +}; + struct fd_range { unsigned int from, to; + enum fd_range_flags flags; }; struct files_struct *dup_fd(struct files_struct *, struct fd_range *) __latent_entropy; -void do_close_on_exec(struct files_struct *); +void close_cloexec_files(struct files_struct *); int iterate_fd(struct files_struct *, unsigned, int (*)(const void *, struct file *, unsigned), const void *); diff --git a/include/linux/file.h b/include/linux/file.h index 27484b444d31..41c3c0be1064 100644 --- a/include/linux/file.h +++ b/include/linux/file.h @@ -12,6 +12,7 @@ #include <linux/errno.h> #include <linux/cleanup.h> #include <linux/err.h> +#include <linux/vfsdebug.h> struct file; @@ -129,117 +130,84 @@ extern unsigned int sysctl_nr_open_min, sysctl_nr_open_max; /* * fd_prepare: Combined fd + file allocation cleanup class. - * @err: Error code to indicate if allocation succeeded. - * @__fd: Allocated fd (may not be accessed directly) - * @__file: Allocated struct file pointer (may not be accessed directly) + * @fd: Allocated fd + * @file: Allocated struct file pointer * * Allocates an fd and a file together. On error paths, automatically cleans * up whichever resource was successfully allocated. Allows flexible file * allocation with different functions per usage. * - * Do not use directly. + * Do not declare directly, use FD_PREPARE(). */ struct fd_prepare { - s32 err; - s32 __fd; /* do not access directly */ - struct file *__file; /* do not access directly */ + int fd; + struct file *file; }; -/* Typedef for fd_prepare cleanup guards. */ -typedef struct fd_prepare class_fd_prepare_t; - -/* - * Accessors for fd_prepare class members. - * _Generic() is used for zero-cost type safety. - */ -#define fd_prepare_fd(_fdf) \ - (_Generic((_fdf), struct fd_prepare: (_fdf).__fd)) - -#define fd_prepare_file(_fdf) \ - (_Generic((_fdf), struct fd_prepare: (_fdf).__file)) - /* Do not use directly. */ -static inline void class_fd_prepare_destructor(const struct fd_prepare *fdf) +static __always_inline void __fd_prepare_cleanup(const struct fd_prepare *fdf) { - if (unlikely(fdf->__fd >= 0)) - put_unused_fd(fdf->__fd); - if (unlikely(!IS_ERR_OR_NULL(fdf->__file))) - fput(fdf->__file); + if (unlikely(fdf->fd >= 0)) { + put_unused_fd(fdf->fd); + fput(fdf->file); + } } /* Do not use directly. */ -static inline int class_fd_prepare_lock_err(const struct fd_prepare *fdf) +static __always_inline struct fd_prepare __fd_prepare(int fd, struct file *file) { - if (unlikely(fdf->err)) - return fdf->err; - if (unlikely(fdf->__fd < 0)) - return fdf->__fd; - if (unlikely(IS_ERR(fdf->__file))) - return PTR_ERR(fdf->__file); - if (unlikely(!fdf->__file)) - return -ENOMEM; - return 0; + if (fd >= 0 && IS_ERR_OR_NULL(file)) { + int err = file ? PTR_ERR(file) : -ENOMEM; + + put_unused_fd(fd); + fd = err; + file = NULL; + } + return (struct fd_prepare){ .fd = fd, .file = file }; } /* - * __FD_PREPARE_INIT - Helper to initialize fd_prepare class. - * @_fd_flags: flags for get_unused_fd_flags() - * @_file_owned: expression that returns struct file * - * - * Returns a struct fd_prepare with fd, file, and err set. - * If fd allocation fails, fd will be negative and err will be set. If - * fd succeeds but file_init_expr fails, file will be ERR_PTR and err - * will be set. The err field is the single source of truth for error - * checking. - */ -#define __FD_PREPARE_INIT(_fd_flags, _file_owned) \ - ({ \ - struct fd_prepare fdf = { \ - .__fd = get_unused_fd_flags((_fd_flags)), \ - }; \ - if (likely(fdf.__fd >= 0)) \ - fdf.__file = (_file_owned); \ - fdf.err = ACQUIRE_ERR(fd_prepare, &fdf); \ - fdf; \ - }) - -/* - * FD_PREPARE - Macro to declare and initialize an fd_prepare variable. + * FD_PREPARE - Declare and initialize an fd_prepare instance. * - * Declares and initializes an fd_prepare variable with automatic - * cleanup. No separate scope required - cleanup happens when variable - * goes out of scope. + * This allocates a new fd and only evaluates @_file_owned if the + * allocation succeeded. Cleanup happens when the variable goes out of + * scope and the guard releases whichever of the descriptor and the file + * was allocated. If fd_publish() was called the fd and file are + * published and cleanup becomes a nop. * - * @_fdf: name of struct fd_prepare variable to define + * @_fdf: name of the const struct fd_prepare pointer to define * @_fd_flags: flags for get_unused_fd_flags() * @_file_owned: struct file to take ownership of (can be expression) */ +#define __FD_PREPARE(_guard, _fdf, _fd_flags, _file_owned) \ + struct fd_prepare _guard __cleanup(__fd_prepare_cleanup) = ({ \ + int __fd = get_unused_fd_flags(_fd_flags); \ + __fd_prepare(__fd, __fd < 0 ? NULL : (_file_owned)); \ + }); \ + const struct fd_prepare *const _fdf = &_guard + #define FD_PREPARE(_fdf, _fd_flags, _file_owned) \ - CLASS_INIT(fd_prepare, _fdf, __FD_PREPARE_INIT(_fd_flags, _file_owned)) + __FD_PREPARE(__UNIQUE_ID(fd_prepare), _fdf, _fd_flags, _file_owned) /* * fd_publish - Publish prepared fd and file to the fd table. - * @_fdf: struct fd_prepare variable + * @fdf: struct fd_prepare pointer defined by FD_PREPARE() */ -#define fd_publish(_fdf) \ - ({ \ - struct fd_prepare *fdp = &(_fdf); \ - VFS_WARN_ON_ONCE(fdp->err); \ - VFS_WARN_ON_ONCE(fdp->__fd < 0); \ - VFS_WARN_ON_ONCE(IS_ERR_OR_NULL(fdp->__file)); \ - fd_install(fdp->__fd, fdp->__file); \ - retain_and_null_ptr(fdp->__file); \ - take_fd(fdp->__fd); \ - }) +static __always_inline int fd_publish(const struct fd_prepare *fdf) +{ + /* Callers only get a const view, the guard itself is writable. */ + struct fd_prepare *guard = (struct fd_prepare *)fdf; + + VFS_WARN_ON_ONCE(guard->fd < 0); + fd_install(guard->fd, guard->file); + return take_fd(guard->fd); +} /* Do not use directly. */ -#define __FD_ADD(_fdf, _fd_flags, _file_owned) \ - ({ \ - FD_PREPARE(_fdf, _fd_flags, _file_owned); \ - s32 ret = _fdf.err; \ - if (likely(!ret)) \ - ret = fd_publish(_fdf); \ - ret; \ +#define __FD_ADD(_fdf, _fd_flags, _file_owned) \ + ({ \ + FD_PREPARE(_fdf, _fd_flags, _file_owned); \ + _fdf->fd < 0 ? _fdf->fd : fd_publish(_fdf); \ }) /* diff --git a/include/linux/fileattr.h b/include/linux/fileattr.h index 58044b598016..09e32b84e02a 100644 --- a/include/linux/fileattr.h +++ b/include/linux/fileattr.h @@ -74,7 +74,7 @@ static inline bool fileattr_has_fsx(const struct file_kattr *fa) } int vfs_fileattr_get(struct dentry *dentry, struct file_kattr *fa); -int vfs_fileattr_set(struct mnt_idmap *idmap, struct dentry *dentry, +int vfs_fileattr_set(const struct mnt_idmap *idmap, struct dentry *dentry, struct file_kattr *fa); int ioctl_getflags(struct file *file, unsigned int __user *argp); int ioctl_setflags(struct file *file, unsigned int __user *argp); diff --git a/include/linux/fs.h b/include/linux/fs.h index f9d1e05e8ae6..3db90996756f 100644 --- a/include/linux/fs.h +++ b/include/linux/fs.h @@ -1436,10 +1436,10 @@ static inline void i_gid_write(struct inode *inode, gid_t gid) * @idmap: idmap of the mount the inode was found from * @inode: inode to map * - * Return: whe inode's i_uid mapped down according to @idmap. + * Return: the inode's i_uid mapped down according to @idmap. * If the inode's i_uid has no mapping INVALID_VFSUID is returned. */ -static inline vfsuid_t i_uid_into_vfsuid(struct mnt_idmap *idmap, +static inline vfsuid_t i_uid_into_vfsuid(const struct mnt_idmap *idmap, const struct inode *inode) { return make_vfsuid(idmap, i_user_ns(inode), inode->i_uid); @@ -1456,7 +1456,7 @@ static inline vfsuid_t i_uid_into_vfsuid(struct mnt_idmap *idmap, * * Return: true if @inode's i_uid field needs to be updated, false if not. */ -static inline bool i_uid_needs_update(struct mnt_idmap *idmap, +static inline bool i_uid_needs_update(const struct mnt_idmap *idmap, const struct iattr *attr, const struct inode *inode) { @@ -1474,7 +1474,7 @@ static inline bool i_uid_needs_update(struct mnt_idmap *idmap, * Safely update @inode's i_uid field translating the vfsuid of any idmapped * mount into the filesystem kuid. */ -static inline void i_uid_update(struct mnt_idmap *idmap, +static inline void i_uid_update(const struct mnt_idmap *idmap, const struct iattr *attr, struct inode *inode) { @@ -1491,7 +1491,7 @@ static inline void i_uid_update(struct mnt_idmap *idmap, * Return: the inode's i_gid mapped down according to @idmap. * If the inode's i_gid has no mapping INVALID_VFSGID is returned. */ -static inline vfsgid_t i_gid_into_vfsgid(struct mnt_idmap *idmap, +static inline vfsgid_t i_gid_into_vfsgid(const struct mnt_idmap *idmap, const struct inode *inode) { return make_vfsgid(idmap, i_user_ns(inode), inode->i_gid); @@ -1508,7 +1508,7 @@ static inline vfsgid_t i_gid_into_vfsgid(struct mnt_idmap *idmap, * * Return: true if @inode's i_gid field needs to be updated, false if not. */ -static inline bool i_gid_needs_update(struct mnt_idmap *idmap, +static inline bool i_gid_needs_update(const struct mnt_idmap *idmap, const struct iattr *attr, const struct inode *inode) { @@ -1526,7 +1526,7 @@ static inline bool i_gid_needs_update(struct mnt_idmap *idmap, * Safely update @inode's i_gid field translating the vfsgid of any idmapped * mount into the filesystem kgid. */ -static inline void i_gid_update(struct mnt_idmap *idmap, +static inline void i_gid_update(const struct mnt_idmap *idmap, const struct iattr *attr, struct inode *inode) { @@ -1544,7 +1544,7 @@ static inline void i_gid_update(struct mnt_idmap *idmap, * an idmapped mount map the caller's fsuid according to @idmap. */ static inline void inode_fsuid_set(struct inode *inode, - struct mnt_idmap *idmap) + const struct mnt_idmap *idmap) { inode->i_uid = mapped_fsuid(idmap, i_user_ns(inode)); } @@ -1558,7 +1558,7 @@ static inline void inode_fsuid_set(struct inode *inode, * an idmapped mount map the caller's fsgid according to @idmap. */ static inline void inode_fsgid_set(struct inode *inode, - struct mnt_idmap *idmap) + const struct mnt_idmap *idmap) { inode->i_gid = mapped_fsgid(idmap, i_user_ns(inode)); } @@ -1575,7 +1575,7 @@ static inline void inode_fsgid_set(struct inode *inode, * Return: true if fsuid and fsgid is mapped, false if not. */ static inline bool fsuidgid_has_mapping(struct super_block *sb, - struct mnt_idmap *idmap) + const struct mnt_idmap *idmap) { struct user_namespace *fs_userns = sb->s_user_ns; kuid_t kuid; @@ -1755,25 +1755,25 @@ static inline bool file_write_not_started(const struct file *file) return sb_write_not_started(file_inode(file)->i_sb); } -bool inode_owner_or_capable(struct mnt_idmap *idmap, +bool inode_owner_or_capable(const struct mnt_idmap *idmap, const struct inode *inode); /* * VFS helper functions.. */ -int vfs_create(struct mnt_idmap *, struct dentry *, umode_t, +int vfs_create(const struct mnt_idmap *, struct dentry *, umode_t, struct delegated_inode *); -struct dentry *vfs_mkdir(struct mnt_idmap *, struct inode *, +struct dentry *vfs_mkdir(const struct mnt_idmap *, struct inode *, struct dentry *, umode_t, struct delegated_inode *); -int vfs_mknod(struct mnt_idmap *, struct inode *, struct dentry *, +int vfs_mknod(const struct mnt_idmap *, struct inode *, struct dentry *, umode_t, dev_t, struct delegated_inode *); -int vfs_symlink(struct mnt_idmap *, struct inode *, +int vfs_symlink(const struct mnt_idmap *, struct inode *, struct dentry *, const char *, struct delegated_inode *); -int vfs_link(struct dentry *, struct mnt_idmap *, struct inode *, +int vfs_link(struct dentry *, const struct mnt_idmap *, struct inode *, struct dentry *, struct delegated_inode *); -int vfs_rmdir(struct mnt_idmap *, struct inode *, struct dentry *, +int vfs_rmdir(const struct mnt_idmap *, struct inode *, struct dentry *, struct delegated_inode *); -int vfs_unlink(struct mnt_idmap *, struct inode *, struct dentry *, +int vfs_unlink(const struct mnt_idmap *, struct inode *, struct dentry *, struct delegated_inode *); /** @@ -1787,7 +1787,7 @@ int vfs_unlink(struct mnt_idmap *, struct inode *, struct dentry *, * @flags: rename flags */ struct renamedata { - struct mnt_idmap *mnt_idmap; + const struct mnt_idmap *mnt_idmap; struct dentry *old_parent; struct dentry *old_dentry; struct dentry *new_parent; @@ -1798,14 +1798,14 @@ struct renamedata { int vfs_rename(struct renamedata *); -static inline int vfs_whiteout(struct mnt_idmap *idmap, +static inline int vfs_whiteout(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *dentry) { return vfs_mknod(idmap, dir, dentry, S_IFCHR | WHITEOUT_MODE, WHITEOUT_DEV, NULL); } -struct file *kernel_tmpfile_open(struct mnt_idmap *idmap, +struct file *kernel_tmpfile_open(const struct mnt_idmap *idmap, const struct path *parentpath, umode_t mode, int open_flag, const struct cred *cred); @@ -1830,12 +1830,12 @@ extern long compat_ptr_ioctl(struct file *file, unsigned int cmd, /* * VFS file helper functions. */ -void inode_init_owner(struct mnt_idmap *idmap, struct inode *inode, +void inode_init_owner(const struct mnt_idmap *idmap, struct inode *inode, const struct inode *dir, umode_t mode); extern bool may_open_dev(const struct path *path); -umode_t mode_strip_sgid(struct mnt_idmap *idmap, +umode_t mode_strip_sgid(const struct mnt_idmap *idmap, const struct inode *dir, umode_t mode); -bool in_group_or_capable(struct mnt_idmap *idmap, +bool in_group_or_capable(const struct mnt_idmap *idmap, const struct inode *inode, vfsgid_t vfsgid); /* @@ -1994,26 +1994,26 @@ enum fs_update_time { struct inode_operations { struct dentry * (*lookup) (struct inode *,struct dentry *, unsigned int); const char * (*get_link) (struct dentry *, struct inode *, struct delayed_call *); - int (*permission) (struct mnt_idmap *, struct inode *, int); + int (*permission) (const struct mnt_idmap *, struct inode *, int); struct posix_acl * (*get_inode_acl)(struct inode *, int, bool); int (*readlink) (struct dentry *, char __user *,int); - int (*create) (struct mnt_idmap *, struct inode *,struct dentry *, + int (*create) (const struct mnt_idmap *, struct inode *,struct dentry *, umode_t); int (*link) (struct dentry *,struct inode *,struct dentry *); int (*unlink) (struct inode *,struct dentry *); - int (*symlink) (struct mnt_idmap *, struct inode *,struct dentry *, + int (*symlink) (const struct mnt_idmap *, struct inode *,struct dentry *, const char *); - struct dentry *(*mkdir) (struct mnt_idmap *, struct inode *, + struct dentry *(*mkdir) (const struct mnt_idmap *, struct inode *, struct dentry *, umode_t); int (*rmdir) (struct inode *,struct dentry *); - int (*mknod) (struct mnt_idmap *, struct inode *,struct dentry *, + int (*mknod) (const struct mnt_idmap *, struct inode *,struct dentry *, umode_t,dev_t); - int (*rename) (struct mnt_idmap *, struct inode *, struct dentry *, + int (*rename) (const struct mnt_idmap *, struct inode *, struct dentry *, struct inode *, struct dentry *, unsigned int); - int (*setattr) (struct mnt_idmap *, struct dentry *, struct iattr *); - int (*getattr) (struct mnt_idmap *, const struct path *, + int (*setattr) (const struct mnt_idmap *, struct dentry *, struct iattr *); + int (*getattr) (const struct mnt_idmap *, const struct path *, struct kstat *, u32, unsigned int); ssize_t (*listxattr) (struct dentry *, char *, size_t); int (*fiemap)(struct inode *, struct fiemap_extent_info *, u64 start, @@ -2024,13 +2024,13 @@ struct inode_operations { int (*atomic_open)(struct inode *, struct dentry *, struct file *, unsigned open_flag, umode_t create_mode); - int (*tmpfile) (struct mnt_idmap *, struct inode *, + int (*tmpfile) (const struct mnt_idmap *, struct inode *, struct file *, umode_t); - struct posix_acl *(*get_acl)(struct mnt_idmap *, struct dentry *, + struct posix_acl *(*get_acl)(const struct mnt_idmap *, struct dentry *, int); - int (*set_acl)(struct mnt_idmap *, struct dentry *, + int (*set_acl)(const struct mnt_idmap *, struct dentry *, struct posix_acl *, int); - int (*fileattr_set)(struct mnt_idmap *idmap, + int (*fileattr_set)(const struct mnt_idmap *idmap, struct dentry *dentry, struct file_kattr *fa); int (*fileattr_get)(struct dentry *dentry, struct file_kattr *fa); struct offset_ctx *(*get_offset_ctx)(struct inode *inode); @@ -2173,7 +2173,7 @@ extern loff_t vfs_dedupe_file_range_one(struct file *src_file, loff_t src_pos, (inode)->i_rdev == WHITEOUT_DEV) #define IS_ANON_FILE(inode) ((inode)->i_flags & S_ANON_INODE) -static inline bool HAS_UNMAPPED_ID(struct mnt_idmap *idmap, +static inline bool HAS_UNMAPPED_ID(const struct mnt_idmap *idmap, struct inode *inode) { return !vfsuid_valid(i_uid_into_vfsuid(idmap, inode)) || @@ -2459,7 +2459,7 @@ struct filename { static_assert(offsetof(struct filename, iname) % sizeof(long) == 0); static_assert(sizeof(struct filename) % 64 == 0); -static inline struct mnt_idmap *file_mnt_idmap(const struct file *file) +static inline const struct mnt_idmap *file_mnt_idmap(const struct file *file) { return mnt_idmap(file->f_path.mnt); } @@ -2483,7 +2483,7 @@ static inline bool is_idmapped_mnt(const struct vfsmount *mnt) } int vfs_truncate(const struct path *, loff_t); -int do_truncate(struct mnt_idmap *, struct dentry *, loff_t start, +int do_truncate(const struct mnt_idmap *, struct dentry *, loff_t start, unsigned int time_attrs, struct file *filp); extern int vfs_fallocate(struct file *file, int mode, loff_t offset, loff_t len); @@ -2707,10 +2707,10 @@ static inline int bmap(struct inode *inode, sector_t *block) } #endif -int notify_change(struct mnt_idmap *, struct dentry *, +int notify_change(const struct mnt_idmap *, struct dentry *, struct iattr *, struct delegated_inode *); -int inode_permission(struct mnt_idmap *, struct inode *, int); -int generic_permission(struct mnt_idmap *, struct inode *, int); +int inode_permission(const struct mnt_idmap *, struct inode *, int); +int generic_permission(const struct mnt_idmap *, struct inode *, int); static inline int file_permission(struct file *file, int mask) { return inode_permission(file_mnt_idmap(file), @@ -2721,12 +2721,12 @@ static inline int path_permission(const struct path *path, int mask) return inode_permission(mnt_idmap(path->mnt), d_inode(path->dentry), mask); } -int __check_sticky(struct mnt_idmap *idmap, struct inode *dir, +int __check_sticky(const struct mnt_idmap *idmap, struct inode *dir, struct inode *inode); -int may_delete_dentry(struct mnt_idmap *idmap, struct inode *dir, +int may_delete_dentry(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *victim, bool isdir); -int may_create_dentry(struct mnt_idmap *idmap, +int may_create_dentry(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *child); static inline bool execute_ok(struct inode *inode) @@ -3045,9 +3045,9 @@ static inline struct inode *new_inode_pseudo(struct super_block *sb) } extern struct inode *new_inode(struct super_block *sb); extern void free_inode_nonrcu(struct inode *inode); -extern int setattr_should_drop_suidgid(struct mnt_idmap *, struct inode *); +extern int setattr_should_drop_suidgid(const struct mnt_idmap *, struct inode *); extern int file_remove_privs(struct file *); -int setattr_should_drop_sgid(struct mnt_idmap *idmap, +int setattr_should_drop_sgid(const struct mnt_idmap *idmap, const struct inode *inode); /* @@ -3204,7 +3204,7 @@ extern int page_symlink(struct inode *inode, const char *symname, int len); extern const struct inode_operations page_symlink_inode_operations; extern void kfree_link(void *); void fill_mg_cmtime(struct kstat *stat, u32 request_mask, struct inode *inode); -void generic_fillattr(struct mnt_idmap *, u32, struct inode *, struct kstat *); +void generic_fillattr(const struct mnt_idmap *, u32, struct inode *, struct kstat *); void generic_fill_statx_attr(struct inode *inode, struct kstat *stat); void generic_fill_statx_atomic_writes(struct kstat *stat, unsigned int unit_min, @@ -3261,9 +3261,9 @@ extern int dcache_dir_open(struct inode *, struct file *); extern int dcache_dir_close(struct inode *, struct file *); extern loff_t dcache_dir_lseek(struct file *, loff_t, int); extern int dcache_readdir(struct file *, struct dir_context *); -extern int simple_setattr(struct mnt_idmap *, struct dentry *, +extern int simple_setattr(const struct mnt_idmap *, struct dentry *, struct iattr *); -extern int simple_getattr(struct mnt_idmap *, const struct path *, +extern int simple_getattr(const struct mnt_idmap *, const struct path *, struct kstat *, u32, unsigned int); extern int simple_statfs(struct dentry *, struct kstatfs *); extern int simple_open(struct inode *inode, struct file *file); @@ -3276,7 +3276,7 @@ void simple_rename_timestamp(struct inode *old_dir, struct dentry *old_dentry, struct inode *new_dir, struct dentry *new_dentry); extern int simple_rename_exchange(struct inode *old_dir, struct dentry *old_dentry, struct inode *new_dir, struct dentry *new_dentry); -extern int simple_rename(struct mnt_idmap *, struct inode *, +extern int simple_rename(const struct mnt_idmap *, struct inode *, struct dentry *, struct inode *, struct dentry *, unsigned int); extern void simple_recursive_removal(struct dentry *, @@ -3397,11 +3397,11 @@ static inline bool generic_ci_validate_strict_name(struct inode *dir, } #endif -int may_setattr(struct mnt_idmap *idmap, struct inode *inode, +int may_setattr(const struct mnt_idmap *idmap, struct inode *inode, unsigned int ia_valid); -int setattr_prepare(struct mnt_idmap *, struct dentry *, struct iattr *); +int setattr_prepare(const struct mnt_idmap *, struct dentry *, struct iattr *); extern int inode_newsize_ok(const struct inode *, loff_t offset); -void setattr_copy(struct mnt_idmap *, struct inode *inode, +void setattr_copy(const struct mnt_idmap *, struct inode *inode, const struct iattr *attr); extern int file_update_time(struct file *file); @@ -3578,7 +3578,7 @@ static inline bool is_sxid(umode_t mode) return mode & (S_ISUID | S_ISGID); } -static inline int check_sticky(struct mnt_idmap *idmap, +static inline int check_sticky(const struct mnt_idmap *idmap, struct inode *dir, struct inode *inode) { if (!(dir->i_mode & S_ISVTX)) @@ -3653,23 +3653,6 @@ extern int vfs_fadvise(struct file *file, loff_t offset, loff_t len, extern int generic_fadvise(struct file *file, loff_t offset, loff_t len, int advice); -static inline bool vfs_empty_path(int dfd, const char __user *path) -{ - char c; - - if (dfd < 0) - return false; - - /* We now allow NULL to be used for empty path. */ - if (!path) - return true; - - if (unlikely(get_user(c, path))) - return false; - - return !c; -} - int generic_atomic_write_valid(struct kiocb *iocb, struct iov_iter *iter); static inline bool extensible_ioctl_valid(unsigned int cmd_a, diff --git a/include/linux/fs_context.h b/include/linux/fs_context.h index 0d6c8a6d7be2..c920aba5177c 100644 --- a/include/linux/fs_context.h +++ b/include/linux/fs_context.h @@ -150,6 +150,10 @@ extern int vfs_parse_fs_param_source(struct fs_context *fc, struct fs_parameter *param); extern void fc_drop_locked(struct fs_context *fc); +extern int get_tree_super(struct fs_context *fc, + int (*test)(struct super_block *, struct fs_context *), + int (*fill_super)(struct super_block *sb, + struct fs_context *fc)); extern int get_tree_nodev(struct fs_context *fc, int (*fill_super)(struct super_block *sb, struct fs_context *fc)); diff --git a/include/linux/fscache-cache.h b/include/linux/fscache-cache.h index 4c91a019972b..ee524c863fa9 100644 --- a/include/linux/fscache-cache.h +++ b/include/linux/fscache-cache.h @@ -67,7 +67,7 @@ struct fscache_cache_ops { /* Change the size of a data object */ void (*resize_cookie)(struct netfs_cache_resources *cres, - loff_t new_size); + uoff_t new_size); /* Invalidate an object */ bool (*invalidate_cookie)(struct fscache_cookie *cookie); diff --git a/include/linux/fscache.h b/include/linux/fscache.h index 58fdb9605425..f2d958bd1f48 100644 --- a/include/linux/fscache.h +++ b/include/linux/fscache.h @@ -112,7 +112,7 @@ struct fscache_cookie { struct list_head proc_link; /* Link in proc list */ struct list_head commit_link; /* Link in commit queue */ struct work_struct work; /* Commit/relinq/withdraw work */ - loff_t object_size; /* Size of the netfs object */ + uoff_t object_size; /* Size of the netfs object */ unsigned long unused_at; /* Time at which unused (jiffies) */ unsigned long flags; #define FSCACHE_COOKIE_RELINQUISHED 0 /* T if cookie has been relinquished */ @@ -147,6 +147,23 @@ struct fscache_cookie { }; }; +enum fscache_extent_type { + FSCACHE_EXTENT_DATA, + FSCACHE_EXTENT_ZERO, +} __mode(byte); + +/* + * Cache occupancy information. + */ +struct fscache_occupancy { + unsigned long long query_from; /* Point to query from */ + unsigned long long query_to; /* Point to query to */ + unsigned long long cached_from[2]; /* Point at which cache extents start */ + unsigned long long cached_to[2]; /* Point at which cache extents end */ + unsigned int granularity; /* Granularity desired */ + enum fscache_extent_type cached_type[2]; /* Type of cache extent */ +}; + /* * slow-path functions for when there is actually caching available, and the * netfs does actually have a valid token @@ -163,22 +180,22 @@ extern struct fscache_cookie *__fscache_acquire_cookie( u8, const void *, size_t, const void *, size_t, - loff_t); + uoff_t); extern void __fscache_use_cookie(struct fscache_cookie *, bool); -extern void __fscache_unuse_cookie(struct fscache_cookie *, const void *, const loff_t *); +extern void __fscache_unuse_cookie(struct fscache_cookie *, const void *, const uoff_t *); extern void __fscache_relinquish_cookie(struct fscache_cookie *, bool); -extern void __fscache_resize_cookie(struct fscache_cookie *, loff_t); -extern void __fscache_invalidate(struct fscache_cookie *, const void *, loff_t, unsigned int); +extern void __fscache_resize_cookie(struct fscache_cookie *, uoff_t); +extern void __fscache_invalidate(struct fscache_cookie *, const void *, uoff_t, unsigned int); extern int __fscache_begin_read_operation(struct netfs_cache_resources *, struct fscache_cookie *); extern int __fscache_begin_write_operation(struct netfs_cache_resources *, struct fscache_cookie *); void __fscache_write_to_cache(struct fscache_cookie *cookie, struct address_space *mapping, - loff_t start, size_t len, loff_t i_size, + uoff_t start, size_t len, uoff_t i_size, netfs_io_terminated_t term_func, void *term_func_priv, bool using_pgpriv2, bool cond); -extern void __fscache_clear_page_bits(struct address_space *, loff_t, size_t); +extern void __fscache_clear_page_bits(struct address_space *, uoff_t, size_t); /** * fscache_acquire_volume - Register a volume as desiring caching services @@ -249,7 +266,7 @@ struct fscache_cookie *fscache_acquire_cookie(struct fscache_volume *volume, size_t index_key_len, const void *aux_data, size_t aux_data_len, - loff_t object_size) + uoff_t object_size) { if (!fscache_volume_valid(volume)) return NULL; @@ -286,7 +303,7 @@ static inline void fscache_use_cookie(struct fscache_cookie *cookie, */ static inline void fscache_unuse_cookie(struct fscache_cookie *cookie, const void *aux_data, - const loff_t *object_size) + const uoff_t *object_size) { if (fscache_cookie_valid(cookie)) __fscache_unuse_cookie(cookie, aux_data, object_size); @@ -327,7 +344,7 @@ static inline void *fscache_get_aux(struct fscache_cookie *cookie) */ static inline void fscache_update_aux(struct fscache_cookie *cookie, - const void *aux_data, const loff_t *object_size) + const void *aux_data, const uoff_t *object_size) { void *p = fscache_get_aux(cookie); @@ -343,7 +360,7 @@ extern atomic_t fscache_n_updates; static inline void __fscache_update_cookie(struct fscache_cookie *cookie, const void *aux_data, - const loff_t *object_size) + const uoff_t *object_size) { #ifdef CONFIG_FSCACHE_STATS atomic_inc(&fscache_n_updates); @@ -369,7 +386,7 @@ void __fscache_update_cookie(struct fscache_cookie *cookie, const void *aux_data */ static inline void fscache_update_cookie(struct fscache_cookie *cookie, const void *aux_data, - const loff_t *object_size) + const uoff_t *object_size) { if (fscache_cookie_enabled(cookie)) __fscache_update_cookie(cookie, aux_data, object_size); @@ -386,7 +403,7 @@ void fscache_update_cookie(struct fscache_cookie *cookie, const void *aux_data, * description. */ static inline -void fscache_resize_cookie(struct fscache_cookie *cookie, loff_t new_size) +void fscache_resize_cookie(struct fscache_cookie *cookie, uoff_t new_size) { if (fscache_cookie_enabled(cookie)) __fscache_resize_cookie(cookie, new_size); @@ -413,7 +430,7 @@ void fscache_resize_cookie(struct fscache_cookie *cookie, loff_t new_size) */ static inline void fscache_invalidate(struct fscache_cookie *cookie, - const void *aux_data, loff_t size, unsigned int flags) + const void *aux_data, uoff_t size, unsigned int flags) { if (fscache_cookie_enabled(cookie)) __fscache_invalidate(cookie, aux_data, size, flags); @@ -502,7 +519,7 @@ static inline void fscache_end_operation(struct netfs_cache_resources *cres) */ static inline int fscache_read(struct netfs_cache_resources *cres, - loff_t start_pos, + uoff_t start_pos, struct iov_iter *iter, enum netfs_read_from_hole read_hole, netfs_io_terminated_t term_func, @@ -561,7 +578,7 @@ int fscache_begin_write_operation(struct netfs_cache_resources *cres, */ static inline int fscache_write(struct netfs_cache_resources *cres, - loff_t start_pos, + uoff_t start_pos, struct iov_iter *iter, netfs_io_terminated_t term_func, void *term_func_priv) @@ -581,7 +598,7 @@ int fscache_write(struct netfs_cache_resources *cres, * waiting. */ static inline void fscache_clear_page_bits(struct address_space *mapping, - loff_t start, size_t len, + uoff_t start, size_t len, bool caching) { if (caching) @@ -615,7 +632,7 @@ static inline void fscache_clear_page_bits(struct address_space *mapping, */ static inline void fscache_write_to_cache(struct fscache_cookie *cookie, struct address_space *mapping, - loff_t start, size_t len, loff_t i_size, + uoff_t start, size_t len, uoff_t i_size, netfs_io_terminated_t term_func, void *term_func_priv, bool using_pgpriv2, bool caching) diff --git a/include/linux/iomap.h b/include/linux/iomap.h index bc7ae6327dbf..59718f73c15a 100644 --- a/include/linux/iomap.h +++ b/include/linux/iomap.h @@ -483,13 +483,35 @@ sector_t iomap_bmap(struct address_space *mapping, sector_t bno, #define IOMAP_IOEND_BOUNDARY (1U << 2) /* is direct I/O */ #define IOMAP_IOEND_DIRECT (1U << 3) +/* generate integrity (PI) information */ +#ifdef CONFIG_BLK_DEV_INTEGRITY +#define IOMAP_IOEND_INTEGRITY (1U << 4) +#else +#define IOMAP_IOEND_INTEGRITY 0 +#endif /* CONFIG_BLK_DEV_INTEGRITY */ /* * Flags that if set on either ioend prevent the merge of two ioends. * (IOMAP_IOEND_BOUNDARY also prevents merges, but only one-way) */ #define IOMAP_IOEND_NOMERGE_FLAGS \ - (IOMAP_IOEND_SHARED | IOMAP_IOEND_UNWRITTEN | IOMAP_IOEND_DIRECT) + (IOMAP_IOEND_SHARED | IOMAP_IOEND_UNWRITTEN | IOMAP_IOEND_DIRECT | \ + IOMAP_IOEND_INTEGRITY) + +/* ioend flags directly implied by iomap flags */ +static inline u16 iomap_ioend_flags(const struct iomap *iomap) +{ + unsigned int flags = 0; + + if (iomap->type == IOMAP_UNWRITTEN) + flags |= IOMAP_IOEND_UNWRITTEN; + if (iomap->flags & IOMAP_F_SHARED) + flags |= IOMAP_IOEND_SHARED; + if (iomap->flags & IOMAP_F_INTEGRITY) + flags |= IOMAP_IOEND_INTEGRITY; + + return flags; +} /* * Structure for writeback I/O completions. @@ -500,6 +522,7 @@ sector_t iomap_bmap(struct address_space *mapping, sector_t bno, struct iomap_ioend { struct list_head io_list; /* next ioend in chain */ u16 io_flags; /* IOMAP_IOEND_* */ + u32 io_bvec_offset; /* offset into first bvec */ struct inode *io_inode; /* file being written to */ size_t io_size; /* size of the extent */ atomic_t io_remaining; /* completetion defer count */ @@ -517,6 +540,13 @@ static inline struct iomap_ioend *iomap_ioend_from_bio(struct bio *bio) return container_of(bio, struct iomap_ioend, io_bio); } +#define BVEC_ITER_IOEND(_ioend) \ +{ \ + .bi_sector = (_ioend)->io_sector, \ + .bi_size = (_ioend)->io_size, \ + .bi_offset = (_ioend)->io_bvec_offset, \ +} + struct iomap_writeback_ops { /* * Performs writeback on the passed in range @@ -565,6 +595,7 @@ void iomap_finish_ioends(struct iomap_ioend *ioend, int error); void iomap_ioend_try_merge(struct iomap_ioend *ioend, struct list_head *more_ioends); void iomap_sort_ioends(struct list_head *ioend_list); +int iomap_ioend_integrity_verify(struct iomap_ioend *ioend); ssize_t iomap_add_to_ioend(struct iomap_writepage_ctx *wpc, struct folio *folio, loff_t pos, loff_t end_pos, unsigned int dirty_len); int iomap_ioend_writeback_submit(struct iomap_writepage_ctx *wpc, int error); @@ -577,6 +608,11 @@ void iomap_finish_folio_write(struct inode *inode, struct folio *folio, int iomap_writeback_folio(struct iomap_writepage_ctx *wpc, struct folio *folio); int iomap_writepages(struct iomap_writepage_ctx *wpc); +void iomap_bounce_read(struct iomap_ioend *orig_ioend, unsigned int minsize, + void (*submit_ioend)(struct iomap_ioend *ioend)); +void iomap_bounce_read_end_io(struct iomap_ioend *ioend, struct bio *orig_bio, + int error); + struct iomap_read_folio_ctx { const struct iomap_read_ops *ops; struct folio *cur_folio; diff --git a/include/linux/lsm_hook_defs.h b/include/linux/lsm_hook_defs.h index 65c9609ec207..c9561564585e 100644 --- a/include/linux/lsm_hook_defs.h +++ b/include/linux/lsm_hook_defs.h @@ -36,6 +36,7 @@ LSM_HOOK(int, 0, binder_transfer_file, const struct cred *from, LSM_HOOK(int, 0, ptrace_access_check, struct task_struct *child, unsigned int mode) LSM_HOOK(int, 0, ptrace_traceme, struct task_struct *parent) +LSM_HOOK(int, 0, mem_foll_force, const struct cred *subject, bool opened_by_owner) LSM_HOOK(int, 0, capget, const struct task_struct *target, kernel_cap_t *effective, kernel_cap_t *inheritable, kernel_cap_t *permitted) LSM_HOOK(int, 0, capset, struct cred *new, const struct cred *old, @@ -94,7 +95,7 @@ LSM_HOOK(int, 0, path_mkdir, const struct path *dir, struct dentry *dentry, LSM_HOOK(int, 0, path_rmdir, const struct path *dir, struct dentry *dentry) LSM_HOOK(int, 0, path_mknod, const struct path *dir, struct dentry *dentry, umode_t mode, unsigned int dev) -LSM_HOOK(void, LSM_RET_VOID, path_post_mknod, struct mnt_idmap *idmap, +LSM_HOOK(void, LSM_RET_VOID, path_post_mknod, const struct mnt_idmap *idmap, struct dentry *dentry) LSM_HOOK(int, 0, path_truncate, const struct path *path) LSM_HOOK(int, 0, path_symlink, const struct path *dir, struct dentry *dentry, @@ -122,7 +123,7 @@ LSM_HOOK(int, 0, inode_init_security_anon, struct inode *inode, const struct qstr *name, const struct inode *context_inode) LSM_HOOK(int, 0, inode_create, struct inode *dir, struct dentry *dentry, umode_t mode) -LSM_HOOK(void, LSM_RET_VOID, inode_post_create_tmpfile, struct mnt_idmap *idmap, +LSM_HOOK(void, LSM_RET_VOID, inode_post_create_tmpfile, const struct mnt_idmap *idmap, struct inode *inode) LSM_HOOK(int, 0, inode_link, struct dentry *old_dentry, struct inode *dir, struct dentry *new_dentry) @@ -140,39 +141,39 @@ LSM_HOOK(int, 0, inode_readlink, struct dentry *dentry) LSM_HOOK(int, 0, inode_follow_link, struct dentry *dentry, struct inode *inode, bool rcu) LSM_HOOK(int, 0, inode_permission, struct inode *inode, int mask) -LSM_HOOK(int, 0, inode_setattr, struct mnt_idmap *idmap, struct dentry *dentry, +LSM_HOOK(int, 0, inode_setattr, const struct mnt_idmap *idmap, struct dentry *dentry, struct iattr *attr) -LSM_HOOK(void, LSM_RET_VOID, inode_post_setattr, struct mnt_idmap *idmap, +LSM_HOOK(void, LSM_RET_VOID, inode_post_setattr, const struct mnt_idmap *idmap, struct dentry *dentry, int ia_valid) LSM_HOOK(int, 0, inode_getattr, const struct path *path) LSM_HOOK(int, 0, inode_xattr_skipcap, const char *name) -LSM_HOOK(int, 0, inode_setxattr, struct mnt_idmap *idmap, +LSM_HOOK(int, 0, inode_setxattr, const struct mnt_idmap *idmap, struct dentry *dentry, const char *name, const void *value, size_t size, int flags) LSM_HOOK(void, LSM_RET_VOID, inode_post_setxattr, struct dentry *dentry, const char *name, const void *value, size_t size, int flags) LSM_HOOK(int, 0, inode_getxattr, struct dentry *dentry, const char *name) LSM_HOOK(int, 0, inode_listxattr, struct dentry *dentry) -LSM_HOOK(int, 0, inode_removexattr, struct mnt_idmap *idmap, +LSM_HOOK(int, 0, inode_removexattr, const struct mnt_idmap *idmap, struct dentry *dentry, const char *name) LSM_HOOK(void, LSM_RET_VOID, inode_post_removexattr, struct dentry *dentry, const char *name) LSM_HOOK(int, 0, inode_file_setattr, struct dentry *dentry, struct file_kattr *fa) LSM_HOOK(int, 0, inode_file_getattr, struct dentry *dentry, struct file_kattr *fa) -LSM_HOOK(int, 0, inode_set_acl, struct mnt_idmap *idmap, +LSM_HOOK(int, 0, inode_set_acl, const struct mnt_idmap *idmap, struct dentry *dentry, const char *acl_name, struct posix_acl *kacl) LSM_HOOK(void, LSM_RET_VOID, inode_post_set_acl, struct dentry *dentry, const char *acl_name, struct posix_acl *kacl) -LSM_HOOK(int, 0, inode_get_acl, struct mnt_idmap *idmap, +LSM_HOOK(int, 0, inode_get_acl, const struct mnt_idmap *idmap, struct dentry *dentry, const char *acl_name) -LSM_HOOK(int, 0, inode_remove_acl, struct mnt_idmap *idmap, +LSM_HOOK(int, 0, inode_remove_acl, const struct mnt_idmap *idmap, struct dentry *dentry, const char *acl_name) -LSM_HOOK(void, LSM_RET_VOID, inode_post_remove_acl, struct mnt_idmap *idmap, +LSM_HOOK(void, LSM_RET_VOID, inode_post_remove_acl, const struct mnt_idmap *idmap, struct dentry *dentry, const char *acl_name) LSM_HOOK(int, 0, inode_need_killpriv, struct dentry *dentry) -LSM_HOOK(int, 0, inode_killpriv, struct mnt_idmap *idmap, +LSM_HOOK(int, 0, inode_killpriv, const struct mnt_idmap *idmap, struct dentry *dentry) -LSM_HOOK(int, -EOPNOTSUPP, inode_getsecurity, struct mnt_idmap *idmap, +LSM_HOOK(int, -EOPNOTSUPP, inode_getsecurity, const struct mnt_idmap *idmap, struct inode *inode, const char *name, void **buffer, bool alloc) LSM_HOOK(int, -EOPNOTSUPP, inode_setsecurity, struct inode *inode, const char *name, const void *value, size_t size, int flags) diff --git a/include/linux/mnt_idmapping.h b/include/linux/mnt_idmapping.h index e71a6070a8f8..78eeef4c2996 100644 --- a/include/linux/mnt_idmapping.h +++ b/include/linux/mnt_idmapping.h @@ -8,8 +8,8 @@ struct mnt_idmap; struct user_namespace; -extern struct mnt_idmap nop_mnt_idmap; -extern struct mnt_idmap invalid_mnt_idmap; +extern const struct mnt_idmap nop_mnt_idmap; +extern const struct mnt_idmap invalid_mnt_idmap; extern struct user_namespace init_user_ns; typedef struct { @@ -121,19 +121,19 @@ static inline bool vfsgid_eq_kgid(vfsgid_t vfsgid, kgid_t kgid) int vfsgid_in_group_p(vfsgid_t vfsgid); -struct mnt_idmap *mnt_idmap_get(struct mnt_idmap *idmap); -void mnt_idmap_put(struct mnt_idmap *idmap); +const struct mnt_idmap *mnt_idmap_get(const struct mnt_idmap *idmap); +void mnt_idmap_put(const struct mnt_idmap *idmap); -vfsuid_t make_vfsuid(struct mnt_idmap *idmap, +vfsuid_t make_vfsuid(const struct mnt_idmap *idmap, struct user_namespace *fs_userns, kuid_t kuid); -vfsgid_t make_vfsgid(struct mnt_idmap *idmap, +vfsgid_t make_vfsgid(const struct mnt_idmap *idmap, struct user_namespace *fs_userns, kgid_t kgid); -kuid_t from_vfsuid(struct mnt_idmap *idmap, +kuid_t from_vfsuid(const struct mnt_idmap *idmap, struct user_namespace *fs_userns, vfsuid_t vfsuid); -kgid_t from_vfsgid(struct mnt_idmap *idmap, +kgid_t from_vfsgid(const struct mnt_idmap *idmap, struct user_namespace *fs_userns, vfsgid_t vfsgid); /** @@ -148,7 +148,7 @@ kgid_t from_vfsgid(struct mnt_idmap *idmap, * * Return: true if @vfsuid has a mapping in the filesystem, false if not. */ -static inline bool vfsuid_has_fsmapping(struct mnt_idmap *idmap, +static inline bool vfsuid_has_fsmapping(const struct mnt_idmap *idmap, struct user_namespace *fs_userns, vfsuid_t vfsuid) { @@ -186,7 +186,7 @@ static inline kuid_t vfsuid_into_kuid(vfsuid_t vfsuid) * * Return: true if @vfsgid has a mapping in the filesystem, false if not. */ -static inline bool vfsgid_has_fsmapping(struct mnt_idmap *idmap, +static inline bool vfsgid_has_fsmapping(const struct mnt_idmap *idmap, struct user_namespace *fs_userns, vfsgid_t vfsgid) { @@ -225,7 +225,7 @@ static inline kgid_t vfsgid_into_kgid(vfsgid_t vfsgid) * * Return: the caller's current fsuid mapped up according to @idmap. */ -static inline kuid_t mapped_fsuid(struct mnt_idmap *idmap, +static inline kuid_t mapped_fsuid(const struct mnt_idmap *idmap, struct user_namespace *fs_userns) { return from_vfsuid(idmap, fs_userns, VFSUIDT_INIT(current_fsuid())); @@ -244,7 +244,7 @@ static inline kuid_t mapped_fsuid(struct mnt_idmap *idmap, * * Return: the caller's current fsgid mapped up according to @idmap. */ -static inline kgid_t mapped_fsgid(struct mnt_idmap *idmap, +static inline kgid_t mapped_fsgid(const struct mnt_idmap *idmap, struct user_namespace *fs_userns) { return from_vfsgid(idmap, fs_userns, VFSGIDT_INIT(current_fsgid())); diff --git a/include/linux/mount.h b/include/linux/mount.h index acfe7ef86a1b..e90ccafef281 100644 --- a/include/linux/mount.h +++ b/include/linux/mount.h @@ -59,10 +59,10 @@ struct vfsmount { struct dentry *mnt_root; /* root of the mounted tree */ struct super_block *mnt_sb; /* pointer to superblock */ int mnt_flags; - struct mnt_idmap *mnt_idmap; + const struct mnt_idmap *mnt_idmap; } __randomize_layout; -static inline struct mnt_idmap *mnt_idmap(const struct vfsmount *mnt) +static inline const struct mnt_idmap *mnt_idmap(const struct vfsmount *mnt) { /* Pairs with smp_store_release() in do_idmap_mount(). */ return READ_ONCE(mnt->mnt_idmap); diff --git a/include/linux/namei.h b/include/linux/namei.h index 86d657b24fc6..c4436e5c2ba6 100644 --- a/include/linux/namei.h +++ b/include/linux/namei.h @@ -32,8 +32,9 @@ enum { MAX_NESTED_LINKS = 8 }; #define LOOKUP_CREATE BIT(17) /* ... in object creation */ #define LOOKUP_EXCL BIT(18) /* ... in target must not exist */ #define LOOKUP_RENAME_TARGET BIT(19) /* ... in destination of rename() */ +#define LOOKUP_SHARED BIT(20) /* Parent lock is held shared */ -/* 4 spare bits for intent */ +/* 3 spare bits for intent */ /* Scoping flags for lookup. */ #define LOOKUP_NO_SYMLINKS BIT(24) /* No symlink crossing. */ @@ -70,24 +71,24 @@ extern struct dentry *try_lookup_noperm(struct qstr *, struct dentry *); extern struct dentry *lookup_noperm(struct qstr *, struct dentry *); extern struct dentry *lookup_noperm_unlocked(struct qstr *, struct dentry *); extern struct dentry *lookup_noperm_positive_unlocked(struct qstr *, struct dentry *); -struct dentry *lookup_one(struct mnt_idmap *, struct qstr *, struct dentry *); -struct dentry *lookup_one_unlocked(struct mnt_idmap *idmap, +struct dentry *lookup_one(const struct mnt_idmap *, struct qstr *, struct dentry *); +struct dentry *lookup_one_unlocked(const struct mnt_idmap *idmap, struct qstr *name, struct dentry *base); -struct dentry *lookup_one_positive_unlocked(struct mnt_idmap *idmap, +struct dentry *lookup_one_positive_unlocked(const struct mnt_idmap *idmap, struct qstr *name, struct dentry *base); -struct dentry *lookup_one_positive_killable(struct mnt_idmap *idmap, +struct dentry *lookup_one_positive_killable(const struct mnt_idmap *idmap, struct qstr *name, struct dentry *base); -struct dentry *start_creating(struct mnt_idmap *idmap, struct dentry *parent, +struct dentry *start_creating(const struct mnt_idmap *idmap, struct dentry *parent, struct qstr *name); -struct dentry *start_removing(struct mnt_idmap *idmap, struct dentry *parent, +struct dentry *start_removing(const struct mnt_idmap *idmap, struct dentry *parent, struct qstr *name); -struct dentry *start_creating_killable(struct mnt_idmap *idmap, +struct dentry *start_creating_killable(const struct mnt_idmap *idmap, struct dentry *parent, struct qstr *name); -struct dentry *start_removing_killable(struct mnt_idmap *idmap, +struct dentry *start_removing_killable(const struct mnt_idmap *idmap, struct dentry *parent, struct qstr *name); struct dentry *start_creating_noperm(struct dentry *parent, struct qstr *name); diff --git a/include/linux/netfs.h b/include/linux/netfs.h index b4dd32863dd4..67e010b6994b 100644 --- a/include/linux/netfs.h +++ b/include/linux/netfs.h @@ -22,6 +22,7 @@ enum netfs_sreq_ref_trace; typedef struct mempool mempool_t; +struct fscache_occupancy; struct folio_queue; /** @@ -62,8 +63,8 @@ struct netfs_inode { struct fscache_cookie *cache; #endif struct list_head wb_queue; /* Queue of processes wanting to do writeback */ - loff_t _remote_i_size; /* Size of the remote file */ - loff_t _zero_point; /* Size after which we assume there's no data + uoff_t _remote_i_size; /* Size of the remote file */ + uoff_t _zero_point; /* Size after which we assume there's no data * on the server */ spinlock_t lock; /* Lock covering wb_queue */ atomic_t io_count; /* Number of outstanding reqs */ @@ -125,6 +126,12 @@ static inline struct netfs_group *netfs_folio_group(struct folio *folio) return priv; } +enum netfs_cache_collect { + NETFS_CACHE_COLLECT_WRITE_GAP, /* Gap in collection, no state either way */ + NETFS_CACHE_COLLECT_WRITE_DATA, /* Currently collecting good writes */ + NETFS_CACHE_COLLECT_WRITE_CANCEL, /* Currently collecting cancelled writes */ +}; + /* * Stream of I/O subrequests going to a particular destination, such as the * server or the local cache. This is mainly intended for writing where we may @@ -142,7 +149,7 @@ struct netfs_io_stream { void (*issue_write)(struct netfs_io_subrequest *subreq); /* Collection tracking */ struct list_head subrequests; /* Contributory I/O operations */ - unsigned long long collected_to; /* Position we've collected results to */ + uoff_t collected_to; /* Position we've collected results to */ size_t transferred; /* The amount transferred from this stream */ unsigned short error; /* Aggregate error for the stream */ enum netfs_io_source source; /* Where to read from/write to */ @@ -152,6 +159,7 @@ struct netfs_io_stream { bool need_retry; /* T if this stream needs retrying */ bool failed; /* T if this stream failed */ bool transferred_valid; /* T is ->transferred is valid */ + enum netfs_cache_collect cache_collect; /* Current writeback cache collect state */ }; /* @@ -161,8 +169,11 @@ struct netfs_cache_resources { const struct netfs_cache_ops *ops; void *cache_priv; void *cache_priv2; - unsigned int debug_id; /* Cookie debug ID */ + uoff_t cache_i_size; /* Initial size of cache file */ + unsigned int cookie_id; /* Cache cookie debug ID */ + unsigned int object_id; /* Cache object debug ID */ unsigned int inval_counter; /* object->inval_counter at begin_op */ + unsigned int dio_size; /* DIO block size */ }; /* @@ -177,7 +188,7 @@ struct netfs_io_subrequest { struct work_struct work; struct list_head rreq_link; /* Link in rreq->subrequests */ struct iov_iter io_iter; /* Iterator for this subrequest */ - unsigned long long start; /* Where to start the I/O */ + uoff_t start; /* Where to start the I/O */ size_t len; /* Size of the I/O */ size_t transferred; /* Amount of data transferred */ refcount_t ref; @@ -196,6 +207,7 @@ struct netfs_io_subrequest { #define NETFS_SREQ_IN_PROGRESS 8 /* Unlocked when the subrequest completes */ #define NETFS_SREQ_NEED_RETRY 9 /* Set if the filesystem requests a retry */ #define NETFS_SREQ_FAILED 10 /* Set if the subreq failed unretryably */ +#define NETFS_SREQ_CANCELLED 11 /* Set if the subreq was cancelled by netfslib */ }; enum netfs_io_origin { @@ -208,7 +220,6 @@ enum netfs_io_origin { NETFS_DIO_READ, /* This is a direct I/O read */ NETFS_WRITEBACK, /* This write was triggered by writepages */ NETFS_WRITEBACK_SINGLE, /* This monolithic write was triggered by writepages */ - NETFS_WRITETHROUGH, /* This write was made by netfs_perform_write() */ NETFS_UNBUFFERED_WRITE, /* This is an unbuffered write */ NETFS_DIO_WRITE, /* This is a direct I/O write */ NETFS_PGPRIV2_COPY_TO_CACHE, /* [DEPRECATED] This is writing read data to the cache */ @@ -243,17 +254,18 @@ struct netfs_io_request { void *netfs_priv; /* Private data for the netfs */ void *netfs_priv2; /* Private data for the netfs */ struct bio_vec *direct_bv; /* DIO buffer list (when handling iovec-iter) */ - unsigned long long submitted; /* Amount submitted for I/O so far */ - unsigned long long len; /* Length of the request */ + uoff_t submitted; /* Amount submitted for I/O so far */ + uoff_t len; /* Length of the request */ size_t transferred; /* Amount to be indicated as transferred */ size_t progress_at; /* Report read progress when hit this much read */ long error; /* 0 or error that occurred */ - unsigned long long i_size; /* Size of the file */ - unsigned long long start; /* Start position */ + uoff_t i_size; /* Size of the file */ + uoff_t start; /* Start position */ atomic64_t issued_to; /* Write issuer folio cursor */ - unsigned long long collected_to; /* Point we've collected to */ - unsigned long long cleaned_to; /* Position we've cleaned folios to */ - unsigned long long abandon_to; /* Position to abandon folios to */ + uoff_t collected_to; /* Point we've collected to */ + uoff_t cache_coll_to; /* Point the cache has collected to */ + uoff_t cleaned_to; /* Position we've cleaned folios to */ + uoff_t abandon_to; /* Position to abandon folios to */ const struct folio *no_unlock_folio; /* Don't unlock this folio after read */ gfp_t gfp; /* GFP flags to use */ unsigned int direct_bv_count; /* Number of elements in direct_bv[] */ @@ -273,14 +285,18 @@ struct netfs_io_request { #define NETFS_RREQ_FAILED 3 /* The request failed */ #define NETFS_RREQ_RETRYING 4 /* Set if we're in the retry path */ #define NETFS_RREQ_SHORT_TRANSFER 5 /* Set if we have a short transfer */ -#define NETFS_RREQ_OFFLOAD_COLLECTION 8 /* Offload collection to workqueue */ -#define NETFS_RREQ_NO_UNLOCK_FOLIO 9 /* Don't unlock no_unlock_folio on completion */ +#define NETFS_RREQ_CACHE_STOP 8 /* Set to stop caching (ENOBUFS or error) */ +#define NETFS_RREQ_CACHE_ERROR 9 /* Set if we got an error from the cache */ #define NETFS_RREQ_CANCEL_CACHING 10 /* Set to cancel caching */ -#define NETFS_RREQ_UPLOAD_TO_SERVER 11 /* Need to write to the server */ -#define NETFS_RREQ_USE_IO_ITER 12 /* Use ->io_iter rather than ->i_pages */ +#define NETFS_RREQ_OFFLOAD_COLLECTION 12 /* Offload collection to workqueue */ +#define NETFS_RREQ_NO_UNLOCK_FOLIO 13 /* Don't unlock no_unlock_folio on completion */ +#define NETFS_RREQ_UPLOAD_TO_SERVER 14 /* Need to write to the server */ +#define NETFS_RREQ_USE_IO_ITER 15 /* Use ->io_iter rather than ->i_pages */ #define NETFS_RREQ_NEED_PUT_RA_REFS 17 /* Need to put the folio refs RA gave us */ +#ifdef CONFIG_NETFS_PGPRIV2 #define NETFS_RREQ_USE_PGPRIV2 31 /* [DEPRECATED] Use PG_private_2 to mark * write to cache on read */ +#endif const struct netfs_request_ops *netfs_ops; }; @@ -299,12 +315,12 @@ struct netfs_request_ops { int (*prepare_read)(struct netfs_io_subrequest *subreq); void (*issue_read)(struct netfs_io_subrequest *subreq); bool (*is_still_valid)(struct netfs_io_request *rreq); - int (*check_write_begin)(struct file *file, loff_t pos, unsigned len, + int (*check_write_begin)(struct file *file, uoff_t pos, unsigned len, struct folio **foliop, void **_fsdata); void (*done)(struct netfs_io_request *rreq); /* Modification handling */ - void (*update_i_size)(struct inode *inode, loff_t i_size); + void (*update_i_size)(struct inode *inode, uoff_t i_size); void (*post_modify)(struct inode *inode); /* Write request handling */ @@ -332,7 +348,7 @@ struct netfs_cache_ops { /* Read data from the cache */ int (*read)(struct netfs_cache_resources *cres, - loff_t start_pos, + uoff_t start_pos, struct iov_iter *iter, enum netfs_read_from_hole read_hole, netfs_io_terminated_t term_func, @@ -340,7 +356,7 @@ struct netfs_cache_ops { /* Write data to the cache */ int (*write)(struct netfs_cache_resources *cres, - loff_t start_pos, + uoff_t start_pos, struct iov_iter *iter, netfs_io_terminated_t term_func, void *term_func_priv); @@ -350,15 +366,14 @@ struct netfs_cache_ops { /* Expand readahead request */ void (*expand_readahead)(struct netfs_cache_resources *cres, - unsigned long long *_start, - unsigned long long *_len, - unsigned long long i_size); + uoff_t *_start, + uoff_t *_len, + uoff_t i_size); /* Prepare a read operation, shortening it to a cached/uncached * boundary as appropriate. */ - enum netfs_io_source (*prepare_read)(struct netfs_io_subrequest *subreq, - unsigned long long i_size); + int (*prepare_read)(struct netfs_io_subrequest *subreq); /* Prepare a write subrequest, working out if we're allowed to do it * and finding out the maximum amount of data to gather before @@ -371,15 +386,24 @@ struct netfs_cache_ops { * actually do. */ int (*prepare_write)(struct netfs_cache_resources *cres, - loff_t *_start, size_t *_len, size_t upper_len, - loff_t i_size, bool no_space_allocated_yet); + uoff_t *_start, size_t *_len, size_t upper_len, + uoff_t i_size, bool no_space_allocated_yet); /* Query the occupancy of the cache in a region, returning where the * next chunk of data starts and how long it is. */ int (*query_occupancy)(struct netfs_cache_resources *cres, - loff_t start, size_t len, size_t granularity, - loff_t *_data_start, size_t *_data_len); + struct fscache_occupancy *occ); + + /* Collect the result of buffered writeback to the cache. This + * includes copying a read to the cache. block_type is one of: + * - NETFS_CACHE_COLLECT_WRITE_DATA for a block of data + * - NETFS_CACHE_COLLECT_WRITE_GAP if a discontiguity was skipped + * - NETFS_CACHE_COLLECT_WRITE_CANCEL for a cancellation gap + */ + void (*collect_write)(struct netfs_io_request *wreq, + uoff_t start, size_t len, + enum netfs_cache_collect block_type); }; /* High-level read API. */ @@ -410,7 +434,7 @@ struct readahead_control; void netfs_readahead(struct readahead_control *); int netfs_read_folio(struct file *, struct folio *); int netfs_write_begin(struct netfs_inode *, struct file *, - struct address_space *, loff_t pos, unsigned int len, + struct address_space *, uoff_t pos, unsigned int len, struct folio **, void **fsdata); int netfs_writepages(struct address_space *mapping, struct writeback_control *wbc); @@ -488,10 +512,10 @@ static inline struct netfs_inode *netfs_inode(struct inode *inode) * cmpxchg8b without the need of the lock prefix). For SMP compiles and 64bit * archs it makes no difference if preempt is enabled or not. */ -static inline unsigned long long netfs_read_remote_i_size(const struct inode *inode) +static inline uoff_t netfs_read_remote_i_size(const struct inode *inode) { const struct netfs_inode *ictx = container_of(inode, struct netfs_inode, inode); - unsigned long long remote_i_size; + uoff_t remote_i_size; #if BITS_PER_LONG==32 && defined(CONFIG_SMP) unsigned int seq; @@ -526,7 +550,7 @@ static inline unsigned long long netfs_read_remote_i_size(const struct inode *in * spinning forever. */ static inline void netfs_write_remote_i_size(struct inode *inode, - unsigned long long remote_i_size) + uoff_t remote_i_size) { struct netfs_inode *ictx = netfs_inode(inode); @@ -563,10 +587,10 @@ static inline void netfs_write_remote_i_size(struct inode *inode, * cmpxchg8b without the need of the lock prefix). For SMP compiles and 64bit * archs it makes no difference if preempt is enabled or not. */ -static inline unsigned long long netfs_read_zero_point(const struct inode *inode) +static inline uoff_t netfs_read_zero_point(const struct inode *inode) { struct netfs_inode *ictx = container_of(inode, struct netfs_inode, inode); - unsigned long long zero_point; + uoff_t zero_point; #if BITS_PER_LONG==32 && defined(CONFIG_SMP) unsigned int seq; @@ -601,7 +625,7 @@ static inline unsigned long long netfs_read_zero_point(const struct inode *inode * forever. */ static inline void netfs_write_zero_point(struct inode *inode, - unsigned long long zero_point) + uoff_t zero_point) { struct netfs_inode *ictx = netfs_inode(inode); @@ -642,9 +666,9 @@ static inline void netfs_write_zero_point(struct inode *inode, * archs it makes no difference if preempt is enabled or not. */ static inline void netfs_read_sizes(const struct inode *inode, - unsigned long long *i_size, - unsigned long long *remote_i_size, - unsigned long long *zero_point) + uoff_t *i_size, + uoff_t *remote_i_size, + uoff_t *zero_point) { const struct netfs_inode *ictx = container_of(inode, struct netfs_inode, inode); #if BITS_PER_LONG==32 && defined(CONFIG_SMP) @@ -690,9 +714,9 @@ static inline void netfs_read_sizes(const struct inode *inode, * forever. */ static inline void netfs_write_sizes(struct inode *inode, - unsigned long long i_size, - unsigned long long remote_i_size, - unsigned long long zero_point) + uoff_t i_size, + uoff_t remote_i_size, + uoff_t zero_point) { struct netfs_inode *ictx = netfs_inode(inode); @@ -760,7 +784,7 @@ static inline void netfs_inode_init(struct netfs_inode *ctx, * Inform the netfs lib that a file got resized so that it can adjust its state. */ static inline void netfs_resize_file(struct netfs_inode *ictx, - unsigned long long new_i_size, + uoff_t new_i_size, bool changed_on_server) { #if BITS_PER_LONG==32 && defined(CONFIG_SMP) diff --git a/include/linux/nfs.h b/include/linux/nfs.h index 0906a0b40c6a..8c2818db43c5 100644 --- a/include/linux/nfs.h +++ b/include/linux/nfs.h @@ -11,59 +11,8 @@ #include <linux/cred.h> #include <linux/sunrpc/auth.h> #include <linux/sunrpc/msg_prot.h> -#include <linux/string.h> -#include <linux/crc32.h> -#include <uapi/linux/nfs.h> - -/* The LOCALIO program is entirely private to Linux and is - * NOT part of the uapi. - */ -#define NFS_LOCALIO_PROGRAM 400122 -#define LOCALIOPROC_NULL 0 -#define LOCALIOPROC_UUID_IS_LOCAL 1 - -/* - * This is the kernel NFS client file handle representation - */ -#define NFS_MAXFHSIZE 128 -struct nfs_fh { - unsigned short size; - unsigned char data[NFS_MAXFHSIZE]; -}; - -/* - * Returns a zero iff the size and data fields match. - * Checks only "size" bytes in the data field. - */ -static inline int nfs_compare_fh(const struct nfs_fh *a, const struct nfs_fh *b) -{ - return a->size != b->size || memcmp(a->data, b->data, a->size) != 0; -} - -static inline void nfs_copy_fh(struct nfs_fh *target, const struct nfs_fh *source) -{ - target->size = source->size; - memcpy(target->data, source->data, source->size); -} - -enum nfs3_stable_how { - NFS_UNSTABLE = 0, - NFS_DATA_SYNC = 1, - NFS_FILE_SYNC = 2, +#include <linux/nfs_fh.h> - /* used by direct.c to mark verf as invalid */ - NFS_INVALID_STABLE_HOW = -1 -}; +#include <uapi/linux/nfs.h> -/** - * nfs_fhandle_hash - calculate the crc32 hash for the filehandle - * @fh - pointer to filehandle - * - * returns a crc32 hash for the filehandle that is compatible with - * the one displayed by "wireshark". - */ -static inline u32 nfs_fhandle_hash(const struct nfs_fh *fh) -{ - return ~crc32_le(0xFFFFFFFF, &fh->data[0], fh->size); -} #endif /* _LINUX_NFS_H */ diff --git a/include/linux/nfs3.h b/include/linux/nfs3.h index 404b8f724fc9..b6539a75edea 100644 --- a/include/linux/nfs3.h +++ b/include/linux/nfs3.h @@ -7,6 +7,49 @@ #include <uapi/linux/nfs3.h> +/* + * NFSv3 error status values. + * See RFC 1813 Section 2.5 + */ +enum { + NFS3ERR_PERM = 1, + NFS3ERR_NOENT = 2, + NFS3ERR_IO = 5, + NFS3ERR_NXIO = 6, + NFS3ERR_ACCES = 13, + NFS3ERR_EXIST = 17, + NFS3ERR_XDEV = 18, + NFS3ERR_NODEV = 19, + NFS3ERR_NOTDIR = 20, + NFS3ERR_ISDIR = 21, + NFS3ERR_INVAL = 22, + NFS3ERR_FBIG = 27, + NFS3ERR_NOSPC = 28, + NFS3ERR_ROFS = 30, + NFS3ERR_MLINK = 31, + NFS3ERR_NAMETOOLONG = 63, + NFS3ERR_NOTEMPTY = 66, + NFS3ERR_DQUOT = 69, + NFS3ERR_STALE = 70, + NFS3ERR_REMOTE = 71, + NFS3ERR_BADHANDLE = 10001, + NFS3ERR_NOT_SYNC = 10002, + NFS3ERR_BAD_COOKIE = 10003, + NFS3ERR_NOTSUPP = 10004, + NFS3ERR_TOOSMALL = 10005, + NFS3ERR_SERVERFAULT = 10006, + NFS3ERR_BADTYPE = 10007, + NFS3ERR_JUKEBOX = 10008, +}; + +enum nfs3_stable_how { + NFS_UNSTABLE = 0, + NFS_DATA_SYNC = 1, + NFS_FILE_SYNC = 2, + + /* used to mark verf as invalid */ + NFS_INVALID_STABLE_HOW = -1 +}; /* Number of 32bit words in post_op_attr */ #define NFS3_POST_OP_ATTR_WORDS 22 diff --git a/include/linux/nfs4.h b/include/linux/nfs4.h index 1a3981c26b23..41b7cdcc674f 100644 --- a/include/linux/nfs4.h +++ b/include/linux/nfs4.h @@ -263,6 +263,12 @@ enum why_no_delegation4 { /* new to v4.1 */ WND4_IS_DIR = 8, }; +enum stable_how4 { + UNSTABLE4 = 0, + DATA_SYNC4 = 1, + FILE_SYNC4 = 2, +}; + enum lock_type4 { NFS4_UNLOCK_LT = 0, NFS4_READ_LT = 1, diff --git a/include/linux/nfs_fh.h b/include/linux/nfs_fh.h new file mode 100644 index 000000000000..49dfc5ec60fe --- /dev/null +++ b/include/linux/nfs_fh.h @@ -0,0 +1,63 @@ +/* SPDX-License-Identifier: GPL-2.0 */ +/* + * struct nfs_fh is an NFS version-agnostic data structure that + * stores an NFS file handle. It is also commonly used in NFS + * related APIs. + */ +#ifndef _LINUX_NFS_FH_H +#define _LINUX_NFS_FH_H + +#include <linux/types.h> +#include <linux/string.h> +#include <linux/crc32.h> + +/* + * The largest file handle size today is an NFSv4 file handle, + * which can be up to 128 octets long. + */ +#define NFS_MAXFHSIZE 128 +struct nfs_fh { + unsigned short size; + unsigned char data[NFS_MAXFHSIZE]; +}; + +/** + * nfs_compare_fh - Compare two NFS file handles + * @a: An NFS file handle to be compared + * @b: An NFS file handle to be compared + * + * Checks only "size" bytes in each data field. + * + * Return: %false if the two file handles are equal, otherwise %true + */ +static inline bool nfs_compare_fh(const struct nfs_fh *a, const struct nfs_fh *b) +{ + return a->size != b->size || memcmp(a->data, b->data, a->size) != 0; +} + +/** + * nfs_copy_fh - Copy an NFS file handle + * @target: Destination file handle + * @source: Source file handle + * + * Copies source->size bytes of file handle data into target. + */ +static inline void nfs_copy_fh(struct nfs_fh *target, const struct nfs_fh *source) +{ + target->size = source->size; + memcpy(target->data, source->data, source->size); +} + +/** + * nfs_fhandle_hash - Calculate the crc32 hash for the filehandle + * @fh: An NFS file handle to hash + * + * Return: a crc32 hash for the filehandle that is compatible with + * the one displayed by "wireshark" + */ +static inline u32 nfs_fhandle_hash(const struct nfs_fh *fh) +{ + return ~crc32_le(0xFFFFFFFF, &fh->data[0], fh->size); +} + +#endif /* _LINUX_NFS_FH_H */ diff --git a/include/linux/nfs_fs.h b/include/linux/nfs_fs.h index b85a73ae7919..d2c716322c6f 100644 --- a/include/linux/nfs_fs.h +++ b/include/linux/nfs_fs.h @@ -437,11 +437,11 @@ extern int nfs_refresh_inode(struct inode *, struct nfs_fattr *); extern int nfs_post_op_update_inode(struct inode *inode, struct nfs_fattr *fattr); extern int nfs_post_op_update_inode_force_wcc(struct inode *inode, struct nfs_fattr *fattr); extern int nfs_post_op_update_inode_force_wcc_locked(struct inode *inode, struct nfs_fattr *fattr); -extern int nfs_getattr(struct mnt_idmap *, const struct path *, +extern int nfs_getattr(const struct mnt_idmap *, const struct path *, struct kstat *, u32, unsigned int); extern void nfs_access_add_cache(struct inode *, struct nfs_access_entry *, const struct cred *); extern void nfs_access_set_mask(struct nfs_access_entry *, u32); -extern int nfs_permission(struct mnt_idmap *, struct inode *, int); +extern int nfs_permission(const struct mnt_idmap *, struct inode *, int); extern int nfs_open(struct inode *, struct file *); extern int nfs_attribute_cache_expired(struct inode *inode); extern int nfs_revalidate_inode(struct inode *inode, unsigned long flags); @@ -450,7 +450,7 @@ extern int nfs_clear_invalid_mapping(struct address_space *mapping); extern bool nfs_mapping_need_revalidate_inode(struct inode *inode); extern int nfs_revalidate_mapping(struct inode *inode, struct address_space *mapping); extern int nfs_revalidate_mapping_rcu(struct inode *inode); -extern int nfs_setattr(struct mnt_idmap *, struct dentry *, struct iattr *); +extern int nfs_setattr(const struct mnt_idmap *, struct dentry *, struct iattr *); extern void nfs_setattr_update_inode(struct inode *inode, struct iattr *attr, struct nfs_fattr *); extern void nfs_setsecurity(struct inode *inode, struct nfs_fattr *fattr); extern struct nfs_open_context *get_nfs_open_context(struct nfs_open_context *ctx); diff --git a/include/linux/nfs_fs_sb.h b/include/linux/nfs_fs_sb.h index 34d294774f8c..416c6f39f31d 100644 --- a/include/linux/nfs_fs_sb.h +++ b/include/linux/nfs_fs_sb.h @@ -74,6 +74,8 @@ struct nfs_client { u64 cl_clientid; /* constant */ nfs4_verifier cl_confirm; /* Clientid verifier */ unsigned long cl_state; + /* bumped on each CB_NOTIFY_DEVICEID CHANGE for this client */ + atomic_t cl_deviceid_change_epoch; spinlock_t cl_lock; @@ -101,6 +103,8 @@ struct nfs_client { /* The flags used for obtaining the clientid during EXCHANGE_ID */ u32 cl_exchange_flags; struct nfs4_session *cl_session; /* shared session */ + /* CB_NOTIFY_DEVICEID DELETE suspects, protected by cl_lock */ + struct list_head cl_deviceid_deletes; bool cl_preserve_clid; struct nfs41_server_owner *cl_serverowner; struct nfs41_server_scope *cl_serverscope; @@ -248,6 +252,10 @@ struct nfs_server { that are supported on this filesystem */ struct pnfs_layoutdriver_type *pnfs_curr_ld; /* Active layout driver */ + unsigned int lg_reply_sz; /* Learned LAYOUTGET reply + buffer size, when the layout + driver's default has proved + too small */ struct rpc_wait_queue roc_rpcwaitq; /* the following fields are protected by nfs_client->cl_lock */ diff --git a/include/linux/nfs_page.h b/include/linux/nfs_page.h index 4b9a35dbc062..c38e4b380be5 100644 --- a/include/linux/nfs_page.h +++ b/include/linux/nfs_page.h @@ -38,6 +38,7 @@ enum { PG_REMOVE, /* page group sync bit in write path */ PG_CONTENDED1, /* Is someone waiting for a lock? */ PG_CONTENDED2, /* Is someone waiting for a lock? */ + PG_PINNED, /* page is pinned by GUP */ }; struct nfs_inode; @@ -58,6 +59,7 @@ struct nfs_page { struct nfs_page *wb_this_page; /* list of reqs for this page */ struct nfs_page *wb_head; /* head pointer for req list */ unsigned short wb_nio; /* Number of I/O attempts */ + unsigned int wb_nr_pinned; /* Number of pinned pages */ }; struct nfs_pgio_mirror; @@ -125,15 +127,17 @@ struct nfs_pageio_descriptor { extern struct nfs_page *nfs_page_create_from_page(struct nfs_open_context *ctx, struct page *page, + bool pinned, unsigned int pgbase, loff_t offset, unsigned int count); extern struct nfs_page *nfs_page_create_from_folio(struct nfs_open_context *ctx, struct folio *folio, + bool pinned, unsigned int offset, unsigned int count); -extern void nfs_release_request(struct nfs_page *); - +void nfs_release_request(struct nfs_page *req); +void nfs_release_request_list(struct list_head *head); extern void nfs_pageio_init(struct nfs_pageio_descriptor *desc, struct inode *inode, diff --git a/include/linux/nfs_ssc.h b/include/linux/nfs_ssc.h index 22265b1ff080..c199ea23e7eb 100644 --- a/include/linux/nfs_ssc.h +++ b/include/linux/nfs_ssc.h @@ -2,80 +2,33 @@ /* * include/linux/nfs_ssc.h * + * NFSv4.2 server-to-server copy, NFS client side APIs + * * Author: Dai Ngo <dai.ngo@oracle.com> * * Copyright (c) 2020, Oracle and/or its affiliates. */ -#include <linux/nfs_fs.h> -#include <linux/sunrpc/svc.h> +#ifndef _LINUX_NFS_SSC_H +#define _LINUX_NFS_SSC_H -extern struct nfs_ssc_client_ops_tbl nfs_ssc_client_tbl; +#include <linux/nfs_fh.h> +#include <linux/nfs4.h> + +struct file; +struct vfsmount; -/* - * NFS_V4 - */ struct nfs4_ssc_client_ops { + struct module *owner; struct file *(*sco_open)(struct vfsmount *ss_mnt, struct nfs_fh *src_fh, nfs4_stateid *stateid); void (*sco_close)(struct file *filep); }; -/* - * NFS_FS - */ -struct nfs_ssc_client_ops { - void (*sco_sb_deactive)(struct super_block *sb); -}; - -struct nfs_ssc_client_ops_tbl { - const struct nfs4_ssc_client_ops *ssc_nfs4_ops; - const struct nfs_ssc_client_ops *ssc_nfs_ops; -}; - extern void nfs42_ssc_register_ops(void); extern void nfs42_ssc_unregister_ops(void); extern void nfs42_ssc_register(const struct nfs4_ssc_client_ops *ops); extern void nfs42_ssc_unregister(const struct nfs4_ssc_client_ops *ops); -#ifdef CONFIG_NFSD_V4_2_INTER_SSC -static inline struct file *nfs42_ssc_open(struct vfsmount *ss_mnt, - struct nfs_fh *src_fh, nfs4_stateid *stateid) -{ - if (nfs_ssc_client_tbl.ssc_nfs4_ops) - return (*nfs_ssc_client_tbl.ssc_nfs4_ops->sco_open)(ss_mnt, src_fh, stateid); - return ERR_PTR(-EIO); -} - -static inline void nfs42_ssc_close(struct file *filep) -{ - if (nfs_ssc_client_tbl.ssc_nfs4_ops) - (*nfs_ssc_client_tbl.ssc_nfs4_ops->sco_close)(filep); -} -#endif - -struct nfsd4_ssc_umount_item { - struct list_head nsui_list; - bool nsui_busy; - /* - * nsui_refcnt inited to 2, 1 on list and 1 for consumer. Entry - * is removed when refcnt drops to 1 and nsui_expire expires. - */ - refcount_t nsui_refcnt; - unsigned long nsui_expire; - struct vfsmount *nsui_vfsmount; - char nsui_ipaddr[RPC_MAX_ADDRBUFLEN + 1]; -}; - -/* - * NFS_FS - */ -extern void nfs_ssc_register(const struct nfs_ssc_client_ops *ops); -extern void nfs_ssc_unregister(const struct nfs_ssc_client_ops *ops); - -static inline void nfs_do_sb_deactive(struct super_block *sb) -{ - if (nfs_ssc_client_tbl.ssc_nfs_ops) - (*nfs_ssc_client_tbl.ssc_nfs_ops->sco_sb_deactive)(sb); -} +#endif /* _LINUX_NFS_SSC_H */ diff --git a/include/linux/nfs_xdr.h b/include/linux/nfs_xdr.h index 7ed8fdb930d6..c0e29b4dfa62 100644 --- a/include/linux/nfs_xdr.h +++ b/include/linux/nfs_xdr.h @@ -1693,6 +1693,7 @@ struct nfs_pgio_header { struct nfs_client *ds_clp; /* pNFS data server */ u32 ds_commit_idx; /* ds index if ds_clp is set */ u32 pgio_mirror_idx;/* mirror index in pgio layer */ + struct nfs4_deviceid_node *ds_dev; /* device node ref held across the I/O */ }; struct nfs_mds_commit_info { @@ -1731,6 +1732,7 @@ struct nfs_commit_data { struct nfs_open_context *context; struct pnfs_layout_segment *lseg; struct nfs_client *ds_clp; /* pNFS data server */ + struct nfs4_deviceid_node *ds_dev; /* device node ref held across the commit */ int ds_commit_index; loff_t lwb; const struct rpc_call_ops *mds_ops; diff --git a/include/linux/nfsd_ssc.h b/include/linux/nfsd_ssc.h new file mode 100644 index 000000000000..7001410f01c2 --- /dev/null +++ b/include/linux/nfsd_ssc.h @@ -0,0 +1,38 @@ +/* SPDX-License-Identifier: GPL-2.0 */ +/* + * include/linux/nfsd_ssc.h + * + * NFSv4.2 server-to-server copy, NFS server side APIs + * + * Author: Dai Ngo <dai.ngo@oracle.com> + * + * Copyright (c) 2020, Oracle and/or its affiliates. + */ + +#ifndef _LINUX_NFSD_SSC_H +#define _LINUX_NFSD_SSC_H + +#include <linux/nfs_fh.h> +#include <linux/nfs4.h> + +struct file; +struct vfsmount; + +#if IS_ENABLED(CONFIG_NFS_V4_2_SSC_HELPER) +struct file *nfsd42_ssc_open(struct vfsmount *ss_mnt, struct nfs_fh *src_fh, + nfs4_stateid *stateid); +void nfsd42_ssc_close(struct file *filp); +#else +static inline struct file *nfsd42_ssc_open(struct vfsmount *ss_mnt, + struct nfs_fh *src_fh, + nfs4_stateid *stateid) +{ + return ERR_PTR(-EIO); +} + +static inline void nfsd42_ssc_close(struct file *filp) +{ +} +#endif + +#endif /* _LINUX_NFSD_SSC_H */ diff --git a/include/linux/nfslocalio.h b/include/linux/nfslocalio.h index 3d91043254e6..8ce4d978a636 100644 --- a/include/linux/nfslocalio.h +++ b/include/linux/nfslocalio.h @@ -13,9 +13,18 @@ #include <linux/uuid.h> #include <linux/sunrpc/clnt.h> #include <linux/sunrpc/svcauth.h> -#include <linux/nfs.h> +#include <linux/nfs_fh.h> + #include <net/net_namespace.h> +/* + * The LOCALIO program is entirely private to Linux and is NOT part of + * the uapi. + */ +#define NFS_LOCALIO_PROGRAM 400122 +#define LOCALIOPROC_NULL 0 +#define LOCALIOPROC_UUID_IS_LOCAL 1 + struct nfs_client; struct nfs_file_localio; diff --git a/include/linux/posix_acl.h b/include/linux/posix_acl.h index 62d497763e25..caf500bed993 100644 --- a/include/linux/posix_acl.h +++ b/include/linux/posix_acl.h @@ -74,20 +74,20 @@ extern int __posix_acl_create(struct posix_acl **, gfp_t, umode_t *); extern int __posix_acl_chmod(struct posix_acl **, gfp_t, umode_t); extern struct posix_acl *get_posix_acl(struct inode *, int); -int set_posix_acl(struct mnt_idmap *, struct dentry *, int, +int set_posix_acl(const struct mnt_idmap *, struct dentry *, int, struct posix_acl *); struct posix_acl *get_cached_acl_rcu(struct inode *inode, int type); struct posix_acl *posix_acl_clone(const struct posix_acl *acl, gfp_t flags); #ifdef CONFIG_FS_POSIX_ACL -int posix_acl_chmod(struct mnt_idmap *, struct dentry *, umode_t); +int posix_acl_chmod(const struct mnt_idmap *, struct dentry *, umode_t); extern int posix_acl_create(struct inode *, umode_t *, struct posix_acl **, struct posix_acl **); -int posix_acl_update_mode(struct mnt_idmap *, struct inode *, umode_t *, +int posix_acl_update_mode(const struct mnt_idmap *, struct inode *, umode_t *, struct posix_acl **); -int simple_set_acl(struct mnt_idmap *, struct dentry *, +int simple_set_acl(const struct mnt_idmap *, struct dentry *, struct posix_acl *, int); extern int simple_acl_create(struct inode *, struct inode *); @@ -96,7 +96,7 @@ void set_cached_acl(struct inode *inode, int type, struct posix_acl *acl); void forget_cached_acl(struct inode *inode, int type); void forget_all_cached_acls(struct inode *inode); int posix_acl_valid(struct user_namespace *, const struct posix_acl *); -int posix_acl_permission(struct mnt_idmap *, struct inode *, +int posix_acl_permission(const struct mnt_idmap *, struct inode *, const struct posix_acl *, int); static inline void cache_no_acl(struct inode *inode) @@ -105,16 +105,16 @@ static inline void cache_no_acl(struct inode *inode) inode->i_default_acl = NULL; } -int vfs_set_acl(struct mnt_idmap *idmap, struct dentry *dentry, +int vfs_set_acl(const struct mnt_idmap *idmap, struct dentry *dentry, const char *acl_name, struct posix_acl *kacl); -struct posix_acl *vfs_get_acl(struct mnt_idmap *idmap, +struct posix_acl *vfs_get_acl(const struct mnt_idmap *idmap, struct dentry *dentry, const char *acl_name); -int vfs_remove_acl(struct mnt_idmap *idmap, struct dentry *dentry, +int vfs_remove_acl(const struct mnt_idmap *idmap, struct dentry *dentry, const char *acl_name); int posix_acl_listxattr(struct inode *inode, char **buffer, ssize_t *remaining_size); #else -static inline int posix_acl_chmod(struct mnt_idmap *idmap, +static inline int posix_acl_chmod(const struct mnt_idmap *idmap, struct dentry *dentry, umode_t mode) { return 0; @@ -141,21 +141,21 @@ static inline void forget_all_cached_acls(struct inode *inode) { } -static inline int vfs_set_acl(struct mnt_idmap *idmap, +static inline int vfs_set_acl(const struct mnt_idmap *idmap, struct dentry *dentry, const char *name, struct posix_acl *acl) { return -EOPNOTSUPP; } -static inline struct posix_acl *vfs_get_acl(struct mnt_idmap *idmap, +static inline struct posix_acl *vfs_get_acl(const struct mnt_idmap *idmap, struct dentry *dentry, const char *acl_name) { return ERR_PTR(-EOPNOTSUPP); } -static inline int vfs_remove_acl(struct mnt_idmap *idmap, +static inline int vfs_remove_acl(const struct mnt_idmap *idmap, struct dentry *dentry, const char *acl_name) { return -EOPNOTSUPP; diff --git a/include/linux/quotaops.h b/include/linux/quotaops.h index f9c0f9d7c9d9..0c64ca674e77 100644 --- a/include/linux/quotaops.h +++ b/include/linux/quotaops.h @@ -20,7 +20,7 @@ static inline struct quota_info *sb_dqopt(struct super_block *sb) } /* i_rwsem must being held */ -static inline bool is_quota_modification(struct mnt_idmap *idmap, +static inline bool is_quota_modification(const struct mnt_idmap *idmap, struct inode *inode, struct iattr *ia) { return ((ia->ia_valid & ATTR_SIZE) || @@ -109,7 +109,7 @@ int dquot_set_dqblk(struct super_block *sb, struct kqid id, struct qc_dqblk *di); int __dquot_transfer(struct inode *inode, struct dquot **transfer_to); -int dquot_transfer(struct mnt_idmap *idmap, struct inode *inode, +int dquot_transfer(const struct mnt_idmap *idmap, struct inode *inode, struct iattr *iattr); static inline struct mem_dqinfo *sb_dqinfo(struct super_block *sb, int type) @@ -229,7 +229,7 @@ static inline void dquot_free_inode(struct inode *inode) { } -static inline int dquot_transfer(struct mnt_idmap *idmap, +static inline int dquot_transfer(const struct mnt_idmap *idmap, struct inode *inode, struct iattr *iattr) { return 0; diff --git a/include/linux/sched.h b/include/linux/sched.h index f45b7d43113c..87ed6705c427 100644 --- a/include/linux/sched.h +++ b/include/linux/sched.h @@ -1817,7 +1817,7 @@ extern struct pid __rcu *cad_pid; * I am cleaning dirty pages from some other bdi. */ #define PF_KTHREAD 0x00200000 /* I am a kernel thread */ #define PF_RANDOMIZE 0x00400000 /* Randomize virtual address space */ -#define PF__HOLE__00800000 0x00800000 +#define PF_NO_NOTIFY_SIGNAL 0x00800000 /* see no_notify_signal_save() */ #define PF__HOLE__01000000 0x01000000 #define PF__HOLE__02000000 0x02000000 #define PF_NO_SETAFFINITY 0x04000000 /* Userland is not allowed to meddle with cpus_mask */ diff --git a/include/linux/sched/signal.h b/include/linux/sched/signal.h index d45a5476b97d..d9419dc902f6 100644 --- a/include/linux/sched/signal.h +++ b/include/linux/sched/signal.h @@ -2,6 +2,7 @@ #ifndef _LINUX_SCHED_SIGNAL_H #define _LINUX_SCHED_SIGNAL_H +#include <linux/cleanup.h> #include <linux/rculist.h> #include <linux/signal.h> #include <linux/sched.h> @@ -79,9 +80,9 @@ struct core_thread { }; struct core_state { - atomic_t nr_threads; - struct core_thread dumper; - struct completion startup; + /* Threads the dumper still waits for. */ + atomic_t threads_remaining; + struct core_thread *tasks; }; /* @@ -384,14 +385,36 @@ static inline int task_sigpending(struct task_struct *p) return unlikely(test_tsk_thread_flag(p,TIF_SIGPENDING)); } +/* Prevent TIF_NOTIFY_SIGNAL from interrupting this task. */ +static inline unsigned int no_notify_signal_save(void) +{ + unsigned int flags = current->flags; + + current->flags |= PF_NO_NOTIFY_SIGNAL; + return flags; +} + +/* Restore the previous PF_NO_NOTIFY_SIGNAL state. */ +static inline void no_notify_signal_restore(unsigned int flags) +{ + current_restore_flags(flags, PF_NO_NOTIFY_SIGNAL); +} + +DEFINE_LOCK_GUARD_0(no_notify_signal, + _T->flags = no_notify_signal_save(), + no_notify_signal_restore(_T->flags), + unsigned int flags) + static inline int signal_pending(struct task_struct *p) { /* * TIF_NOTIFY_SIGNAL isn't really a signal, but it requires the same * behavior in terms of ensuring that we break out of wait loops - * so that notify signal callbacks can be processed. + * so that notify signal callbacks can be processed. Not for a task + * that asked not to be interrupted by it, see no_notify_signal_save(). */ - if (unlikely(test_tsk_thread_flag(p, TIF_NOTIFY_SIGNAL))) + if (unlikely(test_tsk_thread_flag(p, TIF_NOTIFY_SIGNAL)) && + likely(!(READ_ONCE(p->flags) & PF_NO_NOTIFY_SIGNAL))) return 1; return task_sigpending(p); } diff --git a/include/linux/security.h b/include/linux/security.h index 153e9043058f..f7ff72ff956b 100644 --- a/include/linux/security.h +++ b/include/linux/security.h @@ -185,11 +185,11 @@ extern int cap_capset(struct cred *new, const struct cred *old, extern int cap_bprm_creds_from_file(struct linux_binprm *bprm, const struct file *file); int cap_inode_setxattr(struct dentry *dentry, const char *name, const void *value, size_t size, int flags); -int cap_inode_removexattr(struct mnt_idmap *idmap, +int cap_inode_removexattr(const struct mnt_idmap *idmap, struct dentry *dentry, const char *name); int cap_inode_need_killpriv(struct dentry *dentry); -int cap_inode_killpriv(struct mnt_idmap *idmap, struct dentry *dentry); -int cap_inode_getsecurity(struct mnt_idmap *idmap, +int cap_inode_killpriv(const struct mnt_idmap *idmap, struct dentry *dentry); +int cap_inode_getsecurity(const struct mnt_idmap *idmap, struct inode *inode, const char *name, void **buffer, bool alloc); extern int cap_mmap_addr(unsigned long addr); @@ -338,6 +338,7 @@ int security_binder_transfer_file(const struct cred *from, const struct cred *to, const struct file *file); int security_ptrace_access_check(struct task_struct *child, unsigned int mode); int security_ptrace_traceme(struct task_struct *parent); +int security_mem_foll_force(const struct cred *subject, bool opened_by_owner); int security_capget(const struct task_struct *target, kernel_cap_t *effective, kernel_cap_t *inheritable, @@ -405,7 +406,7 @@ int security_inode_init_security_anon(struct inode *inode, const struct qstr *name, const struct inode *context_inode); int security_inode_create(struct inode *dir, struct dentry *dentry, umode_t mode); -void security_inode_post_create_tmpfile(struct mnt_idmap *idmap, +void security_inode_post_create_tmpfile(const struct mnt_idmap *idmap, struct inode *inode); int security_inode_link(struct dentry *old_dentry, struct inode *dir, struct dentry *new_dentry); @@ -422,31 +423,31 @@ int security_inode_readlink(struct dentry *dentry); int security_inode_follow_link(struct dentry *dentry, struct inode *inode, bool rcu); int security_inode_permission(struct inode *inode, int mask); -int security_inode_setattr(struct mnt_idmap *idmap, +int security_inode_setattr(const struct mnt_idmap *idmap, struct dentry *dentry, struct iattr *attr); -void security_inode_post_setattr(struct mnt_idmap *idmap, struct dentry *dentry, +void security_inode_post_setattr(const struct mnt_idmap *idmap, struct dentry *dentry, int ia_valid); int security_inode_getattr(const struct path *path); -int security_inode_setxattr(struct mnt_idmap *idmap, +int security_inode_setxattr(const struct mnt_idmap *idmap, struct dentry *dentry, const char *name, const void *value, size_t size, int flags); -int security_inode_set_acl(struct mnt_idmap *idmap, +int security_inode_set_acl(const struct mnt_idmap *idmap, struct dentry *dentry, const char *acl_name, struct posix_acl *kacl); void security_inode_post_set_acl(struct dentry *dentry, const char *acl_name, struct posix_acl *kacl); -int security_inode_get_acl(struct mnt_idmap *idmap, +int security_inode_get_acl(const struct mnt_idmap *idmap, struct dentry *dentry, const char *acl_name); -int security_inode_remove_acl(struct mnt_idmap *idmap, +int security_inode_remove_acl(const struct mnt_idmap *idmap, struct dentry *dentry, const char *acl_name); -void security_inode_post_remove_acl(struct mnt_idmap *idmap, +void security_inode_post_remove_acl(const struct mnt_idmap *idmap, struct dentry *dentry, const char *acl_name); void security_inode_post_setxattr(struct dentry *dentry, const char *name, const void *value, size_t size, int flags); int security_inode_getxattr(struct dentry *dentry, const char *name); int security_inode_listxattr(struct dentry *dentry); -int security_inode_removexattr(struct mnt_idmap *idmap, +int security_inode_removexattr(const struct mnt_idmap *idmap, struct dentry *dentry, const char *name); void security_inode_post_removexattr(struct dentry *dentry, const char *name); int security_inode_file_setattr(struct dentry *dentry, @@ -454,8 +455,8 @@ int security_inode_file_setattr(struct dentry *dentry, int security_inode_file_getattr(struct dentry *dentry, struct file_kattr *fa); int security_inode_need_killpriv(struct dentry *dentry); -int security_inode_killpriv(struct mnt_idmap *idmap, struct dentry *dentry); -int security_inode_getsecurity(struct mnt_idmap *idmap, +int security_inode_killpriv(const struct mnt_idmap *idmap, struct dentry *dentry); +int security_inode_getsecurity(const struct mnt_idmap *idmap, struct inode *inode, const char *name, void **buffer, bool alloc); int security_inode_setsecurity(struct inode *inode, const char *name, const void *value, size_t size, int flags); @@ -676,6 +677,12 @@ static inline int security_ptrace_traceme(struct task_struct *parent) return cap_ptrace_traceme(parent); } +static inline int security_mem_foll_force(const struct cred *subject, + bool opened_by_owner) +{ + return 0; +} + static inline int security_capget(const struct task_struct *target, kernel_cap_t *effective, kernel_cap_t *inheritable, @@ -910,7 +917,7 @@ static inline int security_inode_create(struct inode *dir, } static inline void -security_inode_post_create_tmpfile(struct mnt_idmap *idmap, struct inode *inode) +security_inode_post_create_tmpfile(const struct mnt_idmap *idmap, struct inode *inode) { } static inline int security_inode_link(struct dentry *old_dentry, @@ -979,7 +986,7 @@ static inline int security_inode_permission(struct inode *inode, int mask) return 0; } -static inline int security_inode_setattr(struct mnt_idmap *idmap, +static inline int security_inode_setattr(const struct mnt_idmap *idmap, struct dentry *dentry, struct iattr *attr) { @@ -987,7 +994,7 @@ static inline int security_inode_setattr(struct mnt_idmap *idmap, } static inline void -security_inode_post_setattr(struct mnt_idmap *idmap, struct dentry *dentry, +security_inode_post_setattr(const struct mnt_idmap *idmap, struct dentry *dentry, int ia_valid) { } @@ -996,14 +1003,14 @@ static inline int security_inode_getattr(const struct path *path) return 0; } -static inline int security_inode_setxattr(struct mnt_idmap *idmap, +static inline int security_inode_setxattr(const struct mnt_idmap *idmap, struct dentry *dentry, const char *name, const void *value, size_t size, int flags) { return cap_inode_setxattr(dentry, name, value, size, flags); } -static inline int security_inode_set_acl(struct mnt_idmap *idmap, +static inline int security_inode_set_acl(const struct mnt_idmap *idmap, struct dentry *dentry, const char *acl_name, struct posix_acl *kacl) @@ -1016,21 +1023,21 @@ static inline void security_inode_post_set_acl(struct dentry *dentry, struct posix_acl *kacl) { } -static inline int security_inode_get_acl(struct mnt_idmap *idmap, +static inline int security_inode_get_acl(const struct mnt_idmap *idmap, struct dentry *dentry, const char *acl_name) { return 0; } -static inline int security_inode_remove_acl(struct mnt_idmap *idmap, +static inline int security_inode_remove_acl(const struct mnt_idmap *idmap, struct dentry *dentry, const char *acl_name) { return 0; } -static inline void security_inode_post_remove_acl(struct mnt_idmap *idmap, +static inline void security_inode_post_remove_acl(const struct mnt_idmap *idmap, struct dentry *dentry, const char *acl_name) { } @@ -1050,7 +1057,7 @@ static inline int security_inode_listxattr(struct dentry *dentry) return 0; } -static inline int security_inode_removexattr(struct mnt_idmap *idmap, +static inline int security_inode_removexattr(const struct mnt_idmap *idmap, struct dentry *dentry, const char *name) { @@ -1078,13 +1085,13 @@ static inline int security_inode_need_killpriv(struct dentry *dentry) return cap_inode_need_killpriv(dentry); } -static inline int security_inode_killpriv(struct mnt_idmap *idmap, +static inline int security_inode_killpriv(const struct mnt_idmap *idmap, struct dentry *dentry) { return cap_inode_killpriv(idmap, dentry); } -static inline int security_inode_getsecurity(struct mnt_idmap *idmap, +static inline int security_inode_getsecurity(const struct mnt_idmap *idmap, struct inode *inode, const char *name, void **buffer, bool alloc) @@ -2085,7 +2092,7 @@ int security_path_mkdir(const struct path *dir, struct dentry *dentry, umode_t m int security_path_rmdir(const struct path *dir, struct dentry *dentry); int security_path_mknod(const struct path *dir, struct dentry *dentry, umode_t mode, unsigned int dev); -void security_path_post_mknod(struct mnt_idmap *idmap, struct dentry *dentry); +void security_path_post_mknod(const struct mnt_idmap *idmap, struct dentry *dentry); int security_path_truncate(const struct path *path); int security_path_symlink(const struct path *dir, struct dentry *dentry, const char *old_name); @@ -2120,7 +2127,7 @@ static inline int security_path_mknod(const struct path *dir, struct dentry *den return 0; } -static inline void security_path_post_mknod(struct mnt_idmap *idmap, +static inline void security_path_post_mknod(const struct mnt_idmap *idmap, struct dentry *dentry) { } diff --git a/include/linux/splice.h b/include/linux/splice.h index 9dec4861d09f..0e6c955dc6ff 100644 --- a/include/linux/splice.h +++ b/include/linux/splice.h @@ -79,8 +79,8 @@ ssize_t add_to_pipe(struct pipe_inode_info *pipe, struct pipe_buffer *buf); ssize_t vfs_splice_read(struct file *in, loff_t *ppos, struct pipe_inode_info *pipe, size_t len, unsigned int flags); -ssize_t splice_direct_to_actor(struct file *file, struct splice_desc *sd, - splice_direct_actor *actor); +ssize_t vfs_splice_to_actor(struct file *file, loff_t pos, size_t count, + splice_direct_actor *actor, void *private); ssize_t do_splice(struct file *in, loff_t *off_in, struct file *out, loff_t *off_out, size_t len, unsigned int flags); ssize_t do_splice_direct(struct file *in, loff_t *ppos, struct file *out, diff --git a/include/linux/sunrpc/svc_xprt.h b/include/linux/sunrpc/svc_xprt.h index da2a2531e110..2af222f3ea2c 100644 --- a/include/linux/sunrpc/svc_xprt.h +++ b/include/linux/sunrpc/svc_xprt.h @@ -37,6 +37,9 @@ struct svc_xprt_class { struct list_head xcl_list; u32 xcl_max_payload; int xcl_ident; + u32 xcl_flags; +/* Set only on classes whose xpo_has_wspace() reads xpt_reserved */ +#define SVC_XPRT_FLAG_WSPACE_RESERVE BIT(0) }; /* @@ -59,7 +62,7 @@ struct svc_xprt { unsigned long xpt_flags; struct svc_serv *xpt_server; /* service for transport */ - atomic_t xpt_reserved; /* space on outq that is rsvd */ + atomic_t xpt_reserved; /* outq space rsvd, UDP only */ atomic_t xpt_nr_rqsts; /* Number of requests */ struct mutex xpt_mutex; /* to serialize sending data */ spinlock_t xpt_lock; /* protects sk_deferred diff --git a/include/linux/uidgid.h b/include/linux/uidgid.h index 2dc767e08f54..02403629b49f 100644 --- a/include/linux/uidgid.h +++ b/include/linux/uidgid.h @@ -130,9 +130,9 @@ static inline bool kgid_has_mapping(struct user_namespace *ns, kgid_t gid) return from_kgid(ns, gid) != (gid_t) -1; } -u32 map_id_down(struct uid_gid_map *map, u32 id); -u32 map_id_up(struct uid_gid_map *map, u32 id); -u32 map_id_range_up(struct uid_gid_map *map, u32 id, u32 count); +u32 map_id_down(const struct uid_gid_map *map, u32 id); +u32 map_id_up(const struct uid_gid_map *map, u32 id); +u32 map_id_range_up(const struct uid_gid_map *map, u32 id, u32 count); #else @@ -182,17 +182,17 @@ static inline bool kgid_has_mapping(struct user_namespace *ns, kgid_t gid) return gid_valid(gid); } -static inline u32 map_id_down(struct uid_gid_map *map, u32 id) +static inline u32 map_id_down(const struct uid_gid_map *map, u32 id) { return id; } -static inline u32 map_id_range_up(struct uid_gid_map *map, u32 id, u32 count) +static inline u32 map_id_range_up(const struct uid_gid_map *map, u32 id, u32 count) { return id; } -static inline u32 map_id_up(struct uid_gid_map *map, u32 id) +static inline u32 map_id_up(const struct uid_gid_map *map, u32 id) { return id; } diff --git a/include/linux/user_namespace.h b/include/linux/user_namespace.h index e38d9e60569f..91232053775d 100644 --- a/include/linux/user_namespace.h +++ b/include/linux/user_namespace.h @@ -29,8 +29,8 @@ struct uid_gid_map { /* 64 bytes -- 1 cache line */ u32 nr_extents; }; struct { - struct uid_gid_extent *forward; - struct uid_gid_extent *reverse; + struct uid_gid_extent *forward __counted_by_ptr(nr_extents); + struct uid_gid_extent *reverse __counted_by_ptr(nr_extents); }; }; }; @@ -207,6 +207,13 @@ extern bool in_userns(const struct user_namespace *ancestor, const struct user_namespace *child); extern bool current_in_userns(const struct user_namespace *target_ns); struct ns_common *ns_get_owner(struct ns_common *ns); + +#if IS_ENABLED(CONFIG_KUNIT) +extern int uid_gid_map_insert_extent(struct uid_gid_map *map, + struct uid_gid_extent *extent); +extern int uid_gid_map_sort(struct uid_gid_map *map); +#endif /* CONFIG_KUNIT */ + #else static inline struct user_namespace *get_user_ns(struct user_namespace *ns) diff --git a/include/linux/wait_bit.h b/include/linux/wait_bit.h index 553d7b23e3ad..af077ed4caf6 100644 --- a/include/linux/wait_bit.h +++ b/include/linux/wait_bit.h @@ -433,6 +433,32 @@ do { \ }) /** + * wait_var_event_state - wait for a variable to be updated and notified + * @var: the address of variable being waited on + * @condition: the condition to wait for + * @state: the task state to sleep in, %TASK_UNINTERRUPTIBLE etc. + * + * Wait for a @condition to be true, only re-checking when a wake up is + * received for the given @var (an arbitrary kernel address which need + * not be directly related to the given condition, but usually is). + * + * Returns 0 if the condition became true, or %-ERESTARTSYS if a signal + * arrived which @state allows to interrupt. + * + * The condition should normally use smp_load_acquire() or a similarly + * ordered access to ensure that any changes to memory made before the + * condition became true will be visible after the wait completes. + */ +#define wait_var_event_state(var, condition, state) \ +({ \ + int __ret = 0; \ + might_sleep(); \ + if (!(condition)) \ + __ret = ___wait_var_event(var, condition, (state), 0, 0, schedule()); \ + __ret; \ +}) + +/** * wait_var_event_any_lock - wait for a variable to be updated under a lock * @var: the address of the variable being waited on * @condition: condition to wait for diff --git a/include/linux/xattr.h b/include/linux/xattr.h index 54ac3cbc133f..4cc4257de084 100644 --- a/include/linux/xattr.h +++ b/include/linux/xattr.h @@ -47,7 +47,7 @@ struct xattr_handler { struct inode *inode, const char *name, void *buffer, size_t size); int (*set)(const struct xattr_handler *, - struct mnt_idmap *idmap, struct dentry *dentry, + const struct mnt_idmap *idmap, struct dentry *dentry, struct inode *inode, const char *name, const void *buffer, size_t size, int flags); }; @@ -77,25 +77,25 @@ struct xattr { }; ssize_t __vfs_getxattr(struct dentry *, struct inode *, const char *, void *, size_t); -ssize_t vfs_getxattr(struct mnt_idmap *, struct dentry *, const char *, +ssize_t vfs_getxattr(const struct mnt_idmap *, struct dentry *, const char *, void *, size_t); ssize_t vfs_listxattr(struct dentry *d, char *list, size_t size); -int __vfs_setxattr(struct mnt_idmap *, struct dentry *, struct inode *, +int __vfs_setxattr(const struct mnt_idmap *, struct dentry *, struct inode *, const char *, const void *, size_t, int); -int __vfs_setxattr_noperm(struct mnt_idmap *, struct dentry *, +int __vfs_setxattr_noperm(const struct mnt_idmap *, struct dentry *, const char *, const void *, size_t, int); -int __vfs_setxattr_locked(struct mnt_idmap *, struct dentry *, +int __vfs_setxattr_locked(const struct mnt_idmap *, struct dentry *, const char *, const void *, size_t, int, struct delegated_inode *); -int vfs_setxattr(struct mnt_idmap *, struct dentry *, const char *, +int vfs_setxattr(const struct mnt_idmap *, struct dentry *, const char *, const void *, size_t, int); -int __vfs_removexattr(struct mnt_idmap *, struct dentry *, const char *); -int __vfs_removexattr_locked(struct mnt_idmap *, struct dentry *, +int __vfs_removexattr(const struct mnt_idmap *, struct dentry *, const char *); +int __vfs_removexattr_locked(const struct mnt_idmap *, struct dentry *, const char *, struct delegated_inode *); -int vfs_removexattr(struct mnt_idmap *, struct dentry *, const char *); +int vfs_removexattr(const struct mnt_idmap *, struct dentry *, const char *); ssize_t generic_listxattr(struct dentry *dentry, char *buffer, size_t buffer_size); -int vfs_getxattr_alloc(struct mnt_idmap *idmap, +int vfs_getxattr_alloc(const struct mnt_idmap *idmap, struct dentry *dentry, const char *name, char **xattr_value, size_t size, gfp_t flags); diff --git a/include/trace/events/cachefiles.h b/include/trace/events/cachefiles.h index e3101410e8b2..a19233e8ae78 100644 --- a/include/trace/events/cachefiles.h +++ b/include/trace/events/cachefiles.h @@ -52,6 +52,8 @@ enum cachefiles_coherency_trace { cachefiles_coherency_check_ok, cachefiles_coherency_check_type, cachefiles_coherency_check_xattr, + cachefiles_coherency_discontiguous, + cachefiles_coherency_remove, cachefiles_coherency_set_fail, cachefiles_coherency_set_ok, cachefiles_coherency_vol_check_cmp, @@ -63,9 +65,11 @@ enum cachefiles_coherency_trace { }; enum cachefiles_trunc_trace { + cachefiles_trunc_clear_padding, cachefiles_trunc_dio_adjust, cachefiles_trunc_expand_tmpfile, cachefiles_trunc_shrink, + cachefiles_trunc_zap, }; enum cachefiles_prepare_read_trace { @@ -80,11 +84,14 @@ enum cachefiles_prepare_read_trace { }; enum cachefiles_error_trace { + cachefiles_trace_alignment_error, + cachefiles_trace_create_nospace, cachefiles_trace_fallocate_error, cachefiles_trace_getxattr_error, cachefiles_trace_link_error, cachefiles_trace_lookup_error, cachefiles_trace_mkdir_error, + cachefiles_trace_mkdir_nospace, cachefiles_trace_notify_change_error, cachefiles_trace_open_error, cachefiles_trace_read_error, @@ -97,6 +104,8 @@ enum cachefiles_error_trace { cachefiles_trace_trunc_error, cachefiles_trace_unlink_error, cachefiles_trace_write_error, + cachefiles_trace_write_nospace, + cachefiles_trace_write_nospace_2, }; #endif @@ -136,6 +145,8 @@ enum cachefiles_error_trace { EM(cachefiles_coherency_check_ok, "OK ") \ EM(cachefiles_coherency_check_type, "BAD type") \ EM(cachefiles_coherency_check_xattr, "BAD xatt") \ + EM(cachefiles_coherency_discontiguous, "--- gap ") \ + EM(cachefiles_coherency_remove, "REMOVE ") \ EM(cachefiles_coherency_set_fail, "SET fail") \ EM(cachefiles_coherency_set_ok, "SET ok ") \ EM(cachefiles_coherency_vol_check_cmp, "VOL BAD cmp ") \ @@ -146,9 +157,11 @@ enum cachefiles_error_trace { E_(cachefiles_coherency_vol_set_ok, "VOL SET ok ") #define cachefiles_trunc_traces \ + EM(cachefiles_trunc_clear_padding, "CLRPAD") \ EM(cachefiles_trunc_dio_adjust, "DIOADJ") \ EM(cachefiles_trunc_expand_tmpfile, "EXPTMP") \ - E_(cachefiles_trunc_shrink, "SHRINK") + EM(cachefiles_trunc_shrink, "SHRINK") \ + E_(cachefiles_trunc_zap, "ZAP ") #define cachefiles_prepare_read_traces \ EM(cachefiles_trace_read_after_eof, "after-eof ") \ @@ -161,11 +174,14 @@ enum cachefiles_error_trace { E_(cachefiles_trace_read_seek_nxio, "seek-enxio") #define cachefiles_error_traces \ + EM(cachefiles_trace_alignment_error, "align") \ + EM(cachefiles_trace_create_nospace, "create-nospace") \ EM(cachefiles_trace_fallocate_error, "fallocate") \ EM(cachefiles_trace_getxattr_error, "getxattr") \ EM(cachefiles_trace_link_error, "link") \ EM(cachefiles_trace_lookup_error, "lookup") \ EM(cachefiles_trace_mkdir_error, "mkdir") \ + EM(cachefiles_trace_mkdir_nospace, "mkdir-nospace") \ EM(cachefiles_trace_notify_change_error, "notify_change") \ EM(cachefiles_trace_open_error, "open") \ EM(cachefiles_trace_read_error, "read") \ @@ -177,7 +193,9 @@ enum cachefiles_error_trace { EM(cachefiles_trace_tmpfile_error, "tmpfile") \ EM(cachefiles_trace_trunc_error, "trunc") \ EM(cachefiles_trace_unlink_error, "unlink") \ - E_(cachefiles_trace_write_error, "write") + EM(cachefiles_trace_write_error, "write") \ + EM(cachefiles_trace_write_nospace, "write-nospace") \ + E_(cachefiles_trace_write_nospace_2, "write-nospace-2") /* @@ -371,12 +389,12 @@ TRACE_EVENT(cachefiles_rename, TRACE_EVENT(cachefiles_coherency, TP_PROTO(struct cachefiles_object *obj, - ino_t ino, + ino_t ino, uoff_t obj_size, const void *disk_aux, enum cachefiles_content content, enum cachefiles_coherency_trace why), - TP_ARGS(obj, ino, disk_aux, content, why), + TP_ARGS(obj, ino, obj_size, disk_aux, content, why), /* Note that obj may be NULL */ TP_STRUCT__entry( @@ -384,6 +402,7 @@ TRACE_EVENT(cachefiles_coherency, __field(enum cachefiles_coherency_trace, why) __field(enum cachefiles_content, content) __field(u64, ino) + __field(u64, obj_size) __field(u64, aux) __field(u64, disk_aux) ), @@ -398,6 +417,7 @@ TRACE_EVENT(cachefiles_coherency, __entry->why = why; __entry->content = content; __entry->ino = ino; + __entry->obj_size = obj_size; __entry->aux = be64_to_cpup((__be64 *)obj->cookie->inline_aux); /* cachefiles_xattr::data is 2-byte aligned but not 8-byte aligned. */ @@ -412,10 +432,11 @@ TRACE_EVENT(cachefiles_coherency, } ), - TP_printk("o=%08x %s B=%llx c=%u aux=%llx dsk=%llx", + TP_printk("o=%08x %s B=%llx oz=%llx c=%u aux=%llx dsk=%llx", __entry->obj, __print_symbolic(__entry->why, cachefiles_coherency_traces), __entry->ino, + __entry->obj_size, __entry->content, __entry->aux, __entry->disk_aux) @@ -449,7 +470,7 @@ TRACE_EVENT(cachefiles_vol_coherency, TRACE_EVENT(cachefiles_prep_read, TP_PROTO(struct cachefiles_object *obj, - loff_t start, + uoff_t start, size_t len, unsigned short flags, enum netfs_io_source source, @@ -464,7 +485,7 @@ TRACE_EVENT(cachefiles_prep_read, __field(enum netfs_io_source, source) __field(enum cachefiles_prepare_read_trace, why) __field(size_t, len) - __field(loff_t, start) + __field(uoff_t, start) __field(unsigned int, netfs_inode) __field(unsigned int, cache_inode) ), @@ -492,16 +513,16 @@ TRACE_EVENT(cachefiles_prep_read, TRACE_EVENT(cachefiles_read, TP_PROTO(struct cachefiles_object *obj, struct inode *backer, - loff_t start, + uoff_t start, size_t len), TP_ARGS(obj, backer, start, len), TP_STRUCT__entry( - __field(unsigned int, obj) - __field(unsigned int, backer) - __field(size_t, len) - __field(loff_t, start) + __field(unsigned int, obj) + __field(unsigned int, backer) + __field(size_t, len) + __field(uoff_t, start) ), TP_fast_assign( @@ -521,16 +542,16 @@ TRACE_EVENT(cachefiles_read, TRACE_EVENT(cachefiles_write, TP_PROTO(struct cachefiles_object *obj, struct inode *backer, - loff_t start, + uoff_t start, size_t len), TP_ARGS(obj, backer, start, len), TP_STRUCT__entry( - __field(unsigned int, obj) - __field(unsigned int, backer) - __field(size_t, len) - __field(loff_t, start) + __field(unsigned int, obj) + __field(unsigned int, backer) + __field(size_t, len) + __field(uoff_t, start) ), TP_fast_assign( @@ -549,7 +570,7 @@ TRACE_EVENT(cachefiles_write, TRACE_EVENT(cachefiles_trunc, TP_PROTO(struct cachefiles_object *obj, struct inode *backer, - loff_t from, loff_t to, enum cachefiles_trunc_trace why), + uoff_t from, uoff_t to, enum cachefiles_trunc_trace why), TP_ARGS(obj, backer, from, to, why), @@ -557,8 +578,8 @@ TRACE_EVENT(cachefiles_trunc, __field(unsigned int, obj) __field(unsigned int, backer) __field(enum cachefiles_trunc_trace, why) - __field(loff_t, from) - __field(loff_t, to) + __field(uoff_t, from) + __field(uoff_t, to) ), TP_fast_assign( @@ -694,6 +715,26 @@ TRACE_EVENT(cachefiles_io_error, __entry->error) ); +TRACE_EVENT(cachefiles_no_space, + TP_PROTO(struct cachefiles_object *obj, enum cachefiles_error_trace trace), + + TP_ARGS(obj, trace), + + TP_STRUCT__entry( + __field(unsigned int, obj) + __field(enum cachefiles_error_trace, trace) + ), + + TP_fast_assign( + __entry->obj = obj ? obj->debug_id : 0; + __entry->trace = trace; + ), + + TP_printk("o=%08x %s", + __entry->obj, + __print_symbolic(__entry->trace, cachefiles_error_traces)) + ); + #endif /* _TRACE_CACHEFILES_H */ /* This part must be outside protection */ diff --git a/include/trace/events/f2fs.h b/include/trace/events/f2fs.h index d53be932df01..df4cd346ae72 100644 --- a/include/trace/events/f2fs.h +++ b/include/trace/events/f2fs.h @@ -1430,6 +1430,50 @@ DEFINE_EVENT(f2fs__folio, f2fs_set_page_dirty, TP_ARGS(folio, type) ); +DECLARE_EVENT_CLASS(f2fs__cached_block, + + TP_PROTO(struct f2fs_cached_block *block, int type), + + TP_ARGS(block, type), + + TP_STRUCT__entry( + __field(dev_t, dev) + __field(pgoff_t, index) + __field(int, type) + __field(int, dirty) + __field(int, uptodate) + ), + + TP_fast_assign( + __entry->dev = block->cache->sbi->sb->s_dev; + __entry->index = block->index; + __entry->type = type; + __entry->dirty = f2fs_cache_test_dirty(block); + __entry->uptodate = f2fs_cache_test_uptodate(block); + ), + + TP_printk("dev = (%d,%d), %s, index = %lu, dirty = %d, uptodate = %d", + show_dev(__entry->dev), + show_block_type(__entry->type), + (unsigned long)__entry->index, + __entry->dirty, + __entry->uptodate) +); + +DEFINE_EVENT(f2fs__cached_block, f2fs_write_cache, + + TP_PROTO(struct f2fs_cached_block *block, int type), + + TP_ARGS(block, type) +); + +DEFINE_EVENT(f2fs__cached_block, f2fs_cache_set_dirty, + + TP_PROTO(struct f2fs_cached_block *block, int type), + + TP_ARGS(block, type) +); + TRACE_EVENT(f2fs_replace_atomic_write_block, TP_PROTO(struct inode *inode, struct inode *cow_inode, pgoff_t index, @@ -1574,6 +1618,33 @@ TRACE_EVENT(f2fs_writepages, __entry->for_sync) ); +TRACE_EVENT(f2fs_write_caches, + + TP_PROTO(struct f2fs_sb_info *sbi, long nr_to_write, long nwritten, int type), + + TP_ARGS(sbi, nr_to_write, nwritten, type), + + TP_STRUCT__entry( + __field(dev_t, dev) + __field(long, nr_to_write) + __field(long, nwritten) + __field(int, type) + ), + + TP_fast_assign( + __entry->dev = sbi->sb->s_dev; + __entry->nr_to_write = nr_to_write; + __entry->nwritten = nwritten; + __entry->type = type; + ), + + TP_printk("dev = (%d,%d), %s, nr_to_write = %ld, nwritten = %ld", + show_dev(__entry->dev), + show_block_type(__entry->type), + __entry->nr_to_write, + __entry->nwritten) +); + TRACE_EVENT(f2fs_readpages, TP_PROTO(struct inode *inode, pgoff_t start, unsigned int nrpage), diff --git a/include/trace/events/fscache.h b/include/trace/events/fscache.h index f1a73aa83fbb..8735d428ebd9 100644 --- a/include/trace/events/fscache.h +++ b/include/trace/events/fscache.h @@ -460,13 +460,13 @@ TRACE_EVENT(fscache_relinquish, ); TRACE_EVENT(fscache_invalidate, - TP_PROTO(struct fscache_cookie *cookie, loff_t new_size), + TP_PROTO(struct fscache_cookie *cookie, uoff_t new_size), TP_ARGS(cookie, new_size), TP_STRUCT__entry( __field(unsigned int, cookie ) - __field(loff_t, new_size ) + __field(uoff_t, new_size ) ), TP_fast_assign( @@ -479,14 +479,14 @@ TRACE_EVENT(fscache_invalidate, ); TRACE_EVENT(fscache_resize, - TP_PROTO(struct fscache_cookie *cookie, loff_t new_size), + TP_PROTO(struct fscache_cookie *cookie, uoff_t new_size), TP_ARGS(cookie, new_size), TP_STRUCT__entry( __field(unsigned int, cookie ) - __field(loff_t, old_size ) - __field(loff_t, new_size ) + __field(uoff_t, old_size ) + __field(uoff_t, new_size ) ), TP_fast_assign( diff --git a/include/trace/events/netfs.h b/include/trace/events/netfs.h index 3fec3e8f91c8..bf1e1f185b05 100644 --- a/include/trace/events/netfs.h +++ b/include/trace/events/netfs.h @@ -30,8 +30,7 @@ EM(netfs_write_trace_dio_write, "DIO-WRITE") \ EM(netfs_write_trace_unbuffered_write, "UNB-WRITE") \ EM(netfs_write_trace_writeback, "WRITEBACK") \ - EM(netfs_write_trace_writeback_single, "WB-SINGLE") \ - E_(netfs_write_trace_writethrough, "WRITETHRU") + E_(netfs_write_trace_writeback_single, "WB-SINGLE") #define netfs_rreq_origins \ EM(NETFS_READAHEAD, "RA") \ @@ -43,13 +42,17 @@ EM(NETFS_DIO_READ, "DR") \ EM(NETFS_WRITEBACK, "WB") \ EM(NETFS_WRITEBACK_SINGLE, "W1") \ - EM(NETFS_WRITETHROUGH, "WT") \ EM(NETFS_UNBUFFERED_WRITE, "UW") \ EM(NETFS_DIO_WRITE, "DW") \ E_(NETFS_PGPRIV2_COPY_TO_CACHE, "2C") #define netfs_rreq_traces \ + EM(netfs_rreq_trace_all_queued, "ALL-Q ") \ EM(netfs_rreq_trace_assess, "ASSESS ") \ + EM(netfs_rreq_trace_cache_cancelled, "CA-CNCL") \ + EM(netfs_rreq_trace_cache_failed, "CA-FAIL") \ + EM(netfs_rreq_trace_cache_fail_collect, "CA-F-CO") \ + EM(netfs_rreq_trace_cache_no_space, "CA-NOSP") \ EM(netfs_rreq_trace_collect, "COLLECT") \ EM(netfs_rreq_trace_complete, "COMPLET") \ EM(netfs_rreq_trace_copy, "COPY ") \ @@ -58,11 +61,14 @@ EM(netfs_rreq_trace_end_copy_to_cache, "END-C2C") \ EM(netfs_rreq_trace_free, "FREE ") \ EM(netfs_rreq_trace_intr, "INTR ") \ + EM(netfs_rreq_trace_inval_cache, "INVL-CA") \ EM(netfs_rreq_trace_ki_complete, "KI-CMPL") \ EM(netfs_rreq_trace_ra_put_ref, "RA-PUT ") \ EM(netfs_rreq_trace_recollect, "RECLLCT") \ EM(netfs_rreq_trace_redirty, "REDIRTY") \ EM(netfs_rreq_trace_resubmit, "RESUBMT") \ + EM(netfs_rreq_trace_retry_begin, "RETRY-BEGIN") \ + EM(netfs_rreq_trace_retry_end, "RETRY-END") \ EM(netfs_rreq_trace_set_abandon, "S-ABNDN") \ EM(netfs_rreq_trace_set_pause, "PAUSE ") \ EM(netfs_rreq_trace_unlock, "UNLOCK ") \ @@ -94,8 +100,10 @@ EM(netfs_sreq_trace_abandoned, "ABNDN") \ EM(netfs_sreq_trace_add_donations, "+DON ") \ EM(netfs_sreq_trace_added, "ADD ") \ + EM(netfs_sreq_trace_cache_nofile, "CA-!F") \ EM(netfs_sreq_trace_cache_nowrite, "CA-NW") \ EM(netfs_sreq_trace_cache_prepare, "CA-PR") \ + EM(netfs_sreq_trace_cache_waitfail, "CA-!W") \ EM(netfs_sreq_trace_cache_write, "CA-WR") \ EM(netfs_sreq_trace_cancel, "CANCL") \ EM(netfs_sreq_trace_clear, "CLEAR") \ @@ -134,12 +142,12 @@ #define netfs_failures \ EM(netfs_fail_check_write_begin, "check-write-begin") \ - EM(netfs_fail_copy_to_cache, "copy-to-cache") \ EM(netfs_fail_dio_read_short, "dio-read-short") \ EM(netfs_fail_dio_read_zero, "dio-read-zero") \ EM(netfs_fail_read, "read") \ EM(netfs_fail_short_read, "short-read") \ EM(netfs_fail_prepare_write, "prep-write") \ + EM(netfs_fail_upload, "upload") \ E_(netfs_fail_write, "write") #define netfs_rreq_ref_traces \ @@ -194,11 +202,11 @@ EM(netfs_folio_trace_alloc_buffer, "alloc-buf") \ EM(netfs_folio_trace_cancel_copy, "cancel-copy") \ EM(netfs_folio_trace_cancel_store, "cancel-store") \ - EM(netfs_folio_trace_clear, "clear") \ - EM(netfs_folio_trace_clear_cc, "clear-cc") \ - EM(netfs_folio_trace_clear_g, "clear-g") \ - EM(netfs_folio_trace_clear_s, "clear-s") \ EM(netfs_folio_trace_end_copy, "end-copy") \ + EM(netfs_folio_trace_endwb, "endwb") \ + EM(netfs_folio_trace_endwb_cc, "endwb-cc") \ + EM(netfs_folio_trace_endwb_g, "endwb-g") \ + EM(netfs_folio_trace_endwb_s, "endwb-s") \ EM(netfs_folio_trace_filled_gaps, "filled-gaps") \ EM(netfs_folio_trace_invalidate_all, "inval-all") \ EM(netfs_folio_trace_invalidate_front, "inval-front") \ @@ -223,9 +231,7 @@ EM(netfs_folio_trace_sched_copy, "sched-copy") \ EM(netfs_folio_trace_store, "store") \ EM(netfs_folio_trace_store_copy, "store-copy") \ - EM(netfs_folio_trace_store_plus, "store+") \ - EM(netfs_folio_trace_wthru, "wthru") \ - E_(netfs_folio_trace_wthru_plus, "wthru+") + E_(netfs_folio_trace_store_plus, "store+") #define netfs_collect_contig_traces \ EM(netfs_contig_trace_collect, "Collect") \ @@ -301,7 +307,7 @@ netfs_folioq_traces; TRACE_EVENT(netfs_read, TP_PROTO(struct netfs_io_request *rreq, - loff_t start, size_t len, + uoff_t start, size_t len, enum netfs_read_trace what), TP_ARGS(rreq, start, len, what), @@ -309,8 +315,9 @@ TRACE_EVENT(netfs_read, TP_STRUCT__entry( __field(unsigned int, rreq) __field(unsigned int, cookie) - __field(loff_t, i_size) - __field(loff_t, start) + __field(unsigned int, object) + __field(uoff_t, i_size) + __field(uoff_t, start) __field(size_t, len) __field(enum netfs_read_trace, what) __field(u64, netfs_inode) @@ -318,7 +325,8 @@ TRACE_EVENT(netfs_read, TP_fast_assign( __entry->rreq = rreq->debug_id; - __entry->cookie = rreq->cache_resources.debug_id; + __entry->cookie = rreq->cache_resources.cookie_id; + __entry->object = rreq->cache_resources.object_id; __entry->i_size = rreq->i_size; __entry->start = start; __entry->len = len; @@ -326,10 +334,10 @@ TRACE_EVENT(netfs_read, __entry->netfs_inode = rreq->inode->i_ino; ), - TP_printk("R=%08x %s c=%08x ni=%llx s=%llx l=%zx sz=%llx", + TP_printk("R=%08x %s c=%08x o=%08x ni=%llx s=%llx l=%zx sz=%llx", __entry->rreq, __print_symbolic(__entry->what, netfs_read_traces), - __entry->cookie, + __entry->cookie, __entry->object, __entry->netfs_inode, __entry->start, __entry->len, __entry->i_size) ); @@ -377,7 +385,7 @@ TRACE_EVENT(netfs_sreq, __field(u8, slot) __field(size_t, len) __field(size_t, transferred) - __field(loff_t, start) + __field(uoff_t, start) ), TP_fast_assign( @@ -418,7 +426,7 @@ TRACE_EVENT(netfs_failure, __field(enum netfs_failure, what) __field(size_t, len) __field(size_t, transferred) - __field(loff_t, start) + __field(uoff_t, start) ), TP_fast_assign( @@ -501,6 +509,7 @@ TRACE_EVENT(netfs_folio, TP_STRUCT__entry( __field(u64, ino) __field(pgoff_t, index) + __field(unsigned long, pfn) __field(unsigned int, nr) __field(enum netfs_folio_trace, why) ), @@ -511,9 +520,11 @@ TRACE_EVENT(netfs_folio, __entry->why = why; __entry->index = folio->index; __entry->nr = folio_nr_pages(folio); + __entry->pfn = folio_pfn(folio); ), - TP_printk("i=%05llx ix=%05lx-%05lx %s", + TP_printk("p=%lx i=%05llx ix=%05lx-%05lx %s", + __entry->pfn, __entry->ino, __entry->index, __entry->index + __entry->nr - 1, __print_symbolic(__entry->why, netfs_folio_traces)) ); @@ -524,10 +535,10 @@ TRACE_EVENT(netfs_write_iter, TP_ARGS(iocb, from), TP_STRUCT__entry( - __field(unsigned long long, start) - __field(size_t, len) - __field(unsigned int, flags) - __field(unsigned int, ino) + __field(uoff_t, start) + __field(size_t, len) + __field(unsigned int, flags) + __field(unsigned int, ino) ), TP_fast_assign( @@ -550,27 +561,27 @@ TRACE_EVENT(netfs_write, TP_STRUCT__entry( __field(unsigned int, wreq) __field(unsigned int, cookie) + __field(unsigned int, object) __field(unsigned int, ino) __field(enum netfs_write_trace, what) - __field(unsigned long long, start) - __field(unsigned long long, len) + __field(uoff_t, start) + __field(uoff_t, len) ), TP_fast_assign( - struct netfs_inode *__ctx = netfs_inode(wreq->inode); - struct fscache_cookie *__cookie = netfs_i_cookie(__ctx); __entry->wreq = wreq->debug_id; - __entry->cookie = __cookie ? __cookie->debug_id : 0; + __entry->cookie = wreq->cache_resources.cookie_id; + __entry->object = wreq->cache_resources.object_id; __entry->ino = wreq->inode->i_ino; __entry->what = what; __entry->start = wreq->start; __entry->len = wreq->len; ), - TP_printk("R=%08x %s c=%08x i=%x by=%llx-%llx", + TP_printk("R=%08x %s c=%08x o=%08x i=%x by=%llx-%llx", __entry->wreq, __print_symbolic(__entry->what, netfs_write_traces), - __entry->cookie, + __entry->cookie, __entry->object, __entry->ino, __entry->start, __entry->start + __entry->len - 1) ); @@ -582,25 +593,26 @@ TRACE_EVENT(netfs_copy2cache, TP_ARGS(rreq, creq), TP_STRUCT__entry( - __field(unsigned int, rreq) - __field(unsigned int, creq) - __field(unsigned int, cookie) - __field(unsigned int, ino) + __field(unsigned int, rreq) + __field(unsigned int, creq) + __field(unsigned int, cookie) + __field(unsigned int, object) + __field(unsigned int, ino) ), TP_fast_assign( - struct netfs_inode *__ctx = netfs_inode(rreq->inode); - struct fscache_cookie *__cookie = netfs_i_cookie(__ctx); __entry->rreq = rreq->debug_id; __entry->creq = creq->debug_id; - __entry->cookie = __cookie ? __cookie->debug_id : 0; + __entry->cookie = rreq->cache_resources.cookie_id; + __entry->object = rreq->cache_resources.object_id; __entry->ino = rreq->inode->i_ino; ), - TP_printk("R=%08x CR=%08x c=%08x i=%x ", + TP_printk("R=%08x CR=%08x c=%08x o=%08x i=%x ", __entry->rreq, __entry->creq, __entry->cookie, + __entry->object, __entry->ino) ); @@ -610,10 +622,10 @@ TRACE_EVENT(netfs_collect, TP_ARGS(wreq), TP_STRUCT__entry( - __field(unsigned int, wreq) - __field(unsigned int, len) - __field(unsigned long long, transferred) - __field(unsigned long long, start) + __field(unsigned int, wreq) + __field(unsigned int, len) + __field(uoff_t, transferred) + __field(uoff_t, start) ), TP_fast_assign( @@ -636,12 +648,12 @@ TRACE_EVENT(netfs_collect_sreq, TP_ARGS(wreq, subreq), TP_STRUCT__entry( - __field(unsigned int, wreq) - __field(unsigned int, subreq) - __field(unsigned int, stream) - __field(unsigned int, len) - __field(unsigned int, transferred) - __field(unsigned long long, start) + __field(unsigned int, wreq) + __field(unsigned int, subreq) + __field(unsigned int, stream) + __field(unsigned int, len) + __field(unsigned int, transferred) + __field(uoff_t, start) ), TP_fast_assign( @@ -660,37 +672,30 @@ TRACE_EVENT(netfs_collect_sreq, TRACE_EVENT(netfs_collect_folio, TP_PROTO(const struct netfs_io_request *wreq, - const struct folio *folio, - unsigned long long fend, - unsigned long long collected_to), + const struct folio *folio), - TP_ARGS(wreq, folio, fend, collected_to), + TP_ARGS(wreq, folio), TP_STRUCT__entry( __field(unsigned int, wreq) __field(unsigned long, index) - __field(unsigned long long, fend) - __field(unsigned long long, cleaned_to) - __field(unsigned long long, collected_to) + __field(unsigned int, nr) ), TP_fast_assign( __entry->wreq = wreq->debug_id; __entry->index = folio->index; - __entry->fend = fend; - __entry->cleaned_to = wreq->cleaned_to; - __entry->collected_to = collected_to; + __entry->nr = folio_nr_pages(folio); ), - TP_printk("R=%08x ix=%05lx r=%llx-%llx t=%llx/%llx", + TP_printk("R=%08x ix=%05lx-%05lx", __entry->wreq, __entry->index, - (unsigned long long)__entry->index * PAGE_SIZE, __entry->fend, - __entry->cleaned_to, __entry->collected_to) + __entry->index + __entry->nr - 1) ); TRACE_EVENT(netfs_collect_state, TP_PROTO(const struct netfs_io_request *wreq, - unsigned long long collected_to, + uoff_t collected_to, unsigned int notes), TP_ARGS(wreq, collected_to, notes), @@ -698,8 +703,8 @@ TRACE_EVENT(netfs_collect_state, TP_STRUCT__entry( __field(unsigned int, wreq) __field(unsigned int, notes) - __field(unsigned long long, collected_to) - __field(unsigned long long, cleaned_to) + __field(uoff_t, collected_to) + __field(uoff_t, cleaned_to) ), TP_fast_assign( @@ -718,7 +723,7 @@ TRACE_EVENT(netfs_collect_state, TRACE_EVENT(netfs_collect_gap, TP_PROTO(const struct netfs_io_request *wreq, const struct netfs_io_stream *stream, - unsigned long long jump_to, char type), + uoff_t jump_to, char type), TP_ARGS(wreq, stream, jump_to, type), @@ -726,8 +731,8 @@ TRACE_EVENT(netfs_collect_gap, __field(unsigned int, wreq) __field(unsigned char, stream) __field(unsigned char, type) - __field(unsigned long long, from) - __field(unsigned long long, to) + __field(uoff_t, from) + __field(uoff_t, to) ), TP_fast_assign( @@ -752,8 +757,8 @@ TRACE_EVENT(netfs_collect_stream, TP_STRUCT__entry( __field(unsigned int, wreq) __field(unsigned char, stream) - __field(unsigned long long, collected_to) - __field(unsigned long long, issued_to) + __field(uoff_t, collected_to) + __field(uoff_t, issued_to) ), TP_fast_assign( diff --git a/include/trace/misc/nfs.h b/include/trace/misc/nfs.h index a394b4d38e18..3146813fc4fe 100644 --- a/include/trace/misc/nfs.h +++ b/include/trace/misc/nfs.h @@ -8,6 +8,7 @@ */ #include <linux/nfs.h> +#include <linux/nfs3.h> #include <linux/nfs4.h> #include <uapi/linux/nfs.h> @@ -358,18 +359,6 @@ TRACE_DEFINE_ENUM(IOMODE_ANY); { IOMODE_RW, "RW" }, \ { IOMODE_ANY, "ANY" }) -#define show_rca_mask(x) \ - __print_flags(x, "|", \ - { BIT(RCA4_TYPE_MASK_RDATA_DLG), "RDATA_DLG" }, \ - { BIT(RCA4_TYPE_MASK_WDATA_DLG), "WDATA_DLG" }, \ - { BIT(RCA4_TYPE_MASK_DIR_DLG), "DIR_DLG" }, \ - { BIT(RCA4_TYPE_MASK_FILE_LAYOUT), "FILE_LAYOUT" }, \ - { BIT(RCA4_TYPE_MASK_BLK_LAYOUT), "BLK_LAYOUT" }, \ - { BIT(RCA4_TYPE_MASK_OBJ_LAYOUT_MIN), "OBJ_LAYOUT_MIN" }, \ - { BIT(RCA4_TYPE_MASK_OBJ_LAYOUT_MAX), "OBJ_LAYOUT_MAX" }, \ - { BIT(RCA4_TYPE_MASK_OTHER_LAYOUT_MIN), "OTHER_LAYOUT_MIN" }, \ - { BIT(RCA4_TYPE_MASK_OTHER_LAYOUT_MAX), "OTHER_LAYOUT_MAX" }) - #define show_nfs4_seq4_status(x) \ __print_flags(x, "|", \ { SEQ4_STATUS_CB_PATH_DOWN, "CB_PATH_DOWN" }, \ diff --git a/include/uapi/linux/btrfs_tree.h b/include/uapi/linux/btrfs_tree.h index cc3b9f7dccaf..47ee52859b45 100644 --- a/include/uapi/linux/btrfs_tree.h +++ b/include/uapi/linux/btrfs_tree.h @@ -230,7 +230,7 @@ * * Stored as an inline ref rather to avoid wasting space on a separate item on * top of the existing extent item. However, unlike the other inline refs, - * there is one one owner ref per extent rather than one per extent. + * there is one owner ref per extent rather than one per extent. * * Because of this, it goes at the front of the list of inline refs, and thus * must have a lower type value than any other inline ref type (to satisfy the @@ -243,7 +243,7 @@ #define BTRFS_EXTENT_DATA_REF_KEY 178 /* - * Obsolete key. Defintion removed in 6.6, value may be reused in the future. + * Obsolete key. Definition removed in 6.6, value may be reused in the future. * * #define BTRFS_EXTENT_REF_V0_KEY 180 */ @@ -1255,13 +1255,16 @@ static inline __u16 btrfs_qgroup_level(__u64 qgroupid) } /* - * is subvolume quota turned on? - */ -#define BTRFS_QGROUP_STATUS_FLAG_ON (1ULL << 0) -/* - * RESCAN is set during the initialization phase + * The following BTRFS_QGROUP_STATUS_BIT_* are for * btrfs_qgroup_status_item::flags. + * + * Is subvolume quota turned on? */ -#define BTRFS_QGROUP_STATUS_FLAG_RESCAN (1ULL << 1) +#define BTRFS_QGROUP_STATUS_BIT_ON (0) +#define BTRFS_QGROUP_STATUS_FLAG_ON (1UL << BTRFS_QGROUP_STATUS_BIT_ON) + +/* RESCAN is set during the initialization phase */ +#define BTRFS_QGROUP_STATUS_BIT_RESCAN (1) +#define BTRFS_QGROUP_STATUS_FLAG_RESCAN (1UL << BTRFS_QGROUP_STATUS_BIT_RESCAN) /* * Some qgroup entries are known to be out of date, * either because the configuration has changed in a way that @@ -1269,14 +1272,16 @@ static inline __u16 btrfs_qgroup_level(__u64 qgroupid) * with a non-qgroup-aware version. * Turning qouta off and on again makes it inconsistent, too. */ -#define BTRFS_QGROUP_STATUS_FLAG_INCONSISTENT (1ULL << 2) +#define BTRFS_QGROUP_STATUS_BIT_INCONSISTENT (2) +#define BTRFS_QGROUP_STATUS_FLAG_INCONSISTENT (1UL << BTRFS_QGROUP_STATUS_BIT_INCONSISTENT) /* * Whether or not this filesystem is using simple quotas. Not exactly the * incompat bit, because we support using simple quotas, disabling it, then * going back to full qgroup quotas. */ -#define BTRFS_QGROUP_STATUS_FLAG_SIMPLE_MODE (1ULL << 3) +#define BTRFS_QGROUP_STATUS_BIT_SIMPLE_MODE (3) +#define BTRFS_QGROUP_STATUS_FLAG_SIMPLE_MODE (1UL << BTRFS_QGROUP_STATUS_BIT_SIMPLE_MODE) #define BTRFS_QGROUP_STATUS_FLAGS_MASK (BTRFS_QGROUP_STATUS_FLAG_ON | \ BTRFS_QGROUP_STATUS_FLAG_RESCAN | \ diff --git a/include/uapi/linux/close_range.h b/include/uapi/linux/close_range.h index 2d804281554c..7da9ed95258a 100644 --- a/include/uapi/linux/close_range.h +++ b/include/uapi/linux/close_range.h @@ -2,11 +2,34 @@ #ifndef _UAPI_LINUX_CLOSE_RANGE_H #define _UAPI_LINUX_CLOSE_RANGE_H -/* Unshare the file descriptor table before closing file descriptors. */ -#define CLOSE_RANGE_UNSHARE (1U << 1) +/* + * A macro of one of these names defined before this header is parsed, by + * a libc or by a program's own fallback, would replace the enumerator. + */ +#undef CLOSE_RANGE_UNSHARE +#undef CLOSE_RANGE_CLOEXEC +#undef CLOSE_RANGE_EXCEPT +#undef CLOSE_RANGE_CLOEXEC_ONLY -/* Set the FD_CLOEXEC bit instead of closing the file descriptor. */ -#define CLOSE_RANGE_CLOEXEC (1U << 2) +enum close_range_flags { + /* Unshare the file descriptor table before closing file descriptors. */ + CLOSE_RANGE_UNSHARE = (1U << 1), + + /* Set the FD_CLOEXEC bit instead of closing the file descriptor. */ + CLOSE_RANGE_CLOEXEC = (1U << 2), + + /* Act on every file descriptor outside of the given range instead. */ + CLOSE_RANGE_EXCEPT = (1U << 3), + + /* Only close file descriptors that have the FD_CLOEXEC bit set. */ + CLOSE_RANGE_CLOEXEC_ONLY = (1U << 4), +}; + +/* Keep #ifdef working and let glibc skip its own definitions. */ +#define CLOSE_RANGE_UNSHARE CLOSE_RANGE_UNSHARE +#define CLOSE_RANGE_CLOEXEC CLOSE_RANGE_CLOEXEC +#define CLOSE_RANGE_EXCEPT CLOSE_RANGE_EXCEPT +#define CLOSE_RANGE_CLOEXEC_ONLY CLOSE_RANGE_CLOEXEC_ONLY #endif /* _UAPI_LINUX_CLOSE_RANGE_H */ diff --git a/include/uapi/linux/coredump.h b/include/uapi/linux/coredump.h index dc3789b78af0..6d0c53b534ea 100644 --- a/include/uapi/linux/coredump.h +++ b/include/uapi/linux/coredump.h @@ -11,12 +11,53 @@ * @COREDUMP_USERSPACE: userspace writes coredump * @COREDUMP_REJECT: don't generate coredump * @COREDUMP_WAIT: wait for coredump server + * @COREDUMP_RECORDS: send the coredump as a sequence of records instead of + * as a plain byte stream, see struct coredump_record_header; + * requires COREDUMP_KERNEL + * @COREDUMP_SPARSE: describe the holes in the coredump as zero records + * instead of transferring them; requires COREDUMP_RECORDS + * @COREDUMP_MEMORY_TYPES: dump the memory types in + * coredump_ack->memory_types instead of the ones + * the task selected; requires COREDUMP_KERNEL */ enum { COREDUMP_KERNEL = (1ULL << 0), COREDUMP_USERSPACE = (1ULL << 1), COREDUMP_REJECT = (1ULL << 2), COREDUMP_WAIT = (1ULL << 3), + COREDUMP_RECORDS = (1ULL << 4), + COREDUMP_SPARSE = (1ULL << 5), + COREDUMP_MEMORY_TYPES = (1ULL << 6), +}; + +/** + * coredump memory types + * @COREDUMP_MEMORY_ANON_PRIVATE: anonymous private memory + * @COREDUMP_MEMORY_ANON_SHARED: anonymous shared memory + * @COREDUMP_MEMORY_FILE_PRIVATE: file-backed private memory + * @COREDUMP_MEMORY_FILE_SHARED: file-backed shared memory + * @COREDUMP_MEMORY_ELF_HEADERS: the first page of a file-backed private + * mapping that starts an ELF file + * @COREDUMP_MEMORY_HUGETLB_PRIVATE: hugetlb private memory + * @COREDUMP_MEMORY_HUGETLB_SHARED: hugetlb shared memory + * @COREDUMP_MEMORY_DAX_PRIVATE: DAX private memory + * @COREDUMP_MEMORY_DAX_SHARED: DAX shared memory + * + * A bitmask of memory types a coredump may request to be included. New + * memory type bits must ensure that they do not steal memory from an + * existing one so a coredump server will continue to get the same + * coredumps even if a new bit is introduced. + */ +enum { + COREDUMP_MEMORY_ANON_PRIVATE = (1ULL << 0), + COREDUMP_MEMORY_ANON_SHARED = (1ULL << 1), + COREDUMP_MEMORY_FILE_PRIVATE = (1ULL << 2), + COREDUMP_MEMORY_FILE_SHARED = (1ULL << 3), + COREDUMP_MEMORY_ELF_HEADERS = (1ULL << 4), + COREDUMP_MEMORY_HUGETLB_PRIVATE = (1ULL << 5), + COREDUMP_MEMORY_HUGETLB_SHARED = (1ULL << 6), + COREDUMP_MEMORY_DAX_PRIVATE = (1ULL << 7), + COREDUMP_MEMORY_DAX_SHARED = (1ULL << 8), }; /** @@ -24,17 +65,19 @@ enum { * @size: size of struct coredump_req * @size_ack: known size of struct coredump_ack on this kernel * @mask: supported features + * @memory_types: the memory types the task selected + * @memory_types_mask: the memory types this kernel knows * * When a coredump happens the kernel will connect to the coredump * socket and send a coredump request to the coredump server. The @size * member is set to the size of struct coredump_req and provides a hint * to userspace how much data can be read. Userspace may use MSG_PEEK to * peek the size of struct coredump_req and then choose to consume it in - * one go. Userspace may also simply read a COREDUMP_ACK_SIZE_VER0 + * one go. Userspace may also simply read a COREDUMP_REQ_SIZE_VER0 * request. If the size the kernel sends is larger userspace simply * discards any remaining data. * - * The coredump_req->mask member is set to the currently know features. + * The coredump_req->mask member is set to the currently known features. * Userspace may only set coredump_ack->mask to the bits raised by the * kernel in coredump_req->mask. * @@ -42,15 +85,27 @@ enum { * struct coredump_ack the kernel knows. Userspace may only send up to * coredump_req->size_ack bytes to the kernel and must set * coredump_ack->size accordingly. + * + * @memory_types is set to the default memory types that are included in + * the coredump. This can be overridden by raising bits in + * coredump_ack->memory_types. + * + * @memory_types_mask contains a bitmask of all memory types the kernel + * knows about. A coredump server may only raise bits in + * coredump_ack->memory_types that are raised in + * coredump_req->memory_types_mask. */ struct coredump_req { __u32 size; __u32 size_ack; __u64 mask; + __u64 memory_types; + __u64 memory_types_mask; }; enum { COREDUMP_REQ_SIZE_VER0 = 16U, /* size of first published struct */ + COREDUMP_REQ_SIZE_VER1 = 32U, /* memory_types and memory_types_mask added */ }; /** @@ -58,6 +113,8 @@ enum { * @size: size of the struct * @spare: unused * @mask: features kernel is supposed to use + * @memory_types: memory types to dump, only with COREDUMP_MEMORY_TYPES + * in @mask * * The @size member must be set to the size of struct coredump_ack. It * may never exceed what the kernel returned in coredump_req->size_ack @@ -67,15 +124,30 @@ enum { * The @mask member must be set to the features the coredump server * wants the kernel to use. Only bits the kernel returned in * coredump_req->mask may be set. + * + * If COREDUMP_MEMORY_TYPES is raised in @mask the kernel dumps the + * memory types set in the @memory_types mask. Zero is valid and dumps + * no memory apart from the mappings that are always dumped. + * + * Note that memory a task excluded via MADV_DONTDUMP is always left + * out. A coredump server wanting to add or drop memory types instead of + * outright replacing it should simply copy coredump_req->memory_types + * and then mask off or raise types as needed. + * + * Note that @memory_types must be zero if COREDUMP_MEMORY_TYPES isn't + * raised. COREDUMP_MEMORY_TYPES requires COREDUMP_KERNEL and an ack of + * at least COREDUMP_ACK_SIZE_VER1 bytes. */ struct coredump_ack { __u32 size; __u32 spare; __u64 mask; + __u64 memory_types; }; enum { COREDUMP_ACK_SIZE_VER0 = 16U, /* size of first published struct */ + COREDUMP_ACK_SIZE_VER1 = 24U, /* memory_types added */ }; /** @@ -83,11 +155,12 @@ enum { * * The kernel will place a single byte on the coredump socket. The * markers notify userspace whether the coredump ack succeeded or - * failed. + * failed. After any marker other than COREDUMP_MARK_REQACK the kernel + * closes the connection and no coredump is generated. * * @COREDUMP_MARK_MINSIZE: the provided coredump_ack size was too small * @COREDUMP_MARK_MAXSIZE: the provided coredump_ack size was too big - * @COREDUMP_MARK_UNSUPPORTED: the provided coredump_ack mask was invalid + * @COREDUMP_MARK_UNSUPPORTED: the provided coredump_ack mask or memory types were invalid * @COREDUMP_MARK_CONFLICTING: the provided coredump_ack mask has conflicting options * @COREDUMP_MARK_REQACK: the coredump request and ack was successful * @__COREDUMP_MARK_MAX: the maximum coredump mark value @@ -101,4 +174,72 @@ enum coredump_mark { __COREDUMP_MARK_MAX = (1U << 31), }; +/** + * enum coredump_record_type - Type of a coredump record + * + * @COREDUMP_RECORD_DATA: the header is followed by ->len bytes of data + * @COREDUMP_RECORD_END: the coredump ends here, the header is not followed + * by any data and no further record is sent + * @COREDUMP_RECORD_ZERO: the header stands for ->len zero bytes and is not + * followed by any data + * @__COREDUMP_RECORD_TYPE_MAX: the maximum coredump record type value + */ +enum coredump_record_type { + COREDUMP_RECORD_DATA = 0U, + COREDUMP_RECORD_END = 1U, + COREDUMP_RECORD_ZERO = 2U, + __COREDUMP_RECORD_TYPE_MAX = (1U << 31), +}; + +/** + * struct coredump_record_header - header of a coredump record + * @size: size of struct coredump_record_header + * @type: one of enum coredump_record_type + * @flags: modifiers for this record + * @offset: offset in the coredump this record starts at + * @len: number of coredump bytes this record accounts for + * + * If the coredump server raises COREDUMP_RECORDS in coredump_ack->mask + * the kernel doesn't send the coredump as a plain byte stream. It sends + * a sequence of records instead. A COREDUMP_RECORD_DATA record is + * followed by @len bytes of actual coredump data. A + * COREDUMP_RECORD_ZERO record is followed by nothing and stands for + * @len zero bytes. A server that didn't raise COREDUMP_SPARSE never + * sees a zero record. Records arrive in order and leave no gaps. So + * @offset is the sum of the @len of all records before it. + * + * The last record is a COREDUMP_RECORD_END record. It is followed by + * nothing. Its @len is zero. Its @offset is the size of the coredump. + * The kernel only sends it once it has written the whole coredump. A + * server that hits end-of-file without having seen an end record must + * treat the coredump as incomplete. + * + * The @size member is set to the size of struct coredump_record_header + * the kernel knows and lets the header grow later. It comes first so it + * can be peeked. Userspace must consume @size bytes and discard + * anything beyond what it knows. It must refuse a @size smaller than + * COREDUMP_RECORD_HEADER_SIZE_VER0. @size covers the header alone. + * @offset and @len count coredump bytes. + * + * The @flags member carries modifiers that change how the record is to + * be interpreted. No flag is defined yet. Userspace must refuse a + * record carrying a flag or a type it doesn't know. Every new record + * type is raised in coredump_req->mask as a feature of its own. A + * server only ever sees the types it asked for. + * + * COREDUMP_RECORDS must be combined with COREDUMP_KERNEL, and + * COREDUMP_SPARSE with COREDUMP_RECORDS. + */ +struct coredump_record_header { + __u32 size; + __u32 type; + __u64 flags; + __u64 offset; + __u64 len; +}; + +enum { + COREDUMP_RECORD_HEADER_SIZE_VER0 = 32U, /* size of first published struct */ +}; + #endif /* _UAPI_LINUX_COREDUMP_H */ diff --git a/include/uapi/linux/fs.h b/include/uapi/linux/fs.h index 34c6f219462a..a46c33692aa2 100644 --- a/include/uapi/linux/fs.h +++ b/include/uapi/linux/fs.h @@ -88,7 +88,7 @@ struct fstrim_range { * We include a length field because some filesystems (vfat) have an identifier * that we do want to expose as a UUID, but doesn't have the standard length. * - * We use a fixed size buffer beacuse this interface will, by fiat, never + * We use a fixed size buffer because this interface will, by fiat, never * support "UUIDs" longer than 16 bytes; we don't want to force all downstream * users to have to deal with that. */ diff --git a/include/uapi/linux/fuse.h b/include/uapi/linux/fuse.h index 7435e09c87fe..10a7f31c4bdf 100644 --- a/include/uapi/linux/fuse.h +++ b/include/uapi/linux/fuse.h @@ -248,6 +248,9 @@ * - add bufpool offset field to fuse_uring_ent_in_out struct * - add FUSE_URING_ZERO_COPY, FUSE_URING_ENT_ZERO_COPY, and * FOPEN_IO_URING_ZERO_COPY flag + * + * 7.47 + * - add FUSE_HAS_SYNCFS opt-in flag for privileged userspace servers */ #ifndef _LINUX_FUSE_H @@ -283,7 +286,7 @@ #define FUSE_KERNEL_VERSION 7 /** Minor version number of this interface */ -#define FUSE_KERNEL_MINOR_VERSION 46 +#define FUSE_KERNEL_MINOR_VERSION 47 /** The node ID of the root inode */ #define FUSE_ROOT_ID 1 @@ -464,6 +467,12 @@ struct fuse_file_lock { * FUSE_REQUEST_TIMEOUT: kernel supports timing out requests. * init_out.request_timeout contains the timeout (in secs) * FUSE_HAS_IO_URING_BUFPOOL: kernel supports io-uring buffer pools + * FUSE_HAS_SYNCFS: server requests that syncfs()/sync() be propagated as + * FUSE_SYNCFS requests. Since an untrusted server can use this + * to stall sync(), it is only honored when /dev/fuse was opened + * with CAP_SYS_ADMIN in the initial user namespace (the same + * privilege that mounting virtiofs or fuseblk requires). + * Insufficiently privileged servers ignore it. */ #define FUSE_ASYNC_READ (1 << 0) #define FUSE_POSIX_LOCKS (1 << 1) @@ -512,6 +521,7 @@ struct fuse_file_lock { #define FUSE_OVER_IO_URING (1ULL << 41) #define FUSE_REQUEST_TIMEOUT (1ULL << 42) #define FUSE_HAS_IO_URING_BUFPOOL (1ULL << 43) +#define FUSE_HAS_SYNCFS (1ULL << 44) /** * CUSE INIT request/reply flags diff --git a/init/Kconfig b/init/Kconfig index b92340aa0d19..5229015e0e3d 100644 --- a/init/Kconfig +++ b/init/Kconfig @@ -1263,6 +1263,17 @@ config USER_NS If unsure, say N. +config USER_NS_MAP_KUNIT_TEST + tristate "KUint test for user namespace map insertion" if !KUNIT_ALL_TESTS + depends on USER_NS && KUNIT + default KUNIT_ALL_TESTS + help + This builds the KUnit test for user namespace uid/gid map insertion. + It validates map insertion, limits, dynamic allocation of the + extended extents array, and mapping sorting functions. + + If unsure, say N. + config PID_NS bool "PID Namespaces" default y diff --git a/init/initramfs.c b/init/initramfs.c index 3cee8b50ad82..ebddde9c8f0d 100644 --- a/init/initramfs.c +++ b/init/initramfs.c @@ -80,7 +80,7 @@ static __initdata struct hash { int ino, minor, major; umode_t mode; struct hash *next; - char name[N_ALIGN(PATH_MAX)]; + char name[]; } *head[32]; static __initdata bool hardlink_seen; @@ -92,7 +92,7 @@ static inline int hash(int major, int minor, int ino) } static char __init *find_link(int major, int minor, int ino, - umode_t mode, char *name) + umode_t mode, const char *name, size_t nlen) { struct hash **p, *q; for (p = head + hash(major, minor, ino); *p; p = &(*p)->next) { @@ -106,14 +106,15 @@ static char __init *find_link(int major, int minor, int ino, continue; return (*p)->name; } - q = kmalloc_obj(struct hash); + + q = kmalloc_flex(struct hash, name, nlen); if (!q) panic_show_mem("can't allocate link hash entry"); q->major = major; q->minor = minor; q->ino = ino; q->mode = mode; - strscpy(q->name, name); + strscpy(q->name, name, nlen); q->next = NULL; *p = q; hardlink_seen = true; @@ -355,7 +356,7 @@ static void __init clean_path(char *path, umode_t fmode) static int __init maybe_link(void) { if (nlink >= 2) { - char *old = find_link(major, minor, ino, mode, collected); + char *old = find_link(major, minor, ino, mode, collected, name_len); if (old) { clean_path(collected, 0); return (init_link(old, collected) < 0) ? -1 : 1; diff --git a/io_uring/io-wq.c b/io_uring/io-wq.c index 2ca223e47d41..2a980e86dd94 100644 --- a/io_uring/io-wq.c +++ b/io_uring/io-wq.c @@ -1324,6 +1324,8 @@ static bool io_task_work_match(struct callback_head *cb, void *data) void io_wq_exit_start(struct io_wq *wq) { set_bit(IO_WQ_BIT_EXIT, &wq->state); + /* Pairs with task_work_add() in io_queue_worker_create(). */ + smp_mb__after_atomic(); } static void io_wq_cancel_tw_create(struct io_wq *wq) diff --git a/io_uring/mock_file.c b/io_uring/mock_file.c index b318ed697998..9f0b4d850c12 100644 --- a/io_uring/mock_file.c +++ b/io_uring/mock_file.c @@ -257,17 +257,17 @@ static int io_create_mock_file(struct io_uring_cmd *cmd, unsigned int issue_flag FD_PREPARE(fdf, O_RDWR | O_CLOEXEC, anon_inode_create_getfile("[io_uring_mock]", fops, mf, O_RDWR | O_CLOEXEC, NULL)); - if (fdf.err) - return fdf.err; + if (fdf->fd < 0) + return fdf->fd; retain_and_null_ptr(mf); - file = fd_prepare_file(fdf); + file = fdf->file; file->f_mode |= FMODE_READ | FMODE_CAN_READ | FMODE_WRITE | FMODE_CAN_WRITE | FMODE_LSEEK; if (mc.flags & IORING_MOCK_CREATE_F_SUPPORT_NOWAIT) file->f_mode |= FMODE_NOWAIT; - mc.out_fd = fd_prepare_fd(fdf); + mc.out_fd = fdf->fd; if (copy_to_user(uarg, &mc, uarg_size)) return -EFAULT; diff --git a/ipc/mqueue.c b/ipc/mqueue.c index d1a1965c9811..322c3dd98c44 100644 --- a/ipc/mqueue.c +++ b/ipc/mqueue.c @@ -614,7 +614,7 @@ out_unlock: return error; } -static int mqueue_create(struct mnt_idmap *idmap, struct inode *dir, +static int mqueue_create(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *dentry, umode_t mode) { return mqueue_create_attr(dentry, mode, NULL); diff --git a/kernel/Makefile b/kernel/Makefile index c64c82c96b40..4ff92e24680c 100644 --- a/kernel/Makefile +++ b/kernel/Makefile @@ -141,6 +141,7 @@ obj-$(CONFIG_WATCH_QUEUE) += watch_queue.o obj-$(CONFIG_RESOURCE_KUNIT_TEST) += resource_kunit.o obj-$(CONFIG_SYSCTL_KUNIT_TEST) += sysctl-test.o +obj-$(CONFIG_USER_NS_MAP_KUNIT_TEST) += tests/user_ns_map_kunit.o CFLAGS_kstack_erase.o += $(DISABLE_KSTACK_ERASE) CFLAGS_kstack_erase.o += $(call cc-option,-mgeneral-regs-only) diff --git a/kernel/bpf/bpf_iter.c b/kernel/bpf/bpf_iter.c index b40eb404adab..d9191df5de22 100644 --- a/kernel/bpf/bpf_iter.c +++ b/kernel/bpf/bpf_iter.c @@ -643,11 +643,11 @@ int bpf_iter_new_fd(struct bpf_link *link) flags = O_RDONLY | O_CLOEXEC; FD_PREPARE(fdf, flags, anon_inode_getfile("bpf_iter", &bpf_iter_fops, NULL, flags)); - if (fdf.err) - return fdf.err; + if (fdf->fd < 0) + return fdf->fd; iter_link = container_of(link, struct bpf_iter_link, link); - err = prepare_seq_file(fd_prepare_file(fdf), iter_link); + err = prepare_seq_file(fdf->file, iter_link); if (err) return err; /* Automatic cleanup handles fput */ diff --git a/kernel/bpf/inode.c b/kernel/bpf/inode.c index 7837968c0842..c6f328e4752e 100644 --- a/kernel/bpf/inode.c +++ b/kernel/bpf/inode.c @@ -176,7 +176,7 @@ static void bpf_dentry_finalize(struct dentry *dentry, struct inode *inode, inode_set_mtime_to_ts(dir, inode_set_ctime_current(dir)); } -static struct dentry *bpf_mkdir(struct mnt_idmap *idmap, struct inode *dir, +static struct dentry *bpf_mkdir(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *dentry, umode_t mode) { struct inode *inode; @@ -424,7 +424,7 @@ bpf_lookup(struct inode *dir, struct dentry *dentry, unsigned flags) return simple_lookup(dir, dentry, flags); } -static int bpf_symlink(struct mnt_idmap *idmap, struct inode *dir, +static int bpf_symlink(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *dentry, const char *target) { struct inode *inode; @@ -874,7 +874,7 @@ enum { }; static int bpf_fs_xattr_set(const struct xattr_handler *handler, - struct mnt_idmap *idmap, struct dentry *unused, + const struct mnt_idmap *idmap, struct dentry *unused, struct inode *inode, const char *name, const void *value, size_t size, int flags) { diff --git a/kernel/bpf/token.c b/kernel/bpf/token.c index e85a179523f0..da915a4f972b 100644 --- a/kernel/bpf/token.c +++ b/kernel/bpf/token.c @@ -169,8 +169,8 @@ int bpf_token_create(union bpf_attr *attr) FD_PREPARE(fdf, O_CLOEXEC, alloc_file_pseudo(inode, path.mnt, BPF_TOKEN_INODE_NAME, O_RDWR, &bpf_token_fops)); - if (fdf.err) - return fdf.err; + if (fdf->fd < 0) + return fdf->fd; token = kzalloc_obj(*token, GFP_USER); if (!token) @@ -190,7 +190,7 @@ int bpf_token_create(union bpf_attr *attr) return err; get_user_ns(token->userns); - fd_prepare_file(fdf)->private_data = no_free_ptr(token); + fdf->file->private_data = no_free_ptr(token); return fd_publish(fdf); } diff --git a/kernel/capability.c b/kernel/capability.c index 90e6ab62f6db..a689dae590ff 100644 --- a/kernel/capability.c +++ b/kernel/capability.c @@ -470,7 +470,7 @@ EXPORT_SYMBOL(file_ns_capable); * Return true if the inode uid and gid are within the namespace. */ bool privileged_wrt_inode_uidgid(struct user_namespace *ns, - struct mnt_idmap *idmap, + const struct mnt_idmap *idmap, const struct inode *inode) { return vfsuid_has_mapping(ns, i_uid_into_vfsuid(idmap, inode)) && @@ -487,7 +487,7 @@ bool privileged_wrt_inode_uidgid(struct user_namespace *ns, * its own user namespace and that the given inode's uid and gid are * mapped into the current user namespace. */ -bool capable_wrt_inode_uidgid(struct mnt_idmap *idmap, +bool capable_wrt_inode_uidgid(const struct mnt_idmap *idmap, const struct inode *inode, int cap) { struct user_namespace *ns = current_user_ns(); diff --git a/kernel/exit.c b/kernel/exit.c index 282328d2b4cf..29e853a36602 100644 --- a/kernel/exit.c +++ b/kernel/exit.c @@ -17,6 +17,7 @@ #include <linux/module.h> #include <linux/capability.h> #include <linux/completion.h> +#include <linux/wait_bit.h> #include <linux/personality.h> #include <linux/tty.h> #include <linux/iocontext.h> @@ -436,15 +437,14 @@ static void coredump_task_exit(struct task_struct *tsk, self.task = tsk; if (self.task->flags & PF_SIGNALED) - self.next = xchg(&core_state->dumper.next, &self); + self.next = xchg(&core_state->tasks, &self); else self.task = NULL; /* * Implies mb(), the result of xchg() must be visible - * to core_state->dumper. + * to the dumper. */ - if (atomic_dec_and_test(&core_state->nr_threads)) - complete(&core_state->startup); + atomic_dec_and_wake_up(&core_state->threads_remaining); for (;;) { set_current_state(TASK_IDLE|TASK_FREEZABLE); @@ -892,7 +892,7 @@ static void synchronize_group_exit(struct task_struct *tsk, long code) * Serialize with any possible pending coredump. * We must hold siglock around checking core_state * and setting PF_POSTCOREDUMP. The core-inducing thread - * will increment ->nr_threads for each thread in the + * will increment ->threads_remaining for each thread in the * group without PF_POSTCOREDUMP set. */ tsk->flags |= PF_POSTCOREDUMP; @@ -978,10 +978,11 @@ void __noreturn do_exit(long code) exit_sem(tsk); exit_shm(tsk); - exit_files(tsk); - exit_fs(tsk); + /* Hang the tty up before the last close of it can clear the session. */ if (group_dead) disassociate_ctty(1); + exit_files(tsk); + exit_fs(tsk); exit_nsproxy_namespaces(tsk); exit_task_work(tsk); exit_thread(tsk); diff --git a/kernel/fork.c b/kernel/fork.c index f8d696c097ed..d458d85d7a65 100644 --- a/kernel/fork.c +++ b/kernel/fork.c @@ -1676,6 +1676,7 @@ static int copy_files(u64 clone_flags, struct task_struct *tsk, if (clone_flags & CLONE_FILES) { atomic_inc(&oldf->count); + tsk->files = oldf; return 0; } @@ -2145,12 +2146,9 @@ __latent_entropy struct task_struct *copy_process( if (args->kthread) p->flags |= PF_KTHREAD; if (args->user_worker) { - /* - * Mark us a user worker, and block any signal that isn't - * fatal or STOP - */ + /* A user worker takes only the signals nobody can block. */ p->flags |= PF_USER_WORKER; - siginitsetinv(&p->blocked, sigmask(SIGKILL)|sigmask(SIGSTOP)); + siginitsetinv(&p->blocked, SIG_KERNEL_ONLY_MASK); } if (args->io_thread) p->flags |= PF_IO_WORKER; @@ -2199,6 +2197,8 @@ __latent_entropy struct task_struct *copy_process( INIT_LIST_HEAD(&p->sibling); rcu_copy_process(p); p->vfork_done = NULL; + /* Set by copy_files(), exit_files() on the error path skips NULL. */ + p->files = NULL; spin_lock_init(&p->alloc_lock); init_sigpending(&p->pending); @@ -2300,7 +2300,7 @@ __latent_entropy struct task_struct *copy_process( goto bad_fork_cleanup_semundo; retval = copy_fs(clone_flags, p, args->umh); if (retval) - goto bad_fork_cleanup_files; + goto bad_fork_cleanup_semundo; retval = copy_sighand(clone_flags, p); if (retval) goto bad_fork_cleanup_fs; @@ -2614,8 +2614,6 @@ bad_fork_cleanup_sighand: __cleanup_sighand(p->sighand); bad_fork_cleanup_fs: exit_fs(p); /* blocking */ -bad_fork_cleanup_files: - exit_files(p); /* blocking */ bad_fork_cleanup_semundo: exit_sem(p); bad_fork_cleanup_security: @@ -2626,6 +2624,8 @@ bad_fork_cleanup_perf: perf_event_free_task(p); bad_fork_sched_cancel_fork: sched_cancel_fork(p); + /* ->release() of a file may need scx_fork_rwsem for write. */ + exit_files(p); /* blocking */ bad_fork_cleanup_policy: lockdep_free_task(p); #ifdef CONFIG_NUMA @@ -2704,6 +2704,10 @@ struct task_struct *create_io_thread(int (*fn)(void *), void *arg, int node) .user_worker = 1, }; + /* A creator past its fatal signal or its coredump point gets no thread. */ + if (current->flags & (PF_SIGNALED | PF_POSTCOREDUMP)) + return ERR_PTR(-EINTR); + return copy_process(NULL, 0, node, &args); } @@ -3212,24 +3216,6 @@ static int unshare_fs(unsigned long unshare_flags, struct fs_struct **new_fsp) } /* - * Unshare file descriptor table if it is being shared - */ -static int unshare_fd(unsigned long unshare_flags, struct files_struct **new_fdp) -{ - struct files_struct *fd = current->files; - - if ((unshare_flags & CLONE_FILES) && - (fd && atomic_read(&fd->count) > 1)) { - fd = dup_fd(fd, NULL); - if (IS_ERR(fd)) - return PTR_ERR(fd); - *new_fdp = fd; - } - - return 0; -} - -/* * unshare allows a process to 'unshare' part of the process * context which was originally shared using clone. copy_* * functions used by kernel_clone() cannot be used here directly @@ -3324,10 +3310,8 @@ int ksys_unshare(unsigned long unshare_flags) if (new_fs) new_fs = switch_fs_struct(new_fs); - if (new_fd) { - guard(task_lock)(current); - swap(current->files, new_fd); - } + if (new_fd) + switch_files_struct(current, no_free_ptr(new_fd)); if (new_cred) { /* Install the new user namespace */ @@ -3360,30 +3344,6 @@ SYSCALL_DEFINE1(unshare, unsigned long, unshare_flags) return ksys_unshare(unshare_flags); } -/* - * Helper to unshare the files of the current task. - * We don't want to expose copy_files internals to - * the exec layer of the kernel. - */ - -int unshare_files(void) -{ - struct task_struct *task = current; - struct files_struct *old, *copy = NULL; - int error; - - error = unshare_fd(CLONE_FILES, ©); - if (error || !copy) - return error; - - old = task->files; - task_lock(task); - task->files = copy; - task_unlock(task); - put_files_struct(old); - return 0; -} - static int sysctl_max_threads(const struct ctl_table *table, int write, void *buffer, size_t *lenp, loff_t *ppos) { diff --git a/kernel/kthread.c b/kernel/kthread.c index a3f95c90456b..cc8cb5d3eab7 100644 --- a/kernel/kthread.c +++ b/kernel/kthread.c @@ -81,7 +81,7 @@ enum KTHREAD_BITS { static inline struct kthread *to_kthread(struct task_struct *k) { - WARN_ON(!(k->flags & PF_KTHREAD)); + WARN_ON(!(READ_ONCE(k->flags) & PF_KTHREAD)); return k->worker_private; } diff --git a/kernel/pid_namespace.c b/kernel/pid_namespace.c index d36afc58ee1d..8bc9edb40b78 100644 --- a/kernel/pid_namespace.c +++ b/kernel/pid_namespace.c @@ -238,9 +238,10 @@ void zap_pid_ns_processes(struct pid_namespace *pid_ns) * kernel_wait4() will also block until our children traced from the * parent namespace are detached and become EXIT_DEAD. */ + /* Task work must not busy-loop the reaper, see signal_pending(). */ + guard(no_notify_signal)(); do { clear_thread_flag(TIF_SIGPENDING); - clear_thread_flag(TIF_NOTIFY_SIGNAL); rc = kernel_wait4(-1, NULL, __WALL, NULL); } while (rc != -ECHILD); diff --git a/kernel/ptrace.c b/kernel/ptrace.c index d041645d9d17..4e9822a87aab 100644 --- a/kernel/ptrace.c +++ b/kernel/ptrace.c @@ -1227,6 +1227,12 @@ int ptrace_request(struct task_struct *child, long request, case PTRACE_SETSIGMASK: { sigset_t new_set; + /* A user worker only ever takes SIGKILL and SIGSTOP. */ + if (child->flags & PF_USER_WORKER) { + ret = -EPERM; + break; + } + if (addr != sizeof(sigset_t)) { ret = -EINVAL; break; diff --git a/kernel/signal.c b/kernel/signal.c index d31ebcb6ed4d..e433de93b430 100644 --- a/kernel/signal.c +++ b/kernel/signal.c @@ -3170,6 +3170,10 @@ static void retarget_shared_pending(struct task_struct *tsk, sigset_t *which) sigset_t retarget; struct task_struct *t; + /* Nobody dequeues them in a dying group, see get_signal(). */ + if (tsk->signal->flags & SIGNAL_GROUP_EXIT) + return; + sigandsets(&retarget, &tsk->signal->shared_pending.signal, which); if (sigisemptyset(&retarget)) return; @@ -3255,6 +3259,16 @@ long do_no_restart_syscall(struct restart_block *param) static void __set_task_blocked(struct task_struct *tsk, const sigset_t *newset) { + sigset_t floor, floored; + + /* A user worker never unblocks anything but SIGKILL and SIGSTOP. */ + if (unlikely(tsk->flags & PF_USER_WORKER)) { + siginitsetinv(&floor, SIG_KERNEL_ONLY_MASK); + sigorsets(&floored, newset, &floor); + WARN_ON_ONCE(!sigequalsets(&floored, newset)); + newset = &floored; + } + if (task_sigpending(tsk) && !thread_group_empty(tsk)) { sigset_t newblocked; /* A set of now blocked but previously unblocked signals. */ diff --git a/kernel/tests/.kunitconfig b/kernel/tests/.kunitconfig new file mode 100644 index 000000000000..b3d1206fd81a --- /dev/null +++ b/kernel/tests/.kunitconfig @@ -0,0 +1,4 @@ +CONFIG_KUNIT=y +CONFIG_NAMESPACES=y +CONFIG_USER_NS=y +CONFIG_USER_NS_MAP_KUNIT_TEST=y diff --git a/kernel/tests/user_ns_map_kunit.c b/kernel/tests/user_ns_map_kunit.c new file mode 100644 index 000000000000..033dccc6a535 --- /dev/null +++ b/kernel/tests/user_ns_map_kunit.c @@ -0,0 +1,98 @@ +// SPDX-License-Identifier: GPL-2.0 +/* + * KUnit test for user namespace map insertion and sorting. + */ + +#define pr_fmt(fmt) "user_namespace: " fmt + +#include <kunit/test.h> +#include <linux/user_namespace.h> + +#define NR_EXTENTS (UID_GID_MAP_MAX_BASE_EXTENTS + 5) + +static void user_ns_map_insert(struct kunit *test) +{ + struct uid_gid_map map; + struct uid_gid_extent extent; + int i, ret; + + memset(&map, 0, sizeof(map)); + + /* Insert up to UID_GID_MAP_MAX_BASE_EXTENTS elements */ + for (i = 0; i < UID_GID_MAP_MAX_BASE_EXTENTS; i++) { + extent.first = i * 10; + extent.lower_first = i * 100; + extent.count = 5; + + ret = uid_gid_map_insert_extent(&map, &extent); + KUNIT_ASSERT_EQ(test, ret, 0); + } + + KUNIT_EXPECT_EQ(test, map.nr_extents, UID_GID_MAP_MAX_BASE_EXTENTS); + + /* Verify the elements ended up in the 'extent' array */ + for (i = 0; i < UID_GID_MAP_MAX_BASE_EXTENTS; i++) { + KUNIT_EXPECT_EQ(test, map.extent[i].first, i * 10); + KUNIT_EXPECT_EQ(test, map.extent[i].lower_first, i * 100); + KUNIT_EXPECT_EQ(test, map.extent[i].count, 5); + } +} + +static void user_ns_map_insert_extended(struct kunit *test) +{ + struct uid_gid_map map; + struct uid_gid_extent extent; + int i, ret; + + memset(&map, 0, sizeof(map)); + + /* Insert more than UID_GID_MAP_MAX_BASE_EXTENTS elements */ + for (i = 0; i < NR_EXTENTS; i++) { + int value = 9 - i; + + extent.first = value * 10; + extent.lower_first = value * 100; + extent.count = 5; + + ret = uid_gid_map_insert_extent(&map, &extent); + KUNIT_ASSERT_EQ(test, ret, 0); + } + + KUNIT_EXPECT_EQ(test, map.nr_extents, NR_EXTENTS); + + /* Now sort the map to set up reverse mapping */ + ret = uid_gid_map_sort(&map); + KUNIT_ASSERT_EQ(test, ret, 0); + KUNIT_ASSERT_NOT_ERR_OR_NULL(test, map.reverse); + + /* Verify the elements are in 'forward' and that sorting is correct */ + for (i = 0; i < map.nr_extents; i++) { + KUNIT_EXPECT_EQ(test, map.forward[i].first, i * 10); + KUNIT_EXPECT_EQ(test, map.forward[i].lower_first, i * 100); + KUNIT_EXPECT_EQ(test, map.forward[i].count, 5); + + KUNIT_EXPECT_EQ(test, map.reverse[i].first, i * 10); + KUNIT_EXPECT_EQ(test, map.reverse[i].lower_first, i * 100); + KUNIT_EXPECT_EQ(test, map.reverse[i].count, 5); + } + + kfree(map.forward); + kfree(map.reverse); +} + +static struct kunit_case user_ns_map_test_cases[] = { + KUNIT_CASE(user_ns_map_insert), + KUNIT_CASE(user_ns_map_insert_extended), + {} +}; + +static struct kunit_suite user_ns_map_test_suite = { + .name = "user_ns_map", + .test_cases = user_ns_map_test_cases, +}; + +kunit_test_suite(user_ns_map_test_suite); + +MODULE_LICENSE("GPL"); +MODULE_DESCRIPTION("KUnit test for user namespace map insertion"); +MODULE_IMPORT_NS("EXPORTED_FOR_KUNIT_TESTING"); diff --git a/kernel/user_namespace.c b/kernel/user_namespace.c index 0bed462e9b2a..1b23d819d398 100644 --- a/kernel/user_namespace.c +++ b/kernel/user_namespace.c @@ -1,5 +1,6 @@ // SPDX-License-Identifier: GPL-2.0-only +#include <kunit/visibility.h> #include <linux/export.h> #include <linux/nsproxy.h> #include <linux/slab.h> @@ -162,9 +163,6 @@ int create_user_ns(struct cred *new) ns_tree_add(ns); return 0; fail_keyring: -#ifdef CONFIG_PERSISTENT_KEYRINGS - key_put(ns->persistent_keyring_register); -#endif ns_common_free(ns); fail_free: kmem_cache_free(user_ns_cachep, ns); @@ -278,8 +276,8 @@ static int cmp_map_id(const void *k, const void *e) * map_id_range_down_max - Find idmap via binary search in ordered idmap array. * Can only be called if number of mappings exceeds UID_GID_MAP_MAX_BASE_EXTENTS. */ -static struct uid_gid_extent * -map_id_range_down_max(unsigned extents, struct uid_gid_map *map, u32 id, u32 count) +static const struct uid_gid_extent * +map_id_range_down_max(unsigned extents, const struct uid_gid_map *map, u32 id, u32 count) { struct idmap_key key; @@ -296,8 +294,8 @@ map_id_range_down_max(unsigned extents, struct uid_gid_map *map, u32 id, u32 cou * Can only be called if number of mappings is equal or less than * UID_GID_MAP_MAX_BASE_EXTENTS. */ -static struct uid_gid_extent * -map_id_range_down_base(unsigned extents, struct uid_gid_map *map, u32 id, u32 count) +static const struct uid_gid_extent * +map_id_range_down_base(unsigned extents, const struct uid_gid_map *map, u32 id, u32 count) { unsigned idx; u32 first, last, id2; @@ -315,9 +313,9 @@ map_id_range_down_base(unsigned extents, struct uid_gid_map *map, u32 id, u32 co return NULL; } -static u32 map_id_range_down(struct uid_gid_map *map, u32 id, u32 count) +static u32 map_id_range_down(const struct uid_gid_map *map, u32 id, u32 count) { - struct uid_gid_extent *extent; + const struct uid_gid_extent *extent; unsigned extents = map->nr_extents; smp_rmb(); @@ -335,7 +333,7 @@ static u32 map_id_range_down(struct uid_gid_map *map, u32 id, u32 count) return id; } -u32 map_id_down(struct uid_gid_map *map, u32 id) +u32 map_id_down(const struct uid_gid_map *map, u32 id) { return map_id_range_down(map, id, 1); } @@ -345,8 +343,8 @@ u32 map_id_down(struct uid_gid_map *map, u32 id) * Can only be called if number of mappings is equal or less than * UID_GID_MAP_MAX_BASE_EXTENTS. */ -static struct uid_gid_extent * -map_id_range_up_base(unsigned extents, struct uid_gid_map *map, u32 id, u32 count) +static const struct uid_gid_extent * +map_id_range_up_base(unsigned extents, const struct uid_gid_map *map, u32 id, u32 count) { unsigned idx; u32 first, last, id2; @@ -368,8 +366,8 @@ map_id_range_up_base(unsigned extents, struct uid_gid_map *map, u32 id, u32 coun * map_id_up_max - Find idmap via binary search in ordered idmap array. * Can only be called if number of mappings exceeds UID_GID_MAP_MAX_BASE_EXTENTS. */ -static struct uid_gid_extent * -map_id_range_up_max(unsigned extents, struct uid_gid_map *map, u32 id, u32 count) +static const struct uid_gid_extent * +map_id_range_up_max(unsigned extents, const struct uid_gid_map *map, u32 id, u32 count) { struct idmap_key key; @@ -381,9 +379,9 @@ map_id_range_up_max(unsigned extents, struct uid_gid_map *map, u32 id, u32 count sizeof(struct uid_gid_extent), cmp_map_id); } -u32 map_id_range_up(struct uid_gid_map *map, u32 id, u32 count) +u32 map_id_range_up(const struct uid_gid_map *map, u32 id, u32 count) { - struct uid_gid_extent *extent; + const struct uid_gid_extent *extent; unsigned extents = map->nr_extents; smp_rmb(); @@ -401,7 +399,7 @@ u32 map_id_range_up(struct uid_gid_map *map, u32 id, u32 count) return id; } -u32 map_id_up(struct uid_gid_map *map, u32 id) +u32 map_id_up(const struct uid_gid_map *map, u32 id) { return map_id_range_up(map, id, 1); } @@ -782,11 +780,13 @@ static bool mappings_overlap(struct uid_gid_map *new_map, } /* - * insert_extent - Safely insert a new idmap extent into struct uid_gid_map. + * uid_gid_map_insert_extent - Safely insert a new idmap extent into + * struct uid_gid_map. * Takes care to allocate a 4K block of memory if the number of mappings exceeds * UID_GID_MAP_MAX_BASE_EXTENTS. */ -static int insert_extent(struct uid_gid_map *map, struct uid_gid_extent *extent) +VISIBLE_IF_KUNIT int uid_gid_map_insert_extent(struct uid_gid_map *map, + struct uid_gid_extent *extent) { struct uid_gid_extent *dest; @@ -809,15 +809,20 @@ static int insert_extent(struct uid_gid_map *map, struct uid_gid_extent *extent) map->reverse = NULL; } - if (map->nr_extents < UID_GID_MAP_MAX_BASE_EXTENTS) - dest = &map->extent[map->nr_extents]; + /* + * nr_extents must be updated before the extent and forward arrays are + * accessed, otherwise KSAN will assert an out-of-bounds error. + */ + map->nr_extents++; + if (map->nr_extents <= UID_GID_MAP_MAX_BASE_EXTENTS) + dest = &map->extent[map->nr_extents - 1]; else - dest = &map->forward[map->nr_extents]; + dest = &map->forward[map->nr_extents - 1]; *dest = *extent; - map->nr_extents++; return 0; } +EXPORT_SYMBOL_IF_KUNIT(uid_gid_map_insert_extent); /* cmp function to sort() forward mappings */ static int cmp_extents_forward(const void *a, const void *b) @@ -850,10 +855,10 @@ static int cmp_extents_reverse(const void *a, const void *b) } /* - * sort_idmaps - Sorts an array of idmap entries. + * uid_gid_map_sort - Sorts an array of idmap entries. * Can only be called if number of mappings exceeds UID_GID_MAP_MAX_BASE_EXTENTS. */ -static int sort_idmaps(struct uid_gid_map *map) +VISIBLE_IF_KUNIT int uid_gid_map_sort(struct uid_gid_map *map) { if (map->nr_extents <= UID_GID_MAP_MAX_BASE_EXTENTS) return 0; @@ -874,6 +879,7 @@ static int sort_idmaps(struct uid_gid_map *map) return 0; } +EXPORT_SYMBOL_IF_KUNIT(uid_gid_map_sort); /** * verify_root_map() - check the uid 0 mapping @@ -1042,7 +1048,7 @@ static ssize_t map_write(struct file *file, const char __user *buf, (next_line != NULL)) goto out; - ret = insert_extent(&new_map, &extent); + ret = uid_gid_map_insert_extent(&new_map, &extent); if (ret < 0) goto out; ret = -EINVAL; @@ -1086,7 +1092,7 @@ static ssize_t map_write(struct file *file, const char __user *buf, * If we want to use binary search for lookup, this clones the extent * array and sorts both copies. */ - ret = sort_idmaps(&new_map); + ret = uid_gid_map_sort(&new_map); if (ret < 0) goto out; diff --git a/kernel/utsname.c b/kernel/utsname.c index ebbfc578a9d3..1ebf87e24607 100644 --- a/kernel/utsname.c +++ b/kernel/utsname.c @@ -81,7 +81,6 @@ struct uts_namespace *copy_utsname(u64 flags, { struct uts_namespace *new_ns; - BUG_ON(!old_ns); get_uts_ns(old_ns); if (!(flags & CLONE_NEWUTS)) diff --git a/mm/secretmem.c b/mm/secretmem.c index 7287a2866897..b3f79ad56790 100644 --- a/mm/secretmem.c +++ b/mm/secretmem.c @@ -238,7 +238,7 @@ const struct address_space_operations secretmem_aops = { .migrate_folio = secretmem_migrate_folio, }; -static int secretmem_setattr(struct mnt_idmap *idmap, +static int secretmem_setattr(const struct mnt_idmap *idmap, struct dentry *dentry, struct iattr *iattr) { struct inode *inode = d_inode(dentry); diff --git a/mm/shmem.c b/mm/shmem.c index 12660bd9aaaa..94dc2981fbba 100644 --- a/mm/shmem.c +++ b/mm/shmem.c @@ -1499,7 +1499,7 @@ void shmem_truncate_range(struct inode *inode, loff_t lstart, uoff_t lend) } EXPORT_SYMBOL_GPL(shmem_truncate_range); -static int shmem_getattr(struct mnt_idmap *idmap, +static int shmem_getattr(const struct mnt_idmap *idmap, const struct path *path, struct kstat *stat, u32 request_mask, unsigned int query_flags) { @@ -1533,7 +1533,7 @@ static int shmem_getattr(struct mnt_idmap *idmap, return 0; } -static int shmem_setattr(struct mnt_idmap *idmap, +static int shmem_setattr(const struct mnt_idmap *idmap, struct dentry *dentry, struct iattr *attr) { struct inode *inode = d_inode(dentry); @@ -3219,7 +3219,7 @@ static struct offset_ctx *shmem_get_offset_ctx(struct inode *inode) return &SHMEM_I(inode)->dir_offsets; } -static struct inode *__shmem_get_inode(struct mnt_idmap *idmap, +static struct inode *__shmem_get_inode(const struct mnt_idmap *idmap, struct super_block *sb, struct inode *dir, umode_t mode, dev_t dev, vma_flags_t flags) @@ -3303,7 +3303,7 @@ static struct inode *__shmem_get_inode(struct mnt_idmap *idmap, } #ifdef CONFIG_TMPFS_QUOTA -static struct inode *shmem_get_inode(struct mnt_idmap *idmap, +static struct inode *shmem_get_inode(const struct mnt_idmap *idmap, struct super_block *sb, struct inode *dir, umode_t mode, dev_t dev, vma_flags_t flags) { @@ -3331,7 +3331,7 @@ errout: return ERR_PTR(err); } #else -static struct inode *shmem_get_inode(struct mnt_idmap *idmap, +static struct inode *shmem_get_inode(const struct mnt_idmap *idmap, struct super_block *sb, struct inode *dir, umode_t mode, dev_t dev, vma_flags_t flags) { @@ -4019,7 +4019,7 @@ static int shmem_statfs(struct dentry *dentry, struct kstatfs *buf) * File creation. Allocate an inode, and we're done.. */ static int -shmem_mknod(struct mnt_idmap *idmap, struct inode *dir, +shmem_mknod(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *dentry, umode_t mode, dev_t dev) { struct inode *inode; @@ -4058,7 +4058,7 @@ out_iput: } static int -shmem_tmpfile(struct mnt_idmap *idmap, struct inode *dir, +shmem_tmpfile(const struct mnt_idmap *idmap, struct inode *dir, struct file *file, umode_t mode) { struct inode *inode; @@ -4086,7 +4086,7 @@ out_iput: return error; } -static struct dentry *shmem_mkdir(struct mnt_idmap *idmap, struct inode *dir, +static struct dentry *shmem_mkdir(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *dentry, umode_t mode) { int error; @@ -4098,7 +4098,7 @@ static struct dentry *shmem_mkdir(struct mnt_idmap *idmap, struct inode *dir, return NULL; } -static int shmem_create(struct mnt_idmap *idmap, struct inode *dir, +static int shmem_create(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *dentry, umode_t mode) { return shmem_mknod(idmap, dir, dentry, mode | S_IFREG, 0); @@ -4171,7 +4171,7 @@ static int shmem_rmdir(struct inode *dir, struct dentry *dentry) return shmem_unlink(dir, dentry); } -static int shmem_whiteout(struct mnt_idmap *idmap, +static int shmem_whiteout(const struct mnt_idmap *idmap, struct inode *old_dir, struct dentry *old_dentry) { struct dentry *whiteout; @@ -4192,7 +4192,7 @@ static int shmem_whiteout(struct mnt_idmap *idmap, * it exists so that the VFS layer correctly free's it when it * gets overwritten. */ -static int shmem_rename2(struct mnt_idmap *idmap, +static int shmem_rename2(const struct mnt_idmap *idmap, struct inode *old_dir, struct dentry *old_dentry, struct inode *new_dir, struct dentry *new_dentry, unsigned int flags) @@ -4248,7 +4248,7 @@ static int shmem_rename2(struct mnt_idmap *idmap, return 0; } -static int shmem_symlink(struct mnt_idmap *idmap, struct inode *dir, +static int shmem_symlink(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *dentry, const char *symname) { int error; @@ -4360,7 +4360,7 @@ static int shmem_fileattr_get(struct dentry *dentry, struct file_kattr *fa) return 0; } -static int shmem_fileattr_set(struct mnt_idmap *idmap, +static int shmem_fileattr_set(const struct mnt_idmap *idmap, struct dentry *dentry, struct file_kattr *fa) { struct inode *inode = d_inode(dentry); @@ -4465,7 +4465,7 @@ static int shmem_xattr_handler_get(const struct xattr_handler *handler, } static int shmem_xattr_handler_set(const struct xattr_handler *handler, - struct mnt_idmap *idmap, + const struct mnt_idmap *idmap, struct dentry *unused, struct inode *inode, const char *name, const void *value, size_t size, int flags) @@ -6004,7 +6004,7 @@ static inline void shmem_unacct_size(unsigned long flags, loff_t size) { } -static inline struct inode *shmem_get_inode(struct mnt_idmap *idmap, +static inline struct inode *shmem_get_inode(const struct mnt_idmap *idmap, struct super_block *sb, struct inode *dir, umode_t mode, dev_t dev, vma_flags_t flags) { diff --git a/mm/shmem_quota.c b/mm/shmem_quota.c index d0b92d6da50f..a3fa4bd93e75 100644 --- a/mm/shmem_quota.c +++ b/mm/shmem_quota.c @@ -30,6 +30,7 @@ #include <linux/slab.h> #include <linux/rbtree.h> #include <linux/shmem_fs.h> +#include <linux/magic.h> #include <linux/quotaops.h> #include <linux/quota.h> @@ -54,6 +55,10 @@ struct quota_id { static int shmem_check_quota_file(struct super_block *sb, int type) { + /* Verify enabling happens on tmpfs superblock */ + if (sb->s_magic != TMPFS_MAGIC) + return 0; + /* There is no real quota file, nothing to do */ return 1; } diff --git a/mm/userfaultfd.c b/mm/userfaultfd.c index 3c7fd39deb13..81967c63fc17 100644 --- a/mm/userfaultfd.c +++ b/mm/userfaultfd.c @@ -4768,12 +4768,12 @@ static int new_userfaultfd(int flags) anon_inode_create_getfile("[userfaultfd]", &userfaultfd_fops, ctx, O_RDONLY | (flags & UFFD_SHARED_FCNTL_FLAGS), NULL)); - if (fdf.err) - return fdf.err; + if (fdf->fd < 0) + return fdf->fd; /* prevent the mm struct to be freed */ mmgrab(ctx->mm); - fd_prepare_file(fdf)->f_mode |= FMODE_NOWAIT; + fdf->file->f_mode |= FMODE_NOWAIT; retain_and_null_ptr(ctx); return fd_publish(fdf); } diff --git a/net/core/scm.c b/net/core/scm.c index f0d44ecdb11f..15c330784a69 100644 --- a/net/core/scm.c +++ b/net/core/scm.c @@ -364,14 +364,14 @@ int scm_recv_one_fd(struct file *f, int __user *ufd, unsigned int flags, return notrunc ? put_user(error, ufd) : error; FD_PREPARE(fdf, flags, get_file(f)); - if (fdf.err) - return fdf.err; + if (fdf->fd < 0) + return fdf->fd; - error = put_user(fd_prepare_fd(fdf), ufd); + error = put_user(fdf->fd, ufd); if (error) return error; - __receive_sock(fd_prepare_file(fdf)); + __receive_sock(fdf->file); return fd_publish(fdf); } diff --git a/net/handshake/netlink.c b/net/handshake/netlink.c index 3fd4fef9bab1..73b8314d9010 100644 --- a/net/handshake/netlink.c +++ b/net/handshake/netlink.c @@ -107,17 +107,17 @@ int handshake_nl_accept_doit(struct sk_buff *skb, struct genl_info *info) req = handshake_req_next(hn, class); if (req) { FD_PREPARE(fdf, O_CLOEXEC, req->hr_file); - if (fdf.err) { + if (fdf->fd < 0) { fput(req->hr_file); /* drop ref from handshake_req_next() */ - err = fdf.err; + err = fdf->fd; goto out_complete; } - err = req->hr_proto->hp_accept(req, info, fd_prepare_fd(fdf)); + err = req->hr_proto->hp_accept(req, info, fdf->fd); if (err) goto out_complete; /* Automatic cleanup handles fput */ - trace_handshake_cmd_accept(net, req, req->hr_sk, fd_prepare_fd(fdf)); + trace_handshake_cmd_accept(net, req, req->hr_sk, fdf->fd); fd_publish(fdf); return 0; } diff --git a/net/kcm/kcmsock.c b/net/kcm/kcmsock.c index 71af69d442f2..962ee2c4acd2 100644 --- a/net/kcm/kcmsock.c +++ b/net/kcm/kcmsock.c @@ -1580,10 +1580,10 @@ static int kcm_ioctl(struct socket *sock, unsigned int cmd, unsigned long arg) struct kcm_clone info; FD_PREPARE(fdf, 0, kcm_clone(sock)); - if (fdf.err) - return fdf.err; + if (fdf->fd < 0) + return fdf->fd; - info.fd = fd_prepare_fd(fdf); + info.fd = fdf->fd; if (copy_to_user((void __user *)arg, &info, sizeof(info))) return -EFAULT; diff --git a/net/socket.c b/net/socket.c index c05d86e63abf..c0ab6a3ad8ed 100644 --- a/net/socket.c +++ b/net/socket.c @@ -422,7 +422,7 @@ static const struct xattr_handler sockfs_xattr_handler = { }; static int sockfs_security_xattr_set(const struct xattr_handler *handler, - struct mnt_idmap *idmap, + const struct mnt_idmap *idmap, struct dentry *dentry, struct inode *inode, const char *suffix, const void *value, size_t size, int flags) @@ -447,7 +447,7 @@ static int sockfs_user_xattr_get(const struct xattr_handler *handler, } static int sockfs_user_xattr_set(const struct xattr_handler *handler, - struct mnt_idmap *idmap, + const struct mnt_idmap *idmap, struct dentry *dentry, struct inode *inode, const char *suffix, const void *value, size_t size, int flags) @@ -670,7 +670,7 @@ static ssize_t sockfs_listxattr(struct dentry *dentry, char *buffer, return used; } -static int sockfs_setattr(struct mnt_idmap *idmap, +static int sockfs_setattr(const struct mnt_idmap *idmap, struct dentry *dentry, struct iattr *iattr) { int err = simple_setattr(&nop_mnt_idmap, dentry, iattr); diff --git a/net/sunrpc/svc_xprt.c b/net/sunrpc/svc_xprt.c index 40040af588fb..77e28dcc4d2a 100644 --- a/net/sunrpc/svc_xprt.c +++ b/net/sunrpc/svc_xprt.c @@ -64,8 +64,10 @@ static LIST_HEAD(svc_xprt_class_list); * - Can be set or cleared at any time. * - After a set, svc_xprt_enqueue must be called to enqueue * the transport for processing. - * - After a clear, the transport must be read/accepted. - * If this succeeds, it must be set again. + * - After clearing XPT_CONN, the transport must be + * accepted. If this succeeds, the bit must be set again. + * - xpo_recvfrom decides when XPT_DATA is cleared; see + * svc_xprt_received. * XPT_CLOSE: * - Can set at any time. It is never cleared. * XPT_DEAD: @@ -218,8 +220,9 @@ EXPORT_SYMBOL_GPL(svc_xprt_init); * The caller must hold the XPT_BUSY bit and must * not thereafter touch transport data. * - * Note: XPT_DATA only gets cleared when a read-attempt finds no (or - * insufficient) data. + * Note: xpo_recvfrom decides when to clear XPT_DATA. A transport may + * leave the bit set until a read attempt finds no (or insufficient) + * data, or clear it as soon as it consumes the last queued receive. */ void svc_xprt_received(struct svc_xprt *xprt) { @@ -475,11 +478,11 @@ static bool svc_xprt_ready(struct svc_xprt *xprt) /* * If another cpu has recently updated xpt_flags, - * sk_sock->flags, xpt_reserved, or xpt_nr_rqsts, we need to - * know about it; otherwise it's possible that both that cpu and - * this one could call svc_xprt_enqueue() without either - * svc_xprt_enqueue() recognizing that the conditions below - * are satisfied, and we could stall indefinitely: + * sk_sock->flags, xpt_reserved (UDP only), or xpt_nr_rqsts, + * we need to know about it; otherwise it's possible that both + * that cpu and this one could call svc_xprt_enqueue() without + * either svc_xprt_enqueue() recognizing that the conditions + * below are satisfied, and we could stall indefinitely: */ smp_rmb(); xpt_flags = READ_ONCE(xprt->xpt_flags); @@ -551,6 +554,10 @@ static struct svc_xprt *svc_xprt_dequeue(struct svc_pool *pool) * to make sure the reply fits. This function reduces that reserved * space to be the amount of space used already, plus @space. * + * The transport's reservation is tracked only on classes that set + * SVC_XPRT_FLAG_WSPACE_RESERVE. On the others, only @rqstp's + * reservation is updated. + * */ void svc_reserve(struct svc_rqst *rqstp, int space) { @@ -559,10 +566,12 @@ void svc_reserve(struct svc_rqst *rqstp, int space) space += rqstp->rq_res.head[0].iov_len; if (xprt && space < rqstp->rq_reserved) { - atomic_sub((rqstp->rq_reserved - space), - &xprt->xpt_reserved); + if (xprt->xpt_class->xcl_flags & SVC_XPRT_FLAG_WSPACE_RESERVE) { + atomic_sub((rqstp->rq_reserved - space), + &xprt->xpt_reserved); + svc_xprt_resource_released(xprt); + } rqstp->rq_reserved = space; - svc_xprt_resource_released(xprt); } } EXPORT_SYMBOL_GPL(svc_reserve); @@ -869,7 +878,8 @@ static void svc_handle_xprt(struct svc_rqst *rqstp, struct svc_xprt *xprt) else len = xprt->xpt_ops->xpo_recvfrom(rqstp); rqstp->rq_reserved = serv->sv_max_mesg; - atomic_add(rqstp->rq_reserved, &xprt->xpt_reserved); + if (xprt->xpt_class->xcl_flags & SVC_XPRT_FLAG_WSPACE_RESERVE) + atomic_add(rqstp->rq_reserved, &xprt->xpt_reserved); if (len <= 0) goto out; diff --git a/net/sunrpc/svcauth_unix.c b/net/sunrpc/svcauth_unix.c index 31a1bc60a5f6..a68fe44d1f62 100644 --- a/net/sunrpc/svcauth_unix.c +++ b/net/sunrpc/svcauth_unix.c @@ -540,6 +540,13 @@ static int unix_gid_parse(struct cache_detail *cd, if (ugp) { struct cache_head *ch; ug.h.flags = 0; + /* + * mountd sends at least the user's primary group on + * success, so an empty list can only mean the lookup + * failed. Keep the credential's own groups instead. + */ + if (gids == 0) + set_bit(CACHE_NEGATIVE, &ug.h.flags); ug.h.expiry_time = expiry; ch = sunrpc_cache_update(cd, &ug.h, &ugp->h, @@ -730,6 +737,8 @@ static int sunrpc_nl_parse_one_unix_gid(struct cache_detail *cd, boot.tv_sec; if (tb[SUNRPC_A_UNIX_GID_NEGATIVE]) { + /* failed lookup: keep the credential's own groups */ + set_bit(CACHE_NEGATIVE, &ug.h.flags); ug.gi = groups_alloc(0); if (!ug.gi) return -ENOMEM; diff --git a/net/sunrpc/svcsock.c b/net/sunrpc/svcsock.c index 50e5e7f5b762..e5459d504b6a 100644 --- a/net/sunrpc/svcsock.c +++ b/net/sunrpc/svcsock.c @@ -8,15 +8,6 @@ * evenly when servicing a single client. May need to modify the * svc_xprt_enqueue procedure... * - * TCP support is largely untested and may be a little slow. The problem - * is that we currently do two separate recvfrom's, one for the 4-byte - * record length, and the second for the actual record. This could possibly - * be improved by always reading a minimum size of around 100 bytes and - * tucking any superfluous bytes away in a temporary store. Still, that - * leaves write requests out in the rain. An alternative may be to peek at - * the first skb in the queue, and if it matches the next TCP sequence - * number, to extract the record marker. Yuck. - * * Copyright (C) 1995, 1996 Olaf Kirch <okir@monad.swb.de> */ @@ -238,136 +229,108 @@ static int svc_one_sock_name(struct svc_sock *svsk, char *buf, int remaining) return len; } -static int -svc_tcp_sock_process_cmsg(struct socket *sock, struct msghdr *msg, - struct cmsghdr *cmsg, int ret) -{ - u8 content_type = tls_get_record_type(sock->sk, cmsg); - u8 level, description; - - switch (content_type) { - case 0: - break; - case TLS_RECORD_TYPE_DATA: - /* TLS sets EOR at the end of each application data - * record, even though there might be more frames - * waiting to be decrypted. - */ - msg->msg_flags &= ~MSG_EOR; - break; - case TLS_RECORD_TYPE_ALERT: - tls_alert_recv(sock->sk, msg, &level, &description); - ret = (level == TLS_ALERT_LEVEL_FATAL) ? - -ENOTCONN : -EAGAIN; - break; - default: - /* discard this record type */ - ret = -EAGAIN; - } - return ret; -} - -static int -svc_tcp_sock_recv_cmsg(struct socket *sock, unsigned int *msg_flags) +/* + * The ->read_sock data path invokes neither security_socket_recvmsg() + * nor the sock:sock_recv_length tracepoint. Dispatch ->recvmsg + * directly so the whole receive path behaves one way. + */ +static int svc_tcp_recv_cmsg(struct socket *sock, int flags, + struct kvec *payload, u8 *type, + unsigned int *msg_flags) { + const struct proto_ops *ops = READ_ONCE(sock->ops); union { struct cmsghdr cmsg; u8 buf[CMSG_SPACE(sizeof(u8))]; - } u; - u8 alert[2]; - struct kvec alert_kvec = { - .iov_base = alert, - .iov_len = sizeof(alert), - }; + } u = {}; struct msghdr msg = { - .msg_flags = *msg_flags, - .msg_control = &u, - .msg_controllen = sizeof(u), + .msg_control = &u, + .msg_controllen = sizeof(u), }; int ret; - iov_iter_kvec(&msg.msg_iter, ITER_DEST, &alert_kvec, 1, - alert_kvec.iov_len); - ret = sock_recvmsg(sock, &msg, MSG_DONTWAIT); - if (ret > 0 && - tls_get_record_type(sock->sk, &u.cmsg) == TLS_RECORD_TYPE_ALERT) { - iov_iter_revert(&msg.msg_iter, ret); - ret = svc_tcp_sock_process_cmsg(sock, &msg, &u.cmsg, -EAGAIN); - } + iov_iter_kvec(&msg.msg_iter, ITER_DEST, payload, 1, payload->iov_len); + ret = ops->recvmsg(sock, &msg, msg_data_left(&msg), flags); + if (ret < 0) + return ret; + *msg_flags = msg.msg_flags; + *type = tls_get_record_type(sock->sk, &u.cmsg); + if (!*type && ret) + return -EBADMSG; return ret; } -static int -svc_tcp_sock_recvmsg(struct svc_sock *svsk, struct msghdr *msg) +static int svc_tcp_recv_ctrl_record(struct svc_sock *svsk) { - int ret; + u8 alert[2], type, level, description; struct socket *sock = svsk->sk_sock; - - ret = sock_recvmsg(sock, msg, MSG_DONTWAIT); - if (msg->msg_flags & MSG_CTRUNC) { - msg->msg_flags &= ~(MSG_CTRUNC | MSG_EOR); - if (ret == 0 || ret == -EIO) - ret = svc_tcp_sock_recv_cmsg(sock, &msg->msg_flags); - } - return ret; -} - -#if ARCH_IMPLEMENTS_FLUSH_DCACHE_PAGE -static void svc_flush_bvec(const struct bio_vec *bvec, size_t size, size_t seek) -{ - struct bvec_iter bi = { - .bi_size = size + seek, + struct kvec recv_kvec = { + .iov_base = alert, + .iov_len = sizeof(alert), }; - struct bio_vec bv; - - bvec_iter_advance(bvec, &bi, seek & PAGE_MASK); - for_each_bvec(bv, bvec, bi, bi) - flush_dcache_page(bv.bv_page); -} -#else -static inline void svc_flush_bvec(const struct bio_vec *bvec, size_t size, - size_t seek) -{ -} -#endif - -/* - * Read from @rqstp's transport socket. The incoming message fills whole - * pages in @rqstp's rq_pages array until the last page of the message - * has been received into a partial page. - */ -static ssize_t svc_tcp_read_msg(struct svc_rqst *rqstp, size_t buflen, - size_t seek) -{ - struct svc_sock *svsk = - container_of(rqstp->rq_xprt, struct svc_sock, sk_xprt); - struct bio_vec *bvec = rqstp->rq_bvec; - struct msghdr msg = { NULL }; - unsigned int i; - ssize_t len; - size_t t; - - clear_bit(XPT_DATA, &svsk->sk_xprt.xpt_flags); + unsigned int msg_flags; + struct msghdr msg = {}; + int ret; - for (i = 0, t = 0; t < buflen; i++, t += PAGE_SIZE) - bvec_set_page(&bvec[i], rqstp->rq_pages[i], PAGE_SIZE, 0); + if (!test_bit(XPT_TLS_SESSION, &svsk->sk_xprt.xpt_flags)) + return 0; - iov_iter_bvec(&msg.msg_iter, ITER_DEST, bvec, i, buflen); - if (seek) { - iov_iter_advance(&msg.msg_iter, seek); - buflen -= seek; + /* A data record can become ready between ->read_sock returning + * and this probe. A plain receive would take two octets of it + * as RPC payload, so peek. + */ + ret = svc_tcp_recv_cmsg(sock, MSG_DONTWAIT | MSG_PEEK, + &recv_kvec, &type, &msg_flags); + if (ret == -EAGAIN || (!ret && !type)) + return 0; + if (ret < 0) + return ret; + if (type == TLS_RECORD_TYPE_DATA) { + /* The peek parks the decrypted record on ctx->rx_list, + * where it draws no further data_ready. Re-arm or the + * RPC hangs until the client times out. + */ + set_bit(XPT_DATA, &svsk->sk_xprt.xpt_flags); + return 0; } - len = svc_tcp_sock_recvmsg(svsk, &msg); - if (len > 0) - svc_flush_bvec(bvec, len, seek); + if (type != TLS_RECORD_TYPE_ALERT) + return -EPROTO; + + ret = svc_tcp_recv_cmsg(sock, MSG_DONTWAIT, &recv_kvec, &type, + &msg_flags); + /* The peek found a record at the head, so an -EAGAIN here is + * spurious. Propagating it strands the record with no later + * announcement, so return -EBADMSG, which closes the transport. + */ + if (ret == -EAGAIN) + return -EBADMSG; + if (ret < 0) + return ret; + + /* An Alert record carries exactly one two-octet message (RFC + * 8446 Section 5.1). recv_kvec caps the receive at two, so a + * longer record produces the same count. MSG_EOR appears only + * once kTLS has drained the whole record. + */ + if (ret != sizeof(alert) || !(msg_flags & MSG_EOR)) + return -EBADMSG; + + iov_iter_kvec(&msg.msg_iter, ITER_DEST, &recv_kvec, 1, + recv_kvec.iov_len); + tls_alert_recv(sock->sk, &msg, &level, &description); - /* If we read a full record, then assume there may be more - * data to read (stream based sockets only!) + /* RFC 8446 Section 6: every alert but a closure alert is + * an error alert. kTLS raises no data_ready for records it + * already holds, so re-arm for what sits behind the alert. */ - if (len == buflen) + switch (description) { + case TLS_ALERT_DESC_CLOSE_NOTIFY: + case TLS_ALERT_DESC_USER_CANCELED: set_bit(XPT_DATA, &svsk->sk_xprt.xpt_flags); - - return len; + return -EAGAIN; + default: + return -ENOTCONN; + } } /* @@ -837,6 +800,7 @@ static struct svc_xprt_class svc_udp_class = { .xcl_ops = &svc_udp_ops, .xcl_max_payload = RPCSVC_MAXPAYLOAD_UDP, .xcl_ident = XPRT_TRANSPORT_UDP, + .xcl_flags = SVC_XPRT_FLAG_WSPACE_RESERVE, }; static void svc_udp_init(struct svc_sock *svsk, struct svc_serv *serv) @@ -992,14 +956,14 @@ failed: return NULL; } -static size_t svc_tcp_restore_pages(struct svc_sock *svsk, - struct svc_rqst *rqstp) +static void svc_tcp_restore_pages(struct svc_sock *svsk, + struct svc_rqst *rqstp) { size_t len = svsk->sk_datalen; unsigned int i, npages; if (!len) - return 0; + return; npages = (len + PAGE_SIZE - 1) >> PAGE_SHIFT; for (i = 0; i < npages; i++) { if (rqstp->rq_pages[i] != NULL) @@ -1009,7 +973,6 @@ static size_t svc_tcp_restore_pages(struct svc_sock *svsk, svsk->sk_pages[i] = NULL; } rqstp->rq_arg.head[0].iov_base = page_address(rqstp->rq_pages[0]); - return len; } static void svc_tcp_save_pages(struct svc_sock *svsk, struct svc_rqst *rqstp) @@ -1048,50 +1011,6 @@ out: svsk->sk_datalen = 0; } -/* - * Receive fragment record header into sk_marker. - */ -static ssize_t svc_tcp_read_marker(struct svc_sock *svsk, - struct svc_rqst *rqstp) -{ - ssize_t want, len; - - /* If we haven't gotten the record length yet, - * get the next four bytes. - */ - if (svsk->sk_tcplen < sizeof(rpc_fraghdr)) { - struct msghdr msg = { NULL }; - struct kvec iov; - - want = sizeof(rpc_fraghdr) - svsk->sk_tcplen; - iov.iov_base = ((char *)&svsk->sk_marker) + svsk->sk_tcplen; - iov.iov_len = want; - iov_iter_kvec(&msg.msg_iter, ITER_DEST, &iov, 1, want); - len = svc_tcp_sock_recvmsg(svsk, &msg); - if (len < 0) - return len; - svsk->sk_tcplen += len; - if (len < want) { - /* call again to read the remaining bytes */ - goto err_short; - } - trace_svcsock_marker(&svsk->sk_xprt, svsk->sk_marker); - if (svc_sock_reclen(svsk) + svsk->sk_datalen > - svsk->sk_xprt.xpt_server->sv_max_mesg) - goto err_too_large; - } - return svc_sock_reclen(svsk); - -err_too_large: - net_notice_ratelimited("svc: %s oversized RPC fragment (%u octets) from %pISpc\n", - svsk->sk_xprt.xpt_server->sv_name, - svc_sock_reclen(svsk), - (struct sockaddr *)&svsk->sk_xprt.xpt_remote); - svc_xprt_deferred_close(&svsk->sk_xprt); -err_short: - return -EAGAIN; -} - static int receive_cb_reply(struct svc_sock *svsk, struct svc_rqst *rqstp) { struct rpc_xprt *bc_xprt = svsk->sk_xprt.xpt_bc_xprt; @@ -1129,20 +1048,162 @@ unlock_eagain: static void svc_tcp_fragment_received(struct svc_sock *svsk) { - /* If we have more data, signal svc_xprt_enqueue() to try again */ svsk->sk_tcplen = 0; svsk->sk_marker = xdr_zero; } +/* + * A non-final fragment carries four octets of marker and may carry + * no payload at all. sk_datalen advances only by the payload, so a + * run of tiny fragments exhausts ->read_sock's byte budget before + * the sv_max_mesg check trips, and a run of empty ones never trips + * it. Cap the fragments per socket-lock hold. The cap leaves the + * record incomplete, and svc_tcp_recvfrom() resumes it on the next + * call. + */ +#define SVC_TCP_MAX_FRAGS 256 + +struct svc_tcp_recv_ctx { + struct svc_rqst *rqstp; + unsigned int frags; + bool complete; +}; + +/* + * Nothing reads the message body before the message is complete, and + * partial receives refill the same pages. Flush once here, after the + * socket lock is released, rather than once per copy in the actor. + */ +static void svc_tcp_flush_pages(struct svc_sock *svsk, + struct svc_rqst *rqstp) +{ + unsigned int pg, pages = DIV_ROUND_UP(svsk->sk_datalen, PAGE_SIZE); + + for (pg = 0; pg < pages; pg++) + flush_dcache_page(rqstp->rq_pages[pg]); +} + +/* + * Mapping the message's unfilled remainder would re-map untouched + * pages on every call, at a cost that grows with the message rather + * than with the octets copied. + */ +static void svc_tcp_recv_iter_init(struct svc_rqst *rqstp, + struct iov_iter *iter, size_t body_off, + size_t len) +{ + unsigned int first = body_off >> PAGE_SHIFT; + size_t seek = offset_in_page(body_off); + unsigned int i, pages = DIV_ROUND_UP(seek + len, PAGE_SIZE); + + for (i = 0; i < pages; i++) + bvec_set_page(&rqstp->rq_bvec[i], rqstp->rq_pages[first + i], + PAGE_SIZE, 0); + + iov_iter_bvec(iter, ITER_DEST, rqstp->rq_bvec, pages, seek + len); + iov_iter_advance(iter, seek); +} + +/* + * ->read_sock actor, called under the socket lock. sk_datalen is both + * the count of body octets received so far and their write offset into + * rq_pages. + */ +static int svc_tcp_recv_actor(read_descriptor_t *desc, struct sk_buff *skb, + unsigned int offset, size_t len) +{ + struct svc_tcp_recv_ctx *ctx = desc->arg.data; + struct svc_rqst *rqstp = ctx->rqstp; + struct svc_sock *svsk = + container_of(rqstp->rq_xprt, struct svc_sock, sk_xprt); + size_t reclen, received, want, take, n; + size_t consumed = 0; + + if (!desc->count) + return 0; + + len = min(len, desc->count); + + if (svsk->sk_tcplen < sizeof(rpc_fraghdr)) { + want = sizeof(rpc_fraghdr) - svsk->sk_tcplen; + n = min(want, len); + + if (skb_copy_bits(skb, offset, + (char *)&svsk->sk_marker + svsk->sk_tcplen, + n)) + goto fault; + svsk->sk_tcplen += n; + offset += n; + len -= n; + consumed += n; + desc->count -= n; + + if (svsk->sk_tcplen < sizeof(rpc_fraghdr)) + return consumed; + + trace_svcsock_marker(&svsk->sk_xprt, svsk->sk_marker); + if (svc_sock_reclen(svsk) + svsk->sk_datalen > + svsk->sk_xprt.xpt_server->sv_max_mesg) { + net_notice_ratelimited("svc: %s oversized RPC fragment (%u octets) from %pISpc\n", + svsk->sk_xprt.xpt_server->sv_name, + svc_sock_reclen(svsk), + (struct sockaddr *)&svsk->sk_xprt.xpt_remote); + desc->error = -EMSGSIZE; + desc->count = 0; + return consumed; + } + } + + reclen = svc_sock_reclen(svsk); + received = svsk->sk_tcplen - sizeof(rpc_fraghdr); + want = reclen - received; + take = min(want, len); + + if (take) { + struct iov_iter iter; + + svc_tcp_recv_iter_init(rqstp, &iter, svsk->sk_datalen, take); + if (skb_copy_datagram_iter(skb, offset, &iter, take)) + goto fault; + svsk->sk_datalen += take; + svsk->sk_tcplen += take; + consumed += take; + desc->count -= take; + } + + if (take == want) { + if (svc_sock_final_rec(svsk)) { + ctx->complete = true; + desc->count = 0; + } else { + svc_tcp_fragment_received(svsk); + if (++ctx->frags >= SVC_TCP_MAX_FRAGS) + desc->count = 0; + } + } + + return consumed; + +fault: + desc->error = -EFAULT; + desc->count = 0; + return consumed; +} + +static bool svc_tcp_at_urg_mark(struct sock *sk) +{ + const struct tcp_sock *tp = tcp_sk(sk); + + return tp->urg_data && tp->urg_seq == tp->copied_seq; +} + /** * svc_tcp_recvfrom - Receive data from a TCP socket * @rqstp: request structure into which to receive an RPC Call * * Called in a loop when XPT_DATA has been set. * - * Read the 4-byte stream record marker, then use the record length - * in that marker to set up exactly the resources needed to receive - * the next RPC message into @rqstp. + * Context: Process context. Takes and releases the socket lock. * * Returns: * On success, the number of bytes in a received RPC Call, or @@ -1157,29 +1218,69 @@ static int svc_tcp_recvfrom(struct svc_rqst *rqstp) struct svc_sock *svsk = container_of(rqstp->rq_xprt, struct svc_sock, sk_xprt); struct svc_serv *serv = svsk->sk_xprt.xpt_server; - size_t want, base; + struct svc_tcp_recv_ctx ctx = { + .rqstp = rqstp, + }; + read_descriptor_t desc = { + .arg.data = &ctx, + .count = serv->sv_max_mesg + sizeof(rpc_fraghdr), + }; + struct socket *sock = svsk->sk_sock; ssize_t len; __be32 *p; __be32 calldir; clear_bit(XPT_DATA, &svsk->sk_xprt.xpt_flags); - len = svc_tcp_read_marker(svsk, rqstp); - if (len < 0) - goto error; + svc_tcp_restore_pages(svsk, rqstp); - base = svc_tcp_restore_pages(svsk, rqstp); - want = len - (svsk->sk_tcplen - sizeof(rpc_fraghdr)); - len = svc_tcp_read_msg(rqstp, base + want, base); - if (len >= 0) { - trace_svcsock_tcp_recv(&svsk->sk_xprt, len); - svsk->sk_tcplen += len; - svsk->sk_datalen += len; + lock_sock(sock->sk); + len = sock->ops->read_sock(sock->sk, &desc, svc_tcp_recv_actor); + /* ->read_sock stops at urgent data and consumes none of it. + * Only recvmsg() clears the condition, and this path calls + * none, so every later read stops at the same octet. An RPC + * stream carries no urgent data, so close the connection. + * + * The read that first reaches the mark consumes the octets + * ahead of it, so the stop does not show up as a zero len. A + * record completed ahead of the mark is returned first. The + * XPT_DATA set below brings the next call back here with + * nothing left to consume. + */ + if (!ctx.complete && svc_tcp_at_urg_mark(sock->sk)) + desc.error = -EPROTO; + release_sock(sock->sk); + + /* ->read_sock returns the octets consumed before an actor + * failure, so a positive len can accompany desc.error. + */ + if (desc.error < 0) { + len = desc.error; + goto err_nuts; } - if (len != want || !svc_sock_final_rec(svsk)) + if (len >= 0) + trace_svcsock_tcp_recv(&svsk->sk_xprt, len); + + if (!ctx.complete) { + if (!desc.count) { + set_bit(XPT_DATA, &svsk->sk_xprt.xpt_flags); + goto err_incomplete; + } + /* A zero return leaves no record at the head to classify. + * -EINVAL means a control record sits there. Screen the + * other errors out first, because a probe calls + * sock_error(), whose xchg clears sk->sk_err as it reads. + */ + if (len <= 0 && len != -EINVAL) + goto err_incomplete; + + len = svc_tcp_recv_ctrl_record(svsk); goto err_incomplete; + } if (svsk->sk_datalen < 8) goto err_nuts; + svc_tcp_flush_pages(svsk, rqstp); + rqstp->rq_arg.len = svsk->sk_datalen; rqstp->rq_arg.page_base = 0; if (rqstp->rq_arg.len <= rqstp->rq_arg.head[0].iov_len) { @@ -1195,6 +1296,12 @@ static int svc_tcp_recvfrom(struct svc_rqst *rqstp) else clear_bit(RQ_LOCAL, &rqstp->rq_flags); + /* Completing one message stops ->read_sock with whatever + * follows still queued, and no path from here re-arms XPT_DATA. + * The queued message would wait for unrelated traffic. + */ + set_bit(XPT_DATA, &svsk->sk_xprt.xpt_flags); + p = (__be32 *)rqstp->rq_arg.head[0].iov_base; calldir = p[1]; if (calldir) @@ -1219,19 +1326,21 @@ err_incomplete: svc_tcp_save_pages(svsk, rqstp); if (len < 0 && len != -EAGAIN) goto err_delete; - if (len == want) - svc_tcp_fragment_received(svsk); - else + if (svsk->sk_tcplen >= sizeof(rpc_fraghdr)) trace_svcsock_tcp_recv_short(&svsk->sk_xprt, svc_sock_reclen(svsk), svsk->sk_tcplen - sizeof(rpc_fraghdr)); + else + trace_svcsock_tcp_recv_eagain(&svsk->sk_xprt, 0); goto err_noclose; error: - if (len != -EAGAIN) - goto err_delete; trace_svcsock_tcp_recv_eagain(&svsk->sk_xprt, 0); goto err_noclose; err_nuts: + /* svc_tcp_save_pages() has not run, so svsk->sk_pages[] is + * empty. A non-zero sk_datalen makes the teardown-time + * svc_tcp_clear_pages() walk empty slots and WARN. + */ svsk->sk_datalen = 0; err_delete: trace_svcsock_tcp_recv_err(&svsk->sk_xprt, len); @@ -1541,6 +1650,9 @@ int svc_addsock(struct svc_serv *serv, struct net *net, const int fd, err = -EISCONN; if (so->state > SS_UNCONNECTED) goto out; + err = -EBUSY; + if (so->sk->sk_user_data) + goto out; err = -ENOENT; if (!try_module_get(THIS_MODULE)) goto out; diff --git a/net/sunrpc/xprtrdma/svc_rdma_recvfrom.c b/net/sunrpc/xprtrdma/svc_rdma_recvfrom.c index fdfed1be97da..d029bcb7a5c0 100644 --- a/net/sunrpc/xprtrdma/svc_rdma_recvfrom.c +++ b/net/sunrpc/xprtrdma/svc_rdma_recvfrom.c @@ -925,8 +925,9 @@ static noinline void svc_rdma_read_complete(struct svc_rqst *rqstp, * %-ENOTCONN if posting failed (connection is lost), * %-EIO if rdma_rw initialization failed (DMA mapping, etc). * - * Called in a loop when XPT_DATA is set. XPT_DATA is cleared only - * when there are no remaining ctxt's to process. + * Called in a loop when XPT_DATA is set. XPT_DATA is cleared as + * soon as both receive queues are empty, so a consumed ctxt does + * not leave a stale bit behind. * * The next ctxt is removed from the "receive" lists. * @@ -960,6 +961,15 @@ int svc_rdma_recvfrom(struct svc_rqst *rqstp) ctxt = svc_rdma_next_recv_ctxt(&rdma_xprt->sc_read_complete_q); if (ctxt) { list_del(&ctxt->rc_list); + /* Producers add to these queues and set XPT_DATA under + * this lock, so the clear cannot race one. The clear can + * drop the XPT_DATA that svc_rdma_send_ctxt_put() sets + * for sc_send_release_list. svc_xprt_release() drains + * that list before this thread looks for more work. + */ + if (list_empty(&rdma_xprt->sc_read_complete_q) && + list_empty(&rdma_xprt->sc_rq_dto_q)) + clear_bit(XPT_DATA, &xprt->xpt_flags); spin_unlock(&rdma_xprt->sc_rq_dto_lock); svc_xprt_received(xprt); svc_rdma_read_complete(rqstp, ctxt); @@ -968,8 +978,8 @@ int svc_rdma_recvfrom(struct svc_rqst *rqstp) ctxt = svc_rdma_next_recv_ctxt(&rdma_xprt->sc_rq_dto_q); if (ctxt) list_del(&ctxt->rc_list); - else - /* No new incoming requests, terminate the loop */ + /* sc_read_complete_q was empty above, under this same lock. */ + if (list_empty(&rdma_xprt->sc_rq_dto_q)) clear_bit(XPT_DATA, &xprt->xpt_flags); spin_unlock(&rdma_xprt->sc_rq_dto_lock); diff --git a/net/sunrpc/xprtrdma/svc_rdma_transport.c b/net/sunrpc/xprtrdma/svc_rdma_transport.c index 927269598ac2..f949601b2144 100644 --- a/net/sunrpc/xprtrdma/svc_rdma_transport.c +++ b/net/sunrpc/xprtrdma/svc_rdma_transport.c @@ -471,7 +471,6 @@ static struct svc_xprt *svc_rdma_accept(struct svc_xprt *xprt) newxprt->sc_max_requests = svcrdma_max_requests; newxprt->sc_max_bc_requests = svcrdma_max_bc_requests; newxprt->sc_recv_batch = RPCRDMA_MAX_RECV_BATCH; - newxprt->sc_fc_credits = cpu_to_be32(newxprt->sc_max_requests); /* Qualify the transport's resource defaults with the * capabilities of this particular device. @@ -492,6 +491,8 @@ static struct svc_xprt *svc_rdma_accept(struct svc_xprt *xprt) newxprt->sc_max_bc_requests = 2; } + newxprt->sc_fc_credits = cpu_to_be32(newxprt->sc_max_requests); + /* Estimate the needed number of rdma_rw contexts. The maximum * Read and Write chunks have one segment each. Each request * can involve one Read chunk and either a Write chunk or Reply diff --git a/net/sunrpc/xprtsock.c b/net/sunrpc/xprtsock.c index 7f60723fa64d..f5a5136327ba 100644 --- a/net/sunrpc/xprtsock.c +++ b/net/sunrpc/xprtsock.c @@ -357,35 +357,6 @@ xs_alloc_sparse_pages(struct xdr_buf *buf, size_t want, gfp_t gfp) } static int -xs_sock_process_cmsg(struct socket *sock, struct msghdr *msg, - unsigned int *msg_flags, struct cmsghdr *cmsg, int ret) -{ - u8 content_type = tls_get_record_type(sock->sk, cmsg); - u8 level, description; - - switch (content_type) { - case 0: - break; - case TLS_RECORD_TYPE_DATA: - /* TLS sets EOR at the end of each application data - * record, even though there might be more frames - * waiting to be decrypted. - */ - *msg_flags &= ~MSG_EOR; - break; - case TLS_RECORD_TYPE_ALERT: - tls_alert_recv(sock->sk, msg, &level, &description); - ret = (level == TLS_ALERT_LEVEL_FATAL) ? - -EACCES : -EAGAIN; - break; - default: - /* discard this record type */ - ret = -EAGAIN; - } - return ret; -} - -static int xs_sock_recv_cmsg(struct socket *sock, unsigned int *msg_flags, int flags) { union { @@ -402,16 +373,42 @@ xs_sock_recv_cmsg(struct socket *sock, unsigned int *msg_flags, int flags) .msg_control = &u, .msg_controllen = sizeof(u), }; + u8 level, description; int ret; iov_iter_kvec(&msg.msg_iter, ITER_DEST, &alert_kvec, 1, alert_kvec.iov_len); ret = sock_recvmsg(sock, &msg, flags); - if (ret > 0) { - if (tls_get_record_type(sock->sk, &u.cmsg) == TLS_RECORD_TYPE_ALERT) - iov_iter_revert(&msg.msg_iter, ret); - ret = xs_sock_process_cmsg(sock, &msg, msg_flags, &u.cmsg, - -EAGAIN); + /* put_cmsg() shrinks msg_controllen, so a short one means + * kTLS filled in u.cmsg. + */ + if (ret >= 0 && msg.msg_controllen < sizeof(u)) { + if (tls_get_record_type(sock->sk, &u.cmsg) != + TLS_RECORD_TYPE_ALERT) + return -EAGAIN; + /* RFC 8446 Section 5.1: a record with an Alert type carries + * exactly one message, and an alert is two octets. + * tls_alert_recv() reads both without checking the length. + * alert_kvec caps the count at two, so a longer record + * fills it as well. kTLS sets MSG_EOR only once the + * record has been drained. + */ + if (ret != sizeof(alert) || !(msg.msg_flags & MSG_EOR)) + return -EACCES; + iov_iter_revert(&msg.msg_iter, ret); + tls_alert_recv(sock->sk, &msg, &level, &description); + /* RFC 8446 Section 6: every alert but a closure alert is + * an error alert, whatever the legacy AlertLevel octet + * says. + */ + switch (description) { + case TLS_ALERT_DESC_CLOSE_NOTIFY: + case TLS_ALERT_DESC_USER_CANCELED: + ret = -EAGAIN; + break; + default: + ret = -EACCES; + } } return ret; } @@ -425,6 +422,10 @@ xs_sock_recvmsg(struct socket *sock, struct msghdr *msg, int flags, size_t seek) ret = sock_recvmsg(sock, msg, flags); /* Handle TLS inband control message lazily */ if (msg->msg_flags & MSG_CTRUNC) { + /* TLS sets EOR at the end of each application data + * record, even though there might be more frames + * waiting to be decrypted. + */ msg->msg_flags &= ~(MSG_CTRUNC | MSG_EOR); if (ret == 0 || ret == -EIO) ret = xs_sock_recv_cmsg(sock, &msg->msg_flags, flags); diff --git a/net/unix/af_unix.c b/net/unix/af_unix.c index 42cffeafc8c1..d3b459310523 100644 --- a/net/unix/af_unix.c +++ b/net/unix/af_unix.c @@ -1361,7 +1361,7 @@ static int unix_bind_bsd(struct sock *sk, struct sockaddr_un *sunaddr, struct unix_sock *u = unix_sk(sk); unsigned int new_hash, old_hash; struct net *net = sock_net(sk); - struct mnt_idmap *idmap; + const struct mnt_idmap *idmap; struct unix_address *addr; struct dentry *dentry; struct path parent; diff --git a/rust/kernel/configfs.rs b/rust/kernel/configfs.rs index cd082b83e9e7..04cea2a1894a 100644 --- a/rust/kernel/configfs.rs +++ b/rust/kernel/configfs.rs @@ -135,8 +135,9 @@ pub struct Subsystem<Data> { // SAFETY: We do not provide any operations on `Subsystem`. unsafe impl<Data> Sync for Subsystem<Data> {} -// SAFETY: Ownership of `Subsystem` can safely be transferred to other threads. -unsafe impl<Data> Send for Subsystem<Data> {} +// SAFETY: Ownership of `Subsystem` can safely be transferred to other threads +// if its data can be transferred as well. +unsafe impl<Data: Send> Send for Subsystem<Data> {} impl<Data> Subsystem<Data> { /// Create an initializer for a [`Subsystem`]. @@ -249,6 +250,13 @@ pub struct Group<Data> { data: Data, } +// SAFETY: We do not provide any operations on `Group`. +unsafe impl<Data> Sync for Group<Data> {} + +// SAFETY: Ownership of `Group` can safely be transferred to other threads if +// its data can be transferred as well. +unsafe impl<Data: Send> Send for Group<Data> {} + impl<Data> Group<Data> { /// Create an initializer for a new group. /// @@ -297,26 +305,36 @@ unsafe impl<Data> HasGroup<Data> for Group<Data> { /// /// `this` must be a valid pointer. /// -/// If `this` does not represent the root group of a configfs subsystem, -/// `this` must be a pointer to a `bindings::config_group` embedded in a -/// `Group<Parent>`. +/// If `this` is the `su_group` field of a `bindings::configfs_subsystem`, that +/// `configfs_subsystem` must be embedded in a `Subsystem<Parent>`. /// -/// Otherwise, `this` must be a pointer to a `bindings::config_group` that -/// is embedded in a `bindings::configfs_subsystem` that is embedded in a -/// `Subsystem<Parent>`. +/// Otherwise, `this` must be a pointer to a `bindings::config_group` embedded +/// in a `Group<Parent>`. unsafe fn get_group_data<'a, Parent>(this: *mut bindings::config_group) -> &'a Parent { // SAFETY: `this` is a valid pointer. - let is_root = unsafe { (*this).cg_subsys.is_null() }; - - if !is_root { - // SAFETY: By C API contact,`this` was returned from a call to - // `make_group`. The pointer is known to be embedded within a - // `Group<Parent>`. - unsafe { &(*Group::<Parent>::container_of(this)).data } - } else { - // SAFETY: By C API contract, `this` is a pointer to the - // `bindings::config_group` field within a `Subsystem<Parent>`. + let subsys = unsafe { (*this).cg_subsys }; + // `link_group()` in `fs/configfs/dir.c` assigns `cg_subsys` for every + // `config_group` attached anywhere in a registered subsystem, including + // the subsystem's own `su_group` (which gets a pointer to itself). The + // only `config_group` with a NULL `cg_subsys` is the configfs root, and + // userspace cannot trigger callbacks on it. The group is therefore the + // subsystem's `su_group` iff it equals `&cg_subsys->su_group`. + // + // SAFETY: For every `config_group` the configfs core dispatches a + // callback on, `cg_subsys` was set by `link_group()` at registration time + // and points to a valid `configfs_subsystem` that outlives the callback. + let is_root = !subsys.is_null() + && core::ptr::eq(this.cast_const(), unsafe { &raw const (*subsys).su_group }); + + if is_root { + // SAFETY: By the above, `this` is the `su_group` field of a + // `configfs_subsystem` that, by function safety requirements, is + // embedded in a `Subsystem<Parent>`. unsafe { &(*Subsystem::container_of(this)).data } + } else { + // SAFETY: By function safety requirements, `this` is a + // `config_group` embedded in a `Group<Parent>`. + unsafe { &(*Group::<Parent>::container_of(this)).data } } } @@ -325,7 +343,9 @@ struct GroupOperationsVTable<Parent, Child>(PhantomData<(Parent, Child)>); impl<Parent, Child> GroupOperationsVTable<Parent, Child> where Parent: GroupOperations<Child = Child>, - Child: 'static, + // We transfer `Arc<Group<Data>>` across a thread boundary in `make_group` + // and `drop_item`. + Arc<Group<Child>>: Send, { /// # Safety /// @@ -344,8 +364,9 @@ where this: *mut bindings::config_group, name: *const kernel::ffi::c_char, ) -> *mut bindings::config_group { - // SAFETY: By function safety requirements of this function, this call - // is safe. + // SAFETY: By function safety requirements, `this` points to a configfs + // group containing `Parent`. The `GroupOperations` bound guarantees + // that `Parent: Sync`, so it is safe to share it with this thread. let parent_data = unsafe { get_group_data(this) }; let group_init = match Parent::make_group( @@ -390,8 +411,9 @@ where this: *mut bindings::config_group, item: *mut bindings::config_item, ) { - // SAFETY: By function safety requirements of this function, this call - // is safe. + // SAFETY: By function safety requirements, `this` points to a configfs + // group containing `Parent`. The `GroupOperations` bound guarantees + // that `Parent: Sync`, so it is safe to share it with this thread. let parent_data = unsafe { get_group_data(this) }; // SAFETY: By function safety requirements, `item` is embedded in a @@ -403,7 +425,9 @@ where if Parent::HAS_DROP_ITEM { // SAFETY: We called `into_raw` to produce `r_child_group_ptr` in - // `make_group`. + // `make_group`. This function may be executing on a different + // thread than `into_raw`. As `Arc<Group<Child>>: Send` this + // ownership transfer is safe. let arc: Arc<Group<Child>> = unsafe { Arc::from_raw(r_child_group_ptr.cast_mut()) }; Parent::drop_item(parent_data, arc.as_arc_borrow()); @@ -434,11 +458,13 @@ struct ItemOperationsVTable<Container, Data>(PhantomData<(Container, Data)>); impl<Data> ItemOperationsVTable<Group<Data>, Data> where Data: 'static, + // We transfer `Arc<Group<Data>>` across a thread boundary in `release`. + Arc<Group<Data>>: Send, { /// # Safety /// - /// `this` must be a pointer to a `bindings::config_group` embedded in a - /// `Group<Parent>`. + /// `this` must be a pointer to the `cg_item` field of a `bindings::config_group` embedded in a + /// valid `Group<Data>`. /// /// This function will destroy the pointee of `this`. The pointee of `this` /// must not be accessed after the function returns. @@ -450,8 +476,9 @@ where // embedded within a `Group<Data>`. let r_group_ptr = unsafe { Group::<Data>::container_of(c_group_ptr) }; - // SAFETY: We called `into_raw` on `r_group_ptr` in - // `make_group`. + // SAFETY: We called `into_raw` on `r_group_ptr` in `make_group`. This + // function may be running on a different thread than the thread that + // called `into_raw`. As `Arc<Group<Data>>: Send`, this is safe. let pin_self: Arc<Group<Data>> = unsafe { Arc::from_raw(r_group_ptr.cast_mut()) }; drop(pin_self); } @@ -483,12 +510,12 @@ impl<Data> ItemOperationsVTable<Subsystem<Data>, Data> { /// /// Implement this trait on structs that embed a [`Subsystem`] or a [`Group`]. #[vtable] -pub trait GroupOperations { +pub trait GroupOperations: Sync { /// The child data object type. /// /// This group will create subgroups (subdirectories) backed by this kind of /// object. - type Child: 'static; + type Child: 'static + Send; /// Creates a new subgroup. /// @@ -555,8 +582,10 @@ where // `config_group`. unsafe { container_of!(item, bindings::config_group, cg_item) }; - // SAFETY: The function safety requirements for this function satisfy - // the conditions for this call. + // SAFETY: By function safety requirements, `c_group` points to a + // configfs group containing `Data`. The `AttributeOperations` bound + // guarantees that `Data: Sync`, so it is safe to share it with this + // thread. let data: &Data = unsafe { get_group_data(c_group) }; // SAFETY: By function safety requirements, `page` is writable for `PAGE_SIZE`. @@ -589,8 +618,10 @@ where // `config_group`. unsafe { container_of!(item, bindings::config_group, cg_item) }; - // SAFETY: The function safety requirements for this function satisfy - // the conditions for this call. + // SAFETY: By function safety requirements, `c_group` points to a + // configfs group containing `Data`. The `AttributeOperations` bound + // guarantees that `Data: Sync`, so it is safe to share it with this + // thread. let data: &Data = unsafe { get_group_data(c_group) }; let ret = O::store( @@ -643,7 +674,7 @@ where pub trait AttributeOperations<const ID: u64 = 0> { /// The type of the object that contains the field that is backing the /// attribute for this operation. - type Data; + type Data: Sync; /// Renders the value of an attribute. /// @@ -749,15 +780,18 @@ macro_rules! impl_item_type { ) -> Self where Data: GroupOperations<Child = Child>, - Child: 'static, + Arc<Group<Child>>: Send, + Arc<Group<Data>>: Send, { Self { item_type: Opaque::new(bindings::config_item_type { ct_owner: owner.as_ptr(), - ct_group_ops: GroupOperationsVTable::<Data, Child>::vtable_ptr().cast_mut(), - ct_item_ops: ItemOperationsVTable::<$tpe, Data>::vtable_ptr().cast_mut(), - ct_attrs: core::ptr::from_ref(attributes).cast_mut().cast(), - ct_bin_attrs: core::ptr::null_mut(), + ct_group_ops: GroupOperationsVTable::<Data, Child>::vtable_ptr(), + ct_item_ops: ItemOperationsVTable::<$tpe, Data>::vtable_ptr(), + __bindgen_anon_1: bindings::config_item_type__bindgen_ty_1 { + ct_attrs_const: core::ptr::from_ref(attributes).cast(), + }, + ct_bin_attrs: core::ptr::null(), }), _p: PhantomData, } @@ -767,14 +801,19 @@ macro_rules! impl_item_type { pub const fn new<const N: usize>( owner: &'static ThisModule, attributes: &'static AttributeList<N, Data>, - ) -> Self { + ) -> Self + where + Arc<Group<Data>>: Send, + { Self { item_type: Opaque::new(bindings::config_item_type { ct_owner: owner.as_ptr(), - ct_group_ops: core::ptr::null_mut(), - ct_item_ops: ItemOperationsVTable::<$tpe, Data>::vtable_ptr().cast_mut(), - ct_attrs: core::ptr::from_ref(attributes).cast_mut().cast(), - ct_bin_attrs: core::ptr::null_mut(), + ct_group_ops: core::ptr::null(), + ct_item_ops: ItemOperationsVTable::<$tpe, Data>::vtable_ptr(), + __bindgen_anon_1: bindings::config_item_type__bindgen_ty_1 { + ct_attrs_const: core::ptr::from_ref(attributes).cast(), + }, + ct_bin_attrs: core::ptr::null(), }), _p: PhantomData, } diff --git a/samples/configfs/configfs_sample.c b/samples/configfs/configfs_sample.c index c1b108ec4ea0..6e3a8e0061ab 100644 --- a/samples/configfs/configfs_sample.c +++ b/samples/configfs/configfs_sample.c @@ -84,7 +84,7 @@ CONFIGFS_ATTR_RO(childless_, showme); CONFIGFS_ATTR(childless_, storeme); CONFIGFS_ATTR_RO(childless_, description); -static struct configfs_attribute *childless_attrs[] = { +static const struct configfs_attribute *const childless_attrs[] = { &childless_attr_showme, &childless_attr_storeme, &childless_attr_description, @@ -92,7 +92,7 @@ static struct configfs_attribute *childless_attrs[] = { }; static const struct config_item_type childless_type = { - .ct_attrs = childless_attrs, + .ct_attrs_const = childless_attrs, .ct_owner = THIS_MODULE, }; @@ -148,7 +148,7 @@ static ssize_t simple_child_storeme_store(struct config_item *item, CONFIGFS_ATTR(simple_child_, storeme); -static struct configfs_attribute *simple_child_attrs[] = { +static const struct configfs_attribute *const simple_child_attrs[] = { &simple_child_attr_storeme, NULL, }; @@ -164,7 +164,7 @@ static const struct configfs_item_operations simple_child_item_ops = { static const struct config_item_type simple_child_type = { .ct_item_ops = &simple_child_item_ops, - .ct_attrs = simple_child_attrs, + .ct_attrs_const = simple_child_attrs, .ct_owner = THIS_MODULE, }; @@ -205,7 +205,7 @@ static ssize_t simple_children_description_show(struct config_item *item, CONFIGFS_ATTR_RO(simple_children_, description); -static struct configfs_attribute *simple_children_attrs[] = { +static const struct configfs_attribute *const simple_children_attrs[] = { &simple_children_attr_description, NULL, }; @@ -230,7 +230,7 @@ static const struct configfs_group_operations simple_children_group_ops = { static const struct config_item_type simple_children_type = { .ct_item_ops = &simple_children_item_ops, .ct_group_ops = &simple_children_group_ops, - .ct_attrs = simple_children_attrs, + .ct_attrs_const = simple_children_attrs, .ct_owner = THIS_MODULE, }; @@ -283,7 +283,7 @@ static ssize_t group_children_description_show(struct config_item *item, CONFIGFS_ATTR_RO(group_children_, description); -static struct configfs_attribute *group_children_attrs[] = { +static const struct configfs_attribute *const group_children_attrs[] = { &group_children_attr_description, NULL, }; @@ -298,7 +298,7 @@ static const struct configfs_group_operations group_children_group_ops = { static const struct config_item_type group_children_type = { .ct_group_ops = &group_children_group_ops, - .ct_attrs = group_children_attrs, + .ct_attrs_const = group_children_attrs, .ct_owner = THIS_MODULE, }; @@ -314,6 +314,126 @@ static struct configfs_subsystem group_children_subsys = { /* ----------------------------------------------------------------- */ /* + * 04-symlink-children + * + * This example has children that are valid sources for symlink(2). A + * child accepts a link to any other config_item and reports how many + * links it currently holds, so ->allow_link() and ->drop_link() are + * observable from userspace. + */ + +struct symlink_child { + struct config_item item; + int nlinks; +}; + +static inline struct symlink_child *to_symlink_child(struct config_item *item) +{ + return container_of(item, struct symlink_child, item); +} + +static ssize_t symlink_child_nlinks_show(struct config_item *item, char *page) +{ + return sprintf(page, "%d\n", to_symlink_child(item)->nlinks); +} + +CONFIGFS_ATTR_RO(symlink_child_, nlinks); + +static struct configfs_attribute *symlink_child_attrs[] = { + &symlink_child_attr_nlinks, + NULL, +}; + +/* + * The VFS holds the source item's directory locked across symlink(2) and + * unlink(2), so ->nlinks needs no lock of its own. + */ +static int symlink_child_allow_link(struct config_item *src, + struct config_item *target) +{ + to_symlink_child(src)->nlinks++; + + return 0; +} + +static void symlink_child_drop_link(struct config_item *src, + struct config_item *target) +{ + to_symlink_child(src)->nlinks--; +} + +static void symlink_child_release(struct config_item *item) +{ + kfree(to_symlink_child(item)); +} + +static const struct configfs_item_operations symlink_child_item_ops = { + .release = symlink_child_release, + .allow_link = symlink_child_allow_link, + .drop_link = symlink_child_drop_link, +}; + +static const struct config_item_type symlink_child_type = { + .ct_item_ops = &symlink_child_item_ops, + .ct_attrs = symlink_child_attrs, + .ct_owner = THIS_MODULE, +}; + +static struct config_item *symlink_children_make_item( + struct config_group *group, const char *name) +{ + struct symlink_child *symlink_child; + + symlink_child = kzalloc_obj(*symlink_child, GFP_KERNEL); + if (!symlink_child) + return ERR_PTR(-ENOMEM); + + config_item_init_type_name(&symlink_child->item, name, + &symlink_child_type); + + return &symlink_child->item; +} + +static ssize_t symlink_children_description_show(struct config_item *item, + char *page) +{ + return sprintf(page, +"[04-symlink-children]\n" +"\n" +"This subsystem allows the creation of child config_items that\n" +"symlink(2) can point at other config_items from. Each child\n" +"reports the number of links it holds.\n"); +} + +CONFIGFS_ATTR_RO(symlink_children_, description); + +static struct configfs_attribute *symlink_children_attrs[] = { + &symlink_children_attr_description, + NULL, +}; + +static const struct configfs_group_operations symlink_children_group_ops = { + .make_item = symlink_children_make_item, +}; + +static const struct config_item_type symlink_children_type = { + .ct_group_ops = &symlink_children_group_ops, + .ct_attrs = symlink_children_attrs, + .ct_owner = THIS_MODULE, +}; + +static struct configfs_subsystem symlink_children_subsys = { + .su_group = { + .cg_item = { + .ci_namebuf = "04-symlink-children", + .ci_type = &symlink_children_type, + }, + }, +}; + +/* ----------------------------------------------------------------- */ + +/* * We're now done with our subsystem definitions. * For convenience in this module, here's a list of them all. It * allows the init function to easily register them. Most modules @@ -324,6 +444,7 @@ static struct configfs_subsystem *example_subsys[] = { &childless_subsys.subsys, &simple_children_subsys, &group_children_subsys, + &symlink_children_subsys, NULL, }; diff --git a/security/apparmor/apparmorfs.c b/security/apparmor/apparmorfs.c index 5b42140e12e8..d1386722537d 100644 --- a/security/apparmor/apparmorfs.c +++ b/security/apparmor/apparmorfs.c @@ -2073,7 +2073,7 @@ fail2: return error; } -static struct dentry *ns_mkdir_op(struct mnt_idmap *idmap, struct inode *dir, +static struct dentry *ns_mkdir_op(const struct mnt_idmap *idmap, struct inode *dir, struct dentry *dentry, umode_t mode) { struct aa_ns *ns, *parent; diff --git a/security/apparmor/lsm.c b/security/apparmor/lsm.c index d502ad0ac26f..c73681d820a0 100644 --- a/security/apparmor/lsm.c +++ b/security/apparmor/lsm.c @@ -397,7 +397,7 @@ static int apparmor_path_rename(const struct path *old_dir, struct dentry *old_d label = begin_current_label_crit_section(&needput); if (!unconfined(label)) { - struct mnt_idmap *idmap = mnt_idmap(old_dir->mnt); + const struct mnt_idmap *idmap = mnt_idmap(old_dir->mnt); vfsuid_t vfsuid; struct path old_path = { .mnt = old_dir->mnt, .dentry = old_dentry }; @@ -485,7 +485,7 @@ static int apparmor_file_open(struct file *file) label = aa_get_newest_cred_label_condref(file->f_cred, &needput); if (!unconfined(label)) { - struct mnt_idmap *idmap = file_mnt_idmap(file); + const struct mnt_idmap *idmap = file_mnt_idmap(file); struct inode *inode = file_inode(file); vfsuid_t vfsuid; struct path_cond cond = { diff --git a/security/commoncap.c b/security/commoncap.c index 3399535808fe..d47ab3022343 100644 --- a/security/commoncap.c +++ b/security/commoncap.c @@ -348,7 +348,7 @@ int cap_inode_need_killpriv(struct dentry *dentry) * * Return: 0 if successful, -ve on error. */ -int cap_inode_killpriv(struct mnt_idmap *idmap, struct dentry *dentry) +int cap_inode_killpriv(const struct mnt_idmap *idmap, struct dentry *dentry) { int error; @@ -417,7 +417,7 @@ static bool is_v3header(int size, const struct vfs_cap_data *cap) * by the integrity subsystem, which really wants the unconverted values - * so that's good. */ -int cap_inode_getsecurity(struct mnt_idmap *idmap, +int cap_inode_getsecurity(const struct mnt_idmap *idmap, struct inode *inode, const char *name, void **buffer, bool alloc) { @@ -566,7 +566,7 @@ static bool validheader(size_t size, const struct vfs_cap_data *cap) * * Return: On success, return the new size; on error, return < 0. */ -int cap_convert_nscap(struct mnt_idmap *idmap, struct dentry *dentry, +int cap_convert_nscap(const struct mnt_idmap *idmap, struct dentry *dentry, const void **ivalue, size_t size) { struct vfs_ns_cap_data *nscap; @@ -672,7 +672,7 @@ static inline int bprm_caps_from_vfs_caps(struct cpu_vfs_cap_data *caps, * permissions. On non-idmapped mounts or if permission checking is to be * performed on the raw inode simply pass @nop_mnt_idmap. */ -int get_vfs_caps_from_disk(struct mnt_idmap *idmap, +int get_vfs_caps_from_disk(const struct mnt_idmap *idmap, const struct dentry *dentry, struct cpu_vfs_cap_data *cpu_caps) { @@ -1063,7 +1063,7 @@ int cap_inode_setxattr(struct dentry *dentry, const char *name, * This is used to make sure security xattrs don't get removed by those who * aren't privileged to remove them. */ -int cap_inode_removexattr(struct mnt_idmap *idmap, +int cap_inode_removexattr(const struct mnt_idmap *idmap, struct dentry *dentry, const char *name) { struct user_namespace *user_ns = dentry->d_sb->s_user_ns; diff --git a/security/integrity/evm/evm_main.c b/security/integrity/evm/evm_main.c index b59e3f121b8a..9bfea858df2d 100644 --- a/security/integrity/evm/evm_main.c +++ b/security/integrity/evm/evm_main.c @@ -481,7 +481,7 @@ static enum integrity_status evm_verify_current_integrity(struct dentry *dentry) * * Returns 1 if passed xattr value differs from current value, 0 otherwise. */ -static int evm_xattr_change(struct mnt_idmap *idmap, +static int evm_xattr_change(const struct mnt_idmap *idmap, struct dentry *dentry, const char *xattr_name, const void *xattr_value, size_t xattr_value_len) { @@ -517,7 +517,7 @@ out: * For posix xattr acls only, permit security.evm, even if it currently * doesn't exist, to be updated unless the EVM signature is immutable. */ -static int evm_protect_xattr(struct mnt_idmap *idmap, +static int evm_protect_xattr(const struct mnt_idmap *idmap, struct dentry *dentry, const char *xattr_name, const void *xattr_value, size_t xattr_value_len) { @@ -607,7 +607,7 @@ out: * userspace from writing HMAC value. Writing 'security.evm' requires * requires CAP_SYS_ADMIN privileges. */ -static int evm_inode_setxattr(struct mnt_idmap *idmap, struct dentry *dentry, +static int evm_inode_setxattr(const struct mnt_idmap *idmap, struct dentry *dentry, const char *xattr_name, const void *xattr_value, size_t xattr_value_len, int flags) { @@ -639,7 +639,7 @@ static int evm_inode_setxattr(struct mnt_idmap *idmap, struct dentry *dentry, * Removing 'security.evm' requires CAP_SYS_ADMIN privileges and that * the current value is valid. */ -static int evm_inode_removexattr(struct mnt_idmap *idmap, struct dentry *dentry, +static int evm_inode_removexattr(const struct mnt_idmap *idmap, struct dentry *dentry, const char *xattr_name) { /* Policy permits modification of the protected xattrs even though @@ -652,7 +652,7 @@ static int evm_inode_removexattr(struct mnt_idmap *idmap, struct dentry *dentry, } #ifdef CONFIG_FS_POSIX_ACL -static int evm_inode_set_acl_change(struct mnt_idmap *idmap, +static int evm_inode_set_acl_change(const struct mnt_idmap *idmap, struct dentry *dentry, const char *name, struct posix_acl *kacl) { @@ -671,7 +671,7 @@ static int evm_inode_set_acl_change(struct mnt_idmap *idmap, return 0; } #else -static inline int evm_inode_set_acl_change(struct mnt_idmap *idmap, +static inline int evm_inode_set_acl_change(const struct mnt_idmap *idmap, struct dentry *dentry, const char *name, struct posix_acl *kacl) @@ -693,7 +693,7 @@ static inline int evm_inode_set_acl_change(struct mnt_idmap *idmap, * * Return: zero on success, -EPERM on failure. */ -static int evm_inode_set_acl(struct mnt_idmap *idmap, struct dentry *dentry, +static int evm_inode_set_acl(const struct mnt_idmap *idmap, struct dentry *dentry, const char *acl_name, struct posix_acl *kacl) { enum integrity_status evm_status; @@ -745,7 +745,7 @@ static int evm_inode_set_acl(struct mnt_idmap *idmap, struct dentry *dentry, * * Return: zero on success, -EPERM on failure. */ -static int evm_inode_remove_acl(struct mnt_idmap *idmap, struct dentry *dentry, +static int evm_inode_remove_acl(const struct mnt_idmap *idmap, struct dentry *dentry, const char *acl_name) { return evm_inode_set_acl(idmap, dentry, acl_name, NULL); @@ -926,14 +926,14 @@ static void evm_inode_post_removexattr(struct dentry *dentry, * Update the 'security.evm' xattr with the EVM HMAC re-calculated after * removing posix acls. */ -static inline void evm_inode_post_remove_acl(struct mnt_idmap *idmap, +static inline void evm_inode_post_remove_acl(const struct mnt_idmap *idmap, struct dentry *dentry, const char *acl_name) { evm_inode_post_removexattr(dentry, acl_name); } -static int evm_attr_change(struct mnt_idmap *idmap, +static int evm_attr_change(const struct mnt_idmap *idmap, struct dentry *dentry, struct iattr *attr) { struct inode *inode = d_backing_inode(dentry); @@ -956,7 +956,7 @@ static int evm_attr_change(struct mnt_idmap *idmap, * Permit update of file attributes when files have a valid EVM signature, * except in the case of them having an immutable portable signature. */ -static int evm_inode_setattr(struct mnt_idmap *idmap, struct dentry *dentry, +static int evm_inode_setattr(const struct mnt_idmap *idmap, struct dentry *dentry, struct iattr *attr) { unsigned int ia_valid = attr->ia_valid; @@ -1008,7 +1008,7 @@ static int evm_inode_setattr(struct mnt_idmap *idmap, struct dentry *dentry, * This function is called from notify_change(), which expects the caller * to lock the inode's i_mutex. */ -static void evm_inode_post_setattr(struct mnt_idmap *idmap, +static void evm_inode_post_setattr(const struct mnt_idmap *idmap, struct dentry *dentry, int ia_valid) { if (!evm_revalidate_status(NULL)) @@ -1140,7 +1140,7 @@ static void evm_file_release(struct file *file) iint->flags &= ~EVM_NEW_FILE; } -static void evm_post_path_mknod(struct mnt_idmap *idmap, struct dentry *dentry) +static void evm_post_path_mknod(const struct mnt_idmap *idmap, struct dentry *dentry) { struct inode *inode = d_backing_inode(dentry); struct evm_iint_cache *iint = evm_iint_inode(inode); diff --git a/security/integrity/ima/ima.h b/security/integrity/ima/ima.h index 10214f73ca1e..b502854f28ee 100644 --- a/security/integrity/ima/ima.h +++ b/security/integrity/ima/ima.h @@ -423,7 +423,7 @@ static inline void ima_process_queued_keys(void) {} #endif /* CONFIG_IMA_QUEUE_EARLY_BOOT_KEYS */ /* LIM API function definitions */ -int ima_get_action(struct mnt_idmap *idmap, struct inode *inode, +int ima_get_action(const struct mnt_idmap *idmap, struct inode *inode, const struct cred *cred, struct lsm_prop *prop, int mask, enum ima_hooks func, int *pcr, struct ima_template_desc **template_desc, @@ -437,7 +437,7 @@ void ima_store_measurement(struct ima_iint_cache *iint, struct file *file, struct evm_ima_xattr_data *xattr_value, int xattr_len, const struct modsig *modsig, int pcr, struct ima_template_desc *template_desc); -int process_buffer_measurement(struct mnt_idmap *idmap, +int process_buffer_measurement(const struct mnt_idmap *idmap, struct inode *inode, const void *buf, int size, const char *eventname, enum ima_hooks func, int pcr, const char *func_data, @@ -454,7 +454,7 @@ void ima_free_template_entry(struct ima_template_entry *entry); const char *ima_d_path(const struct path *path, char **pathbuf, char *filename); /* IMA policy related functions */ -int ima_match_policy(struct mnt_idmap *idmap, struct inode *inode, +int ima_match_policy(const struct mnt_idmap *idmap, struct inode *inode, const struct cred *cred, struct lsm_prop *prop, enum ima_hooks func, int mask, int flags, int *pcr, struct ima_template_desc **template_desc, @@ -489,7 +489,7 @@ int ima_appraise_measurement(enum ima_hooks func, struct ima_iint_cache *iint, struct evm_ima_xattr_data *xattr_value, int xattr_len, const struct modsig *modsig, bool bprm_is_check); -int ima_must_appraise(struct mnt_idmap *idmap, struct inode *inode, +int ima_must_appraise(const struct mnt_idmap *idmap, struct inode *inode, int mask, enum ima_hooks func); void ima_update_xattr(struct ima_iint_cache *iint, struct file *file); enum integrity_status ima_get_cache_status(struct ima_iint_cache *iint, @@ -519,7 +519,7 @@ static inline int ima_appraise_measurement(enum ima_hooks func, return INTEGRITY_UNKNOWN; } -static inline int ima_must_appraise(struct mnt_idmap *idmap, +static inline int ima_must_appraise(const struct mnt_idmap *idmap, struct inode *inode, int mask, enum ima_hooks func) { diff --git a/security/integrity/ima/ima_api.c b/security/integrity/ima/ima_api.c index 122d127e108d..3c5a23b4a2ff 100644 --- a/security/integrity/ima/ima_api.c +++ b/security/integrity/ima/ima_api.c @@ -188,7 +188,7 @@ err_out: * Returns IMA_MEASURE, IMA_APPRAISE mask. * */ -int ima_get_action(struct mnt_idmap *idmap, struct inode *inode, +int ima_get_action(const struct mnt_idmap *idmap, struct inode *inode, const struct cred *cred, struct lsm_prop *prop, int mask, enum ima_hooks func, int *pcr, struct ima_template_desc **template_desc, diff --git a/security/integrity/ima/ima_appraise.c b/security/integrity/ima/ima_appraise.c index b280488e15fc..63b0cb957c65 100644 --- a/security/integrity/ima/ima_appraise.c +++ b/security/integrity/ima/ima_appraise.c @@ -71,7 +71,7 @@ bool is_ima_appraise_enabled(void) * * Return 1 to appraise or hash */ -int ima_must_appraise(struct mnt_idmap *idmap, struct inode *inode, +int ima_must_appraise(const struct mnt_idmap *idmap, struct inode *inode, int mask, enum ima_hooks func) { struct lsm_prop prop; @@ -634,7 +634,7 @@ void ima_update_xattr(struct ima_iint_cache *iint, struct file *file) * This function is called from notify_change(), which expects the caller * to lock the inode's i_mutex. */ -static void ima_inode_post_setattr(struct mnt_idmap *idmap, +static void ima_inode_post_setattr(const struct mnt_idmap *idmap, struct dentry *dentry, int ia_valid) { struct inode *inode = d_backing_inode(dentry); @@ -759,7 +759,7 @@ static int validate_hash_algo(struct dentry *dentry, return -EACCES; } -static int ima_inode_setxattr(struct mnt_idmap *idmap, struct dentry *dentry, +static int ima_inode_setxattr(const struct mnt_idmap *idmap, struct dentry *dentry, const char *xattr_name, const void *xattr_value, size_t xattr_value_len, int flags) { @@ -792,7 +792,7 @@ static int ima_inode_setxattr(struct mnt_idmap *idmap, struct dentry *dentry, return result; } -static int ima_inode_set_acl(struct mnt_idmap *idmap, struct dentry *dentry, +static int ima_inode_set_acl(const struct mnt_idmap *idmap, struct dentry *dentry, const char *acl_name, struct posix_acl *kacl) { if (evm_revalidate_status(acl_name)) @@ -801,7 +801,7 @@ static int ima_inode_set_acl(struct mnt_idmap *idmap, struct dentry *dentry, return 0; } -static int ima_inode_removexattr(struct mnt_idmap *idmap, struct dentry *dentry, +static int ima_inode_removexattr(const struct mnt_idmap *idmap, struct dentry *dentry, const char *xattr_name) { int result, digsig = -1; @@ -817,7 +817,7 @@ static int ima_inode_removexattr(struct mnt_idmap *idmap, struct dentry *dentry, return result; } -static int ima_inode_remove_acl(struct mnt_idmap *idmap, struct dentry *dentry, +static int ima_inode_remove_acl(const struct mnt_idmap *idmap, struct dentry *dentry, const char *acl_name) { return ima_inode_set_acl(idmap, dentry, acl_name, NULL); diff --git a/security/integrity/ima/ima_main.c b/security/integrity/ima/ima_main.c index ab1e53b3210d..7ac38a98b1f9 100644 --- a/security/integrity/ima/ima_main.c +++ b/security/integrity/ima/ima_main.c @@ -846,7 +846,7 @@ EXPORT_SYMBOL_GPL(ima_inode_hash); * Skip calling process_measurement(), but indicate which newly, created * tmpfiles are in policy. */ -static void ima_post_create_tmpfile(struct mnt_idmap *idmap, +static void ima_post_create_tmpfile(const struct mnt_idmap *idmap, struct inode *inode) { @@ -879,7 +879,7 @@ static void ima_post_create_tmpfile(struct mnt_idmap *idmap, * Mark files created via the mknodat syscall as new, so that the * file data can be written later. */ -static void ima_post_path_mknod(struct mnt_idmap *idmap, struct dentry *dentry) +static void ima_post_path_mknod(const struct mnt_idmap *idmap, struct dentry *dentry) { struct ima_iint_cache *iint; struct inode *inode = dentry->d_inode; @@ -1095,7 +1095,7 @@ static int ima_post_load_data(char *buf, loff_t size, * has been written to the passed location but not added to a measurement entry, * a negative value otherwise. */ -int process_buffer_measurement(struct mnt_idmap *idmap, +int process_buffer_measurement(const struct mnt_idmap *idmap, struct inode *inode, const void *buf, int size, const char *eventname, enum ima_hooks func, int pcr, const char *func_data, diff --git a/security/integrity/ima/ima_policy.c b/security/integrity/ima/ima_policy.c index 68d9a5e6c232..b34d9621929b 100644 --- a/security/integrity/ima/ima_policy.c +++ b/security/integrity/ima/ima_policy.c @@ -580,7 +580,7 @@ static bool ima_match_rule_data(struct ima_rule_entry *rule, * Returns true on rule match, false on failure. */ static bool ima_match_rules(struct ima_rule_entry *rule, - struct mnt_idmap *idmap, + const struct mnt_idmap *idmap, struct inode *inode, const struct cred *cred, struct lsm_prop *prop, enum ima_hooks func, int mask, const char *func_data) @@ -762,7 +762,7 @@ static int get_subaction(struct ima_rule_entry *rule, enum ima_hooks func) * list when walking it. Reads are many orders of magnitude more numerous * than writes so ima_match_policy() is classical RCU candidate. */ -int ima_match_policy(struct mnt_idmap *idmap, struct inode *inode, +int ima_match_policy(const struct mnt_idmap *idmap, struct inode *inode, const struct cred *cred, struct lsm_prop *prop, enum ima_hooks func, int mask, int flags, int *pcr, struct ima_template_desc **template_desc, diff --git a/security/security.c b/security/security.c index 2ee276ab15c5..f476b0736c9a 100644 --- a/security/security.c +++ b/security/security.c @@ -596,6 +596,31 @@ int security_ptrace_traceme(struct task_struct *parent) } /** + * security_mem_foll_force() - Check if FOLL_FORCE is allowed + * @subject: credentials using which /proc/$pid/mem was opened + * @opened_by_owner: whether checks on open() were bypassed because the opener + * has the same MM as the target + * + * Check if FOLL_FORCE is allowed for accessing process memory through + * /proc/$pid/mem. opened_by_owner signals whether the opener's MM was the same + * as the target MM, meaning the security_ptrace_access_check() hook was + * bypassed on open(). + * (Current current->mm does not matter for this; for example, if write() is + * called on an FD that was received from another process which obtained it with + * open("/proc/self/mem"), @opened_by_owner is still true.) + * + * Note that this hook is only designed to be useful in the opened_by_owner + * case, where the subject credentials effectively also describe the object. + * + * Return: Returns 0 if permission is granted. + */ +int security_mem_foll_force(const struct cred *subject, + bool opened_by_owner) +{ + return call_int_hook(mem_foll_force, subject, opened_by_owner); +} + +/** * security_capget() - Get the capability sets for a process * @target: target process * @effective: effective capability set @@ -1425,7 +1450,7 @@ EXPORT_SYMBOL(security_path_mknod); * * Update inode security field after a regular file has been created. */ -void security_path_post_mknod(struct mnt_idmap *idmap, struct dentry *dentry) +void security_path_post_mknod(const struct mnt_idmap *idmap, struct dentry *dentry) { if (unlikely(IS_PRIVATE(d_backing_inode(dentry)))) return; @@ -1638,7 +1663,7 @@ EXPORT_SYMBOL_GPL(security_inode_create); * * Update inode security data after a tmpfile has been created. */ -void security_inode_post_create_tmpfile(struct mnt_idmap *idmap, +void security_inode_post_create_tmpfile(const struct mnt_idmap *idmap, struct inode *inode) { if (unlikely(IS_PRIVATE(inode))) @@ -1855,7 +1880,7 @@ int security_inode_permission(struct inode *inode, int mask) * * Return: Returns 0 if permission is granted. */ -int security_inode_setattr(struct mnt_idmap *idmap, +int security_inode_setattr(const struct mnt_idmap *idmap, struct dentry *dentry, struct iattr *attr) { if (unlikely(IS_PRIVATE(d_backing_inode(dentry)))) @@ -1872,7 +1897,7 @@ EXPORT_SYMBOL_GPL(security_inode_setattr); * * Update inode security field after successful setting file attributes. */ -void security_inode_post_setattr(struct mnt_idmap *idmap, struct dentry *dentry, +void security_inode_post_setattr(const struct mnt_idmap *idmap, struct dentry *dentry, int ia_valid) { if (unlikely(IS_PRIVATE(d_backing_inode(dentry)))) @@ -1921,7 +1946,7 @@ int security_inode_getattr(const struct path *path) * * Return: Returns 0 if permission is granted. */ -int security_inode_setxattr(struct mnt_idmap *idmap, +int security_inode_setxattr(const struct mnt_idmap *idmap, struct dentry *dentry, const char *name, const void *value, size_t size, int flags) { @@ -1953,7 +1978,7 @@ int security_inode_setxattr(struct mnt_idmap *idmap, * * Return: Returns 0 if permission is granted. */ -int security_inode_set_acl(struct mnt_idmap *idmap, +int security_inode_set_acl(const struct mnt_idmap *idmap, struct dentry *dentry, const char *acl_name, struct posix_acl *kacl) { @@ -1990,7 +2015,7 @@ void security_inode_post_set_acl(struct dentry *dentry, const char *acl_name, * * Return: Returns 0 if permission is granted. */ -int security_inode_get_acl(struct mnt_idmap *idmap, +int security_inode_get_acl(const struct mnt_idmap *idmap, struct dentry *dentry, const char *acl_name) { if (unlikely(IS_PRIVATE(d_backing_inode(dentry)))) @@ -2009,7 +2034,7 @@ int security_inode_get_acl(struct mnt_idmap *idmap, * * Return: Returns 0 if permission is granted. */ -int security_inode_remove_acl(struct mnt_idmap *idmap, +int security_inode_remove_acl(const struct mnt_idmap *idmap, struct dentry *dentry, const char *acl_name) { if (unlikely(IS_PRIVATE(d_backing_inode(dentry)))) @@ -2026,7 +2051,7 @@ int security_inode_remove_acl(struct mnt_idmap *idmap, * Update inode security data after successfully removing posix acls on * @dentry in @idmap. The posix acls are identified by @acl_name. */ -void security_inode_post_remove_acl(struct mnt_idmap *idmap, +void security_inode_post_remove_acl(const struct mnt_idmap *idmap, struct dentry *dentry, const char *acl_name) { if (unlikely(IS_PRIVATE(d_backing_inode(dentry)))) @@ -2108,7 +2133,7 @@ int security_inode_listxattr(struct dentry *dentry) * * Return: Returns 0 if permission is granted. */ -int security_inode_removexattr(struct mnt_idmap *idmap, +int security_inode_removexattr(const struct mnt_idmap *idmap, struct dentry *dentry, const char *name) { int rc; @@ -2197,7 +2222,7 @@ int security_inode_need_killpriv(struct dentry *dentry) * Return: Return 0 on success. If error is returned, then the operation * causing setuid bit removal is failed. */ -int security_inode_killpriv(struct mnt_idmap *idmap, +int security_inode_killpriv(const struct mnt_idmap *idmap, struct dentry *dentry) { return call_int_hook(inode_killpriv, idmap, dentry); @@ -2219,7 +2244,7 @@ int security_inode_killpriv(struct mnt_idmap *idmap, * * Return: Returns size of buffer on success. */ -int security_inode_getsecurity(struct mnt_idmap *idmap, +int security_inode_getsecurity(const struct mnt_idmap *idmap, struct inode *inode, const char *name, void **buffer, bool alloc) { diff --git a/security/selinux/hooks.c b/security/selinux/hooks.c index 3f4a6ddee322..73496690ab52 100644 --- a/security/selinux/hooks.c +++ b/security/selinux/hooks.c @@ -2165,6 +2165,36 @@ static int selinux_ptrace_traceme(struct task_struct *parent) SECCLASS_PROCESS, PROCESS__PTRACE, NULL); } +/** + * selinux_mem_foll_force() - Determine whether /proc/$pid/mem can use FOLL_FORCE + * @subject: credentials using which /proc/$pid/mem was opened + * @opened_by_owner: whether checks on open() were bypassed because the opener + * has the same MM as the target + * + * Decide whether it should be possible to read non-readable VMAs and write + * non-writable VMAs via /proc/self/mem. + * The @opened_by_owner case only applies to systems configured with + * PROC_MEM_FORCE_ALWAYS, and only happens on accesses that are not visible to + * selinux_ptrace_access_check() because of the introspection exceptions in + * may_access_mm() and __ptrace_may_access(). + * + * This allows a process to overwrite read-only code in its own address space. + * + * Creating an audit record on denial doesn't make sense here, since we can't + * tell whether FOLL_FORCE matters for the accessed VMAs. + */ +static int selinux_mem_foll_force(const struct cred *subject, bool opened_by_owner) +{ + struct av_decision avd; + u32 sid; + + if (!opened_by_owner) + return 0; + sid = cred_sid(subject); + + return avc_has_perm_noaudit(sid, sid, SECCLASS_PROCESS, PROCESS__PTRACE, 0, &avd); +} + static int selinux_capget(const struct task_struct *target, kernel_cap_t *effective, kernel_cap_t *inheritable, kernel_cap_t *permitted) { @@ -3303,7 +3333,7 @@ static int selinux_inode_permission(struct inode *inode, int requested) return rc; } -static int selinux_inode_setattr(struct mnt_idmap *idmap, struct dentry *dentry, +static int selinux_inode_setattr(const struct mnt_idmap *idmap, struct dentry *dentry, struct iattr *iattr) { const struct cred *cred = current_cred(); @@ -3373,7 +3403,7 @@ static int selinux_inode_xattr_skipcap(const char *name) return !strcmp(name, XATTR_NAME_SELINUX); } -static int selinux_inode_setxattr(struct mnt_idmap *idmap, +static int selinux_inode_setxattr(const struct mnt_idmap *idmap, struct dentry *dentry, const char *name, const void *value, size_t size, int flags) { @@ -3459,20 +3489,20 @@ static int selinux_inode_setxattr(struct mnt_idmap *idmap, &ad); } -static int selinux_inode_set_acl(struct mnt_idmap *idmap, +static int selinux_inode_set_acl(const struct mnt_idmap *idmap, struct dentry *dentry, const char *acl_name, struct posix_acl *kacl) { return dentry_has_perm(current_cred(), dentry, FILE__SETATTR); } -static int selinux_inode_get_acl(struct mnt_idmap *idmap, +static int selinux_inode_get_acl(const struct mnt_idmap *idmap, struct dentry *dentry, const char *acl_name) { return dentry_has_perm(current_cred(), dentry, FILE__GETATTR); } -static int selinux_inode_remove_acl(struct mnt_idmap *idmap, +static int selinux_inode_remove_acl(const struct mnt_idmap *idmap, struct dentry *dentry, const char *acl_name) { return dentry_has_perm(current_cred(), dentry, FILE__SETATTR); @@ -3532,7 +3562,7 @@ static int selinux_inode_listxattr(struct dentry *dentry) return dentry_has_perm(cred, dentry, FILE__GETATTR); } -static int selinux_inode_removexattr(struct mnt_idmap *idmap, +static int selinux_inode_removexattr(const struct mnt_idmap *idmap, struct dentry *dentry, const char *name) { /* if not a selinux xattr, only check the ordinary setattr perm */ @@ -3612,7 +3642,7 @@ static int selinux_path_notify(const struct path *path, u64 mask, * * Permission check is handled by selinux_inode_getxattr hook. */ -static int selinux_inode_getsecurity(struct mnt_idmap *idmap, +static int selinux_inode_getsecurity(const struct mnt_idmap *idmap, struct inode *inode, const char *name, void **buffer, bool alloc) { @@ -7664,6 +7694,7 @@ static struct security_hook_list selinux_hooks[] __ro_after_init = { LSM_HOOK_INIT(ptrace_access_check, selinux_ptrace_access_check), LSM_HOOK_INIT(ptrace_traceme, selinux_ptrace_traceme), + LSM_HOOK_INIT(mem_foll_force, selinux_mem_foll_force), LSM_HOOK_INIT(capget, selinux_capget), LSM_HOOK_INIT(capset, selinux_capset), LSM_HOOK_INIT(capable, selinux_capable), diff --git a/security/selinux/selinuxfs.c b/security/selinux/selinuxfs.c index 545a6f89f9e7..dbe3a221caab 100644 --- a/security/selinux/selinuxfs.c +++ b/security/selinux/selinuxfs.c @@ -1777,7 +1777,7 @@ static struct dentry *sel_make_dir(struct dentry *dir, const char *name, return sel_attach(dir, name, inode); } -static int reject_all(struct mnt_idmap *idmap, struct inode *inode, int mask) +static int reject_all(const struct mnt_idmap *idmap, struct inode *inode, int mask) { return -EPERM; // no access for anyone, root or no root. } diff --git a/security/smack/smack_lsm.c b/security/smack/smack_lsm.c index 8e88ac65fd7f..bb78569b3d5b 100644 --- a/security/smack/smack_lsm.c +++ b/security/smack/smack_lsm.c @@ -1270,7 +1270,7 @@ static int smack_inode_permission(struct inode *inode, int mask) * * Returns 0 if access is permitted, an error code otherwise */ -static int smack_inode_setattr(struct mnt_idmap *idmap, struct dentry *dentry, +static int smack_inode_setattr(const struct mnt_idmap *idmap, struct dentry *dentry, struct iattr *iattr) { struct smk_audit_info ad; @@ -1348,7 +1348,7 @@ static int smack_inode_xattr_skipcap(const char *name) * * Returns 0 if access is permitted, an error code otherwise */ -static int smack_inode_setxattr(struct mnt_idmap *idmap, +static int smack_inode_setxattr(const struct mnt_idmap *idmap, struct dentry *dentry, const char *name, const void *value, size_t size, int flags) { @@ -1482,7 +1482,7 @@ static int smack_inode_getxattr(struct dentry *dentry, const char *name) * * Returns 0 if access is permitted, an error code otherwise */ -static int smack_inode_removexattr(struct mnt_idmap *idmap, +static int smack_inode_removexattr(const struct mnt_idmap *idmap, struct dentry *dentry, const char *name) { struct inode_smack *isp; @@ -1543,7 +1543,7 @@ static int smack_inode_removexattr(struct mnt_idmap *idmap, * * Returns 0 if access is permitted, an error code otherwise */ -static int smack_inode_set_acl(struct mnt_idmap *idmap, +static int smack_inode_set_acl(const struct mnt_idmap *idmap, struct dentry *dentry, const char *acl_name, struct posix_acl *kacl) { @@ -1566,7 +1566,7 @@ static int smack_inode_set_acl(struct mnt_idmap *idmap, * * Returns 0 if access is permitted, an error code otherwise */ -static int smack_inode_get_acl(struct mnt_idmap *idmap, +static int smack_inode_get_acl(const struct mnt_idmap *idmap, struct dentry *dentry, const char *acl_name) { struct smk_audit_info ad; @@ -1588,7 +1588,7 @@ static int smack_inode_get_acl(struct mnt_idmap *idmap, * * Returns 0 if access is permitted, an error code otherwise */ -static int smack_inode_remove_acl(struct mnt_idmap *idmap, +static int smack_inode_remove_acl(const struct mnt_idmap *idmap, struct dentry *dentry, const char *acl_name) { struct smk_audit_info ad; @@ -1612,7 +1612,7 @@ static int smack_inode_remove_acl(struct mnt_idmap *idmap, * * Returns the size of the attribute or an error code */ -static int smack_inode_getsecurity(struct mnt_idmap *idmap, +static int smack_inode_getsecurity(const struct mnt_idmap *idmap, struct inode *inode, const char *name, void **buffer, bool alloc) { diff --git a/tools/include/uapi/linux/coredump.h b/tools/include/uapi/linux/coredump.h index dc3789b78af0..6d0c53b534ea 100644 --- a/tools/include/uapi/linux/coredump.h +++ b/tools/include/uapi/linux/coredump.h @@ -11,12 +11,53 @@ * @COREDUMP_USERSPACE: userspace writes coredump * @COREDUMP_REJECT: don't generate coredump * @COREDUMP_WAIT: wait for coredump server + * @COREDUMP_RECORDS: send the coredump as a sequence of records instead of + * as a plain byte stream, see struct coredump_record_header; + * requires COREDUMP_KERNEL + * @COREDUMP_SPARSE: describe the holes in the coredump as zero records + * instead of transferring them; requires COREDUMP_RECORDS + * @COREDUMP_MEMORY_TYPES: dump the memory types in + * coredump_ack->memory_types instead of the ones + * the task selected; requires COREDUMP_KERNEL */ enum { COREDUMP_KERNEL = (1ULL << 0), COREDUMP_USERSPACE = (1ULL << 1), COREDUMP_REJECT = (1ULL << 2), COREDUMP_WAIT = (1ULL << 3), + COREDUMP_RECORDS = (1ULL << 4), + COREDUMP_SPARSE = (1ULL << 5), + COREDUMP_MEMORY_TYPES = (1ULL << 6), +}; + +/** + * coredump memory types + * @COREDUMP_MEMORY_ANON_PRIVATE: anonymous private memory + * @COREDUMP_MEMORY_ANON_SHARED: anonymous shared memory + * @COREDUMP_MEMORY_FILE_PRIVATE: file-backed private memory + * @COREDUMP_MEMORY_FILE_SHARED: file-backed shared memory + * @COREDUMP_MEMORY_ELF_HEADERS: the first page of a file-backed private + * mapping that starts an ELF file + * @COREDUMP_MEMORY_HUGETLB_PRIVATE: hugetlb private memory + * @COREDUMP_MEMORY_HUGETLB_SHARED: hugetlb shared memory + * @COREDUMP_MEMORY_DAX_PRIVATE: DAX private memory + * @COREDUMP_MEMORY_DAX_SHARED: DAX shared memory + * + * A bitmask of memory types a coredump may request to be included. New + * memory type bits must ensure that they do not steal memory from an + * existing one so a coredump server will continue to get the same + * coredumps even if a new bit is introduced. + */ +enum { + COREDUMP_MEMORY_ANON_PRIVATE = (1ULL << 0), + COREDUMP_MEMORY_ANON_SHARED = (1ULL << 1), + COREDUMP_MEMORY_FILE_PRIVATE = (1ULL << 2), + COREDUMP_MEMORY_FILE_SHARED = (1ULL << 3), + COREDUMP_MEMORY_ELF_HEADERS = (1ULL << 4), + COREDUMP_MEMORY_HUGETLB_PRIVATE = (1ULL << 5), + COREDUMP_MEMORY_HUGETLB_SHARED = (1ULL << 6), + COREDUMP_MEMORY_DAX_PRIVATE = (1ULL << 7), + COREDUMP_MEMORY_DAX_SHARED = (1ULL << 8), }; /** @@ -24,17 +65,19 @@ enum { * @size: size of struct coredump_req * @size_ack: known size of struct coredump_ack on this kernel * @mask: supported features + * @memory_types: the memory types the task selected + * @memory_types_mask: the memory types this kernel knows * * When a coredump happens the kernel will connect to the coredump * socket and send a coredump request to the coredump server. The @size * member is set to the size of struct coredump_req and provides a hint * to userspace how much data can be read. Userspace may use MSG_PEEK to * peek the size of struct coredump_req and then choose to consume it in - * one go. Userspace may also simply read a COREDUMP_ACK_SIZE_VER0 + * one go. Userspace may also simply read a COREDUMP_REQ_SIZE_VER0 * request. If the size the kernel sends is larger userspace simply * discards any remaining data. * - * The coredump_req->mask member is set to the currently know features. + * The coredump_req->mask member is set to the currently known features. * Userspace may only set coredump_ack->mask to the bits raised by the * kernel in coredump_req->mask. * @@ -42,15 +85,27 @@ enum { * struct coredump_ack the kernel knows. Userspace may only send up to * coredump_req->size_ack bytes to the kernel and must set * coredump_ack->size accordingly. + * + * @memory_types is set to the default memory types that are included in + * the coredump. This can be overridden by raising bits in + * coredump_ack->memory_types. + * + * @memory_types_mask contains a bitmask of all memory types the kernel + * knows about. A coredump server may only raise bits in + * coredump_ack->memory_types that are raised in + * coredump_req->memory_types_mask. */ struct coredump_req { __u32 size; __u32 size_ack; __u64 mask; + __u64 memory_types; + __u64 memory_types_mask; }; enum { COREDUMP_REQ_SIZE_VER0 = 16U, /* size of first published struct */ + COREDUMP_REQ_SIZE_VER1 = 32U, /* memory_types and memory_types_mask added */ }; /** @@ -58,6 +113,8 @@ enum { * @size: size of the struct * @spare: unused * @mask: features kernel is supposed to use + * @memory_types: memory types to dump, only with COREDUMP_MEMORY_TYPES + * in @mask * * The @size member must be set to the size of struct coredump_ack. It * may never exceed what the kernel returned in coredump_req->size_ack @@ -67,15 +124,30 @@ enum { * The @mask member must be set to the features the coredump server * wants the kernel to use. Only bits the kernel returned in * coredump_req->mask may be set. + * + * If COREDUMP_MEMORY_TYPES is raised in @mask the kernel dumps the + * memory types set in the @memory_types mask. Zero is valid and dumps + * no memory apart from the mappings that are always dumped. + * + * Note that memory a task excluded via MADV_DONTDUMP is always left + * out. A coredump server wanting to add or drop memory types instead of + * outright replacing it should simply copy coredump_req->memory_types + * and then mask off or raise types as needed. + * + * Note that @memory_types must be zero if COREDUMP_MEMORY_TYPES isn't + * raised. COREDUMP_MEMORY_TYPES requires COREDUMP_KERNEL and an ack of + * at least COREDUMP_ACK_SIZE_VER1 bytes. */ struct coredump_ack { __u32 size; __u32 spare; __u64 mask; + __u64 memory_types; }; enum { COREDUMP_ACK_SIZE_VER0 = 16U, /* size of first published struct */ + COREDUMP_ACK_SIZE_VER1 = 24U, /* memory_types added */ }; /** @@ -83,11 +155,12 @@ enum { * * The kernel will place a single byte on the coredump socket. The * markers notify userspace whether the coredump ack succeeded or - * failed. + * failed. After any marker other than COREDUMP_MARK_REQACK the kernel + * closes the connection and no coredump is generated. * * @COREDUMP_MARK_MINSIZE: the provided coredump_ack size was too small * @COREDUMP_MARK_MAXSIZE: the provided coredump_ack size was too big - * @COREDUMP_MARK_UNSUPPORTED: the provided coredump_ack mask was invalid + * @COREDUMP_MARK_UNSUPPORTED: the provided coredump_ack mask or memory types were invalid * @COREDUMP_MARK_CONFLICTING: the provided coredump_ack mask has conflicting options * @COREDUMP_MARK_REQACK: the coredump request and ack was successful * @__COREDUMP_MARK_MAX: the maximum coredump mark value @@ -101,4 +174,72 @@ enum coredump_mark { __COREDUMP_MARK_MAX = (1U << 31), }; +/** + * enum coredump_record_type - Type of a coredump record + * + * @COREDUMP_RECORD_DATA: the header is followed by ->len bytes of data + * @COREDUMP_RECORD_END: the coredump ends here, the header is not followed + * by any data and no further record is sent + * @COREDUMP_RECORD_ZERO: the header stands for ->len zero bytes and is not + * followed by any data + * @__COREDUMP_RECORD_TYPE_MAX: the maximum coredump record type value + */ +enum coredump_record_type { + COREDUMP_RECORD_DATA = 0U, + COREDUMP_RECORD_END = 1U, + COREDUMP_RECORD_ZERO = 2U, + __COREDUMP_RECORD_TYPE_MAX = (1U << 31), +}; + +/** + * struct coredump_record_header - header of a coredump record + * @size: size of struct coredump_record_header + * @type: one of enum coredump_record_type + * @flags: modifiers for this record + * @offset: offset in the coredump this record starts at + * @len: number of coredump bytes this record accounts for + * + * If the coredump server raises COREDUMP_RECORDS in coredump_ack->mask + * the kernel doesn't send the coredump as a plain byte stream. It sends + * a sequence of records instead. A COREDUMP_RECORD_DATA record is + * followed by @len bytes of actual coredump data. A + * COREDUMP_RECORD_ZERO record is followed by nothing and stands for + * @len zero bytes. A server that didn't raise COREDUMP_SPARSE never + * sees a zero record. Records arrive in order and leave no gaps. So + * @offset is the sum of the @len of all records before it. + * + * The last record is a COREDUMP_RECORD_END record. It is followed by + * nothing. Its @len is zero. Its @offset is the size of the coredump. + * The kernel only sends it once it has written the whole coredump. A + * server that hits end-of-file without having seen an end record must + * treat the coredump as incomplete. + * + * The @size member is set to the size of struct coredump_record_header + * the kernel knows and lets the header grow later. It comes first so it + * can be peeked. Userspace must consume @size bytes and discard + * anything beyond what it knows. It must refuse a @size smaller than + * COREDUMP_RECORD_HEADER_SIZE_VER0. @size covers the header alone. + * @offset and @len count coredump bytes. + * + * The @flags member carries modifiers that change how the record is to + * be interpreted. No flag is defined yet. Userspace must refuse a + * record carrying a flag or a type it doesn't know. Every new record + * type is raised in coredump_req->mask as a feature of its own. A + * server only ever sees the types it asked for. + * + * COREDUMP_RECORDS must be combined with COREDUMP_KERNEL, and + * COREDUMP_SPARSE with COREDUMP_RECORDS. + */ +struct coredump_record_header { + __u32 size; + __u32 type; + __u64 flags; + __u64 offset; + __u64 len; +}; + +enum { + COREDUMP_RECORD_HEADER_SIZE_VER0 = 32U, /* size of first published struct */ +}; + #endif /* _UAPI_LINUX_COREDUMP_H */ diff --git a/tools/testing/selftests/Makefile b/tools/testing/selftests/Makefile index 273853937c25..79a00e9ee46d 100644 --- a/tools/testing/selftests/Makefile +++ b/tools/testing/selftests/Makefile @@ -34,10 +34,12 @@ TARGETS += exec TARGETS += fchmodat2 TARGETS += filesystems TARGETS += filesystems/binderfs +TARGETS += filesystems/configfs TARGETS += filesystems/epoll TARGETS += filesystems/eventfd TARGETS += filesystems/failfs TARGETS += filesystems/fat +TARGETS += filesystems/file_stressor TARGETS += filesystems/openat2 TARGETS += filesystems/open_tree_ns TARGETS += filesystems/overlayfs @@ -50,6 +52,8 @@ TARGETS += filesystems/empty_mntns TARGETS += filesystems/fsmount_ns TARGETS += filesystems/fscontext_ns TARGETS += filesystems/xattr +TARGETS += filesystems/mntns_unbindable +TARGETS += filesystems/umount_propagation TARGETS += firmware TARGETS += fpu TARGETS += ftrace diff --git a/tools/testing/selftests/clone3/clone3_clear_sighand.c b/tools/testing/selftests/clone3/clone3_clear_sighand.c index de0c9d62015d..0ad62b80ada2 100644 --- a/tools/testing/selftests/clone3/clone3_clear_sighand.c +++ b/tools/testing/selftests/clone3/clone3_clear_sighand.c @@ -50,12 +50,12 @@ static void test_clone3_clear_sighand(void) * Check that CLONE_CLEAR_SIGHAND and CLONE_SIGHAND are mutually * exclusive. */ - args.flags |= CLONE_CLEAR_SIGHAND | CLONE_SIGHAND; + args.flags |= CLONE_VM | CLONE_CLEAR_SIGHAND | CLONE_SIGHAND; args.exit_signal = SIGCHLD; pid = sys_clone3(&args, sizeof(args)); - if (pid > 0) + if (pid != -1 || errno != EINVAL) ksft_exit_fail_msg( - "clone3(CLONE_CLEAR_SIGHAND | CLONE_SIGHAND) succeeded\n"); + "clone3(CLONE_CLEAR_SIGHAND | CLONE_SIGHAND) did not fail with EINVAL\n"); act.sa_handler = nop_handler; ret = sigemptyset(&act.sa_mask); diff --git a/tools/testing/selftests/core/close_range_test.c b/tools/testing/selftests/core/close_range_test.c index f14eca63f20c..20ecb65e529b 100644 --- a/tools/testing/selftests/core/close_range_test.c +++ b/tools/testing/selftests/core/close_range_test.c @@ -36,6 +36,14 @@ static inline int sys_close_range(unsigned int fd, unsigned int max_fd, return syscall(__NR_close_range, fd, max_fd, flags); } +static void clear_cloexec(const int *fds, size_t n) +{ + size_t i; + + for (i = 0; i < n; i++) + fcntl(fds[i], F_SETFD, 0); +} + TEST(core_close_range) { int i, ret; @@ -236,6 +244,50 @@ TEST(close_range_unshare_capped) EXPECT_EQ(0, WEXITSTATUS(status)); } +TEST(close_range_unshare_hole) +{ + int i, status; + pid_t pid; + struct __clone_args args = { + .flags = CLONE_FILES, + .exit_signal = SIGCHLD, + }; + + /* Fill the first two words of the table. */ + for (i = 3; i < 128; i++) + ASSERT_GE(dup2(0, i), 0); + + pid = sys_clone3(&args, sizeof(args)); + ASSERT_GE(pid, 0); + + if (pid == 0) { + /* Punch a hole into the second word, behind a full first one. */ + if (sys_close_range(70, 80, CLOSE_RANGE_UNSHARE)) + exit(EXIT_FAILURE); + + for (i = 3; i < 128; i++) { + bool closed = i >= 70 && i <= 80; + + if (closed == (fcntl(i, F_GETFD) != -1)) + exit(EXIT_FAILURE); + } + + /* A stale full bit on word 1 would hand out 128, not 70. */ + if (dup(0) != 70) + exit(EXIT_FAILURE); + + exit(EXIT_SUCCESS); + } + + EXPECT_EQ(waitpid(pid, &status, 0), pid); + EXPECT_EQ(true, WIFEXITED(status)); + EXPECT_EQ(0, WEXITSTATUS(status)); + + /* The shared table the child unshared from is untouched. */ + for (i = 3; i < 128; i++) + EXPECT_NE(-1, fcntl(i, F_GETFD)); +} + TEST(close_range_cloexec) { int i, ret; @@ -593,6 +645,905 @@ TEST(close_range_cloexec_unshare_syzbot) EXPECT_EQ(close(fd3), 0); } +TEST(close_range_except) +{ + int i, ret, status; + pid_t pid; + int open_fds[101]; + struct __clone_args args = { + .exit_signal = SIGCHLD, + }; + + for (i = 0; i < ARRAY_SIZE(open_fds); i++) { + int fd; + + fd = open("/dev/null", O_RDONLY); + ASSERT_GE(fd, 0) { + if (errno == ENOENT) + SKIP(return, "Skipping test since /dev/null does not exist"); + } + + open_fds[i] = fd; + } + + /* A range covering everything keeps everything. */ + ret = sys_close_range(0, UINT_MAX, CLOSE_RANGE_EXCEPT); + if (ret < 0) { + if (errno == ENOSYS) + SKIP(return, "close_range() syscall not supported"); + if (errno == EINVAL) + SKIP(return, "close_range() doesn't support CLOSE_RANGE_EXCEPT"); + } + ASSERT_EQ(0, ret); + + for (i = 0; i < ARRAY_SIZE(open_fds); i++) + EXPECT_NE(-1, fcntl(open_fds[i], F_GETFD)); + + /* The bounds are checked before the range is turned around. */ + EXPECT_EQ(-1, sys_close_range(open_fds[20], open_fds[10], + CLOSE_RANGE_EXCEPT)); + EXPECT_EQ(EINVAL, errno); + + /* Everything above open_fds[50] goes. */ + ASSERT_EQ(0, sys_close_range(0, open_fds[50], CLOSE_RANGE_EXCEPT)); + + for (i = 0; i < ARRAY_SIZE(open_fds); i++) + EXPECT_EQ(i <= 50, fcntl(open_fds[i], F_GETFD) != -1); + + /* A window in the middle takes stdio with it, so do that in a fork. */ + pid = sys_clone3(&args, sizeof(args)); + ASSERT_GE(pid, 0); + + if (pid == 0) { + ret = sys_close_range(open_fds[10], open_fds[20], + CLOSE_RANGE_EXCEPT); + if (ret) + exit(EXIT_FAILURE); + + for (i = 0; i <= 50; i++) { + bool kept = i >= 10 && i <= 20; + + if (kept != (fcntl(open_fds[i], F_GETFD) != -1)) + exit(EXIT_FAILURE); + } + + if (fcntl(STDERR_FILENO, F_GETFD) != -1) + exit(EXIT_FAILURE); + + exit(EXIT_SUCCESS); + } + + EXPECT_EQ(waitpid(pid, &status, 0), pid); + EXPECT_EQ(true, WIFEXITED(status)); + EXPECT_EQ(0, WEXITSTATUS(status)); + + /* The fork had a table of its own. */ + for (i = 0; i <= 50; i++) + EXPECT_NE(-1, fcntl(open_fds[i], F_GETFD)); +} + +TEST(close_range_except_cloexec) +{ + int i, ret; + int open_fds[101]; + + for (i = 0; i < ARRAY_SIZE(open_fds); i++) { + int fd; + + fd = open("/dev/null", O_RDONLY); + ASSERT_GE(fd, 0) { + if (errno == ENOENT) + SKIP(return, "Skipping test since /dev/null does not exist"); + } + + open_fds[i] = fd; + } + + ret = sys_close_range(open_fds[10], open_fds[20], + CLOSE_RANGE_CLOEXEC | CLOSE_RANGE_EXCEPT); + if (ret < 0) { + if (errno == ENOSYS) + SKIP(return, "close_range() syscall not supported"); + if (errno == EINVAL) + SKIP(return, "close_range() doesn't support CLOSE_RANGE_EXCEPT"); + } + ASSERT_EQ(0, ret); + + for (i = 0; i < ARRAY_SIZE(open_fds); i++) { + bool inside = i >= 10 && i <= 20; + int flags = fcntl(open_fds[i], F_GETFD); + + EXPECT_NE(-1, flags); + EXPECT_EQ(inside ? 0 : FD_CLOEXEC, flags & FD_CLOEXEC); + } + + /* stdio sits outside of the window too. */ + EXPECT_EQ(FD_CLOEXEC, fcntl(STDERR_FILENO, F_GETFD) & FD_CLOEXEC); + + /* A window that starts at 0 marks only what lies above it. */ + clear_cloexec(open_fds, ARRAY_SIZE(open_fds)); + ASSERT_EQ(0, fcntl(STDERR_FILENO, F_SETFD, 0)); + ASSERT_EQ(0, sys_close_range(0, open_fds[20], + CLOSE_RANGE_CLOEXEC | CLOSE_RANGE_EXCEPT)); + + for (i = 0; i < ARRAY_SIZE(open_fds); i++) + EXPECT_EQ(i <= 20 ? 0 : FD_CLOEXEC, + fcntl(open_fds[i], F_GETFD) & FD_CLOEXEC); + EXPECT_EQ(0, fcntl(STDERR_FILENO, F_GETFD) & FD_CLOEXEC); + + /* One at the top marks only what lies below it, stdio included. */ + clear_cloexec(open_fds, ARRAY_SIZE(open_fds)); + ASSERT_EQ(0, sys_close_range(open_fds[80], UINT_MAX, + CLOSE_RANGE_CLOEXEC | CLOSE_RANGE_EXCEPT)); + + for (i = 0; i < ARRAY_SIZE(open_fds); i++) + EXPECT_EQ(i < 80 ? FD_CLOEXEC : 0, + fcntl(open_fds[i], F_GETFD) & FD_CLOEXEC); + EXPECT_EQ(FD_CLOEXEC, fcntl(STDERR_FILENO, F_GETFD) & FD_CLOEXEC); + + /* One that cannot hold a descriptor marks everything. */ + clear_cloexec(open_fds, ARRAY_SIZE(open_fds)); + ASSERT_EQ(0, fcntl(STDERR_FILENO, F_SETFD, 0)); + ASSERT_EQ(0, sys_close_range(UINT_MAX, UINT_MAX, + CLOSE_RANGE_CLOEXEC | CLOSE_RANGE_EXCEPT)); + + for (i = 0; i < ARRAY_SIZE(open_fds); i++) + EXPECT_EQ(FD_CLOEXEC, fcntl(open_fds[i], F_GETFD) & FD_CLOEXEC); + EXPECT_EQ(FD_CLOEXEC, fcntl(STDERR_FILENO, F_GETFD) & FD_CLOEXEC); +} + +TEST(close_range_except_cloexec_unshare) +{ + int i, ret, status; + pid_t pid; + int open_fds[101]; + struct __clone_args args = { + .flags = CLONE_FILES, + .exit_signal = SIGCHLD, + }; + + for (i = 0; i < ARRAY_SIZE(open_fds); i++) { + int fd; + + fd = open("/dev/null", O_RDONLY); + ASSERT_GE(fd, 0) { + if (errno == ENOENT) + SKIP(return, "Skipping test since /dev/null does not exist"); + } + + open_fds[i] = fd; + } + + /* A range covering everything marks nothing. */ + ret = sys_close_range(0, UINT_MAX, + CLOSE_RANGE_CLOEXEC | CLOSE_RANGE_EXCEPT); + if (ret < 0) { + if (errno == ENOSYS) + SKIP(return, "close_range() syscall not supported"); + if (errno == EINVAL) + SKIP(return, "close_range() doesn't support CLOSE_RANGE_EXCEPT"); + } + ASSERT_EQ(0, ret); + + for (i = 0; i < ARRAY_SIZE(open_fds); i++) + EXPECT_EQ(0, fcntl(open_fds[i], F_GETFD) & FD_CLOEXEC); + + pid = sys_clone3(&args, sizeof(args)); + ASSERT_GE(pid, 0); + + if (pid == 0) { + ret = sys_close_range(open_fds[10], open_fds[20], + CLOSE_RANGE_UNSHARE | CLOSE_RANGE_CLOEXEC | + CLOSE_RANGE_EXCEPT); + if (ret) + exit(EXIT_FAILURE); + + for (i = 0; i < ARRAY_SIZE(open_fds); i++) { + bool inside = i >= 10 && i <= 20; + int flags = fcntl(open_fds[i], F_GETFD); + + if (flags == -1) + exit(EXIT_FAILURE); + if ((flags & FD_CLOEXEC) != (inside ? 0 : FD_CLOEXEC)) + exit(EXIT_FAILURE); + } + + exit(EXIT_SUCCESS); + } + + EXPECT_EQ(waitpid(pid, &status, 0), pid); + EXPECT_EQ(true, WIFEXITED(status)); + EXPECT_EQ(0, WEXITSTATUS(status)); + + /* The shared table the child unshared from is untouched. */ + for (i = 0; i < ARRAY_SIZE(open_fds); i++) + EXPECT_EQ(0, fcntl(open_fds[i], F_GETFD) & FD_CLOEXEC); +} + +TEST(close_range_except_bounds) +{ + int i, c, ret, status; + pid_t pid; + int open_fds[101]; + struct __clone_args args = { + .exit_signal = SIGCHLD, + }; + + for (i = 0; i < ARRAY_SIZE(open_fds); i++) { + int fd; + + fd = open("/dev/null", O_RDONLY); + ASSERT_GE(fd, 0) { + if (errno == ENOENT) + SKIP(return, "Skipping test since /dev/null does not exist"); + } + + open_fds[i] = fd; + } + + /* A range covering everything keeps everything. */ + ret = sys_close_range(0, UINT_MAX, CLOSE_RANGE_EXCEPT); + if (ret < 0) { + if (errno == ENOSYS) + SKIP(return, "close_range() syscall not supported"); + if (errno == EINVAL) + SKIP(return, "close_range() doesn't support CLOSE_RANGE_EXCEPT"); + } + ASSERT_EQ(0, ret); + + struct { + unsigned int fd, max_fd, flags; + } cases[] = { + /* A window at the top drops everything below it. */ + { open_fds[50], UINT_MAX, CLOSE_RANGE_EXCEPT }, + /* One that cannot hold a descriptor keeps nothing. */ + { UINT_MAX, UINT_MAX, CLOSE_RANGE_EXCEPT }, + /* One of a single descriptor keeps just that. */ + { open_fds[30], open_fds[30], CLOSE_RANGE_EXCEPT }, + /* The unshare form on a table that is not shared acts in place. */ + { open_fds[10], open_fds[20], + CLOSE_RANGE_UNSHARE | CLOSE_RANGE_EXCEPT }, + }; + + /* Each of them takes stdio with it, so do that in a fork. */ + for (c = 0; c < ARRAY_SIZE(cases); c++) { + pid = sys_clone3(&args, sizeof(args)); + ASSERT_GE(pid, 0); + + if (pid == 0) { + ret = sys_close_range(cases[c].fd, cases[c].max_fd, + cases[c].flags); + if (ret) + exit(EXIT_FAILURE); + + for (i = 0; i < ARRAY_SIZE(open_fds); i++) { + unsigned int fd = open_fds[i]; + bool kept = fd >= cases[c].fd && + fd <= cases[c].max_fd; + + if (kept != (fcntl(fd, F_GETFD) != -1)) + exit(EXIT_FAILURE); + } + + if (fcntl(STDERR_FILENO, F_GETFD) != -1) + exit(EXIT_FAILURE); + + exit(EXIT_SUCCESS); + } + + EXPECT_EQ(waitpid(pid, &status, 0), pid); + EXPECT_EQ(true, WIFEXITED(status)); + EXPECT_EQ(0, WEXITSTATUS(status)); + } + + /* Each fork had a table of its own. */ + for (i = 0; i < ARRAY_SIZE(open_fds); i++) + EXPECT_NE(-1, fcntl(open_fds[i], F_GETFD)); +} + +TEST(close_range_except_unshare) +{ + int i, ret, status; + pid_t pid; + int open_fds[200]; + struct __clone_args args = { + .flags = CLONE_FILES, + .exit_signal = SIGCHLD, + }; + + for (i = 0; i < ARRAY_SIZE(open_fds); i++) { + int fd; + + /* Odd slots are close-on-exec, which makes no difference here. */ + fd = open("/dev/null", O_RDONLY | (i % 2 ? O_CLOEXEC : 0)); + ASSERT_GE(fd, 0) { + if (errno == ENOENT) + SKIP(return, "Skipping test since /dev/null does not exist"); + } + + open_fds[i] = fd; + } + + /* A range covering everything keeps everything. */ + ret = sys_close_range(0, UINT_MAX, CLOSE_RANGE_EXCEPT); + if (ret < 0) { + if (errno == ENOSYS) + SKIP(return, "close_range() syscall not supported"); + if (errno == EINVAL) + SKIP(return, "close_range() doesn't support CLOSE_RANGE_EXCEPT"); + } + ASSERT_EQ(0, ret); + + pid = sys_clone3(&args, sizeof(args)); + ASSERT_GE(pid, 0); + + if (pid == 0) { + /* The window sits near the top, so the clone is sized off its end. */ + ret = sys_close_range(open_fds[150], open_fds[160], + CLOSE_RANGE_UNSHARE | CLOSE_RANGE_EXCEPT); + if (ret) + exit(EXIT_FAILURE); + + for (i = 0; i < ARRAY_SIZE(open_fds); i++) { + bool kept = i >= 150 && i <= 160; + + if (kept != (fcntl(open_fds[i], F_GETFD) != -1)) + exit(EXIT_FAILURE); + } + + if (fcntl(STDERR_FILENO, F_GETFD) != -1) + exit(EXIT_FAILURE); + + /* What was left behind is handed out again, from the bottom. */ + if (dup(open_fds[150]) != 0) + exit(EXIT_FAILURE); + + exit(EXIT_SUCCESS); + } + + EXPECT_EQ(waitpid(pid, &status, 0), pid); + EXPECT_EQ(true, WIFEXITED(status)); + EXPECT_EQ(0, WEXITSTATUS(status)); + + /* A window at the bottom keeps just that, stdio included. */ + pid = sys_clone3(&args, sizeof(args)); + ASSERT_GE(pid, 0); + + if (pid == 0) { + ret = sys_close_range(0, open_fds[10], + CLOSE_RANGE_UNSHARE | CLOSE_RANGE_EXCEPT); + if (ret) + exit(EXIT_FAILURE); + + for (i = 0; i < ARRAY_SIZE(open_fds); i++) { + if ((i <= 10) != (fcntl(open_fds[i], F_GETFD) != -1)) + exit(EXIT_FAILURE); + } + + if (fcntl(STDERR_FILENO, F_GETFD) == -1) + exit(EXIT_FAILURE); + + /* The first slot left behind is the next one handed out. */ + if (dup(0) != open_fds[10] + 1) + exit(EXIT_FAILURE); + + exit(EXIT_SUCCESS); + } + + EXPECT_EQ(waitpid(pid, &status, 0), pid); + EXPECT_EQ(true, WIFEXITED(status)); + EXPECT_EQ(0, WEXITSTATUS(status)); + + /* A window that cannot hold a descriptor keeps nothing. */ + pid = sys_clone3(&args, sizeof(args)); + ASSERT_GE(pid, 0); + + if (pid == 0) { + ret = sys_close_range(UINT_MAX, UINT_MAX, + CLOSE_RANGE_UNSHARE | CLOSE_RANGE_EXCEPT); + if (ret) + exit(EXIT_FAILURE); + + for (i = 0; i < ARRAY_SIZE(open_fds); i++) + if (fcntl(open_fds[i], F_GETFD) != -1) + exit(EXIT_FAILURE); + + if (fcntl(STDERR_FILENO, F_GETFD) != -1) + exit(EXIT_FAILURE); + + exit(EXIT_SUCCESS); + } + + EXPECT_EQ(waitpid(pid, &status, 0), pid); + EXPECT_EQ(true, WIFEXITED(status)); + EXPECT_EQ(0, WEXITSTATUS(status)); + + /* The shared table the child unshared from is untouched. */ + for (i = 0; i < ARRAY_SIZE(open_fds); i++) + EXPECT_NE(-1, fcntl(open_fds[i], F_GETFD)); +} + +TEST(close_range_cloexec_only) +{ + int i, ret; + int open_fds[101]; + + for (i = 0; i < ARRAY_SIZE(open_fds); i++) { + int fd; + + /* Odd slots are close-on-exec, even ones are not. */ + fd = open("/dev/null", O_RDONLY | (i % 2 ? O_CLOEXEC : 0)); + ASSERT_GE(fd, 0) { + if (errno == ENOENT) + SKIP(return, "Skipping test since /dev/null does not exist"); + } + + open_fds[i] = fd; + } + + ret = sys_close_range(open_fds[10], open_fds[20], + CLOSE_RANGE_CLOEXEC_ONLY); + if (ret < 0) { + if (errno == ENOSYS) + SKIP(return, "close_range() syscall not supported"); + if (errno == EINVAL) + SKIP(return, "close_range() doesn't support CLOSE_RANGE_CLOEXEC_ONLY"); + } + ASSERT_EQ(0, ret); + + for (i = 0; i < ARRAY_SIZE(open_fds); i++) { + bool closed = i % 2 && i >= 10 && i <= 20; + + EXPECT_EQ(!closed, fcntl(open_fds[i], F_GETFD) != -1); + } + + /* A range above the table closes nothing. */ + ASSERT_EQ(0, sys_close_range(UINT_MAX, UINT_MAX, + CLOSE_RANGE_CLOEXEC_ONLY)); + + for (i = 0; i < ARRAY_SIZE(open_fds); i++) { + bool closed = i % 2 && i >= 10 && i <= 20; + + EXPECT_EQ(!closed, fcntl(open_fds[i], F_GETFD) != -1); + } + + /* Do what an exec would do to the rest. */ + ASSERT_EQ(0, sys_close_range(0, UINT_MAX, CLOSE_RANGE_CLOEXEC_ONLY)); + + for (i = 0; i < ARRAY_SIZE(open_fds); i++) + EXPECT_EQ(!(i % 2), fcntl(open_fds[i], F_GETFD) != -1); +} + +TEST(close_range_cloexec_only_except) +{ + int i, ret; + int open_fds[101]; + + for (i = 0; i < ARRAY_SIZE(open_fds); i++) { + int fd; + + fd = open("/dev/null", O_RDONLY | (i % 2 ? O_CLOEXEC : 0)); + ASSERT_GE(fd, 0) { + if (errno == ENOENT) + SKIP(return, "Skipping test since /dev/null does not exist"); + } + + open_fds[i] = fd; + } + + ret = sys_close_range(open_fds[10], open_fds[20], + CLOSE_RANGE_CLOEXEC_ONLY | CLOSE_RANGE_EXCEPT); + if (ret < 0) { + if (errno == ENOSYS) + SKIP(return, "close_range() syscall not supported"); + if (errno == EINVAL) + SKIP(return, "close_range() doesn't support CLOSE_RANGE_CLOEXEC_ONLY"); + } + ASSERT_EQ(0, ret); + + for (i = 0; i < ARRAY_SIZE(open_fds); i++) { + bool kept = !(i % 2) || (i >= 10 && i <= 20); + int flags = i % 2 ? FD_CLOEXEC : 0; + + /* The kept ones keep their flag, so exec still drops them. */ + EXPECT_EQ(kept ? flags : -1, fcntl(open_fds[i], F_GETFD)); + } + + /* A range that cannot hold an open descriptor keeps nothing. */ + ASSERT_EQ(0, sys_close_range(UINT_MAX, UINT_MAX, + CLOSE_RANGE_CLOEXEC_ONLY | + CLOSE_RANGE_EXCEPT)); + + for (i = 0; i < ARRAY_SIZE(open_fds); i++) + EXPECT_EQ(!(i % 2), fcntl(open_fds[i], F_GETFD) != -1); +} + +TEST(close_range_cloexec_only_except_bounds) +{ + int i, c, ret, status; + pid_t pid; + int open_fds[101]; + struct __clone_args args = { + .exit_signal = SIGCHLD, + }; + + for (i = 0; i < ARRAY_SIZE(open_fds); i++) { + int fd; + + fd = open("/dev/null", O_RDONLY | (i % 2 ? O_CLOEXEC : 0)); + ASSERT_GE(fd, 0) { + if (errno == ENOENT) + SKIP(return, "Skipping test since /dev/null does not exist"); + } + + open_fds[i] = fd; + } + + /* A range covering everything keeps everything. */ + ret = sys_close_range(0, UINT_MAX, + CLOSE_RANGE_CLOEXEC_ONLY | CLOSE_RANGE_EXCEPT); + if (ret < 0) { + if (errno == ENOSYS) + SKIP(return, "close_range() syscall not supported"); + if (errno == EINVAL) + SKIP(return, "close_range() doesn't support CLOSE_RANGE_CLOEXEC_ONLY"); + } + ASSERT_EQ(0, ret); + + struct { + unsigned int fd, max_fd; + } cases[] = { + /* A window at the top keeps the marked ones in it. */ + { open_fds[80], UINT_MAX }, + /* One at the bottom keeps the marked ones in it. */ + { 0, open_fds[20] }, + /* One that cannot hold a descriptor keeps none of them. */ + { UINT_MAX, UINT_MAX }, + }; + + for (c = 0; c < ARRAY_SIZE(cases); c++) { + pid = sys_clone3(&args, sizeof(args)); + ASSERT_GE(pid, 0); + + if (pid == 0) { + ret = sys_close_range(cases[c].fd, cases[c].max_fd, + CLOSE_RANGE_CLOEXEC_ONLY | + CLOSE_RANGE_EXCEPT); + if (ret) + exit(EXIT_FAILURE); + + for (i = 0; i < ARRAY_SIZE(open_fds); i++) { + unsigned int fd = open_fds[i]; + bool kept = !(i % 2) || (fd >= cases[c].fd && + fd <= cases[c].max_fd); + int flags = i % 2 ? FD_CLOEXEC : 0; + + if (fcntl(fd, F_GETFD) != (kept ? flags : -1)) + exit(EXIT_FAILURE); + } + + /* stdio is neither marked nor gone. */ + if (fcntl(STDERR_FILENO, F_GETFD) & FD_CLOEXEC) + exit(EXIT_FAILURE); + + exit(EXIT_SUCCESS); + } + + EXPECT_EQ(waitpid(pid, &status, 0), pid); + EXPECT_EQ(true, WIFEXITED(status)); + EXPECT_EQ(0, WEXITSTATUS(status)); + } + + /* Each fork had a table of its own. */ + for (i = 0; i < ARRAY_SIZE(open_fds); i++) + EXPECT_NE(-1, fcntl(open_fds[i], F_GETFD)); +} + +TEST(close_range_cloexec_only_unshare) +{ + int i, ret, status; + pid_t pid; + int open_fds[101]; + struct __clone_args args = { + .flags = CLONE_FILES, + .exit_signal = SIGCHLD, + }; + + for (i = 0; i < ARRAY_SIZE(open_fds); i++) { + int fd; + + fd = open("/dev/null", O_RDONLY | (i % 2 ? O_CLOEXEC : 0)); + ASSERT_GE(fd, 0) { + if (errno == ENOENT) + SKIP(return, "Skipping test since /dev/null does not exist"); + } + + open_fds[i] = fd; + } + + /* A range covering everything keeps everything. */ + ret = sys_close_range(0, UINT_MAX, + CLOSE_RANGE_CLOEXEC_ONLY | CLOSE_RANGE_EXCEPT); + if (ret < 0) { + if (errno == ENOSYS) + SKIP(return, "close_range() syscall not supported"); + if (errno == EINVAL) + SKIP(return, "close_range() doesn't support CLOSE_RANGE_CLOEXEC_ONLY"); + } + ASSERT_EQ(0, ret); + + for (i = 0; i < ARRAY_SIZE(open_fds); i++) + ASSERT_NE(-1, fcntl(open_fds[i], F_GETFD)); + + pid = sys_clone3(&args, sizeof(args)); + ASSERT_GE(pid, 0); + + if (pid == 0) { + ret = sys_close_range(open_fds[10], open_fds[20], + CLOSE_RANGE_UNSHARE | + CLOSE_RANGE_CLOEXEC_ONLY); + if (ret) + exit(EXIT_FAILURE); + + for (i = 0; i < ARRAY_SIZE(open_fds); i++) { + bool closed = i % 2 && i >= 10 && i <= 20; + + if (closed == (fcntl(open_fds[i], F_GETFD) != -1)) + exit(EXIT_FAILURE); + } + + exit(EXIT_SUCCESS); + } + + EXPECT_EQ(waitpid(pid, &status, 0), pid); + EXPECT_EQ(true, WIFEXITED(status)); + EXPECT_EQ(0, WEXITSTATUS(status)); + + /* A range at the top keeps the descriptors without the flag in it. */ + pid = sys_clone3(&args, sizeof(args)); + ASSERT_GE(pid, 0); + + if (pid == 0) { + ret = sys_close_range(open_fds[50], UINT_MAX, + CLOSE_RANGE_UNSHARE | + CLOSE_RANGE_CLOEXEC_ONLY); + if (ret) + exit(EXIT_FAILURE); + + for (i = 0; i < ARRAY_SIZE(open_fds); i++) { + bool closed = i % 2 && i >= 50; + + if (closed == (fcntl(open_fds[i], F_GETFD) != -1)) + exit(EXIT_FAILURE); + } + + /* The first slot left behind is the next one handed out. */ + if (dup(0) != open_fds[51]) + exit(EXIT_FAILURE); + + exit(EXIT_SUCCESS); + } + + EXPECT_EQ(waitpid(pid, &status, 0), pid); + EXPECT_EQ(true, WIFEXITED(status)); + EXPECT_EQ(0, WEXITSTATUS(status)); + + /* The shared table the child unshared from is untouched. */ + for (i = 0; i < ARRAY_SIZE(open_fds); i++) + EXPECT_NE(-1, fcntl(open_fds[i], F_GETFD)); +} + +TEST(close_range_cloexec_only_except_unshare) +{ + int i, ret, status; + pid_t pid; + int open_fds[101]; + struct __clone_args args = { + .flags = CLONE_FILES, + .exit_signal = SIGCHLD, + }; + + for (i = 0; i < ARRAY_SIZE(open_fds); i++) { + int fd; + + fd = open("/dev/null", O_RDONLY | (i % 2 ? O_CLOEXEC : 0)); + ASSERT_GE(fd, 0) { + if (errno == ENOENT) + SKIP(return, "Skipping test since /dev/null does not exist"); + } + + open_fds[i] = fd; + } + + /* A range covering everything keeps everything. */ + ret = sys_close_range(0, UINT_MAX, + CLOSE_RANGE_CLOEXEC_ONLY | CLOSE_RANGE_EXCEPT); + if (ret < 0) { + if (errno == ENOSYS) + SKIP(return, "close_range() syscall not supported"); + if (errno == EINVAL) + SKIP(return, "close_range() doesn't support CLOSE_RANGE_CLOEXEC_ONLY"); + } + ASSERT_EQ(0, ret); + + for (i = 0; i < ARRAY_SIZE(open_fds); i++) + ASSERT_NE(-1, fcntl(open_fds[i], F_GETFD)); + + pid = sys_clone3(&args, sizeof(args)); + ASSERT_GE(pid, 0); + + if (pid == 0) { + ret = sys_close_range(open_fds[10], open_fds[20], + CLOSE_RANGE_UNSHARE | + CLOSE_RANGE_CLOEXEC_ONLY | + CLOSE_RANGE_EXCEPT); + if (ret) + exit(EXIT_FAILURE); + + for (i = 0; i < ARRAY_SIZE(open_fds); i++) { + bool kept = !(i % 2) || (i >= 10 && i <= 20); + int flags = i % 2 ? FD_CLOEXEC : 0; + + if (fcntl(open_fds[i], F_GETFD) != (kept ? flags : -1)) + exit(EXIT_FAILURE); + } + + exit(EXIT_SUCCESS); + } + + EXPECT_EQ(waitpid(pid, &status, 0), pid); + EXPECT_EQ(true, WIFEXITED(status)); + EXPECT_EQ(0, WEXITSTATUS(status)); + + /* A window that cannot hold a descriptor keeps none of the marked. */ + pid = sys_clone3(&args, sizeof(args)); + ASSERT_GE(pid, 0); + + if (pid == 0) { + ret = sys_close_range(UINT_MAX, UINT_MAX, + CLOSE_RANGE_UNSHARE | + CLOSE_RANGE_CLOEXEC_ONLY | + CLOSE_RANGE_EXCEPT); + if (ret) + exit(EXIT_FAILURE); + + for (i = 0; i < ARRAY_SIZE(open_fds); i++) + if ((i % 2) == (fcntl(open_fds[i], F_GETFD) != -1)) + exit(EXIT_FAILURE); + + if (fcntl(STDERR_FILENO, F_GETFD) == -1) + exit(EXIT_FAILURE); + + exit(EXIT_SUCCESS); + } + + EXPECT_EQ(waitpid(pid, &status, 0), pid); + EXPECT_EQ(true, WIFEXITED(status)); + EXPECT_EQ(0, WEXITSTATUS(status)); + + /* One that covers everything keeps everything, in a clone too. */ + pid = sys_clone3(&args, sizeof(args)); + ASSERT_GE(pid, 0); + + if (pid == 0) { + ret = sys_close_range(0, UINT_MAX, + CLOSE_RANGE_UNSHARE | + CLOSE_RANGE_CLOEXEC_ONLY | + CLOSE_RANGE_EXCEPT); + if (ret) + exit(EXIT_FAILURE); + + for (i = 0; i < ARRAY_SIZE(open_fds); i++) + if (fcntl(open_fds[i], F_GETFD) != (i % 2 ? FD_CLOEXEC : 0)) + exit(EXIT_FAILURE); + + exit(EXIT_SUCCESS); + } + + EXPECT_EQ(waitpid(pid, &status, 0), pid); + EXPECT_EQ(true, WIFEXITED(status)); + EXPECT_EQ(0, WEXITSTATUS(status)); + + /* The shared table the child unshared from is untouched. */ + for (i = 0; i < ARRAY_SIZE(open_fds); i++) + EXPECT_NE(-1, fcntl(open_fds[i], F_GETFD)); +} + +TEST(close_range_cloexec_only_except_unshare_sizing) +{ + int i, ret, status; + pid_t pid; + int open_fds[200]; + struct __clone_args args = { + .flags = CLONE_FILES, + .exit_signal = SIGCHLD, + }; + + /* All close-on-exec, so the kept range alone sizes the clone. */ + for (i = 0; i < ARRAY_SIZE(open_fds); i++) { + int fd; + + fd = open("/dev/null", O_RDONLY | O_CLOEXEC); + ASSERT_GE(fd, 0) { + if (errno == ENOENT) + SKIP(return, "Skipping test since /dev/null does not exist"); + } + + open_fds[i] = fd; + } + + ret = sys_close_range(0, UINT_MAX, + CLOSE_RANGE_CLOEXEC_ONLY | CLOSE_RANGE_EXCEPT); + if (ret < 0) { + if (errno == ENOSYS) + SKIP(return, "close_range() syscall not supported"); + if (errno == EINVAL) + SKIP(return, "close_range() doesn't support CLOSE_RANGE_CLOEXEC_ONLY"); + } + ASSERT_EQ(0, ret); + + pid = sys_clone3(&args, sizeof(args)); + ASSERT_GE(pid, 0); + + if (pid == 0) { + ret = sys_close_range(open_fds[150], open_fds[160], + CLOSE_RANGE_UNSHARE | + CLOSE_RANGE_CLOEXEC_ONLY | + CLOSE_RANGE_EXCEPT); + if (ret) + exit(EXIT_FAILURE); + + for (i = 0; i < ARRAY_SIZE(open_fds); i++) { + bool kept = i >= 150 && i <= 160; + + if (kept != (fcntl(open_fds[i], F_GETFD) != -1)) + exit(EXIT_FAILURE); + } + + /* Nothing set close-on-exec on stdio. */ + if (fcntl(STDERR_FILENO, F_GETFD) == -1) + exit(EXIT_FAILURE); + + exit(EXIT_SUCCESS); + } + + EXPECT_EQ(waitpid(pid, &status, 0), pid); + EXPECT_EQ(true, WIFEXITED(status)); + EXPECT_EQ(0, WEXITSTATUS(status)); +} + +TEST(close_range_cloexec_only_einval) +{ + int ret; + + /* A range covering everything keeps everything, so this only probes. */ + ret = sys_close_range(0, UINT_MAX, + CLOSE_RANGE_CLOEXEC_ONLY | CLOSE_RANGE_EXCEPT); + if (ret < 0) { + if (errno == ENOSYS) + SKIP(return, "close_range() syscall not supported"); + if (errno == EINVAL) + SKIP(return, "close_range() doesn't support CLOSE_RANGE_CLOEXEC_ONLY"); + } + ASSERT_EQ(0, ret); + + EXPECT_EQ(-1, sys_close_range(3, UINT_MAX, CLOSE_RANGE_CLOEXEC | + CLOSE_RANGE_CLOEXEC_ONLY)); + EXPECT_EQ(EINVAL, errno); + + /* The other flags do not make the pair acceptable. */ + EXPECT_EQ(-1, sys_close_range(3, UINT_MAX, CLOSE_RANGE_UNSHARE | + CLOSE_RANGE_CLOEXEC | + CLOSE_RANGE_CLOEXEC_ONLY | + CLOSE_RANGE_EXCEPT)); + EXPECT_EQ(EINVAL, errno); + + /* The bounds are checked with the new flag too. */ + EXPECT_EQ(-1, sys_close_range(4, 3, CLOSE_RANGE_CLOEXEC_ONLY | + CLOSE_RANGE_EXCEPT)); + EXPECT_EQ(EINVAL, errno); +} + TEST(close_range_bitmap_corruption) { pid_t pid; diff --git a/tools/testing/selftests/coredump/.gitignore b/tools/testing/selftests/coredump/.gitignore index 097f52db0be9..e32f6e9006f6 100644 --- a/tools/testing/selftests/coredump/.gitignore +++ b/tools/testing/selftests/coredump/.gitignore @@ -2,3 +2,5 @@ stackdump_test coredump_socket_test coredump_socket_protocol_test +coredump_signal_test +coredump_worker_test diff --git a/tools/testing/selftests/coredump/Makefile b/tools/testing/selftests/coredump/Makefile index dece1a31d561..c4cdab7b7d0b 100644 --- a/tools/testing/selftests/coredump/Makefile +++ b/tools/testing/selftests/coredump/Makefile @@ -3,7 +3,12 @@ CFLAGS += -Wall -O0 -g $(KHDR_INCLUDES) $(TOOLS_INCLUDES) TEST_GEN_PROGS := stackdump_test \ coredump_socket_test \ - coredump_socket_protocol_test + coredump_socket_protocol_test \ + coredump_notify_signal_test \ + coredump_signal_test \ + coredump_worker_test +# Spawned by the kernel as the |helper, not a test of its own. +TEST_GEN_FILES := coredump_notify_signal_helper TEST_FILES := stackdump include ../lib.mk @@ -11,3 +16,7 @@ include ../lib.mk $(OUTPUT)/stackdump_test: coredump_test_helpers.c $(OUTPUT)/coredump_socket_test: coredump_test_helpers.c $(OUTPUT)/coredump_socket_protocol_test: coredump_test_helpers.c +$(OUTPUT)/coredump_notify_signal_test: coredump_test_helpers.c +$(OUTPUT)/coredump_notify_signal_helper: coredump_test_helpers.c +$(OUTPUT)/coredump_signal_test: coredump_test_helpers.c +$(OUTPUT)/coredump_worker_test: coredump_test_helpers.c diff --git a/tools/testing/selftests/coredump/coredump_notify_signal.h b/tools/testing/selftests/coredump/coredump_notify_signal.h new file mode 100644 index 000000000000..92868a43425c --- /dev/null +++ b/tools/testing/selftests/coredump/coredump_notify_signal.h @@ -0,0 +1,29 @@ +/* SPDX-License-Identifier: GPL-2.0 */ + +#ifndef __COREDUMP_NOTIFY_SIGNAL_H +#define __COREDUMP_NOTIFY_SIGNAL_H + +#include <stdbool.h> +#include <sys/types.h> + +/* + * Define a bunch of constants we need. We create a situation where the + * NT_FILE note blows past 200K. That's way beyond the default 64K + * pipe ring and past the ~36K an af_unix skb holds. So we force a write + * to come back short. + */ +#define NOTIFY_SIGNAL_MAP_COUNT 4000 +#define NOTIFY_SIGNAL_ANON_BYTES (4UL << 20) +#define NOTIFY_SIGNAL_STALL_US 200000 +#define NOTIFY_SIGNAL_MAPFILE "/tmp/coredump.notify_signal.mapfile" +#define NOTIFY_SIGNAL_TRIGGER "/tmp/coredump.notify_signal.trigger" +#define NOTIFY_SIGNAL_CORE_FILE "/tmp/coredump.notify_signal.core" +#define NOTIFY_SIGNAL_CORE_TMPFILE "/tmp/coredump.notify_signal.core.tmp" +#define NOTIFY_SIGNAL_SOCKET "/tmp/coredump.notify_signal.socket" + +void crashing_child_notify_signal(void); +bool coredump_io_uring_available(void); +ssize_t recv_coredump_notify_signal(int fd, int fd_out, bool arm); +long long coredump_expected_size(const char *path); + +#endif /* __COREDUMP_NOTIFY_SIGNAL_H */ diff --git a/tools/testing/selftests/coredump/coredump_notify_signal_helper.c b/tools/testing/selftests/coredump/coredump_notify_signal_helper.c new file mode 100644 index 000000000000..849f5c1ea736 --- /dev/null +++ b/tools/testing/selftests/coredump/coredump_notify_signal_helper.c @@ -0,0 +1,46 @@ +// SPDX-License-Identifier: GPL-2.0 + +/* + * The |helper half of coredump_notify_signal_test. The kernel spawns this + * with the coredump on stdin, so it cannot be part of the test binary. + * It saves the dump and, once the notes have started, trips the fifo the + * crashing task is polling so TIF_NOTIFY_SIGNAL is raised while the note + * write is in flight. + */ + +#include <fcntl.h> +#include <stdio.h> +#include <stdlib.h> +#include <unistd.h> + +#include "coredump_notify_signal.h" + +int main(int argc, char *argv[]) +{ + int fd_core_file; + ssize_t ret; + + fd_core_file = open(NOTIFY_SIGNAL_CORE_TMPFILE, + O_WRONLY | O_CREAT | O_TRUNC | O_CLOEXEC, 0600); + if (fd_core_file < 0) { + fprintf(stderr, "%s: open failed: %m\n", argv[0]); + return EXIT_FAILURE; + } + + ret = recv_coredump_notify_signal(STDIN_FILENO, fd_core_file, true); + close(fd_core_file); + if (ret < 0) + goto err; + + /* The test polls for this name, so only create it once it is whole. */ + if (rename(NOTIFY_SIGNAL_CORE_TMPFILE, NOTIFY_SIGNAL_CORE_FILE)) { + fprintf(stderr, "%s: rename failed: %m\n", argv[0]); + goto err; + } + + return EXIT_SUCCESS; + +err: + unlink(NOTIFY_SIGNAL_CORE_TMPFILE); + return EXIT_FAILURE; +} diff --git a/tools/testing/selftests/coredump/coredump_notify_signal_test.c b/tools/testing/selftests/coredump/coredump_notify_signal_test.c new file mode 100644 index 000000000000..4a98ab141c41 --- /dev/null +++ b/tools/testing/selftests/coredump/coredump_notify_signal_test.c @@ -0,0 +1,245 @@ +// SPDX-License-Identifier: GPL-2.0 + +/* + * A coredump is meant to be interrupted by SIGKILL and by the freezer and + * by nothing else. dump_interrupted() says so, but the blocking waits + * underneath it test signal_pending(), which is also true for + * TIF_NOTIFY_SIGNAL. A crashing task that has an io_uring completion land + * on it mid-dump therefore keeps dumping while every wait it enters bails + * out at once, and the dump is silently cut short. Nothing reports it: + * binfmt_elf sets has_dumped before it writes anything, so WCOREDUMP() + * says the dump worked. + * + * A coredump note is the only dump_emit() that exceeds what the transport + * takes in one go, so it is the one write that is certain to block. Both + * tests crash a child holding enough file backed mappings for its NT_FILE + * note to run past that, arm an io_uring poll on it, and trip the poll + * while the note is being written. What comes out has to be the whole + * dump. + */ + +#include <fcntl.h> +#include <limits.h> +#include <sys/socket.h> +#include <sys/stat.h> +#include <sys/un.h> +#include <sys/wait.h> +#include <unistd.h> + +#include "coredump_test.h" + +FIXTURE_SETUP(coredump) +{ + FILE *file; + int ret; + + self->pid_coredump_server = -ESRCH; + self->fd_tmpfs_detached = -1; + file = fopen("/proc/sys/kernel/core_pattern", "r"); + ASSERT_NE(NULL, file); + + ret = fread(self->original_core_pattern, 1, + sizeof(self->original_core_pattern), file); + ASSERT_TRUE(ret || feof(file)); + ASSERT_LT(ret, sizeof(self->original_core_pattern)); + + self->original_core_pattern[ret] = '\0'; + self->fd_tmpfs_detached = create_detached_tmpfs(); + ASSERT_GE(self->fd_tmpfs_detached, 0); + + ret = fclose(file); + ASSERT_EQ(0, ret); + + /* A stale core file from a killed previous run would fake a pass. */ + unlink(NOTIFY_SIGNAL_CORE_FILE); + /* And a stale socket would fail the server's bind. */ + unlink(NOTIFY_SIGNAL_SOCKET); + unlink(NOTIFY_SIGNAL_TRIGGER); + ASSERT_EQ(mkfifo(NOTIFY_SIGNAL_TRIGGER, 0600), 0); +} + +FIXTURE_TEARDOWN(coredump) +{ + const char *reason; + FILE *file; + int ret, status; + + if (self->pid_coredump_server > 0) { + kill(self->pid_coredump_server, SIGTERM); + waitpid(self->pid_coredump_server, &status, 0); + } + unlink(NOTIFY_SIGNAL_CORE_FILE); + unlink(NOTIFY_SIGNAL_CORE_TMPFILE); + unlink(NOTIFY_SIGNAL_SOCKET); + unlink(NOTIFY_SIGNAL_TRIGGER); + unlink(NOTIFY_SIGNAL_MAPFILE); + + file = fopen("/proc/sys/kernel/core_pattern", "w"); + if (!file) { + reason = "Unable to open core_pattern"; + goto fail; + } + + ret = fprintf(file, "%s", self->original_core_pattern); + if (ret < 0) { + reason = "Unable to write to core_pattern"; + goto fail; + } + + ret = fclose(file); + if (ret) { + reason = "Unable to close core_pattern"; + goto fail; + } + + if (self->fd_tmpfs_detached >= 0) { + ret = close(self->fd_tmpfs_detached); + if (ret < 0) { + reason = "Unable to close detached tmpfs"; + goto fail; + } + self->fd_tmpfs_detached = -1; + } + + return; +fail: + /* This should never happen */ + fprintf(stderr, "Failed to cleanup coredump test: %s\n", reason); +} + +/* Check that what the helper or the server saved is the whole dump. */ +static void check_whole_coredump(struct __test_metadata *const _metadata) +{ + long long expected; + struct stat st; + + expected = coredump_expected_size(NOTIFY_SIGNAL_CORE_FILE); + ASSERT_GT(expected, 0); + ASSERT_EQ(stat(NOTIFY_SIGNAL_CORE_FILE, &st), 0); + ASSERT_EQ((long long)st.st_size, expected); +} + +/* + * The dump goes to a |helper, so the reader is a separate program the + * kernel spawns. It saves what it received to NOTIFY_SIGNAL_CORE_FILE. + */ +TEST_F(coredump, notify_signal_pipe) +{ + char pattern[PATH_MAX], helper[PATH_MAX], *p; + struct stat st; + int status, i; + pid_t pid; + ssize_t n; + + if (!coredump_io_uring_available()) + SKIP(return, "io_uring not available"); + + n = readlink("/proc/self/exe", helper, sizeof(helper) - 1); + ASSERT_GT(n, 0); + helper[n] = '\0'; + p = strstr(helper, "coredump_notify_signal_test"); + ASSERT_NE(p, NULL); + ASSERT_LE((size_t)(p - helper) + sizeof("coredump_notify_signal_helper"), + sizeof(helper)); + strcpy(p, "coredump_notify_signal_helper"); + if (access(helper, X_OK)) + SKIP(return, "coredump_notify_signal_helper not built"); + + ASSERT_LT(snprintf(pattern, sizeof(pattern), "|%s", helper), + (int)sizeof(pattern)); + ASSERT_TRUE(set_core_pattern(pattern)); + + pid = fork(); + ASSERT_GE(pid, 0); + if (pid == 0) + crashing_child_notify_signal(); + + ASSERT_EQ(waitpid(pid, &status, 0), pid); + ASSERT_TRUE(WIFSIGNALED(status)); + + /* The kernel does not wait for the helper, so poll for it. */ + for (i = 0; i < 100; i++) { + if (!stat(NOTIFY_SIGNAL_CORE_FILE, &st) && st.st_size) + break; + usleep(100000); + } + + check_whole_coredump(_metadata); +} + +/* The same thing with the dump going to a coredump socket. */ +TEST_F(coredump, notify_signal_socket) +{ + pid_t pid, pid_coredump_server; + int ipc_sockets[2], status; + char pattern[PATH_MAX]; + char c; + + if (!coredump_io_uring_available()) + SKIP(return, "io_uring not available"); + + ASSERT_EQ(socketpair(AF_UNIX, SOCK_STREAM | SOCK_CLOEXEC, 0, + ipc_sockets), 0); + ASSERT_LT(snprintf(pattern, sizeof(pattern), "@%s", + NOTIFY_SIGNAL_SOCKET), (int)sizeof(pattern)); + ASSERT_TRUE(set_core_pattern(pattern)); + + pid_coredump_server = fork(); + ASSERT_GE(pid_coredump_server, 0); + if (pid_coredump_server == 0) { + int fd_server = -1, fd_coredump = -1, fd_core_file = -1; + int exit_code = EXIT_FAILURE; + + close(ipc_sockets[0]); + + fd_server = create_and_listen_unix_socket(NOTIFY_SIGNAL_SOCKET); + if (fd_server < 0) + goto out; + if (write_nointr(ipc_sockets[1], "1", 1) < 0) + goto out; + close(ipc_sockets[1]); + + fd_coredump = accept4(fd_server, NULL, NULL, SOCK_CLOEXEC); + if (fd_coredump < 0) + goto out; + + fd_core_file = open(NOTIFY_SIGNAL_CORE_FILE, + O_WRONLY | O_CREAT | O_TRUNC | O_CLOEXEC, + 0600); + if (fd_core_file < 0) + goto out; + + if (recv_coredump_notify_signal(fd_coredump, fd_core_file, + true) < 0) + goto out; + + exit_code = EXIT_SUCCESS; +out: + if (fd_core_file >= 0) + close(fd_core_file); + if (fd_coredump >= 0) + close(fd_coredump); + if (fd_server >= 0) + close(fd_server); + _exit(exit_code); + } + self->pid_coredump_server = pid_coredump_server; + + EXPECT_EQ(close(ipc_sockets[1]), 0); + ASSERT_EQ(read_nointr(ipc_sockets[0], &c, 1), 1); + EXPECT_EQ(close(ipc_sockets[0]), 0); + + pid = fork(); + ASSERT_GE(pid, 0); + if (pid == 0) + crashing_child_notify_signal(); + + ASSERT_EQ(waitpid(pid, &status, 0), pid); + ASSERT_TRUE(WIFSIGNALED(status)); + + wait_and_check_coredump_server(pid_coredump_server, _metadata, self); + + check_whole_coredump(_metadata); +} + +TEST_HARNESS_MAIN diff --git a/tools/testing/selftests/coredump/coredump_signal_test.c b/tools/testing/selftests/coredump/coredump_signal_test.c new file mode 100644 index 000000000000..fdf48b144781 --- /dev/null +++ b/tools/testing/selftests/coredump/coredump_signal_test.c @@ -0,0 +1,238 @@ +// SPDX-License-Identifier: GPL-2.0 + +#include <fcntl.h> +#include <pthread.h> +#include <signal.h> +#include <sys/mman.h> +#include <sys/socket.h> +#include <sys/stat.h> +#include <sys/syscall.h> +#include <sys/un.h> +#include <unistd.h> + +#include "coredump_test.h" + +/* Big enough to fill the socket buffer many times over. */ +#define CRASH_MAPPING_SIZE (32 * 1024 * 1024) + +FIXTURE_SETUP(coredump) +{ + FILE *file; + int ret; + + self->pid_coredump_server = -ESRCH; + self->fd_tmpfs_detached = -1; + file = fopen("/proc/sys/kernel/core_pattern", "r"); + ASSERT_NE(NULL, file); + + ret = fread(self->original_core_pattern, 1, sizeof(self->original_core_pattern), file); + ASSERT_TRUE(ret || feof(file)); + ASSERT_LT(ret, sizeof(self->original_core_pattern)); + + self->original_core_pattern[ret] = '\0'; + + ret = fclose(file); + ASSERT_EQ(0, ret); +} + +FIXTURE_TEARDOWN(coredump) +{ + const char *reason; + FILE *file; + int ret, status; + + if (self->pid_coredump_server > 0) { + kill(self->pid_coredump_server, SIGTERM); + waitpid(self->pid_coredump_server, &status, 0); + } + unlink("/tmp/coredump.file"); + unlink("/tmp/coredump.socket"); + + file = fopen("/proc/sys/kernel/core_pattern", "w"); + if (!file) { + reason = "Unable to open core_pattern"; + goto fail; + } + + ret = fprintf(file, "%s", self->original_core_pattern); + if (ret < 0) { + reason = "Unable to write to core_pattern"; + goto fail; + } + + ret = fclose(file); + if (ret) { + reason = "Unable to close core_pattern"; + goto fail; + } + + return; +fail: + /* This should never happen */ + fprintf(stderr, "Failed to cleanup coredump test: %s\n", reason); +} + +static volatile int waiter_ready; + +static void usr1_handler(int sig) +{ +} + +/* Sleeps in sigtimedwait() with SIGUSR1 unblocked only inside the kernel. */ +static void *sigwaiter(void *arg) +{ + sigset_t set; + siginfo_t info; + + sigemptyset(&set); + sigaddset(&set, SIGUSR1); + __atomic_store_n(&waiter_ready, 1, __ATOMIC_RELEASE); + for (;;) + sigtimedwait(&set, &info, NULL); + return NULL; +} + +/* + * Crash with a shared SIGUSR1 still queued for this thread. The vfork() + * child queues it while we sleep killably and then the SIGSEGV that is + * dequeued first. The zap wakes the sibling out of sigtimedwait() and its + * mask restore retargets SIGUSR1 to the dumper. + */ +static void crashing_child_retarget(bool queue_shared) +{ + sigset_t all, old; + pthread_t thread; + pid_t pid, tid; + char *p; + + p = mmap(NULL, CRASH_MAPPING_SIZE, PROT_READ | PROT_WRITE, + MAP_PRIVATE | MAP_ANONYMOUS, -1, 0); + if (p == MAP_FAILED) + _exit(EXIT_FAILURE); + memset(p, 0x5a, CRASH_MAPPING_SIZE); + + signal(SIGUSR1, usr1_handler); + sigfillset(&all); + pthread_sigmask(SIG_BLOCK, &all, &old); + /* One waiter only, retarget stops at the first thread not blocking it. */ + if (pthread_create(&thread, NULL, sigwaiter, NULL)) + _exit(EXIT_FAILURE); + pthread_sigmask(SIG_SETMASK, &old, NULL); + while (!__atomic_load_n(&waiter_ready, __ATOMIC_ACQUIRE)) + usleep(1000); + usleep(50 * 1000); + + pid = getpid(); + tid = syscall(SYS_gettid); + if (vfork() == 0) { + if (queue_shared) + syscall(SYS_kill, pid, SIGUSR1); + syscall(SYS_tgkill, pid, tid, SIGSEGV); + syscall(SYS_exit, 0); + } + + /* Not reached, the pending SIGSEGV dumps core. */ + for (;;) + pause(); +} + +/* + * Dump the crashing child into /tmp/coredump.file through a server that + * holds the read back so the dumper blocks on the full socket buffer. + */ +static void run_coredump(struct __test_metadata *const _metadata, + FIXTURE_DATA(coredump) *self, bool queue_shared) +{ + pid_t pid, pid_coredump_server; + int ipc_sockets[2]; + int status; + char c; + + unlink("/tmp/coredump.file"); + unlink("/tmp/coredump.socket"); + ASSERT_TRUE(set_core_pattern("@/tmp/coredump.socket")); + ASSERT_EQ(socketpair(AF_UNIX, SOCK_STREAM | SOCK_CLOEXEC, 0, ipc_sockets), 0); + + pid_coredump_server = fork(); + ASSERT_GE(pid_coredump_server, 0); + if (pid_coredump_server == 0) { + int fd_server = -1, fd_coredump = -1, fd_core_file = -1; + int exit_code = EXIT_FAILURE; + + close(ipc_sockets[0]); + + fd_server = create_and_listen_unix_socket("/tmp/coredump.socket"); + if (fd_server < 0) + goto out; + + if (write_nointr(ipc_sockets[1], "1", 1) < 0) + goto out; + close(ipc_sockets[1]); + + fd_coredump = accept4(fd_server, NULL, NULL, SOCK_CLOEXEC); + if (fd_coredump < 0) { + fprintf(stderr, "%s: accept4 failed: %m\n", __func__); + goto out; + } + + /* Let the dumper run into the full socket buffer first. */ + sleep(1); + + fd_core_file = creat("/tmp/coredump.file", 0644); + if (fd_core_file < 0) { + fprintf(stderr, "%s: creat failed: %m\n", __func__); + goto out; + } + + if (recv_coredump_bytes(fd_coredump, fd_core_file) < 0) + goto out; + + exit_code = EXIT_SUCCESS; +out: + if (fd_core_file >= 0) + close(fd_core_file); + if (fd_coredump >= 0) + close(fd_coredump); + if (fd_server >= 0) + close(fd_server); + _exit(exit_code); + } + self->pid_coredump_server = pid_coredump_server; + + EXPECT_EQ(close(ipc_sockets[1]), 0); + ASSERT_EQ(read_nointr(ipc_sockets[0], &c, 1), 1); + EXPECT_EQ(close(ipc_sockets[0]), 0); + + pid = fork(); + ASSERT_GE(pid, 0); + if (pid == 0) + crashing_child_retarget(queue_shared); + + waitpid(pid, &status, 0); + ASSERT_TRUE(WIFSIGNALED(status)); + ASSERT_EQ(WTERMSIG(status), SIGSEGV); + ASSERT_TRUE(WCOREDUMP(status)); + + wait_and_check_coredump_server(pid_coredump_server, _metadata, self); +} + +static void check_coredump_complete(struct __test_metadata *const _metadata) +{ + int fd; + + fd = open("/tmp/coredump.file", O_RDONLY | O_CLOEXEC); + ASSERT_GE(fd, 0); + ASSERT_TRUE(check_coredump_extent(fd)); + close(fd); +} + +TEST_F(coredump, retarget_shared_pending) +{ + run_coredump(_metadata, self, false); + check_coredump_complete(_metadata); + + run_coredump(_metadata, self, true); + check_coredump_complete(_metadata); +} + +TEST_HARNESS_MAIN diff --git a/tools/testing/selftests/coredump/coredump_socket_protocol_test.c b/tools/testing/selftests/coredump/coredump_socket_protocol_test.c index d9fa6239b5a9..f5c9bad87546 100644 --- a/tools/testing/selftests/coredump/coredump_socket_protocol_test.c +++ b/tools/testing/selftests/coredump/coredump_socket_protocol_test.c @@ -151,9 +151,7 @@ TEST_F(coredump, socket_request_kernel) goto out; } - if (!check_coredump_req(&req, COREDUMP_ACK_SIZE_VER0, - COREDUMP_KERNEL | COREDUMP_USERSPACE | - COREDUMP_REJECT | COREDUMP_WAIT)) { + if (!check_coredump_req(&req)) { fprintf(stderr, "socket_request_kernel: check_coredump_req failed\n"); goto out; } @@ -301,9 +299,7 @@ TEST_F(coredump, socket_request_userspace) goto out; } - if (!check_coredump_req(&req, COREDUMP_ACK_SIZE_VER0, - COREDUMP_KERNEL | COREDUMP_USERSPACE | - COREDUMP_REJECT | COREDUMP_WAIT)) { + if (!check_coredump_req(&req)) { fprintf(stderr, "socket_request_userspace: check_coredump_req failed\n"); goto out; } @@ -441,9 +437,7 @@ TEST_F(coredump, socket_request_reject) goto out; } - if (!check_coredump_req(&req, COREDUMP_ACK_SIZE_VER0, - COREDUMP_KERNEL | COREDUMP_USERSPACE | - COREDUMP_REJECT | COREDUMP_WAIT)) { + if (!check_coredump_req(&req)) { fprintf(stderr, "socket_request_reject: check_coredump_req failed\n"); goto out; } @@ -514,93 +508,89 @@ out: wait_and_check_coredump_server(pid_coredump_server, _metadata, self); } -TEST_F(coredump, socket_request_invalid_flag_combination) +/* An ack the kernel must refuse and how. */ +struct refused_ack { + /* The ack, and how many bytes of it the server sends before it hangs up. */ + struct coredump_ack ack; + size_t bytes; + /* The marker the kernel answers with, or none if @no_marker. */ + enum coredump_mark mark; + bool no_marker; +}; + +/* Send @refused, expect the kernel to refuse it and hang up. */ +static void check_refused_ack(struct __test_metadata *const _metadata, + FIXTURE_DATA(coredump) *self, + const struct refused_ack *refused) { - int pidfd, ret, status; + int pidfd, status; pid_t pid, pid_coredump_server; struct pidfd_info info = {}; int ipc_sockets[2]; char c; + ASSERT_EQ(socketpair(AF_UNIX, SOCK_STREAM | SOCK_CLOEXEC, 0, ipc_sockets), 0); ASSERT_TRUE(set_core_pattern("@@/tmp/coredump.socket")); - ret = socketpair(AF_UNIX, SOCK_STREAM | SOCK_CLOEXEC, 0, ipc_sockets); - ASSERT_EQ(ret, 0); - pid_coredump_server = fork(); ASSERT_GE(pid_coredump_server, 0); if (pid_coredump_server == 0) { - struct coredump_req req = {}; int fd_server = -1, fd_coredump = -1, fd_peer_pidfd = -1; int exit_code = EXIT_FAILURE; + struct coredump_req req = {}; close(ipc_sockets[0]); fd_server = create_and_listen_unix_socket("/tmp/coredump.socket"); - if (fd_server < 0) { - fprintf(stderr, "socket_request_invalid_flag_combination: create_and_listen_unix_socket failed: %m\n"); + if (fd_server < 0) goto out; - } - if (write_nointr(ipc_sockets[1], "1", 1) < 0) { - fprintf(stderr, "socket_request_invalid_flag_combination: write_nointr to ipc socket failed: %m\n"); + if (write_nointr(ipc_sockets[1], "1", 1) < 0) goto out; - } close(ipc_sockets[1]); fd_coredump = accept4(fd_server, NULL, NULL, SOCK_CLOEXEC); - if (fd_coredump < 0) { - fprintf(stderr, "socket_request_invalid_flag_combination: accept4 failed: %m\n"); + if (fd_coredump < 0) goto out; - } fd_peer_pidfd = get_peer_pidfd(fd_coredump); - if (fd_peer_pidfd < 0) { - fprintf(stderr, "socket_request_invalid_flag_combination: get_peer_pidfd failed\n"); + if (fd_peer_pidfd < 0) goto out; - } - if (!get_pidfd_info(fd_peer_pidfd, &info)) { - fprintf(stderr, "socket_request_invalid_flag_combination: get_pidfd_info failed\n"); + /* The task shows as dumping while it waits for the ack. */ + if (!get_pidfd_info(fd_peer_pidfd, &info)) goto out; - } - if (!(info.mask & PIDFD_INFO_COREDUMP)) { - fprintf(stderr, "socket_request_invalid_flag_combination: PIDFD_INFO_COREDUMP not set in mask\n"); + if (!(info.mask & PIDFD_INFO_COREDUMP) || + !(info.coredump_mask & PIDFD_COREDUMPED)) { + fprintf(stderr, "Peer isn't marked as dumping\n"); goto out; } - if (!(info.coredump_mask & PIDFD_COREDUMPED)) { - fprintf(stderr, "socket_request_invalid_flag_combination: PIDFD_COREDUMPED not set in coredump_mask\n"); + if (!read_coredump_req(fd_coredump, &req)) goto out; - } - if (!read_coredump_req(fd_coredump, &req)) { - fprintf(stderr, "socket_request_invalid_flag_combination: read_coredump_req failed\n"); + if (!check_coredump_req(&req)) goto out; - } - if (!check_coredump_req(&req, COREDUMP_ACK_SIZE_VER0, - COREDUMP_KERNEL | COREDUMP_USERSPACE | - COREDUMP_REJECT | COREDUMP_WAIT)) { - fprintf(stderr, "socket_request_invalid_flag_combination: check_coredump_req failed\n"); + if (!send_coredump_ack_bytes(fd_coredump, &refused->ack, + refused->bytes)) goto out; - } - if (!send_coredump_ack(fd_coredump, &req, - COREDUMP_KERNEL | COREDUMP_REJECT | COREDUMP_WAIT, 0)) { - fprintf(stderr, "socket_request_invalid_flag_combination: send_coredump_ack failed\n"); + /* Nothing more to say. A server that died looks the same. */ + if (shutdown(fd_coredump, SHUT_WR)) goto out; - } - if (!read_marker(fd_coredump, COREDUMP_MARK_CONFLICTING)) { - fprintf(stderr, "socket_request_invalid_flag_combination: read_marker COREDUMP_MARK_CONFLICTING failed\n"); + if (!refused->no_marker && + !read_marker(fd_coredump, refused->mark)) + goto out; + + /* The kernel hangs up after a refusal, marker or not. */ + if (!read_hangup(fd_coredump)) goto out; - } exit_code = EXIT_SUCCESS; - fprintf(stderr, "socket_request_invalid_flag_combination: completed successfully\n"); out: if (fd_peer_pidfd >= 0) close(fd_peer_pidfd); @@ -635,368 +625,72 @@ out: wait_and_check_coredump_server(pid_coredump_server, _metadata, self); } -TEST_F(coredump, socket_request_unknown_flag) +/* Ack @ack_mask, expect the kernel to refuse it as conflicting. */ +static void check_conflicting_ack(struct __test_metadata *const _metadata, + FIXTURE_DATA(coredump) *self, __u64 ack_mask) { - int pidfd, ret, status; - pid_t pid, pid_coredump_server; - struct pidfd_info info = {}; - int ipc_sockets[2]; - char c; - - ASSERT_TRUE(set_core_pattern("@@/tmp/coredump.socket")); - - ret = socketpair(AF_UNIX, SOCK_STREAM | SOCK_CLOEXEC, 0, ipc_sockets); - ASSERT_EQ(ret, 0); - - pid_coredump_server = fork(); - ASSERT_GE(pid_coredump_server, 0); - if (pid_coredump_server == 0) { - struct coredump_req req = {}; - int fd_server = -1, fd_coredump = -1, fd_peer_pidfd = -1; - int exit_code = EXIT_FAILURE; - - close(ipc_sockets[0]); - - fd_server = create_and_listen_unix_socket("/tmp/coredump.socket"); - if (fd_server < 0) { - fprintf(stderr, "socket_request_unknown_flag: create_and_listen_unix_socket failed: %m\n"); - goto out; - } - - if (write_nointr(ipc_sockets[1], "1", 1) < 0) { - fprintf(stderr, "socket_request_unknown_flag: write_nointr to ipc socket failed: %m\n"); - goto out; - } - - close(ipc_sockets[1]); - - fd_coredump = accept4(fd_server, NULL, NULL, SOCK_CLOEXEC); - if (fd_coredump < 0) { - fprintf(stderr, "socket_request_unknown_flag: accept4 failed: %m\n"); - goto out; - } - - fd_peer_pidfd = get_peer_pidfd(fd_coredump); - if (fd_peer_pidfd < 0) { - fprintf(stderr, "socket_request_unknown_flag: get_peer_pidfd failed\n"); - goto out; - } - - if (!get_pidfd_info(fd_peer_pidfd, &info)) { - fprintf(stderr, "socket_request_unknown_flag: get_pidfd_info failed\n"); - goto out; - } - - if (!(info.mask & PIDFD_INFO_COREDUMP)) { - fprintf(stderr, "socket_request_unknown_flag: PIDFD_INFO_COREDUMP not set in mask\n"); - goto out; - } - - if (!(info.coredump_mask & PIDFD_COREDUMPED)) { - fprintf(stderr, "socket_request_unknown_flag: PIDFD_COREDUMPED not set in coredump_mask\n"); - goto out; - } - - if (!read_coredump_req(fd_coredump, &req)) { - fprintf(stderr, "socket_request_unknown_flag: read_coredump_req failed\n"); - goto out; - } - - if (!check_coredump_req(&req, COREDUMP_ACK_SIZE_VER0, - COREDUMP_KERNEL | COREDUMP_USERSPACE | - COREDUMP_REJECT | COREDUMP_WAIT)) { - fprintf(stderr, "socket_request_unknown_flag: check_coredump_req failed\n"); - goto out; - } - - if (!send_coredump_ack(fd_coredump, &req, (1ULL << 63), 0)) { - fprintf(stderr, "socket_request_unknown_flag: send_coredump_ack failed\n"); - goto out; - } - - if (!read_marker(fd_coredump, COREDUMP_MARK_UNSUPPORTED)) { - fprintf(stderr, "socket_request_unknown_flag: read_marker COREDUMP_MARK_UNSUPPORTED failed\n"); - goto out; - } - - exit_code = EXIT_SUCCESS; - fprintf(stderr, "socket_request_unknown_flag: completed successfully\n"); -out: - if (fd_peer_pidfd >= 0) - close(fd_peer_pidfd); - if (fd_coredump >= 0) - close(fd_coredump); - if (fd_server >= 0) - close(fd_server); - _exit(exit_code); - } - self->pid_coredump_server = pid_coredump_server; - - EXPECT_EQ(close(ipc_sockets[1]), 0); - ASSERT_EQ(read_nointr(ipc_sockets[0], &c, 1), 1); - EXPECT_EQ(close(ipc_sockets[0]), 0); - - pid = fork(); - ASSERT_GE(pid, 0); - if (pid == 0) - crashing_child(); - - pidfd = sys_pidfd_open(pid, 0); - ASSERT_GE(pidfd, 0); - - waitpid(pid, &status, 0); - ASSERT_TRUE(WIFSIGNALED(status)); - ASSERT_FALSE(WCOREDUMP(status)); + struct refused_ack refused = { + .ack = { + .size = sizeof(struct coredump_ack), + .mask = ack_mask, + }, + .bytes = sizeof(struct coredump_ack), + .mark = COREDUMP_MARK_CONFLICTING, + }; + + check_refused_ack(_metadata, self, &refused); +} - ASSERT_TRUE(get_pidfd_info(pidfd, &info)); - ASSERT_GT((info.mask & PIDFD_INFO_COREDUMP), 0); - ASSERT_GT((info.coredump_mask & PIDFD_COREDUMPED), 0); +/* More than one of KERNEL, USERSPACE and REJECT. */ +TEST_F(coredump, socket_request_invalid_flag_combination) +{ + check_conflicting_ack(_metadata, self, + COREDUMP_KERNEL | COREDUMP_REJECT | COREDUMP_WAIT); +} - wait_and_check_coredump_server(pid_coredump_server, _metadata, self); +/* A flag the kernel didn't advertise in coredump_req->mask. */ +TEST_F(coredump, socket_request_unknown_flag) +{ + struct refused_ack refused = { + .ack = { + .size = sizeof(struct coredump_ack), + .mask = 1ULL << 63, + }, + .bytes = sizeof(struct coredump_ack), + .mark = COREDUMP_MARK_UNSUPPORTED, + }; + + check_refused_ack(_metadata, self, &refused); } +/* An ack smaller than the first published struct. */ TEST_F(coredump, socket_request_invalid_size_small) { - int pidfd, ret, status; - pid_t pid, pid_coredump_server; - struct pidfd_info info = {}; - int ipc_sockets[2]; - char c; - - ASSERT_TRUE(set_core_pattern("@@/tmp/coredump.socket")); - - ret = socketpair(AF_UNIX, SOCK_STREAM | SOCK_CLOEXEC, 0, ipc_sockets); - ASSERT_EQ(ret, 0); - - pid_coredump_server = fork(); - ASSERT_GE(pid_coredump_server, 0); - if (pid_coredump_server == 0) { - struct coredump_req req = {}; - int fd_server = -1, fd_coredump = -1, fd_peer_pidfd = -1; - int exit_code = EXIT_FAILURE; - - close(ipc_sockets[0]); - - fd_server = create_and_listen_unix_socket("/tmp/coredump.socket"); - if (fd_server < 0) { - fprintf(stderr, "socket_request_invalid_size_small: create_and_listen_unix_socket failed: %m\n"); - goto out; - } - - if (write_nointr(ipc_sockets[1], "1", 1) < 0) { - fprintf(stderr, "socket_request_invalid_size_small: write_nointr to ipc socket failed: %m\n"); - goto out; - } - - close(ipc_sockets[1]); - - fd_coredump = accept4(fd_server, NULL, NULL, SOCK_CLOEXEC); - if (fd_coredump < 0) { - fprintf(stderr, "socket_request_invalid_size_small: accept4 failed: %m\n"); - goto out; - } - - fd_peer_pidfd = get_peer_pidfd(fd_coredump); - if (fd_peer_pidfd < 0) { - fprintf(stderr, "socket_request_invalid_size_small: get_peer_pidfd failed\n"); - goto out; - } - - if (!get_pidfd_info(fd_peer_pidfd, &info)) { - fprintf(stderr, "socket_request_invalid_size_small: get_pidfd_info failed\n"); - goto out; - } - - if (!(info.mask & PIDFD_INFO_COREDUMP)) { - fprintf(stderr, "socket_request_invalid_size_small: PIDFD_INFO_COREDUMP not set in mask\n"); - goto out; - } - - if (!(info.coredump_mask & PIDFD_COREDUMPED)) { - fprintf(stderr, "socket_request_invalid_size_small: PIDFD_COREDUMPED not set in coredump_mask\n"); - goto out; - } - - if (!read_coredump_req(fd_coredump, &req)) { - fprintf(stderr, "socket_request_invalid_size_small: read_coredump_req failed\n"); - goto out; - } - - if (!check_coredump_req(&req, COREDUMP_ACK_SIZE_VER0, - COREDUMP_KERNEL | COREDUMP_USERSPACE | - COREDUMP_REJECT | COREDUMP_WAIT)) { - fprintf(stderr, "socket_request_invalid_size_small: check_coredump_req failed\n"); - goto out; - } - - if (!send_coredump_ack(fd_coredump, &req, - COREDUMP_REJECT | COREDUMP_WAIT, - COREDUMP_ACK_SIZE_VER0 / 2)) { - fprintf(stderr, "socket_request_invalid_size_small: send_coredump_ack failed\n"); - goto out; - } - - if (!read_marker(fd_coredump, COREDUMP_MARK_MINSIZE)) { - fprintf(stderr, "socket_request_invalid_size_small: read_marker COREDUMP_MARK_MINSIZE failed\n"); - goto out; - } - - exit_code = EXIT_SUCCESS; - fprintf(stderr, "socket_request_invalid_size_small: completed successfully\n"); -out: - if (fd_peer_pidfd >= 0) - close(fd_peer_pidfd); - if (fd_coredump >= 0) - close(fd_coredump); - if (fd_server >= 0) - close(fd_server); - _exit(exit_code); - } - self->pid_coredump_server = pid_coredump_server; - - EXPECT_EQ(close(ipc_sockets[1]), 0); - ASSERT_EQ(read_nointr(ipc_sockets[0], &c, 1), 1); - EXPECT_EQ(close(ipc_sockets[0]), 0); - - pid = fork(); - ASSERT_GE(pid, 0); - if (pid == 0) - crashing_child(); - - pidfd = sys_pidfd_open(pid, 0); - ASSERT_GE(pidfd, 0); - - waitpid(pid, &status, 0); - ASSERT_TRUE(WIFSIGNALED(status)); - ASSERT_FALSE(WCOREDUMP(status)); - - ASSERT_TRUE(get_pidfd_info(pidfd, &info)); - ASSERT_GT((info.mask & PIDFD_INFO_COREDUMP), 0); - ASSERT_GT((info.coredump_mask & PIDFD_COREDUMPED), 0); - - wait_and_check_coredump_server(pid_coredump_server, _metadata, self); + struct refused_ack refused = { + .ack = { + .size = COREDUMP_ACK_SIZE_VER0 / 2, + .mask = COREDUMP_REJECT | COREDUMP_WAIT, + }, + .bytes = COREDUMP_ACK_SIZE_VER0 / 2, + .mark = COREDUMP_MARK_MINSIZE, + }; + + check_refused_ack(_metadata, self, &refused); } +/* An ack bigger than the kernel said it accepts. */ TEST_F(coredump, socket_request_invalid_size_large) { - int pidfd, ret, status; - pid_t pid, pid_coredump_server; - struct pidfd_info info = {}; - int ipc_sockets[2]; - char c; - - ASSERT_TRUE(set_core_pattern("@@/tmp/coredump.socket")); - - ret = socketpair(AF_UNIX, SOCK_STREAM | SOCK_CLOEXEC, 0, ipc_sockets); - ASSERT_EQ(ret, 0); - - pid_coredump_server = fork(); - ASSERT_GE(pid_coredump_server, 0); - if (pid_coredump_server == 0) { - struct coredump_req req = {}; - int fd_server = -1, fd_coredump = -1, fd_peer_pidfd = -1; - int exit_code = EXIT_FAILURE; - - close(ipc_sockets[0]); - - fd_server = create_and_listen_unix_socket("/tmp/coredump.socket"); - if (fd_server < 0) { - fprintf(stderr, "socket_request_invalid_size_large: create_and_listen_unix_socket failed: %m\n"); - goto out; - } - - if (write_nointr(ipc_sockets[1], "1", 1) < 0) { - fprintf(stderr, "socket_request_invalid_size_large: write_nointr to ipc socket failed: %m\n"); - goto out; - } - - close(ipc_sockets[1]); - - fd_coredump = accept4(fd_server, NULL, NULL, SOCK_CLOEXEC); - if (fd_coredump < 0) { - fprintf(stderr, "socket_request_invalid_size_large: accept4 failed: %m\n"); - goto out; - } - - fd_peer_pidfd = get_peer_pidfd(fd_coredump); - if (fd_peer_pidfd < 0) { - fprintf(stderr, "socket_request_invalid_size_large: get_peer_pidfd failed\n"); - goto out; - } - - if (!get_pidfd_info(fd_peer_pidfd, &info)) { - fprintf(stderr, "socket_request_invalid_size_large: get_pidfd_info failed\n"); - goto out; - } - - if (!(info.mask & PIDFD_INFO_COREDUMP)) { - fprintf(stderr, "socket_request_invalid_size_large: PIDFD_INFO_COREDUMP not set in mask\n"); - goto out; - } - - if (!(info.coredump_mask & PIDFD_COREDUMPED)) { - fprintf(stderr, "socket_request_invalid_size_large: PIDFD_COREDUMPED not set in coredump_mask\n"); - goto out; - } - - if (!read_coredump_req(fd_coredump, &req)) { - fprintf(stderr, "socket_request_invalid_size_large: read_coredump_req failed\n"); - goto out; - } - - if (!check_coredump_req(&req, COREDUMP_ACK_SIZE_VER0, - COREDUMP_KERNEL | COREDUMP_USERSPACE | - COREDUMP_REJECT | COREDUMP_WAIT)) { - fprintf(stderr, "socket_request_invalid_size_large: check_coredump_req failed\n"); - goto out; - } - - if (!send_coredump_ack(fd_coredump, &req, - COREDUMP_REJECT | COREDUMP_WAIT, - COREDUMP_ACK_SIZE_VER0 + PAGE_SIZE)) { - fprintf(stderr, "socket_request_invalid_size_large: send_coredump_ack failed\n"); - goto out; - } - - if (!read_marker(fd_coredump, COREDUMP_MARK_MAXSIZE)) { - fprintf(stderr, "socket_request_invalid_size_large: read_marker COREDUMP_MARK_MAXSIZE failed\n"); - goto out; - } - - exit_code = EXIT_SUCCESS; - fprintf(stderr, "socket_request_invalid_size_large: completed successfully\n"); -out: - if (fd_peer_pidfd >= 0) - close(fd_peer_pidfd); - if (fd_coredump >= 0) - close(fd_coredump); - if (fd_server >= 0) - close(fd_server); - _exit(exit_code); - } - self->pid_coredump_server = pid_coredump_server; - - EXPECT_EQ(close(ipc_sockets[1]), 0); - ASSERT_EQ(read_nointr(ipc_sockets[0], &c, 1), 1); - EXPECT_EQ(close(ipc_sockets[0]), 0); - - pid = fork(); - ASSERT_GE(pid, 0); - if (pid == 0) - crashing_child(); - - pidfd = sys_pidfd_open(pid, 0); - ASSERT_GE(pidfd, 0); - - waitpid(pid, &status, 0); - ASSERT_TRUE(WIFSIGNALED(status)); - ASSERT_FALSE(WCOREDUMP(status)); - - ASSERT_TRUE(get_pidfd_info(pidfd, &info)); - ASSERT_GT((info.mask & PIDFD_INFO_COREDUMP), 0); - ASSERT_GT((info.coredump_mask & PIDFD_COREDUMPED), 0); - - wait_and_check_coredump_server(pid_coredump_server, _metadata, self); + struct refused_ack refused = { + .ack = { + .size = COREDUMP_ACK_SIZE_VER0 + PAGE_SIZE, + .mask = COREDUMP_REJECT | COREDUMP_WAIT, + }, + .bytes = COREDUMP_ACK_SIZE_VER0 + PAGE_SIZE, + .mark = COREDUMP_MARK_MAXSIZE, + }; + + check_refused_ack(_metadata, self, &refused); } /* @@ -1355,9 +1049,7 @@ TEST_F_TIMEOUT(coredump, socket_multiple_crashing_coredumps, 500) goto out; } - if (!check_coredump_req(&req, COREDUMP_ACK_SIZE_VER0, - COREDUMP_KERNEL | COREDUMP_USERSPACE | - COREDUMP_REJECT | COREDUMP_WAIT)) { + if (!check_coredump_req(&req)) { fprintf(stderr, "check_coredump_req failed for fd %d\n", fd_coredump); goto out; } @@ -1509,9 +1201,7 @@ TEST_F_TIMEOUT(coredump, socket_multiple_crashing_coredumps_epoll_workers, 500) fprintf(stderr, "socket_multiple_crashing_coredumps_epoll_workers: read_coredump_req failed\n"); goto out; } - if (!check_coredump_req(&req, COREDUMP_ACK_SIZE_VER0, - COREDUMP_KERNEL | COREDUMP_USERSPACE | - COREDUMP_REJECT | COREDUMP_WAIT)) { + if (!check_coredump_req(&req)) { fprintf(stderr, "socket_multiple_crashing_coredumps_epoll_workers: check_coredump_req failed\n"); goto out; } @@ -1591,4 +1281,1335 @@ out: wait_and_check_coredump_server(pid_coredump_server, _metadata, self); } +/* + * Reassemble a record stream and check that what comes out is an ELF + * core file. The records themselves are validated by recv_coredump_records(). + */ +TEST_F(coredump, socket_request_sparse_reassemble) +{ + int fd_core_file, pidfd, status; + pid_t pid, pid_coredump_server; + struct pidfd_info info = {}; + int ipc_sockets[2]; + char c; + + ASSERT_EQ(socketpair(AF_UNIX, SOCK_STREAM | SOCK_CLOEXEC, 0, ipc_sockets), 0); + ASSERT_TRUE(set_core_pattern("@@/tmp/coredump.socket")); + + pid_coredump_server = fork(); + ASSERT_GE(pid_coredump_server, 0); + if (pid_coredump_server == 0) { + int fd_server = -1, fd_coredump = -1, fd_peer_pidfd = -1; + int fd_file = -1; + int exit_code = EXIT_FAILURE; + struct coredump_req req = {}; + + close(ipc_sockets[0]); + + fd_server = create_and_listen_unix_socket("/tmp/coredump.socket"); + if (fd_server < 0) + goto out; + + if (write_nointr(ipc_sockets[1], "1", 1) < 0) + goto out; + + close(ipc_sockets[1]); + + fd_coredump = accept4(fd_server, NULL, NULL, SOCK_CLOEXEC); + if (fd_coredump < 0) + goto out; + + fd_peer_pidfd = get_peer_pidfd(fd_coredump); + if (fd_peer_pidfd < 0) + goto out; + + fd_file = creat("/tmp/coredump.file", 0644); + if (fd_file < 0) + goto out; + + if (!read_coredump_req(fd_coredump, &req)) + goto out; + + if (!check_coredump_req(&req)) + goto out; + + if (!send_coredump_ack(fd_coredump, &req, + COREDUMP_KERNEL | COREDUMP_RECORDS | + COREDUMP_SPARSE | COREDUMP_WAIT, 0)) + goto out; + + if (!read_marker(fd_coredump, COREDUMP_MARK_REQACK)) + goto out; + + if (recv_coredump_records(fd_coredump, fd_file, NULL, NULL, -1) < 0) + goto out; + + exit_code = EXIT_SUCCESS; +out: + if (fd_file >= 0) + close(fd_file); + if (fd_peer_pidfd >= 0) + close(fd_peer_pidfd); + if (fd_coredump >= 0) + close(fd_coredump); + if (fd_server >= 0) + close(fd_server); + _exit(exit_code); + } + self->pid_coredump_server = pid_coredump_server; + + EXPECT_EQ(close(ipc_sockets[1]), 0); + ASSERT_EQ(read_nointr(ipc_sockets[0], &c, 1), 1); + EXPECT_EQ(close(ipc_sockets[0]), 0); + + pid = fork(); + ASSERT_GE(pid, 0); + if (pid == 0) + crashing_child(); + + pidfd = sys_pidfd_open(pid, 0); + ASSERT_GE(pidfd, 0); + + waitpid(pid, &status, 0); + ASSERT_TRUE(WIFSIGNALED(status)); + ASSERT_TRUE(WCOREDUMP(status)); + + ASSERT_TRUE(get_pidfd_info(pidfd, &info)); + ASSERT_GT((info.mask & PIDFD_INFO_COREDUMP), 0); + ASSERT_GT((info.coredump_mask & PIDFD_COREDUMPED), 0); + + wait_and_check_coredump_server(pid_coredump_server, _metadata, self); + + /* What the records reassemble into has to be an ELF core file. */ + fd_core_file = open("/tmp/coredump.file", O_RDONLY | O_CLOEXEC); + ASSERT_GE(fd_core_file, 0); + ASSERT_TRUE(is_elf_core(fd_core_file)); + EXPECT_EQ(close(fd_core_file), 0); +} + +/* + * Crash a child with a mostly-unpopulated mapping and reassemble its + * record stream, reporting what crossed the socket and the coredump + * size the records describe. With @kill_peer the server kills the task + * once the coredump is under way so the kernel has to cut it short. + */ +static void check_record_dump(struct __test_metadata *const _metadata, + FIXTURE_DATA(coredump) *self, __u64 ack_mask, + bool kill_peer, ssize_t *received, + off_t *coredump_size) +{ + bool truncated = false; + int pidfd, status; + pid_t pid, pid_coredump_server; + struct pidfd_info info = {}; + int ipc_sockets[2]; + int pipefds[2]; + char c; + + ASSERT_EQ(socketpair(AF_UNIX, SOCK_STREAM | SOCK_CLOEXEC, 0, ipc_sockets), 0); + ASSERT_EQ(pipe(pipefds), 0); + ASSERT_TRUE(set_core_pattern("@@/tmp/coredump.socket")); + + pid_coredump_server = fork(); + ASSERT_GE(pid_coredump_server, 0); + if (pid_coredump_server == 0) { + int fd_server = -1, fd_coredump = -1, fd_peer_pidfd = -1; + int fd_file = -1; + int exit_code = EXIT_FAILURE; + struct coredump_req req = {}; + bool is_truncated = false; + off_t size = 0; + ssize_t ret; + + close(ipc_sockets[0]); + close(pipefds[0]); + + fd_server = create_and_listen_unix_socket("/tmp/coredump.socket"); + if (fd_server < 0) + goto out; + + if (write_nointr(ipc_sockets[1], "1", 1) < 0) + goto out; + + close(ipc_sockets[1]); + + fd_coredump = accept4(fd_server, NULL, NULL, SOCK_CLOEXEC); + if (fd_coredump < 0) + goto out; + + fd_peer_pidfd = get_peer_pidfd(fd_coredump); + if (fd_peer_pidfd < 0) + goto out; + + /* + * The reassembled coredump is bigger than the mapping the + * child made, so keep it on the detached tmpfs and sparse. + */ + fd_file = open_coredump_tmpfile(self->fd_tmpfs_detached); + if (fd_file < 0) + goto out; + + if (!read_coredump_req(fd_coredump, &req)) + goto out; + + if (!check_coredump_req(&req)) + goto out; + + if (!send_coredump_ack(fd_coredump, &req, ack_mask, 0)) + goto out; + + if (!read_marker(fd_coredump, COREDUMP_MARK_REQACK)) + goto out; + + ret = recv_coredump_records(fd_coredump, fd_file, &size, &is_truncated, + kill_peer ? fd_peer_pidfd : -1); + if (ret < 0) + goto out; + + if (write_nointr(pipefds[1], &ret, sizeof(ret)) != sizeof(ret)) + goto out; + if (write_nointr(pipefds[1], &size, sizeof(size)) != sizeof(size)) + goto out; + if (write_nointr(pipefds[1], &is_truncated, + sizeof(is_truncated)) != sizeof(is_truncated)) + goto out; + + exit_code = EXIT_SUCCESS; +out: + close(pipefds[1]); + if (fd_file >= 0) + close(fd_file); + if (fd_peer_pidfd >= 0) + close(fd_peer_pidfd); + if (fd_coredump >= 0) + close(fd_coredump); + if (fd_server >= 0) + close(fd_server); + _exit(exit_code); + } + self->pid_coredump_server = pid_coredump_server; + + EXPECT_EQ(close(ipc_sockets[1]), 0); + EXPECT_EQ(close(pipefds[1]), 0); + ASSERT_EQ(read_nointr(ipc_sockets[0], &c, 1), 1); + EXPECT_EQ(close(ipc_sockets[0]), 0); + + pid = fork(); + ASSERT_GE(pid, 0); + if (pid == 0) + crashing_child_sparse(SPARSE_MAPPING_SIZE); + + pidfd = sys_pidfd_open(pid, 0); + ASSERT_GE(pidfd, 0); + + waitpid(pid, &status, 0); + ASSERT_TRUE(WIFSIGNALED(status)); + + ASSERT_EQ(read_nointr(pipefds[0], received, sizeof(*received)), + sizeof(*received)); + ASSERT_EQ(read_nointr(pipefds[0], coredump_size, sizeof(*coredump_size)), + sizeof(*coredump_size)); + ASSERT_EQ(read_nointr(pipefds[0], &truncated, sizeof(truncated)), + sizeof(truncated)); + EXPECT_EQ(close(pipefds[0]), 0); + + wait_and_check_coredump_server(pid_coredump_server, _metadata, self); + + if (kill_peer) { + /* The kernel gave up partway, so no end record closed the stream. */ + ASSERT_TRUE(truncated); + ASSERT_FALSE(WCOREDUMP(status)); + ASSERT_LT(*coredump_size, (off_t)SPARSE_MAPPING_SIZE); + return; + } + + ASSERT_FALSE(truncated); + ASSERT_TRUE(WCOREDUMP(status)); + + ASSERT_TRUE(get_pidfd_info(pidfd, &info)); + ASSERT_GT((info.mask & PIDFD_INFO_COREDUMP), 0); + ASSERT_GT((info.coredump_mask & PIDFD_COREDUMPED), 0); + + /* The mapping is in the coredump, holes included. */ + ASSERT_GT(*coredump_size, (off_t)SPARSE_MAPPING_SIZE); +} + +/* + * A mapping that has been written to is dumped whole, including the parts + * of it that were never faulted in. With COREDUMP_SPARSE the holes stay + * off the wire. + */ +TEST_F(coredump, socket_request_sparse_hole) +{ + off_t coredump_size = 0; + ssize_t received = 0; + + check_record_dump(_metadata, self, + COREDUMP_KERNEL | COREDUMP_RECORDS | + COREDUMP_SPARSE | COREDUMP_WAIT, + false, &received, &coredump_size); + + /* The holes didn't have to go over the socket. */ + ASSERT_LT(received, coredump_size / 8); +} + +/* + * COREDUMP_RECORDS alone splits the stream into records but elides + * nothing: the holes cross the socket as data records. + */ +TEST_F(coredump, socket_request_records_hole) +{ + off_t coredump_size = 0; + ssize_t received = 0; + + check_record_dump(_metadata, self, + COREDUMP_KERNEL | COREDUMP_RECORDS | COREDUMP_WAIT, + false, &received, &coredump_size); + + /* Records alone elide nothing, so everything crossed the socket. */ + ASSERT_GT(received, coredump_size); +} + +/* + * A coredump the kernel gives up on halfway still ends in an end record, + * and that record says the coredump is incomplete. COREDUMP_SPARSE is left + * out on purpose: the holes have to cross the socket so the coredump is + * far larger than the socket buffer and the kernel is still writing it + * when the kill lands. + */ +TEST_F(coredump, socket_request_records_truncated) +{ + off_t coredump_size = 0; + ssize_t received = 0; + + check_record_dump(_metadata, self, + COREDUMP_KERNEL | COREDUMP_RECORDS | COREDUMP_WAIT, + true, &received, &coredump_size); + + /* The end record crossed the socket even though the task was killed. */ + ASSERT_GT(received, 0); +} + +/* + * A coredump server that uploads to a blob store can't upload a sparse + * file. It doesn't have to: it streams the data records into the object + * as they arrive, leaves the holes out, and uploads the corrected + * program header table last. What it ends up with is an ordinary ELF + * core file that describes the same memory as the coredump the records + * came from, minus the holes. + */ +TEST_F(coredump, socket_request_sparse_blob_upload) +{ + int fd_core_file, pidfd, status; + pid_t pid, pid_coredump_server; + struct pidfd_info info = {}; + off_t coredump_size = 0; + ssize_t received = 0; + int ipc_sockets[2]; + int pipefds[2]; + struct stat st; + char c; + + ASSERT_EQ(socketpair(AF_UNIX, SOCK_STREAM | SOCK_CLOEXEC, 0, ipc_sockets), 0); + ASSERT_EQ(pipe(pipefds), 0); + ASSERT_TRUE(set_core_pattern("@@/tmp/coredump.socket")); + + pid_coredump_server = fork(); + ASSERT_GE(pid_coredump_server, 0); + if (pid_coredump_server == 0) { + int fd_server = -1, fd_coredump = -1, fd_peer_pidfd = -1; + int fd_object = -1, fd_reference = -1; + int exit_code = EXIT_FAILURE; + struct coredump_req req = {}; + off_t size = 0; + ssize_t ret; + + close(ipc_sockets[0]); + close(pipefds[0]); + + fd_server = create_and_listen_unix_socket("/tmp/coredump.socket"); + if (fd_server < 0) + goto out; + + if (write_nointr(ipc_sockets[1], "1", 1) < 0) + goto out; + + close(ipc_sockets[1]); + + fd_coredump = accept4(fd_server, NULL, NULL, SOCK_CLOEXEC); + if (fd_coredump < 0) + goto out; + + fd_peer_pidfd = get_peer_pidfd(fd_coredump); + if (fd_peer_pidfd < 0) + goto out; + + /* The object is a plain file. It never sees a hole. */ + fd_object = open("/tmp/coredump.file", + O_RDWR | O_CREAT | O_TRUNC | O_CLOEXEC, 0600); + if (fd_object < 0) + goto out; + + /* + * The coredump with its holes still in it is bigger than + * the mapping the child made, so keep it on the detached + * tmpfs and sparse. + */ + fd_reference = open_coredump_tmpfile(self->fd_tmpfs_detached); + if (fd_reference < 0) + goto out; + + if (!read_coredump_req(fd_coredump, &req)) + goto out; + + if (!check_coredump_req(&req)) + goto out; + + if (!send_coredump_ack(fd_coredump, &req, + COREDUMP_KERNEL | COREDUMP_RECORDS | + COREDUMP_SPARSE | COREDUMP_WAIT, 0)) + goto out; + + if (!read_marker(fd_coredump, COREDUMP_MARK_REQACK)) + goto out; + + ret = recv_coredump_compact(fd_coredump, fd_object, + fd_reference, &size); + if (ret < 0) + goto out; + + if (check_compact_coredump(fd_object, fd_reference)) + goto out; + + if (write_nointr(pipefds[1], &ret, sizeof(ret)) != sizeof(ret)) + goto out; + if (write_nointr(pipefds[1], &size, sizeof(size)) != sizeof(size)) + goto out; + + exit_code = EXIT_SUCCESS; +out: + close(pipefds[1]); + if (fd_reference >= 0) + close(fd_reference); + if (fd_object >= 0) + close(fd_object); + if (fd_peer_pidfd >= 0) + close(fd_peer_pidfd); + if (fd_coredump >= 0) + close(fd_coredump); + if (fd_server >= 0) + close(fd_server); + _exit(exit_code); + } + self->pid_coredump_server = pid_coredump_server; + + EXPECT_EQ(close(ipc_sockets[1]), 0); + EXPECT_EQ(close(pipefds[1]), 0); + ASSERT_EQ(read_nointr(ipc_sockets[0], &c, 1), 1); + EXPECT_EQ(close(ipc_sockets[0]), 0); + + pid = fork(); + ASSERT_GE(pid, 0); + if (pid == 0) + crashing_child_sparse(SPARSE_MAPPING_SIZE); + + pidfd = sys_pidfd_open(pid, 0); + ASSERT_GE(pidfd, 0); + + waitpid(pid, &status, 0); + ASSERT_TRUE(WIFSIGNALED(status)); + ASSERT_TRUE(WCOREDUMP(status)); + + ASSERT_EQ(read_nointr(pipefds[0], &received, sizeof(received)), + sizeof(received)); + ASSERT_EQ(read_nointr(pipefds[0], &coredump_size, sizeof(coredump_size)), + sizeof(coredump_size)); + EXPECT_EQ(close(pipefds[0]), 0); + + ASSERT_TRUE(get_pidfd_info(pidfd, &info)); + ASSERT_GT((info.mask & PIDFD_INFO_COREDUMP), 0); + ASSERT_GT((info.coredump_mask & PIDFD_COREDUMPED), 0); + + wait_and_check_coredump_server(pid_coredump_server, _metadata, self); + + /* The mapping is in the coredump, holes included. */ + ASSERT_GT(coredump_size, (off_t)SPARSE_MAPPING_SIZE); + + /* The object isn't sparse and doesn't carry them. */ + ASSERT_EQ(stat("/tmp/coredump.file", &st), 0); + ASSERT_LT(st.st_size, coredump_size / 8); + + /* And a debugger still sees an ordinary ELF core file. */ + fd_core_file = open("/tmp/coredump.file", O_RDONLY | O_CLOEXEC); + ASSERT_GE(fd_core_file, 0); + ASSERT_TRUE(is_elf_core(fd_core_file)); + EXPECT_EQ(close(fd_core_file), 0); +} + +/* COREDUMP_RECORDS applies to a coredump the kernel writes, nothing else. */ +TEST_F(coredump, socket_request_records_without_kernel) +{ + check_conflicting_ack(_metadata, self, COREDUMP_USERSPACE | COREDUMP_RECORDS); +} + +/* A zero record can't exist outside a record stream. */ +TEST_F(coredump, socket_request_sparse_without_records) +{ + check_conflicting_ack(_metadata, self, COREDUMP_KERNEL | COREDUMP_SPARSE); +} + +/* What the server reports back about the coredump it decided to take. */ +struct stream_choice { + bool sparse; + ssize_t received; + off_t size; + ssize_t vm_size; +}; + +/* + * The kernel blocks in the coredump request until the ack arrives, so a + * coredump server gets to look at the task before it commits to a + * stream. Take the record stream only for a task whose mappings are + * worth it and the plain byte stream for everything else. + */ +static void check_stream_choice(struct __test_metadata *const _metadata, + FIXTURE_DATA(coredump) *self, bool big, + struct stream_choice *choice) +{ + int pidfd, status; + pid_t pid, pid_coredump_server; + struct pidfd_info info = {}; + int ipc_sockets[2]; + int pipefds[2]; + char c; + + ASSERT_EQ(socketpair(AF_UNIX, SOCK_STREAM | SOCK_CLOEXEC, 0, ipc_sockets), 0); + ASSERT_EQ(pipe(pipefds), 0); + ASSERT_TRUE(set_core_pattern("@@/tmp/coredump.socket")); + + pid_coredump_server = fork(); + ASSERT_GE(pid_coredump_server, 0); + if (pid_coredump_server == 0) { + int fd_server = -1, fd_coredump = -1, fd_peer_pidfd = -1; + int fd_file = -1; + int exit_code = EXIT_FAILURE; + struct coredump_req req = {}; + struct stream_choice got = {}; + __u64 mask; + + close(ipc_sockets[0]); + close(pipefds[0]); + + fd_server = create_and_listen_unix_socket("/tmp/coredump.socket"); + if (fd_server < 0) + goto out; + + if (write_nointr(ipc_sockets[1], "1", 1) < 0) + goto out; + + close(ipc_sockets[1]); + + fd_coredump = accept4(fd_server, NULL, NULL, SOCK_CLOEXEC); + if (fd_coredump < 0) + goto out; + + fd_peer_pidfd = get_peer_pidfd(fd_coredump); + if (fd_peer_pidfd < 0) + goto out; + + /* + * The reassembled coredump is bigger than the mapping the + * child made, so keep it on the detached tmpfs and sparse. + */ + fd_file = open_coredump_tmpfile(self->fd_tmpfs_detached); + if (fd_file < 0) + goto out; + + if (!read_coredump_req(fd_coredump, &req)) + goto out; + + if (!check_coredump_req(&req)) + goto out; + + /* + * Nothing is on the wire yet and the kernel is waiting for + * the ack, so there is all the time in the world to look at + * the task and decide what to ask it for. + */ + got.vm_size = peer_vm_size(fd_peer_pidfd); + if (got.vm_size < 0) + goto out; + got.sparse = got.vm_size >= SPARSE_STREAM_THRESHOLD; + + fprintf(stderr, "Peer maps %zd bytes, asking for %s\n", + got.vm_size, + got.sparse ? "a sparse record stream" : "a byte stream"); + + mask = COREDUMP_KERNEL | COREDUMP_WAIT; + if (got.sparse) + mask |= COREDUMP_RECORDS | COREDUMP_SPARSE; + + if (!send_coredump_ack(fd_coredump, &req, mask, 0)) + goto out; + + if (!read_marker(fd_coredump, COREDUMP_MARK_REQACK)) + goto out; + + if (got.sparse) { + got.received = recv_coredump_records(fd_coredump, fd_file, + &got.size, NULL, -1); + } else { + got.received = recv_coredump_bytes(fd_coredump, fd_file); + got.size = got.received; + } + if (got.received < 0) + goto out; + + /* Either way a debugger has to see an ordinary core file. */ + if (!is_elf_core(fd_file)) + goto out; + + if (write_nointr(pipefds[1], &got, sizeof(got)) != sizeof(got)) + goto out; + + exit_code = EXIT_SUCCESS; +out: + close(pipefds[1]); + if (fd_file >= 0) + close(fd_file); + if (fd_peer_pidfd >= 0) + close(fd_peer_pidfd); + if (fd_coredump >= 0) + close(fd_coredump); + if (fd_server >= 0) + close(fd_server); + _exit(exit_code); + } + self->pid_coredump_server = pid_coredump_server; + + EXPECT_EQ(close(ipc_sockets[1]), 0); + EXPECT_EQ(close(pipefds[1]), 0); + ASSERT_EQ(read_nointr(ipc_sockets[0], &c, 1), 1); + EXPECT_EQ(close(ipc_sockets[0]), 0); + + pid = fork(); + ASSERT_GE(pid, 0); + if (pid == 0) + crashing_child_sparse(big ? SPARSE_MAPPING_SIZE : PAGE_SIZE); + + pidfd = sys_pidfd_open(pid, 0); + ASSERT_GE(pidfd, 0); + + waitpid(pid, &status, 0); + ASSERT_TRUE(WIFSIGNALED(status)); + ASSERT_TRUE(WCOREDUMP(status)); + + ASSERT_EQ(read_nointr(pipefds[0], choice, sizeof(*choice)), + sizeof(*choice)); + EXPECT_EQ(close(pipefds[0]), 0); + + ASSERT_TRUE(get_pidfd_info(pidfd, &info)); + ASSERT_GT((info.mask & PIDFD_INFO_COREDUMP), 0); + ASSERT_GT((info.coredump_mask & PIDFD_COREDUMPED), 0); + + wait_and_check_coredump_server(pid_coredump_server, _metadata, self); +} + +/* A task with little mapped isn't worth a record stream. */ +TEST_F(coredump, socket_request_stream_choice_small) +{ + struct stream_choice choice = {}; + + check_stream_choice(_metadata, self, false, &choice); + + ASSERT_LT(choice.vm_size, (ssize_t)SPARSE_STREAM_THRESHOLD); + ASSERT_FALSE(choice.sparse); + ASSERT_GT(choice.received, 0); +} + +/* A task sitting on a big mapping is. */ +TEST_F(coredump, socket_request_stream_choice_large) +{ + struct stream_choice choice = {}; + + check_stream_choice(_metadata, self, true, &choice); + + ASSERT_GE(choice.vm_size, (ssize_t)SPARSE_STREAM_THRESHOLD); + ASSERT_TRUE(choice.sparse); + ASSERT_GT(choice.size, (off_t)SPARSE_MAPPING_SIZE); + + /* The holes didn't have to go over the socket. */ + ASSERT_LT(choice.received, choice.size / 8); +} + +/* What a coredump server was built with. */ +struct server_build { + /* sizeof(struct coredump_req) and sizeof(struct coredump_ack) back then. */ + size_t req_size; + size_t ack_size; + /* The features it raises if the kernel offers them. */ + __u64 wants; + /* Its policy: what it drops from and adds to the task's selection. */ + __u64 drop; + __u64 add; +}; + +/* A server from when the structs were first published: kernel-written dumps. */ +static const struct server_build server_build_ver0 = { + .req_size = COREDUMP_REQ_SIZE_VER0, + .ack_size = COREDUMP_ACK_SIZE_VER0, + .wants = COREDUMP_KERNEL, +}; + +/* A server built against this header: no shared memory, always the ELF headers. */ +static const struct server_build server_build_ver1 = { + .req_size = sizeof(struct coredump_req), + .ack_size = sizeof(struct coredump_ack), + .wants = COREDUMP_KERNEL | COREDUMP_RECORDS | COREDUMP_SPARSE | + COREDUMP_MEMORY_TYPES, + .drop = COREDUMP_MEMORY_ANON_SHARED | COREDUMP_MEMORY_FILE_SHARED, + .add = COREDUMP_MEMORY_ELF_HEADERS, +}; + +/* + * Build the ack the way a server does: from what the kernel offers, what + * this build implements, and what fits in the ack the kernel accepts. + * Fields the build never read are zero and never consulted. + */ +static void negotiate(const struct coredump_req *req, + const struct server_build *build, + struct coredump_ack *ack) +{ + __u64 offered = req->mask & build->wants; + + memset(ack, 0, sizeof(*ack)); + ack->size = build->ack_size < req->size_ack ? build->ack_size : req->size_ack; + /* These builds only ever have the kernel write the coredump. */ + ack->mask = COREDUMP_KERNEL; + + /* Sparse needs records, records need the kernel to write. */ + if (offered & COREDUMP_RECORDS) { + ack->mask |= COREDUMP_RECORDS; + if (offered & COREDUMP_SPARSE) + ack->mask |= COREDUMP_SPARSE; + } + + /* The memory types need an ack that carries them. */ + if ((offered & COREDUMP_MEMORY_TYPES) && ack->size >= COREDUMP_ACK_SIZE_VER1) { + ack->mask |= COREDUMP_MEMORY_TYPES; + /* Start from the task's selection; only advertised types pass. */ + ack->memory_types = (req->memory_types & ~build->drop) | build->add; + ack->memory_types &= req->memory_types_mask; + } +} + +/* What a memory types test asks of the kernel and what it expects back. */ +struct memory_choice { + /* Memory types the crashing child selects, or FILTER_TASK_INHERIT. */ + __u64 task_filter; + /* Negotiate the ack as this server build, NULL to send it as given. */ + const struct server_build *build; + /* The ack, or what the negotiation must arrive at. */ + __u64 mask; + __u64 memory_types; + size_t size_ack; + /* The shared mapping is in the coredump with all of its memory. */ + bool shared_dumped; + /* No memory at all. Pull a page from /proc/<pid>/mem instead. */ + bool skeleton; +}; + +/* A skeleton still carries the vdso and friends, nothing bigger. */ +#define SKELETON_DATA_PAGES 16 + +/* + * The crashing child maps shared anonymous memory and tells the server + * where. The server acks with @choice and checks whether that mapping's + * segment in the coredump carries its memory. + */ +static void check_memory_dump(struct __test_metadata *const _metadata, + FIXTURE_DATA(coredump) *self, + const struct memory_choice *choice) +{ + int pidfd, status; + pid_t pid, pid_coredump_server; + struct pidfd_info info = {}; + int ipc_sockets[2]; + int addr_pipe[2]; + char c; + + ASSERT_EQ(socketpair(AF_UNIX, SOCK_STREAM | SOCK_CLOEXEC, 0, ipc_sockets), 0); + ASSERT_EQ(pipe(addr_pipe), 0); + ASSERT_TRUE(set_core_pattern("@@/tmp/coredump.socket")); + + pid_coredump_server = fork(); + ASSERT_GE(pid_coredump_server, 0); + if (pid_coredump_server == 0) { + int fd_server = -1, fd_coredump = -1, fd_peer_pidfd = -1; + int fd_file = -1; + int exit_code = EXIT_FAILURE; + struct coredump_req req = {}; + struct coredump_ack ack = { + .size = choice->size_ack, + .mask = choice->mask, + .memory_types = choice->memory_types, + }; + /* How much of the request this server reads. */ + size_t req_size = choice->build ? choice->build->req_size : sizeof(req); + __u64 task_filter; + ElfW(Phdr) segment; + ssize_t received; + off_t size; + char *addr; + + close(ipc_sockets[0]); + close(addr_pipe[1]); + + fd_server = create_and_listen_unix_socket("/tmp/coredump.socket"); + if (fd_server < 0) + goto out; + + if (write_nointr(ipc_sockets[1], "1", 1) < 0) + goto out; + + close(ipc_sockets[1]); + + fd_coredump = accept4(fd_server, NULL, NULL, SOCK_CLOEXEC); + if (fd_coredump < 0) + goto out; + + fd_peer_pidfd = get_peer_pidfd(fd_coredump); + if (fd_peer_pidfd < 0) + goto out; + + fd_file = open_coredump_tmpfile(self->fd_tmpfs_detached); + if (fd_file < 0) + goto out; + + if (!read_coredump_req_sized(fd_coredump, &req, req_size)) + goto out; + + if (!peer_coredump_filter(fd_peer_pidfd, &task_filter)) + goto out; + + /* A build from before the memory types never read that far. */ + if (req_size >= COREDUMP_REQ_SIZE_VER1) { + if (!check_coredump_req(&req)) + goto out; + + /* The request reports the memory types the task selected. */ + if (req.memory_types != task_filter) { + fprintf(stderr, "Request reports 0x%llx, task selected 0x%llx\n", + (unsigned long long)req.memory_types, + (unsigned long long)task_filter); + goto out; + } + } + + if (choice->task_filter != FILTER_TASK_INHERIT && + task_filter != choice->task_filter) { + fprintf(stderr, "Task selected 0x%llx, child asked for 0x%llx\n", + (unsigned long long)task_filter, + (unsigned long long)choice->task_filter); + goto out; + } + + /* The child sent the address of its mapping before it crashed. */ + if (read_nointr(addr_pipe[0], &addr, sizeof(addr)) != sizeof(addr)) + goto out; + + /* A server build negotiates its ack and must arrive at the choice. */ + if (choice->build) { + negotiate(&req, choice->build, &ack); + + if (ack.size != choice->size_ack || ack.mask != choice->mask || + ack.memory_types != choice->memory_types) { + fprintf(stderr, + "Negotiated %u bytes, mask 0x%llx, types 0x%llx\n", + ack.size, (unsigned long long)ack.mask, + (unsigned long long)ack.memory_types); + goto out; + } + } + + if (!send_coredump_ack_types(fd_coredump, &req, ack.mask, + ack.memory_types, ack.size)) + goto out; + + if (!read_marker(fd_coredump, COREDUMP_MARK_REQACK)) + goto out; + + if (ack.mask & COREDUMP_RECORDS) + received = recv_coredump_records(fd_coredump, fd_file, + &size, NULL, -1); + else + received = recv_coredump_bytes(fd_coredump, fd_file); + if (received < 0) + goto out; + + if (!is_elf_core(fd_file)) + goto out; + + /* A dump ending in holes or empty segments must still be whole. */ + if (!check_coredump_extent(fd_file)) + goto out; + + if (!find_coredump_segment(fd_file, (__u64)(uintptr_t)addr, &segment)) + goto out; + + if (segment.p_memsz != MEMORY_MAPPING_SIZE) { + fprintf(stderr, "Segment spans %llu bytes, the mapping %u\n", + (unsigned long long)segment.p_memsz, + MEMORY_MAPPING_SIZE); + goto out; + } + + if (segment.p_filesz != (choice->shared_dumped ? segment.p_memsz : 0)) { + fprintf(stderr, "Segment carries %llu bytes, expected %s of them\n", + (unsigned long long)segment.p_filesz, + choice->shared_dumped ? "all" : "none"); + goto out; + } + + if (choice->skeleton) { + __u64 data, notes, data_max; + char buf[PAGE_SIZE]; + + if (!sum_coredump_segments(fd_file, &data, ¬es)) + goto out; + + data_max = SKELETON_DATA_PAGES * sysconf(_SC_PAGESIZE); + if (!notes || data > data_max) { + fprintf(stderr, "Skeleton has %llu note and %llu memory bytes\n", + (unsigned long long)notes, + (unsigned long long)data); + goto out; + } + + /* The task is parked in COREDUMP_WAIT with its memory. */ + if (peer_read_mem(fd_peer_pidfd, (__u64)(uintptr_t)addr, + buf, sizeof(buf)) != sizeof(buf)) + goto out; + + if (buf[0] != 'x') { + fprintf(stderr, "Pulled memory lacks the child's mark\n"); + goto out; + } + + fprintf(stderr, "Skeleton of %zd bytes, pulled %zu bytes of memory\n", + received, sizeof(buf)); + } + + exit_code = EXIT_SUCCESS; +out: + close(addr_pipe[0]); + if (fd_file >= 0) + close(fd_file); + if (fd_peer_pidfd >= 0) + close(fd_peer_pidfd); + if (fd_coredump >= 0) + close(fd_coredump); + if (fd_server >= 0) + close(fd_server); + _exit(exit_code); + } + self->pid_coredump_server = pid_coredump_server; + + EXPECT_EQ(close(ipc_sockets[1]), 0); + EXPECT_EQ(close(addr_pipe[0]), 0); + ASSERT_EQ(read_nointr(ipc_sockets[0], &c, 1), 1); + EXPECT_EQ(close(ipc_sockets[0]), 0); + + pid = fork(); + ASSERT_GE(pid, 0); + if (pid == 0) + crashing_child_memory(choice->task_filter, addr_pipe[1]); + EXPECT_EQ(close(addr_pipe[1]), 0); + + pidfd = sys_pidfd_open(pid, 0); + ASSERT_GE(pidfd, 0); + + waitpid(pid, &status, 0); + ASSERT_TRUE(WIFSIGNALED(status)); + ASSERT_TRUE(WCOREDUMP(status)); + + ASSERT_TRUE(get_pidfd_info(pidfd, &info)); + ASSERT_GT((info.mask & PIDFD_INFO_COREDUMP), 0); + ASSERT_GT((info.coredump_mask & PIDFD_COREDUMPED), 0); + + wait_and_check_coredump_server(pid_coredump_server, _metadata, self); +} + +/* Without COREDUMP_MEMORY_TYPES the task's own selection decides. */ +TEST_F(coredump, socket_request_memory_types_task_includes) +{ + struct memory_choice choice = { + .task_filter = COREDUMP_MEMORY_ANON_PRIVATE | + COREDUMP_MEMORY_ANON_SHARED, + .mask = COREDUMP_KERNEL, + .shared_dumped = true, + }; + + check_memory_dump(_metadata, self, &choice); +} + +TEST_F(coredump, socket_request_memory_types_task_excludes) +{ + struct memory_choice choice = { + .task_filter = 0, + .mask = COREDUMP_KERNEL, + .shared_dumped = false, + }; + + check_memory_dump(_metadata, self, &choice); +} + +/* The server drops a memory type the task would have dumped. */ +TEST_F(coredump, socket_request_memory_types_restricts) +{ + struct memory_choice choice = { + .task_filter = COREDUMP_MEMORY_ANON_PRIVATE | + COREDUMP_MEMORY_ANON_SHARED, + .mask = COREDUMP_KERNEL | COREDUMP_MEMORY_TYPES, + .memory_types = COREDUMP_MEMORY_ANON_PRIVATE, + .shared_dumped = false, + }; + + check_memory_dump(_metadata, self, &choice); +} + +/* The server adds a memory type the task had excluded. */ +TEST_F(coredump, socket_request_memory_types_widens) +{ + struct memory_choice choice = { + .task_filter = 0, + .mask = COREDUMP_KERNEL | COREDUMP_MEMORY_TYPES, + .memory_types = COREDUMP_MEMORY_ANON_PRIVATE | + COREDUMP_MEMORY_ANON_SHARED, + .shared_dumped = true, + }; + + check_memory_dump(_metadata, self, &choice); +} + +/* The memory types decide what goes into a record stream just the same. */ +TEST_F(coredump, socket_request_memory_types_records) +{ + struct memory_choice choice = { + .task_filter = FILTER_TASK_INHERIT, + .mask = COREDUMP_KERNEL | COREDUMP_RECORDS | COREDUMP_SPARSE | + COREDUMP_MEMORY_TYPES, + .memory_types = COREDUMP_MEMORY_ANON_PRIVATE, + .shared_dumped = false, + }; + + check_memory_dump(_metadata, self, &choice); +} + +/* + * An empty selection leaves a skeleton: every program header and every note + * but no memory. A server that wants to pick the memory itself reads it + * from /proc/<pid>/mem while the task waits for it to finish. + */ +TEST_F(coredump, socket_request_memory_types_skeleton) +{ + struct memory_choice choice = { + .task_filter = FILTER_TASK_INHERIT, + .mask = COREDUMP_KERNEL | COREDUMP_WAIT | COREDUMP_MEMORY_TYPES, + .memory_types = 0, + .shared_dumped = false, + .skeleton = true, + }; + + check_memory_dump(_metadata, self, &choice); +} + +/* A memory type the kernel didn't advertise in memory_types_mask. */ +TEST_F(coredump, socket_request_memory_types_unknown_bit) +{ + struct refused_ack refused = { + .ack = { + .size = sizeof(struct coredump_ack), + .mask = COREDUMP_KERNEL | COREDUMP_MEMORY_TYPES, + .memory_types = 1ULL << 63, + }, + .bytes = sizeof(struct coredump_ack), + .mark = COREDUMP_MARK_UNSUPPORTED, + }; + + check_refused_ack(_metadata, self, &refused); +} + +/* The memory types must be zero unless COREDUMP_MEMORY_TYPES is raised. */ +TEST_F(coredump, socket_request_memory_types_stale_field) +{ + struct refused_ack refused = { + .ack = { + .size = sizeof(struct coredump_ack), + .mask = COREDUMP_KERNEL, + .memory_types = COREDUMP_MEMORY_ANON_PRIVATE, + }, + .bytes = sizeof(struct coredump_ack), + .mark = COREDUMP_MARK_UNSUPPORTED, + }; + + check_refused_ack(_metadata, self, &refused); +} + +/* COREDUMP_MEMORY_TYPES needs an ack that has the memory types. */ +TEST_F(coredump, socket_request_memory_types_short_ack) +{ + struct refused_ack refused = { + .ack = { + .size = COREDUMP_ACK_SIZE_VER0, + .mask = COREDUMP_KERNEL | COREDUMP_MEMORY_TYPES, + }, + .bytes = COREDUMP_ACK_SIZE_VER0, + .mark = COREDUMP_MARK_MINSIZE, + }; + + check_refused_ack(_metadata, self, &refused); +} + +/* The memory types select what the kernel writes, nothing else. */ +TEST_F(coredump, socket_request_memory_types_without_kernel) +{ + check_conflicting_ack(_metadata, self, COREDUMP_USERSPACE | COREDUMP_MEMORY_TYPES); +} + +/* + * A server built with the first structs reads the request it knows, + * discards the rest and acks with the ack it knows. It raises nothing + * it wasn't built for and the kernel dumps what the task selected. + */ +TEST_F(coredump, socket_request_negotiate_ver0) +{ + struct memory_choice choice = { + .task_filter = COREDUMP_MEMORY_ANON_PRIVATE | + COREDUMP_MEMORY_ANON_SHARED, + .build = &server_build_ver0, + .mask = COREDUMP_KERNEL, + .memory_types = 0, + .size_ack = COREDUMP_ACK_SIZE_VER0, + .shared_dumped = true, + }; + + check_memory_dump(_metadata, self, &choice); +} + +/* + * A server built against this header takes every feature the kernel + * offers, drops shared memory from what the task selected and adds the + * ELF headers. + */ +TEST_F(coredump, socket_request_negotiate_ver1) +{ + struct memory_choice choice = { + .task_filter = COREDUMP_MEMORY_ANON_PRIVATE | + COREDUMP_MEMORY_ANON_SHARED, + .build = &server_build_ver1, + .mask = COREDUMP_KERNEL | COREDUMP_RECORDS | COREDUMP_SPARSE | + COREDUMP_MEMORY_TYPES, + .memory_types = COREDUMP_MEMORY_ANON_PRIVATE | + COREDUMP_MEMORY_ELF_HEADERS, + .size_ack = COREDUMP_ACK_SIZE_VER1, + .shared_dumped = false, + }; + + check_memory_dump(_metadata, self, &choice); +} + +/* An ack that picks none of KERNEL, USERSPACE and REJECT. */ +TEST_F(coredump, socket_request_no_mode) +{ + struct refused_ack refused = { + .ack = { + .size = sizeof(struct coredump_ack), + .mask = COREDUMP_WAIT, + }, + .bytes = sizeof(struct coredump_ack), + .mark = COREDUMP_MARK_CONFLICTING, + }; + + check_refused_ack(_metadata, self, &refused); +} + +/* @spare must be zero, like every field that isn't in use. */ +TEST_F(coredump, socket_request_spare) +{ + struct refused_ack refused = { + .ack = { + .size = sizeof(struct coredump_ack), + .spare = 1, + .mask = COREDUMP_KERNEL, + }, + .bytes = sizeof(struct coredump_ack), + .mark = COREDUMP_MARK_UNSUPPORTED, + }; + + check_refused_ack(_metadata, self, &refused); +} + +/* An ack size is a byte count. One that ends inside a field is valid. */ +#define ACK_SIZE_BETWEEN (COREDUMP_ACK_SIZE_VER0 + sizeof(__u32)) + +/* Any size from VER0 up to what the kernel accepts works without memory types. */ +TEST_F(coredump, socket_request_ack_size_between) +{ + struct memory_choice choice = { + .task_filter = COREDUMP_MEMORY_ANON_PRIVATE | + COREDUMP_MEMORY_ANON_SHARED, + .mask = COREDUMP_KERNEL, + .size_ack = ACK_SIZE_BETWEEN, + .shared_dumped = true, + }; + + check_memory_dump(_metadata, self, &choice); +} + +/* The memory types need the whole field, not the part that happens to fit. */ +TEST_F(coredump, socket_request_memory_types_ack_size_between) +{ + struct refused_ack refused = { + .ack = { + .size = ACK_SIZE_BETWEEN, + .mask = COREDUMP_KERNEL | COREDUMP_MEMORY_TYPES, + }, + .bytes = ACK_SIZE_BETWEEN, + .mark = COREDUMP_MARK_MINSIZE, + }; + + check_refused_ack(_metadata, self, &refused); +} + +/* A server that hangs up without acking gets no marker and no coredump. */ +TEST_F(coredump, socket_request_server_hangs_up) +{ + struct refused_ack refused = { + .bytes = 0, + .no_marker = true, + }; + + check_refused_ack(_metadata, self, &refused); +} + +/* A server that hangs up in the middle of its ack looks the same. */ +TEST_F(coredump, socket_request_ack_truncated) +{ + struct refused_ack refused = { + .ack = { + .size = COREDUMP_ACK_SIZE_VER0, + .mask = COREDUMP_KERNEL, + }, + .bytes = COREDUMP_ACK_SIZE_VER0 / 2, + .no_marker = true, + }; + + check_refused_ack(_metadata, self, &refused); +} + +/* + * The kernels a server built against this header can't meet here: + * negotiate() against their requests, no coredump involved. + */ + +/* The request of a kernel with the first structs and features. */ +static const struct coredump_req req_ver0 = { + .size = COREDUMP_REQ_SIZE_VER0, + .size_ack = COREDUMP_ACK_SIZE_VER0, + .mask = COREDUMP_KERNEL | COREDUMP_USERSPACE | + COREDUMP_REJECT | COREDUMP_WAIT, +}; + +/* The request of this kernel. */ +static const struct coredump_req req_ver1 = { + .size = COREDUMP_REQ_SIZE_VER1, + .size_ack = COREDUMP_ACK_SIZE_VER1, + .mask = COREDUMP_KERNEL | COREDUMP_USERSPACE | + COREDUMP_REJECT | COREDUMP_WAIT | + COREDUMP_RECORDS | COREDUMP_SPARSE | + COREDUMP_MEMORY_TYPES, + .memory_types = COREDUMP_MEMORY_ANON_PRIVATE | + COREDUMP_MEMORY_ANON_SHARED, + .memory_types_mask = TEST_MEMORY_ALL, +}; + +/* A kernel with the first structs gets the first ack and nothing newer. */ +TEST(negotiate_ver0_kernel) +{ + struct coredump_ack ack; + + negotiate(&req_ver0, &server_build_ver1, &ack); + ASSERT_EQ(ack.size, COREDUMP_ACK_SIZE_VER0); + ASSERT_EQ(ack.mask, COREDUMP_KERNEL); + ASSERT_EQ(ack.memory_types, 0); +} + +/* A kernel with records and sparse but the first structs: both, no types. */ +TEST(negotiate_sparse_kernel) +{ + struct coredump_req req = req_ver0; + struct coredump_ack ack; + + req.mask |= COREDUMP_RECORDS | COREDUMP_SPARSE; + negotiate(&req, &server_build_ver1, &ack); + ASSERT_EQ(ack.size, COREDUMP_ACK_SIZE_VER0); + ASSERT_EQ(ack.mask, COREDUMP_KERNEL | COREDUMP_RECORDS | COREDUMP_SPARSE); + ASSERT_EQ(ack.memory_types, 0); +} + +/* Records without sparse: sparse isn't raised on its own. */ +TEST(negotiate_records_without_sparse) +{ + struct coredump_req req = req_ver0; + struct coredump_ack ack; + + req.mask |= COREDUMP_RECORDS; + negotiate(&req, &server_build_ver1, &ack); + ASSERT_EQ(ack.mask, COREDUMP_KERNEL | COREDUMP_RECORDS); +} + +/* + * A feature whose ack field lies past what the kernel accepts can't be + * raised. No kernel offers the memory types without the room for them, so a + * request that does stands in for a feature newer than this header. + */ +TEST(negotiate_types_need_room) +{ + struct coredump_req req = req_ver0; + struct coredump_ack ack; + + req.mask |= COREDUMP_MEMORY_TYPES; + negotiate(&req, &server_build_ver1, &ack); + ASSERT_EQ(ack.size, COREDUMP_ACK_SIZE_VER0); + ASSERT_EQ(ack.mask, COREDUMP_KERNEL); + ASSERT_EQ(ack.memory_types, 0); +} + +/* This kernel: the policy applied to the task's selection. */ +TEST(negotiate_ver1_kernel) +{ + struct coredump_ack ack; + + negotiate(&req_ver1, &server_build_ver1, &ack); + ASSERT_EQ(ack.size, COREDUMP_ACK_SIZE_VER1); + ASSERT_EQ(ack.mask, COREDUMP_KERNEL | COREDUMP_RECORDS | COREDUMP_SPARSE | + COREDUMP_MEMORY_TYPES); + ASSERT_EQ(ack.memory_types, COREDUMP_MEMORY_ANON_PRIVATE | + COREDUMP_MEMORY_ELF_HEADERS); +} + +/* A kernel that doesn't know a type the policy adds isn't asked for it. */ +TEST(negotiate_unknown_type) +{ + struct coredump_req req = req_ver1; + struct coredump_ack ack; + + req.memory_types_mask &= ~(__u64)COREDUMP_MEMORY_ELF_HEADERS; + negotiate(&req, &server_build_ver1, &ack); + ASSERT_EQ(ack.mask, COREDUMP_KERNEL | COREDUMP_RECORDS | COREDUMP_SPARSE | + COREDUMP_MEMORY_TYPES); + ASSERT_EQ(ack.memory_types, COREDUMP_MEMORY_ANON_PRIVATE); +} + TEST_HARNESS_MAIN diff --git a/tools/testing/selftests/coredump/coredump_test.h b/tools/testing/selftests/coredump/coredump_test.h index ed47f01fa53c..1c154f513700 100644 --- a/tools/testing/selftests/coredump/coredump_test.h +++ b/tools/testing/selftests/coredump/coredump_test.h @@ -3,18 +3,10 @@ #ifndef __COREDUMP_TEST_H #define __COREDUMP_TEST_H -#include <stdbool.h> -#include <sys/types.h> -#include <linux/coredump.h> - #include "../kselftest_harness.h" -#include "../pidfd/pidfd.h" - -#ifndef PAGE_SIZE -#define PAGE_SIZE 4096 -#endif +#include "coredump_notify_signal.h" -#define NUM_THREAD_SPAWN 128 +#include "coredump_test_helpers.h" /* Coredump fixture */ FIXTURE(coredump) @@ -24,15 +16,6 @@ FIXTURE(coredump) int fd_tmpfs_detached; }; -/* Shared helper function declarations */ -void *do_nothing(void *arg); -void crashing_child(void); -int create_detached_tmpfs(void); -int create_and_listen_unix_socket(const char *path); -bool set_core_pattern(const char *pattern); -int get_peer_pidfd(int fd); -bool get_pidfd_info(int fd_peer_pidfd, struct pidfd_info *info); - /* Inline helper that uses harness types */ static inline void wait_and_check_coredump_server(pid_t pid_coredump_server, struct __test_metadata *const _metadata, @@ -45,15 +28,4 @@ static inline void wait_and_check_coredump_server(pid_t pid_coredump_server, ASSERT_EQ(WEXITSTATUS(status), 0); } -/* Protocol helper function declarations */ -ssize_t recv_marker(int fd); -bool read_marker(int fd, enum coredump_mark mark); -bool read_coredump_req(int fd, struct coredump_req *req); -bool send_coredump_ack(int fd, const struct coredump_req *req, - __u64 mask, size_t size_ack); -bool check_coredump_req(const struct coredump_req *req, size_t min_size, - __u64 required_mask); -int open_coredump_tmpfile(int fd_tmpfs_detached); -void process_coredump_worker(int fd_coredump, int fd_peer_pidfd, int fd_core_file); - #endif /* __COREDUMP_TEST_H */ diff --git a/tools/testing/selftests/coredump/coredump_test_helpers.c b/tools/testing/selftests/coredump/coredump_test_helpers.c index 2a20faf9cb0a..91607e09127f 100644 --- a/tools/testing/selftests/coredump/coredump_test_helpers.c +++ b/tools/testing/selftests/coredump/coredump_test_helpers.c @@ -1,11 +1,18 @@ // SPDX-License-Identifier: GPL-2.0 #include <assert.h> +#include <elf.h> +#include <endian.h> #include <errno.h> #include <fcntl.h> #include <limits.h> +#include <link.h> +#include <linux/stddef.h> +#include <linux/io_uring.h> +#include <linux/swab.h> #include <linux/coredump.h> #include <linux/fs.h> +#include <poll.h> #include <pthread.h> #include <stdbool.h> #include <stdio.h> @@ -13,31 +20,26 @@ #include <string.h> #include <sys/epoll.h> #include <sys/ioctl.h> +#include <sys/mman.h> #include <sys/socket.h> +#include <sys/stat.h> +#include <sys/syscall.h> #include <sys/types.h> #include <sys/un.h> #include <sys/wait.h> #include <unistd.h> #include "../filesystems/wrappers.h" -#include "../pidfd/pidfd.h" +#include "coredump_notify_signal.h" -/* Forward declarations to avoid including harness header */ -struct __test_metadata; +#include "coredump_test_helpers.h" -/* Match the fixture definition from coredump_test.h */ -struct _fixture_coredump_data { - char original_core_pattern[256]; - pid_t pid_coredump_server; - int fd_tmpfs_detached; -}; - -#ifndef PAGE_SIZE -#define PAGE_SIZE 4096 +#if __ELF_NATIVE_CLASS == 64 +#define COREDUMP_ELFCLASS ELFCLASS64 +#else +#define COREDUMP_ELFCLASS ELFCLASS32 #endif -#define NUM_THREAD_SPAWN 128 - void *do_nothing(void *arg) { (void)arg; @@ -59,6 +61,1193 @@ void crashing_child(void) i = *(volatile int *)NULL; } +void crashing_child_sparse(size_t size) +{ + char *p; + + /* + * Touch the first and the last page. This will cause the whole mapping + * to be dumped because it has been written to. Everything between + * those two pages is a hole though. + */ + p = mmap(NULL, size, PROT_READ | PROT_WRITE, + MAP_PRIVATE | MAP_ANONYMOUS | MAP_NORESERVE, -1, 0); + if (p != MAP_FAILED) { + p[0] = 'x'; + p[size - 1] = 'x'; + } + + /* crash on purpose */ + *(volatile int *)NULL = 0; +} + +/* Select @types through the caller's own /proc/self/coredump_filter. */ +static bool set_coredump_filter(__u64 types) +{ + char buf[32]; + int fd, len; + bool ok; + + fd = open("/proc/self/coredump_filter", O_WRONLY | O_CLOEXEC); + if (fd < 0) + return false; + + len = snprintf(buf, sizeof(buf), "0x%llx", (unsigned long long)types); + ok = write_nointr(fd, buf, len) == len; + close(fd); + return ok; +} + +/* + * Map shared anonymous memory, touch it, tell the server where it is and + * crash. A @task_filter other than FILTER_TASK_INHERIT is selected first. + */ +void crashing_child_memory(__u64 task_filter, int fd_addr) +{ + char *p; + + if (task_filter != FILTER_TASK_INHERIT && !set_coredump_filter(task_filter)) + _exit(EXIT_FAILURE); + + p = mmap(NULL, MEMORY_MAPPING_SIZE, PROT_READ | PROT_WRITE, + MAP_SHARED | MAP_ANONYMOUS, -1, 0); + if (p == MAP_FAILED) + _exit(EXIT_FAILURE); + p[0] = 'x'; + + if (write_nointr(fd_addr, &p, sizeof(p)) != sizeof(p)) + _exit(EXIT_FAILURE); + close(fd_addr); + + /* crash on purpose */ + *(volatile int *)NULL = 0; +} + +/* Sink a reassembled record stream is handed to, record by record. */ +struct coredump_record_sink { + /* @len bytes of coredump data that belong at @offset. */ + int (*data)(void *ctx, const void *buf, size_t len, __u64 offset); + /* @len zero bytes that belong at @offset. */ + int (*zero)(void *ctx, __u64 offset, __u64 len); + void *ctx; +}; + +/* Read @len bytes off the socket and hand them to @sink, if there is one. */ +static ssize_t recv_record_bytes(int fd_coredump, __u64 len, + const struct coredump_record_sink *sink, + __u64 offset) +{ + ssize_t received = 0; + + while (len) { + char buffer[PAGE_SIZE]; + size_t chunk = len < sizeof(buffer) ? len : sizeof(buffer); + ssize_t ret; + + ret = recv(fd_coredump, buffer, chunk, MSG_WAITALL); + if (ret <= 0) { + fprintf(stderr, "%s: short read %zd: %m\n", + __func__, ret); + return -1; + } + + if (sink && sink->data(sink->ctx, buffer, ret, offset + received)) + return -1; + + received += ret; + len -= ret; + } + + return received; +} + +/* Put the data where the records say it goes and leave the holes alone. */ +static int file_sink_data(void *ctx, const void *buf, size_t len, __u64 offset) +{ + int fd = *(int *)ctx; + + if (pwrite(fd, buf, len, offset) != (ssize_t)len) { + fprintf(stderr, "%s: pwrite failed: %m\n", __func__); + return -1; + } + + return 0; +} + +static int file_sink_zero(void *ctx, __u64 offset, __u64 len) +{ + /* Nothing has to be written for a hole. */ + return 0; +} + +/* + * Read a coredump strea and funnel it into @sink. Allow to pass in a + * @fd_peer_pidfd to simulate coredump truncation by killing it after having + * received a coredump record. + */ +static ssize_t __recv_coredump_records(int fd_coredump, + const struct coredump_record_sink *sink, + off_t *coredump_size, bool *truncated, + int fd_peer_pidfd) +{ + ssize_t received = 0; + off_t size = 0; + bool is_truncated = false; + bool ended = false; + char trailing; + + while (!ended) { + struct coredump_record_header record = {}; + size_t known_size; + ssize_t ret; + + /* Peek the header size the way read_coredump_req() does. */ + ret = recv(fd_coredump, &record, sizeof(record.size), + MSG_PEEK | MSG_WAITALL); + if (ret == 0) { + /* Nothing closed the stream, so the coredump was cut short. */ + if (truncated) { + is_truncated = true; + break; + } + fprintf(stderr, "%s: stream ended without an end record\n", + __func__); + return -1; + } + if (ret != sizeof(record.size)) { + fprintf(stderr, "%s: short record peek %zd: %m\n", + __func__, ret); + return -1; + } + + if (record.size < COREDUMP_RECORD_HEADER_SIZE_VER0) { + fprintf(stderr, "%s: header size %u below minimum %u\n", + __func__, record.size, + COREDUMP_RECORD_HEADER_SIZE_VER0); + return -1; + } + + /* Consume as much of the header as we know about. */ + known_size = record.size < sizeof(record) ? record.size : sizeof(record); + ret = recv(fd_coredump, &record, known_size, MSG_WAITALL); + if (ret != (ssize_t)known_size) { + fprintf(stderr, "%s: short record read %zd: %m\n", + __func__, ret); + return -1; + } + received += ret; + + /* + * A flag changes what the record means, so refuse one we + * don't know rather than guess. + */ + if (record.flags) { + fprintf(stderr, "%s: unknown header flags 0x%llx\n", + __func__, (unsigned long long)record.flags); + return -1; + } + + /* Discard any part of the header we have no use for. */ + ret = recv_record_bytes(fd_coredump, record.size - known_size, + NULL, 0); + if (ret < 0) + return -1; + received += ret; + + /* Records are sent in order and they don't leave gaps. */ + if (record.offset != (__u64)size) { + fprintf(stderr, "%s: record at %llu, expected %llu\n", + __func__, (unsigned long long)record.offset, + (unsigned long long)size); + return -1; + } + + switch (record.type) { + case COREDUMP_RECORD_ZERO: + /* A hole. It comes with no data and needs none. */ + if (sink->zero(sink->ctx, record.offset, record.len)) + return -1; + break; + case COREDUMP_RECORD_DATA: + ret = recv_record_bytes(fd_coredump, record.len, sink, + record.offset); + if (ret < 0) + return -1; + received += ret; + if (fd_peer_pidfd >= 0) { + if (sys_pidfd_send_signal(fd_peer_pidfd, SIGKILL, + NULL, 0)) { + fprintf(stderr, "%s: kill failed: %m\n", + __func__); + return -1; + } + fd_peer_pidfd = -1; + } + break; + case COREDUMP_RECORD_END: + /* The coredump ends here and nothing follows it. */ + if (record.len) { + fprintf(stderr, "%s: end record covers %llu bytes\n", + __func__, + (unsigned long long)record.len); + return -1; + } + ended = true; + break; + default: + fprintf(stderr, "%s: unknown record type %u\n", + __func__, record.type); + return -1; + } + + size += record.len; + } + + /* The end record is the last thing on the wire. */ + if (recv(fd_coredump, &trailing, sizeof(trailing), MSG_DONTWAIT) > 0) { + fprintf(stderr, "%s: data after the end record\n", __func__); + return -1; + } + + if (truncated) + *truncated = is_truncated; + + *coredump_size = size; + + fprintf(stderr, "Received %zd bytes for a %s coredump of %llu bytes\n", + received, is_truncated ? "truncated" : "complete", + (unsigned long long)size); + return received; +} + +/* Reassemble a record stream into the coredump it describes. */ +ssize_t recv_coredump_records(int fd_coredump, int fd_core_file, + off_t *coredump_size, bool *truncated, + int fd_peer_pidfd) +{ + struct coredump_record_sink sink = { + .data = file_sink_data, + .zero = file_sink_zero, + .ctx = &fd_core_file, + }; + ssize_t received; + off_t size = 0; + + received = __recv_coredump_records(fd_coredump, &sink, &size, truncated, + fd_peer_pidfd); + if (received < 0) + return -1; + + /* + * Nothing is written for a hole, so grow the file to the size the + * records describe in case the coredump ended in one. + */ + if (ftruncate(fd_core_file, size) < 0) { + fprintf(stderr, "%s: ftruncate to %llu failed: %m\n", + __func__, (unsigned long long)size); + return -1; + } + + if (coredump_size) + *coredump_size = size; + + return received; +} + +/* The ELF header of a native core file. */ +static bool is_core_ehdr(const ElfW(Ehdr) *ehdr) +{ + return !memcmp(ehdr->e_ident, ELFMAG, SELFMAG) && + ehdr->e_ident[EI_CLASS] == COREDUMP_ELFCLASS && + ehdr->e_type == ET_CORE; +} + +/* Whatever the server ends up with has to be an ELF core file. */ +bool is_elf_core(int fd) +{ + ElfW(Ehdr) ehdr; + + if (pread(fd, &ehdr, sizeof(ehdr), 0) != sizeof(ehdr)) { + fprintf(stderr, "%s: short read: %m\n", __func__); + return false; + } + + if (!is_core_ehdr(&ehdr)) { + fprintf(stderr, "%s: not an ELF core file\n", __func__); + return false; + } + + return true; +} + +/* + * A coredump server that uploads to a blob store can't upload a sparse + * file and can't seek in the object it is uploading. It streams the data + * records into the object as they arrive, remembers the holes it left + * out, and uploads the program header table that describes the result + * last. What comes out is an ordinary ELF core file without the holes. + */ + +/* A run of the coredump the object doesn't carry. */ +struct compact_hole { + __u64 offset; + __u64 len; +}; + +/* A program header of the object and where its bytes sat in the coredump. */ +struct compact_piece { + ElfW(Phdr) phdr; + __u64 src; +}; + +struct compact_ctx { + int fd_body; /* the object's payload, append only */ + int fd_reference; /* the coredump with its holes, for the test */ + unsigned char *head; /* everything ahead of the segment data */ + size_t head_len; + size_t head_cap; + __u64 data_offset; /* where the segment data starts, 0 while unknown */ + struct compact_hole *holes; + size_t nr_holes; + size_t holes_cap; + __u64 body_len; +}; + +/* Write @len bytes out, short writes and all. */ +static int compact_write(int fd, const void *buf, size_t len) +{ + const unsigned char *pos = buf; + + while (len) { + ssize_t ret = write(fd, pos, len); + + if (ret <= 0) { + fprintf(stderr, "%s: write failed: %m\n", __func__); + return -1; + } + + pos += ret; + len -= ret; + } + + return 0; +} + +/* Keep @len bytes of the head, or @len zeroes if @buf is NULL. */ +static int compact_head_append(struct compact_ctx *ctx, const void *buf, + size_t len) +{ + if (ctx->head_len + len > ctx->head_cap) { + size_t cap = ctx->head_cap ? ctx->head_cap : PAGE_SIZE; + unsigned char *head; + + while (cap < ctx->head_len + len) + cap *= 2; + + head = realloc(ctx->head, cap); + if (!head) { + fprintf(stderr, "%s: out of memory\n", __func__); + return -1; + } + ctx->head = head; + ctx->head_cap = cap; + } + + if (buf) + memcpy(ctx->head + ctx->head_len, buf, len); + else + memset(ctx->head + ctx->head_len, 0, len); + ctx->head_len += len; + + return 0; +} + +/* Remember a hole so the program header table can account for it later. */ +static int compact_keep_hole(struct compact_ctx *ctx, __u64 offset, __u64 len) +{ + if (ctx->nr_holes == ctx->holes_cap) { + size_t cap = ctx->holes_cap ? ctx->holes_cap * 2 : 64; + struct compact_hole *holes; + + holes = realloc(ctx->holes, cap * sizeof(*holes)); + if (!holes) { + fprintf(stderr, "%s: out of memory\n", __func__); + return -1; + } + ctx->holes = holes; + ctx->holes_cap = cap; + } + + ctx->holes[ctx->nr_holes].offset = offset; + ctx->holes[ctx->nr_holes].len = len; + ctx->nr_holes++; + + return 0; +} + +/* The segment data starts where the first PT_LOAD points. */ +static int compact_probe(struct compact_ctx *ctx) +{ + const ElfW(Ehdr) *ehdr = (const ElfW(Ehdr) *)ctx->head; + const ElfW(Phdr) *phdr; + size_t i; + + if (ctx->data_offset || ctx->head_len < sizeof(*ehdr)) + return 0; + + if (!is_core_ehdr(ehdr)) { + fprintf(stderr, "%s: not an ELF core file\n", __func__); + return -1; + } + + if (ehdr->e_phoff != sizeof(*ehdr) || + ehdr->e_phentsize != sizeof(ElfW(Phdr)) || + ehdr->e_phnum == 0 || ehdr->e_phnum == PN_XNUM) { + fprintf(stderr, "%s: unhandled program header table\n", __func__); + return -1; + } + + if (ctx->head_len < ehdr->e_phoff + + (size_t)ehdr->e_phnum * ehdr->e_phentsize) + return 0; + + phdr = (const ElfW(Phdr) *)(ctx->head + ehdr->e_phoff); + for (i = 0; i < ehdr->e_phnum; i++) { + if (phdr[i].p_type != PT_LOAD) + continue; + if (!ctx->data_offset || phdr[i].p_offset < ctx->data_offset) + ctx->data_offset = phdr[i].p_offset; + } + + if (!ctx->data_offset) { + fprintf(stderr, "%s: coredump without a single segment\n", + __func__); + return -1; + } + + return 0; +} + +/* + * Take whatever of [@offset, @offset + @len) still belongs to the head. + * @buf is NULL for a hole. Returns how much was taken. + */ +static ssize_t compact_head_take(struct compact_ctx *ctx, const void *buf, + __u64 offset, __u64 len) +{ + __u64 chunk; + + if (!len || (ctx->data_offset && offset >= ctx->data_offset)) + return 0; + + chunk = len; + if (ctx->data_offset && offset + chunk > ctx->data_offset) + chunk = ctx->data_offset - offset; + + if (offset != ctx->head_len) { + fprintf(stderr, "%s: head has a gap at %llu\n", __func__, + (unsigned long long)offset); + return -1; + } + + if (compact_head_append(ctx, buf, chunk)) + return -1; + + return chunk; +} + +static int compact_data(void *arg, const void *buf, size_t len, __u64 offset) +{ + struct compact_ctx *ctx = arg; + const unsigned char *pos = buf; + ssize_t head; + + /* Only the test needs a coredump with the holes still in it. */ + if (pwrite(ctx->fd_reference, pos, len, offset) != (ssize_t)len) { + fprintf(stderr, "%s: pwrite failed: %m\n", __func__); + return -1; + } + + /* The head has to be rewritten at the end, so hold on to it. */ + head = compact_head_take(ctx, pos, offset, len); + if (head < 0) + return -1; + if (head && compact_probe(ctx)) + return -1; + + pos += head; + len -= head; + if (!len) + return 0; + + /* Everything else goes into the object as it arrives. */ + if (compact_write(ctx->fd_body, pos, len)) + return -1; + ctx->body_len += len; + + return 0; +} + +static int compact_zero(void *arg, __u64 offset, __u64 len) +{ + struct compact_ctx *ctx = arg; + ssize_t head; + + /* A hole in the head is alignment padding. Write it out. */ + head = compact_head_take(ctx, NULL, offset, len); + if (head < 0) + return -1; + + offset += head; + len -= head; + if (!len) + return 0; + + /* This is what the object doesn't have to carry. */ + return compact_keep_hole(ctx, offset, len); +} + +/* Where @offset ends up in the object once the holes ahead of it are gone. */ +static __u64 compact_offset(const struct compact_ctx *ctx, __u64 body_start, + __u64 offset) +{ + __u64 elided = 0; + size_t i; + + for (i = 0; i < ctx->nr_holes; i++) { + __u64 len = ctx->holes[i].len; + + if (ctx->holes[i].offset >= offset) + break; + if (ctx->holes[i].offset + len > offset) + len = offset - ctx->holes[i].offset; + elided += len; + } + + return body_start + (offset - ctx->data_offset) - elided; +} + +/* A run of segment data that made it into the object. */ +static void compact_add_data(struct compact_piece *pieces, size_t *nr, + const ElfW(Phdr) *phdr, __u64 start, __u64 end) +{ + struct compact_piece *piece = &pieces[(*nr)++]; + + piece->phdr = *phdr; + piece->phdr.p_vaddr = phdr->p_vaddr + (start - phdr->p_offset); + piece->phdr.p_paddr = 0; + piece->phdr.p_filesz = end - start; + piece->phdr.p_memsz = end - start; + piece->src = start; +} + +/* + * A run of @len bytes the object doesn't carry. It grows the piece in + * front of it if this segment already has one, because everything a + * segment covers past p_filesz is zeroes anyway. + */ +static void compact_add_zero(struct compact_piece *pieces, size_t *nr, + size_t first, const ElfW(Phdr) *phdr, __u64 vaddr, + __u64 len) +{ + struct compact_piece *piece; + + if (*nr > first) { + pieces[*nr - 1].phdr.p_memsz += len; + return; + } + + piece = &pieces[(*nr)++]; + piece->phdr = *phdr; + piece->phdr.p_vaddr = vaddr; + piece->phdr.p_paddr = 0; + piece->phdr.p_filesz = 0; + piece->phdr.p_memsz = len; + piece->src = 0; +} + +/* Split the segments at the holes and write out what the object became. */ +static int compact_build(struct compact_ctx *ctx, int fd_object) +{ + __u64 note_offset = 0, note_len = 0, note_new; + __u64 align = 0, head_len, body_start, pos; + size_t nr_old, nr_new = 0, note_piece = 0, i; + struct compact_piece *pieces; + char buffer[PAGE_SIZE]; + const ElfW(Phdr) *old; + ElfW(Ehdr) ehdr; + int ret = -1; + + if (!ctx->data_offset) { + fprintf(stderr, "%s: coredump without segment data\n", __func__); + return -1; + } + + memcpy(&ehdr, ctx->head, sizeof(ehdr)); + if (ehdr.e_shoff) { + fprintf(stderr, "%s: section headers are not handled\n", + __func__); + return -1; + } + + old = (const ElfW(Phdr) *)(ctx->head + ehdr.e_phoff); + nr_old = ehdr.e_phnum; + + pieces = calloc(nr_old + 2 * ctx->nr_holes + 1, sizeof(*pieces)); + if (!pieces) { + fprintf(stderr, "%s: out of memory\n", __func__); + return -1; + } + + for (i = 0; i < nr_old; i++) { + ElfW(Phdr) phdr = old[i]; + __u64 end = phdr.p_offset + phdr.p_filesz; + __u64 cur = phdr.p_offset; + size_t first = nr_new, h; + + /* The notes move because the table in front of them grows. */ + if (phdr.p_type == PT_NOTE) { + if (note_len) { + fprintf(stderr, "%s: more than one note segment\n", + __func__); + goto out; + } + note_offset = phdr.p_offset; + note_len = phdr.p_filesz; + note_piece = nr_new; + pieces[nr_new].phdr = phdr; + pieces[nr_new++].src = 0; + continue; + } + + if (phdr.p_type != PT_LOAD) { + if (phdr.p_filesz && phdr.p_offset < ctx->data_offset) { + fprintf(stderr, "%s: segment %zu is in the head\n", + __func__, i); + goto out; + } + pieces[nr_new].phdr = phdr; + pieces[nr_new++].src = phdr.p_offset; + continue; + } + + if (!align) + align = phdr.p_align; + + for (h = 0; h < ctx->nr_holes && cur < end; h++) { + __u64 start = ctx->holes[h].offset; + __u64 stop = start + ctx->holes[h].len; + + if (stop <= cur) + continue; + if (start >= end) + break; + + /* A hole can span more than this one segment. */ + if (start < cur) + start = cur; + if (stop > end) + stop = end; + + if (start > cur) { + compact_add_data(pieces, &nr_new, &phdr, cur, + start); + cur = start; + } + compact_add_zero(pieces, &nr_new, first, &phdr, + phdr.p_vaddr + (cur - phdr.p_offset), + stop - cur); + cur = stop; + } + + if (cur < end) + compact_add_data(pieces, &nr_new, &phdr, cur, end); + + /* Whatever the kernel didn't dump of this mapping. */ + if (phdr.p_memsz > phdr.p_filesz) + compact_add_zero(pieces, &nr_new, first, &phdr, + phdr.p_vaddr + phdr.p_filesz, + phdr.p_memsz - phdr.p_filesz); + } + + if (!note_len || note_offset + note_len > ctx->head_len) { + fprintf(stderr, "%s: notes aren't where they should be\n", + __func__); + goto out; + } + + if (nr_new >= PN_XNUM) { + fprintf(stderr, "%s: %zu program headers don't fit\n", __func__, + nr_new); + goto out; + } + + if (!align || (align & (align - 1))) + align = sysconf(_SC_PAGESIZE); + + note_new = sizeof(ehdr) + (__u64)nr_new * sizeof(ElfW(Phdr)); + head_len = note_new + note_len; + body_start = (head_len + align - 1) & ~(align - 1); + + for (i = 0; i < nr_new; i++) { + struct compact_piece *piece = &pieces[i]; + + if (i == note_piece) + piece->phdr.p_offset = note_new; + else if (piece->phdr.p_filesz) + piece->phdr.p_offset = compact_offset(ctx, body_start, + piece->src); + else + piece->phdr.p_offset = 0; + } + + /* Only now is the head known. That's why it is uploaded last. */ + ehdr.e_phnum = nr_new; + if (compact_write(fd_object, &ehdr, sizeof(ehdr))) + goto out; + + for (i = 0; i < nr_new; i++) + if (compact_write(fd_object, &pieces[i].phdr, + sizeof(pieces[i].phdr))) + goto out; + + if (compact_write(fd_object, ctx->head + note_offset, note_len)) + goto out; + + /* Keep the segments aligned the way a debugger expects them. */ + memset(buffer, 0, sizeof(buffer)); + for (pos = head_len; pos < body_start; ) { + __u64 chunk = body_start - pos; + + if (chunk > sizeof(buffer)) + chunk = sizeof(buffer); + if (compact_write(fd_object, buffer, chunk)) + goto out; + pos += chunk; + } + + /* Putting the parts together is the blob store's job. Do it here. */ + for (pos = 0; pos < ctx->body_len; ) { + ssize_t chunk = pread(ctx->fd_body, buffer, sizeof(buffer), pos); + + if (chunk <= 0) { + fprintf(stderr, "%s: short read %zd: %m\n", __func__, + chunk); + goto out; + } + if (compact_write(fd_object, buffer, chunk)) + goto out; + pos += chunk; + } + + fprintf(stderr, "Object is %llu bytes in %zu program headers, %zu holes left out\n", + (unsigned long long)(body_start + ctx->body_len), nr_new, + ctx->nr_holes); + ret = 0; +out: + free(pieces); + return ret; +} + +/* + * Reassemble a record stream into an ELF core file that has no holes in + * it, the way a coredump server that uploads to a blob store has to. If + * @fd_reference is valid it gets the coredump the records describe, + * holes and all, so the test can compare the two. + */ +ssize_t recv_coredump_compact(int fd_coredump, int fd_object, int fd_reference, + off_t *coredump_size) +{ + struct compact_ctx ctx = { + .fd_body = -1, + .fd_reference = fd_reference, + }; + struct coredump_record_sink sink = { + .data = compact_data, + .zero = compact_zero, + .ctx = &ctx, + }; + ssize_t received; + off_t size = 0; + FILE *body; + + body = tmpfile(); + if (!body) { + fprintf(stderr, "%s: tmpfile failed: %m\n", __func__); + return -1; + } + ctx.fd_body = fileno(body); + + /* An upload is appended to. Make sure nothing here can seek. */ + if (fcntl(ctx.fd_body, F_SETFL, O_APPEND)) { + fprintf(stderr, "%s: F_SETFL failed: %m\n", __func__); + received = -1; + goto out; + } + + received = __recv_coredump_records(fd_coredump, &sink, &size, NULL, -1); + if (received < 0) + goto out; + + /* + * Nothing is written for a hole, so grow the reference to the size + * the records describe in case the coredump ended in one. + */ + if (ftruncate(fd_reference, size) < 0) { + fprintf(stderr, "%s: ftruncate to %llu failed: %m\n", + __func__, (unsigned long long)size); + received = -1; + goto out; + } + + if (compact_build(&ctx, fd_object)) { + received = -1; + goto out; + } + + if (coredump_size) + *coredump_size = size; +out: + fclose(body); + free(ctx.head); + free(ctx.holes); + return received; +} + +/* Read the ELF header and the program header table of @fd. */ +static ElfW(Phdr) *read_phdrs(int fd, size_t *nr) +{ + ElfW(Ehdr) ehdr; + ElfW(Phdr) *phdr; + size_t size; + + if (pread(fd, &ehdr, sizeof(ehdr), 0) != sizeof(ehdr)) { + fprintf(stderr, "%s: no ELF header: %m\n", __func__); + return NULL; + } + + if (!is_core_ehdr(&ehdr) || !ehdr.e_phnum || + ehdr.e_phentsize != sizeof(*phdr)) { + fprintf(stderr, "%s: not an ELF core file\n", __func__); + return NULL; + } + + size = (size_t)ehdr.e_phnum * ehdr.e_phentsize; + phdr = malloc(size); + if (!phdr) { + fprintf(stderr, "%s: out of memory\n", __func__); + return NULL; + } + + if (pread(fd, phdr, size, ehdr.e_phoff) != (ssize_t)size) { + fprintf(stderr, "%s: short program header table: %m\n", __func__); + free(phdr); + return NULL; + } + + *nr = ehdr.e_phnum; + return phdr; +} + +/* The segment @vaddr falls into. */ +static const ElfW(Phdr) *find_segment(const ElfW(Phdr) *phdr, size_t nr, + __u64 vaddr) +{ + size_t i; + + for (i = 0; i < nr; i++) { + if (phdr[i].p_type != PT_LOAD) + continue; + if (vaddr >= phdr[i].p_vaddr && + vaddr < phdr[i].p_vaddr + phdr[i].p_memsz) + return &phdr[i]; + } + + return NULL; +} + +/* The PT_LOAD segment @vaddr falls into. */ +bool find_coredump_segment(int fd, __u64 vaddr, ElfW(Phdr) *segment) +{ + const ElfW(Phdr) *found; + ElfW(Phdr) *phdr; + size_t nr; + + phdr = read_phdrs(fd, &nr); + if (!phdr) + return false; + + found = find_segment(phdr, nr, vaddr); + if (found) + *segment = *found; + else + fprintf(stderr, "%s: no segment for 0x%llx\n", __func__, + (unsigned long long)vaddr); + + free(phdr); + return found; +} + +/* How many bytes the PT_LOAD and the PT_NOTE segments of @fd carry. */ +bool sum_coredump_segments(int fd, __u64 *data, __u64 *notes) +{ + ElfW(Phdr) *phdr; + size_t nr, i; + + phdr = read_phdrs(fd, &nr); + if (!phdr) + return false; + + *data = 0; + *notes = 0; + for (i = 0; i < nr; i++) { + if (phdr[i].p_type == PT_LOAD) + *data += phdr[i].p_filesz; + else if (phdr[i].p_type == PT_NOTE) + *notes += phdr[i].p_filesz; + } + + free(phdr); + return true; +} + +/* The coredump in @fd is at least as long as every segment it declares. */ +bool check_coredump_extent(int fd) +{ + ElfW(Phdr) *phdr; + struct stat st; + size_t nr, i; + bool ok = true; + + if (fstat(fd, &st)) { + fprintf(stderr, "%s: fstat: %m\n", __func__); + return false; + } + + phdr = read_phdrs(fd, &nr); + if (!phdr) + return false; + + for (i = 0; i < nr; i++) { + if (phdr[i].p_offset + phdr[i].p_filesz <= (__u64)st.st_size) + continue; + fprintf(stderr, "%s: segment %zu ends at %llu, the coredump at %llu\n", + __func__, i, + (unsigned long long)(phdr[i].p_offset + phdr[i].p_filesz), + (unsigned long long)st.st_size); + ok = false; + } + + free(phdr); + return ok; +} + +/* The next stretch of memory the segments cover, split ones merged back. */ +static bool next_range(const ElfW(Phdr) *phdr, size_t nr, size_t *i, + __u64 *start, __u64 *end) +{ + while (*i < nr && phdr[*i].p_type != PT_LOAD) + (*i)++; + + if (*i >= nr) + return false; + + *start = phdr[*i].p_vaddr; + *end = phdr[*i].p_vaddr + phdr[*i].p_memsz; + (*i)++; + + while (*i < nr) { + if (phdr[*i].p_type != PT_LOAD) { + (*i)++; + continue; + } + if (phdr[*i].p_vaddr != *end) + break; + *end = phdr[*i].p_vaddr + phdr[*i].p_memsz; + (*i)++; + } + + return true; +} + +/* Compare @len bytes at @offset against @len bytes at @offset_ref. */ +static int compare_range(int fd, __u64 offset, int fd_ref, __u64 offset_ref, + __u64 len) +{ + char buffer[PAGE_SIZE], buffer_ref[PAGE_SIZE]; + + while (len) { + size_t chunk = len < sizeof(buffer) ? len : sizeof(buffer); + + if (pread(fd, buffer, chunk, offset) != (ssize_t)chunk || + pread(fd_ref, buffer_ref, chunk, offset_ref) != (ssize_t)chunk) { + fprintf(stderr, "%s: short read at %llu: %m\n", + __func__, (unsigned long long)offset); + return -1; + } + + if (memcmp(buffer, buffer_ref, chunk)) { + fprintf(stderr, "%s: %llu differs from %llu\n", __func__, + (unsigned long long)offset, + (unsigned long long)offset_ref); + return -1; + } + + offset += chunk; + offset_ref += chunk; + len -= chunk; + } + + return 0; +} + +/* The @len bytes at @offset the object left out have to have been zeroes. */ +static int check_zero_range(int fd, __u64 offset, __u64 len) +{ + static const char zeroes[PAGE_SIZE]; + char buffer[PAGE_SIZE]; + + while (len) { + size_t chunk = len < sizeof(buffer) ? len : sizeof(buffer); + + if (pread(fd, buffer, chunk, offset) != (ssize_t)chunk) { + fprintf(stderr, "%s: short read at %llu: %m\n", + __func__, (unsigned long long)offset); + return -1; + } + + if (memcmp(buffer, zeroes, chunk)) { + fprintf(stderr, "%s: %llu isn't a hole\n", __func__, + (unsigned long long)offset); + return -1; + } + + offset += chunk; + len -= chunk; + } + + return 0; +} + +/* + * The object has to describe the same memory as the coredump it was built + * from, and it has to describe it correctly. + */ +int check_compact_coredump(int fd_object, int fd_reference) +{ + ElfW(Phdr) *object = NULL, *reference = NULL; + size_t nr_object, nr_reference, i; + size_t io = 0, ir = 0; + int ret = -1; + + object = read_phdrs(fd_object, &nr_object); + reference = read_phdrs(fd_reference, &nr_reference); + if (!object || !reference) + goto out; + + /* Nothing may have been dropped and nothing may have been added. */ + for (;;) { + __u64 start = 0, end = 0, start_ref = 0, end_ref = 0; + bool has, has_ref; + + has = next_range(object, nr_object, &io, &start, &end); + has_ref = next_range(reference, nr_reference, &ir, &start_ref, + &end_ref); + if (!has && !has_ref) + break; + + if (has != has_ref || start != start_ref || end != end_ref) { + fprintf(stderr, "%s: object covers 0x%llx-0x%llx, coredump 0x%llx-0x%llx\n", + __func__, (unsigned long long)start, + (unsigned long long)end, + (unsigned long long)start_ref, + (unsigned long long)end_ref); + goto out; + } + } + + for (i = 0; i < nr_object; i++) { + const ElfW(Phdr) *segment; + __u64 offset, dumped; + + if (object[i].p_type != PT_LOAD || !object[i].p_memsz) + continue; + + segment = find_segment(reference, nr_reference, + object[i].p_vaddr); + if (!segment) { + fprintf(stderr, "%s: 0x%llx isn't in the coredump\n", + __func__, + (unsigned long long)object[i].p_vaddr); + goto out; + } + + offset = object[i].p_vaddr - segment->p_vaddr; + dumped = offset < segment->p_filesz ? + segment->p_filesz - offset : 0; + + /* What the object carries is what the coredump had. */ + if (object[i].p_filesz > dumped) { + fprintf(stderr, "%s: object carries %llu bytes the coredump doesn't have\n", + __func__, + (unsigned long long)(object[i].p_filesz - dumped)); + goto out; + } + + if (compare_range(fd_object, object[i].p_offset, fd_reference, + segment->p_offset + offset, + object[i].p_filesz)) + goto out; + + /* And what it left out was a hole. */ + if (object[i].p_memsz > object[i].p_filesz && + dumped > object[i].p_filesz) { + __u64 left_out = dumped - object[i].p_filesz; + + if (left_out > object[i].p_memsz - object[i].p_filesz) + left_out = object[i].p_memsz - object[i].p_filesz; + + if (check_zero_range(fd_reference, + segment->p_offset + offset + + object[i].p_filesz, left_out)) + goto out; + } + } + + ret = 0; +out: + free(object); + free(reference); + return ret; +} + +/* Read a plain coredump byte stream to end-of-file. */ +ssize_t recv_coredump_bytes(int fd_coredump, int fd_core_file) +{ + ssize_t received = 0; + + for (;;) { + char buffer[PAGE_SIZE]; + ssize_t ret = read_nointr(fd_coredump, buffer, sizeof(buffer)); + + if (ret < 0) { + fprintf(stderr, "%s: read failed: %m\n", __func__); + return -1; + } + if (ret == 0) + break; + + if (write_nointr(fd_core_file, buffer, ret) != ret) { + fprintf(stderr, "%s: write failed: %m\n", __func__); + return -1; + } + received += ret; + } + + fprintf(stderr, "Received %zd bytes of coredump\n", received); + return received; +} + int create_detached_tmpfs(void) { int fd_context, fd_tmpfs; @@ -101,6 +1290,7 @@ int create_and_listen_unix_socket(const char *path) return fd; out: + fprintf(stderr, "%s: %s: %m\n", __func__, path); if (fd >= 0) close(fd); return -1; @@ -153,8 +1343,95 @@ bool get_pidfd_info(int fd_peer_pidfd, struct pidfd_info *info) return true; } +/* + * How much the peer has mapped. The task is parked in the coredump + * handshake, so its mm is still there to be looked at. + */ +ssize_t peer_vm_size(int fd_peer_pidfd) +{ + struct pidfd_info info = {}; + unsigned long pages; + char path[64]; + FILE *f; + + if (!get_pidfd_info(fd_peer_pidfd, &info)) + return -1; + + snprintf(path, sizeof(path), "/proc/%d/statm", info.pid); + f = fopen(path, "r"); + if (!f) { + fprintf(stderr, "%s: %s: %m\n", __func__, path); + return -1; + } + + if (fscanf(f, "%lu", &pages) != 1) { + fprintf(stderr, "%s: %s: no size\n", __func__, path); + fclose(f); + return -1; + } + fclose(f); + + return (ssize_t)pages * sysconf(_SC_PAGESIZE); +} + /* Protocol helper functions */ +/* The peer's /proc/<pid>/coredump_filter, which is in memory types. */ +bool peer_coredump_filter(int fd_peer_pidfd, __u64 *memory_types) +{ + struct pidfd_info info = {}; + unsigned long value; + char path[64]; + FILE *f; + int ret; + + if (!get_pidfd_info(fd_peer_pidfd, &info)) + return false; + + snprintf(path, sizeof(path), "/proc/%d/coredump_filter", info.pid); + f = fopen(path, "r"); + if (!f) { + fprintf(stderr, "%s: %s: %m\n", __func__, path); + return false; + } + + ret = fscanf(f, "%lx", &value); + fclose(f); + if (ret != 1) { + fprintf(stderr, "%s: %s: no value\n", __func__, path); + return false; + } + + *memory_types = value; + return true; +} + +/* Read @len bytes at @addr from the peer's /proc/<pid>/mem. */ +ssize_t peer_read_mem(int fd_peer_pidfd, __u64 addr, void *buf, size_t len) +{ + struct pidfd_info info = {}; + char path[64]; + ssize_t ret; + int fd; + + if (!get_pidfd_info(fd_peer_pidfd, &info)) + return -1; + + snprintf(path, sizeof(path), "/proc/%d/mem", info.pid); + fd = open(path, O_RDONLY | O_CLOEXEC); + if (fd < 0) { + fprintf(stderr, "%s: %s: %m\n", __func__, path); + return -1; + } + + ret = pread(fd, buf, len, addr); + if (ret < 0) + fprintf(stderr, "%s: %s at 0x%llx: %m\n", __func__, path, + (unsigned long long)addr); + close(fd); + return ret; +} + ssize_t recv_marker(int fd) { enum coredump_mark mark = COREDUMP_MARK_REQACK; @@ -197,10 +1474,34 @@ bool read_marker(int fd, enum coredump_mark mark) return ret == mark; } -bool read_coredump_req(int fd, struct coredump_req *req) +/* + * The kernel hung up without sending anything more: end of stream, or a + * reset if it refused the ack on its peeked size and never read it. + */ +bool read_hangup(int fd) { ssize_t ret; - size_t field_size, user_size, ack_size, kernel_size, remaining_size; + char c; + + ret = recv(fd, &c, sizeof(c), MSG_WAITALL); + if (ret == 0) { + fprintf(stderr, "Kernel closed the connection\n"); + return true; + } + if (ret < 0 && errno == ECONNRESET) { + fprintf(stderr, "Kernel closed the connection with the ack unread\n"); + return true; + } + + fprintf(stderr, "%s: expected a hangup, got %zd: %m\n", __func__, ret); + return false; +} + +/* Read the request as a server built with a @user_size byte struct does. */ +bool read_coredump_req_sized(int fd, struct coredump_req *req, size_t user_size) +{ + ssize_t ret; + size_t field_size, known_size, kernel_size, remaining_size; memset(req, 0, sizeof(*req)); field_size = sizeof(req->size); @@ -208,37 +1509,36 @@ bool read_coredump_req(int fd, struct coredump_req *req) /* Peek the size of the coredump request. */ ret = recv(fd, req, field_size, MSG_PEEK | MSG_WAITALL); if (ret != field_size) { - fprintf(stderr, "read_coredump_req: peek failed (got %zd, expected %zu): %m\n", + fprintf(stderr, "%s: peek failed (got %zd, expected %zu): %m\n", __func__, ret, field_size); return false; } kernel_size = req->size; - if (kernel_size < COREDUMP_ACK_SIZE_VER0) { - fprintf(stderr, "read_coredump_req: kernel_size %zu < min %d\n", - kernel_size, COREDUMP_ACK_SIZE_VER0); + if (kernel_size < COREDUMP_REQ_SIZE_VER0) { + fprintf(stderr, "%s: kernel_size %zu < min %d\n", __func__, + kernel_size, COREDUMP_REQ_SIZE_VER0); return false; } if (kernel_size >= PAGE_SIZE) { - fprintf(stderr, "read_coredump_req: kernel_size %zu >= PAGE_SIZE %d\n", + fprintf(stderr, "%s: kernel_size %zu >= PAGE_SIZE %d\n", __func__, kernel_size, PAGE_SIZE); return false; } - /* Use the minimum of user and kernel size to read the full request. */ - user_size = sizeof(struct coredump_req); - ack_size = user_size < kernel_size ? user_size : kernel_size; - ret = recv(fd, req, ack_size, MSG_WAITALL); - if (ret != ack_size) + /* Consume as much of the request as we know about. */ + known_size = user_size < kernel_size ? user_size : kernel_size; + ret = recv(fd, req, known_size, MSG_WAITALL); + if (ret != known_size) return false; fprintf(stderr, "Read coredump request with size %u and mask 0x%llx\n", req->size, (unsigned long long)req->mask); - if (user_size > kernel_size) - remaining_size = user_size - kernel_size; - else + if (kernel_size > user_size) remaining_size = kernel_size - user_size; + else + remaining_size = 0; if (PAGE_SIZE <= remaining_size) return false; @@ -250,7 +1550,7 @@ bool read_coredump_req(int fd, struct coredump_req *req) if (remaining_size) { char buffer[PAGE_SIZE]; - ret = recv(fd, buffer, sizeof(buffer), MSG_WAITALL); + ret = recv(fd, buffer, remaining_size, MSG_WAITALL); if (ret != remaining_size) return false; fprintf(stderr, "Discarded %zu bytes of data after coredump request\n", remaining_size); @@ -259,8 +1559,13 @@ bool read_coredump_req(int fd, struct coredump_req *req) return true; } -bool send_coredump_ack(int fd, const struct coredump_req *req, - __u64 mask, size_t size_ack) +bool read_coredump_req(int fd, struct coredump_req *req) +{ + return read_coredump_req_sized(fd, req, sizeof(*req)); +} + +/* Send @len bytes of @ack as they are, more than the struct if asked to. */ +bool send_coredump_ack_bytes(int fd, const struct coredump_ack *ack, size_t len) { ssize_t ret; /* @@ -272,30 +1577,74 @@ bool send_coredump_ack(int fd, const struct coredump_req *req, char buffer[PAGE_SIZE]; } large_ack = {}; + if (len > sizeof(large_ack)) + return false; + + large_ack.ack = *ack; + ret = send(fd, &large_ack, len, MSG_NOSIGNAL); + if (ret != len) { + fprintf(stderr, "%s: short send %zd: %m\n", __func__, ret); + return false; + } + + fprintf(stderr, "Sent %zu bytes of coredump ack: size %u, mask 0x%llx, types 0x%llx\n", + len, ack->size, (unsigned long long)ack->mask, + (unsigned long long)ack->memory_types); + return true; +} + +bool send_coredump_ack_types(int fd, const struct coredump_req *req, + __u64 mask, __u64 memory_types, size_t size_ack) +{ + struct coredump_ack ack = { + .mask = mask, + .memory_types = memory_types, + }; + if (!size_ack) size_ack = sizeof(struct coredump_ack) < req->size_ack ? sizeof(struct coredump_ack) : req->size_ack; - large_ack.ack.mask = mask; - large_ack.ack.size = size_ack; - ret = send(fd, &large_ack, size_ack, MSG_NOSIGNAL); - if (ret != size_ack) - return false; + ack.size = size_ack; + return send_coredump_ack_bytes(fd, &ack, size_ack); +} - fprintf(stderr, "Sent coredump ack with size %zu and mask 0x%llx\n", - size_ack, (unsigned long long)mask); - return true; +bool send_coredump_ack(int fd, const struct coredump_req *req, + __u64 mask, size_t size_ack) +{ + return send_coredump_ack_types(fd, req, mask, 0, size_ack); } -bool check_coredump_req(const struct coredump_req *req, size_t min_size, - __u64 required_mask) +/* Every option the kernel is expected to advertise in coredump_req->mask. */ +#define TEST_REQ_MASK_ALL \ + (COREDUMP_KERNEL | COREDUMP_USERSPACE | \ + COREDUMP_REJECT | COREDUMP_WAIT | \ + COREDUMP_RECORDS | COREDUMP_SPARSE | COREDUMP_MEMORY_TYPES) + +bool check_coredump_req(const struct coredump_req *req) { - if (req->size < min_size) + if (req->size != COREDUMP_REQ_SIZE_VER1) { + fprintf(stderr, "%s: size %u, expected %d\n", + __func__, req->size, COREDUMP_REQ_SIZE_VER1); return false; - if ((req->mask & required_mask) != required_mask) + } + if (req->size_ack != COREDUMP_ACK_SIZE_VER1) { + fprintf(stderr, "%s: size_ack %u, expected %d\n", + __func__, req->size_ack, COREDUMP_ACK_SIZE_VER1); + return false; + } + if (req->mask != TEST_REQ_MASK_ALL) { + fprintf(stderr, "%s: mask 0x%llx, expected 0x%llx\n", + __func__, (unsigned long long)req->mask, + (unsigned long long)TEST_REQ_MASK_ALL); return false; - if (req->mask & ~required_mask) + } + if (req->memory_types_mask != TEST_MEMORY_ALL) { + fprintf(stderr, "%s: memory_types_mask 0x%llx, expected 0x%llx\n", + __func__, (unsigned long long)req->memory_types_mask, + (unsigned long long)TEST_MEMORY_ALL); return false; + } return true; } @@ -381,3 +1730,304 @@ out: close(fd_coredump); _exit(exit_code); } + +/* + * TIF_NOTIFY_SIGNAL coredump helpers. + * + * __dump_emit() takes anything short of a full write as the end of the + * dump, so every emit that blocks on a full transport can lose the rest + * of it. The NT_FILE note is the one emit that is certain to block, + * because it is the only one larger than the transport, so these helpers + * build a note large enough for that and then raise TIF_NOTIFY_SIGNAL on + * the dumping task while that write is in flight. + */ + +static int io_uring_setup_raw(unsigned int entries, struct io_uring_params *p) +{ + return syscall(__NR_io_uring_setup, entries, p); +} + +/* io_uring reads poll32_events back through swahw32() on big endian. */ +static __u32 notify_poll_mask(__u32 events) +{ +#if defined(__BYTE_ORDER) && __BYTE_ORDER == __BIG_ENDIAN + return __swahw32(events); +#else + return events; +#endif +} + +bool coredump_io_uring_available(void) +{ + struct io_uring_params p = {}; + int fd; + + fd = io_uring_setup_raw(1, &p); + if (fd < 0) + return false; + close(fd); + return true; +} + +/* + * Arm a poll on @fd. io_uring leaves ctx->notify_method at TWA_SIGNAL + * unless the ring asks for SQPOLL or COOP_TASKRUN, so the completion runs + * set_notify_signal() against the task that submitted it. That is us, and + * we are about to become the coredumping task. + */ +static int arm_poll_notify(int trigger_fd) +{ + struct io_uring_params p = {}; + unsigned int *sq_tail, *sq_array; + struct io_uring_sqe *sqes; + size_t sqring_sz; + void *sq; + int ring; + + ring = io_uring_setup_raw(8, &p); + if (ring < 0) + return -1; + + sqring_sz = p.sq_off.array + p.sq_entries * sizeof(unsigned int); + sq = mmap(NULL, sqring_sz, PROT_READ | PROT_WRITE, + MAP_SHARED | MAP_POPULATE, ring, IORING_OFF_SQ_RING); + if (sq == MAP_FAILED) + return -1; + + sqes = mmap(NULL, p.sq_entries * sizeof(*sqes), PROT_READ | PROT_WRITE, + MAP_SHARED | MAP_POPULATE, ring, IORING_OFF_SQES); + if (sqes == MAP_FAILED) + return -1; + + sq_tail = (unsigned int *)((char *)sq + p.sq_off.tail); + sq_array = (unsigned int *)((char *)sq + p.sq_off.array); + + memset(&sqes[0], 0, sizeof(sqes[0])); + sqes[0].opcode = IORING_OP_POLL_ADD; + sqes[0].fd = trigger_fd; + sqes[0].poll32_events = notify_poll_mask(POLLIN); + + sq_array[0] = 0; + __atomic_store_n(sq_tail, 1, __ATOMIC_RELEASE); + + if (syscall(__NR_io_uring_enter, ring, 1, 0, 0, NULL, 0) < 0) + return -1; + + /* Deliberately leaked, we are about to crash. */ + return 0; +} + +/* + * Adjacent mappings with identical flags and contiguous file offsets are + * merged into one VMA, which would collapse NT_FILE back to nothing, so + * alternate the protection to keep every mapping an entry of its own. + * Not PROT_EXEC, /tmp is often mounted noexec. Nothing is ever written + * through these so they get no anon_vma and stay out of the dump itself. + */ +static int make_file_mappings(void) +{ + long pgsz = sysconf(_SC_PAGESIZE); + int fd, i; + + fd = open(NOTIFY_SIGNAL_MAPFILE, + O_RDWR | O_CREAT | O_TRUNC | O_CLOEXEC, 0600); + if (fd < 0) + return 0; + if (ftruncate(fd, (off_t)NOTIFY_SIGNAL_MAP_COUNT * pgsz)) { + close(fd); + return 0; + } + + for (i = 0; i < NOTIFY_SIGNAL_MAP_COUNT; i++) { + int prot = (i & 1) ? PROT_READ : (PROT_READ | PROT_WRITE); + + if (mmap(NULL, pgsz, prot, MAP_PRIVATE, fd, + (off_t)i * pgsz) == MAP_FAILED) + break; + } + close(fd); + return i; +} + +void crashing_child_notify_signal(void) +{ + long pgsz = sysconf(_SC_PAGESIZE); + unsigned char *p; + int trigger_fd; + unsigned long off; + + /* Open the read side first so the reader's open() cannot block. */ + trigger_fd = open(NOTIFY_SIGNAL_TRIGGER, O_RDONLY | O_NONBLOCK | O_CLOEXEC); + + /* Exit rather than crash: a dump without the poll armed proves nothing. */ + if (trigger_fd < 0) + _exit(EXIT_FAILURE); + + if (make_file_mappings() < NOTIFY_SIGNAL_MAP_COUNT) + _exit(EXIT_FAILURE); + + p = mmap(NULL, NOTIFY_SIGNAL_ANON_BYTES, PROT_READ | PROT_WRITE, + MAP_PRIVATE | MAP_ANONYMOUS, -1, 0); + if (p == MAP_FAILED) + _exit(EXIT_FAILURE); + for (off = 0; off < NOTIFY_SIGNAL_ANON_BYTES; off += pgsz) + p[off] = 1; + + if (arm_poll_notify(trigger_fd)) + _exit(EXIT_FAILURE); + + /* crash on purpose */ + *(volatile int *)NULL = 0; + _exit(EXIT_FAILURE); +} + +static int pull_trigger(void) +{ + int fd; + + fd = open(NOTIFY_SIGNAL_TRIGGER, O_WRONLY | O_NONBLOCK | O_CLOEXEC); + if (fd < 0) { + fprintf(stderr, "%s: open failed: %m\n", __func__); + return -1; + } + if (write(fd, "x", 1) != 1) { + fprintf(stderr, "%s: write failed: %m\n", __func__); + close(fd); + return -1; + } + close(fd); + return 0; +} + +/* + * phdr[0] is the PT_NOTE entry: elf_core_dump() emits it right after the + * ELF header. Its p_offset is where the notes begin. + */ +static long long note_offset(const unsigned char *hdr) +{ + ElfW(Phdr) ph; + ElfW(Ehdr) eh; + + memcpy(&eh, hdr, sizeof(eh)); + if (memcmp(eh.e_ident, ELFMAG, SELFMAG) || eh.e_type != ET_CORE) + return -1; + memcpy(&ph, hdr + sizeof(eh), sizeof(ph)); + if (ph.p_type != PT_NOTE) + return -1; + return (long long)ph.p_offset; +} + +/* + * Drain a coredump off @fd, counting what arrives. Once the notes have + * started the kernel is inside the one big note write, so poke the fifo + * the crashing task is polling and stop reading, which keeps the + * transport full and the write blocked with a partial count when the + * wakeup lands. Then carry on to end of file. + * + * What arrives is written to @fd_out when that is not negative. + * Returns the number of bytes received, or -1. Failing to trip the fifo + * is an error too: a dump that was never interrupted proves nothing. + */ +ssize_t recv_coredump_notify_signal(int fd, int fd_out, bool arm) +{ + unsigned char hdr[sizeof(ElfW(Ehdr)) + sizeof(ElfW(Phdr))]; + static char buf[64 << 10]; + long pgsz = sysconf(_SC_PAGESIZE); + long long note_off = 0; + size_t hdrlen = 0; + ssize_t total = 0; + bool armed = false; + + for (;;) { + ssize_t n = read(fd, buf, sizeof(buf)); + + if (n < 0) { + if (errno == EINTR) + continue; + return -1; + } + if (n == 0) + break; + if (fd_out >= 0 && write(fd_out, buf, n) != n) + return -1; + + if (hdrlen < sizeof(hdr)) { + size_t want = sizeof(hdr) - hdrlen; + + if (want > (size_t)n) + want = (size_t)n; + memcpy(hdr + hdrlen, buf, want); + hdrlen += want; + if (hdrlen == sizeof(hdr)) + note_off = note_offset(hdr); + } + + total += n; + + if (arm && !armed && note_off > 0 && + total > note_off + (long long)pgsz) { + if (pull_trigger()) + return -1; + armed = true; + usleep(NOTIFY_SIGNAL_STALL_US); + continue; + } + } + + if (arm && !armed) + return -1; + + return total; +} + +/* + * How large the dump was meant to be. The ELF header and the program + * headers are the first thing emitted, so even a truncated dump says how + * far it should have run: the end is max(p_offset + p_filesz). + * + * That end is exact even when the last segment ends in a hole: + * coredump_write() flushes the pending cprm->to_skip with a final one + * byte emit and __dump_skip() writes zeroes for transports that cannot + * seek, so a whole dump carries every byte the headers promise. + */ +long long coredump_expected_size(const char *path) +{ + ElfW(Phdr) *phdr = NULL; + long long expected = 0; + unsigned int nphdr, i; + ElfW(Ehdr) eh; + int fd; + + fd = open(path, O_RDONLY | O_CLOEXEC); + if (fd < 0) + return -1; + if (read(fd, &eh, sizeof(eh)) != sizeof(eh)) + goto err; + if (memcmp(eh.e_ident, ELFMAG, SELFMAG) || eh.e_type != ET_CORE) + goto err; + if (!eh.e_phnum || eh.e_phentsize != sizeof(*phdr)) + goto err; + + nphdr = eh.e_phnum; + phdr = calloc(nphdr, sizeof(*phdr)); + if (!phdr) + goto err; + if (pread(fd, phdr, (size_t)nphdr * sizeof(*phdr), (off_t)eh.e_phoff) != + (ssize_t)((size_t)nphdr * sizeof(*phdr))) + goto err; + + for (i = 0; i < nphdr; i++) { + long long end = (long long)phdr[i].p_offset + + (long long)phdr[i].p_filesz; + if (end > expected) + expected = end; + } + + free(phdr); + close(fd); + return expected; +err: + free(phdr); + close(fd); + return -1; +} diff --git a/tools/testing/selftests/coredump/coredump_test_helpers.h b/tools/testing/selftests/coredump/coredump_test_helpers.h new file mode 100644 index 000000000000..3f2f87837558 --- /dev/null +++ b/tools/testing/selftests/coredump/coredump_test_helpers.h @@ -0,0 +1,79 @@ +/* SPDX-License-Identifier: GPL-2.0 */ + +#ifndef __COREDUMP_TEST_HELPERS_H +#define __COREDUMP_TEST_HELPERS_H + +#include <link.h> +#include <stdbool.h> +#include <sys/types.h> +#include <linux/coredump.h> + +#include "../pidfd/pidfd.h" + +#ifndef PAGE_SIZE +#define PAGE_SIZE 4096 +#endif + +#define NUM_THREAD_SPAWN 128 + +/* Size of the mostly unpopulated mapping the sparse coredump test maps. */ +#define SPARSE_MAPPING_SIZE (256 * 1024 * 1024) + +/* A task mapping at least this much is worth a record stream. */ +#define SPARSE_STREAM_THRESHOLD (SPARSE_MAPPING_SIZE / 2) + +/* Size of the shared anonymous mapping the memory types tests map. */ +#define MEMORY_MAPPING_SIZE (4 * 1024 * 1024) + +/* Leave the coredump_filter the crashing child inherited alone. */ +#define FILTER_TASK_INHERIT ((__u64)-1) + +/* Every memory type the kernel is expected to advertise. */ +#define TEST_MEMORY_ALL \ + (COREDUMP_MEMORY_ANON_PRIVATE | COREDUMP_MEMORY_ANON_SHARED | \ + COREDUMP_MEMORY_FILE_PRIVATE | COREDUMP_MEMORY_FILE_SHARED | \ + COREDUMP_MEMORY_ELF_HEADERS | \ + COREDUMP_MEMORY_HUGETLB_PRIVATE | COREDUMP_MEMORY_HUGETLB_SHARED | \ + COREDUMP_MEMORY_DAX_PRIVATE | COREDUMP_MEMORY_DAX_SHARED) + +/* Shared helper function declarations */ +void *do_nothing(void *arg); +void crashing_child(void); +void crashing_child_sparse(size_t size); +void crashing_child_memory(__u64 task_filter, int fd_addr); +bool find_coredump_segment(int fd, __u64 vaddr, ElfW(Phdr) *segment); +bool sum_coredump_segments(int fd, __u64 *data, __u64 *notes); +bool check_coredump_extent(int fd); +bool peer_coredump_filter(int fd_peer_pidfd, __u64 *memory_types); +ssize_t peer_read_mem(int fd_peer_pidfd, __u64 addr, void *buf, size_t len); +ssize_t recv_coredump_records(int fd_coredump, int fd_core_file, + off_t *coredump_size, bool *truncated, + int fd_peer_pidfd); +ssize_t recv_coredump_compact(int fd_coredump, int fd_object, int fd_reference, + off_t *coredump_size); +ssize_t recv_coredump_bytes(int fd_coredump, int fd_core_file); +ssize_t peer_vm_size(int fd_peer_pidfd); +bool is_elf_core(int fd); +int check_compact_coredump(int fd_object, int fd_reference); +int create_detached_tmpfs(void); +int create_and_listen_unix_socket(const char *path); +bool set_core_pattern(const char *pattern); +int get_peer_pidfd(int fd); +bool get_pidfd_info(int fd_peer_pidfd, struct pidfd_info *info); + +/* Protocol helper function declarations */ +ssize_t recv_marker(int fd); +bool read_marker(int fd, enum coredump_mark mark); +bool read_hangup(int fd); +bool read_coredump_req(int fd, struct coredump_req *req); +bool read_coredump_req_sized(int fd, struct coredump_req *req, size_t user_size); +bool send_coredump_ack(int fd, const struct coredump_req *req, + __u64 mask, size_t size_ack); +bool send_coredump_ack_types(int fd, const struct coredump_req *req, + __u64 mask, __u64 memory_types, size_t size_ack); +bool send_coredump_ack_bytes(int fd, const struct coredump_ack *ack, size_t len); +bool check_coredump_req(const struct coredump_req *req); +int open_coredump_tmpfile(int fd_tmpfs_detached); +void process_coredump_worker(int fd_coredump, int fd_peer_pidfd, int fd_core_file); + +#endif /* __COREDUMP_TEST_HELPERS_H */ diff --git a/tools/testing/selftests/coredump/coredump_worker_test.c b/tools/testing/selftests/coredump/coredump_worker_test.c new file mode 100644 index 000000000000..9a4270b65a6e --- /dev/null +++ b/tools/testing/selftests/coredump/coredump_worker_test.c @@ -0,0 +1,447 @@ +// SPDX-License-Identifier: GPL-2.0 + +/* + * A user worker as the coredumping thread. + * + * io-wq workers and SQPOLL threads are threads of the process that never + * return to userspace. They block every signal but SIGKILL and SIGSTOP, + * but a tracer can replace that mask with PTRACE_SETSIGMASK and inject a + * coredump signal. get_signal() then runs vfs_coredump() in the worker. + * The worker's own exit bookkeeping runs only after the dump, so a + * zapped sibling that waits for it in its exit path deadlocks with the + * dumper and the whole thread group is stuck in D state. + * + * Inject SIGSEGV into a chosen thread and require that the thread group + * is gone in bounded time, either because the dump completed or because + * SIGKILL still works. A failure leaves the stuck process behind. + */ +#include <ctype.h> +#include <errno.h> +#include <dirent.h> +#include <fcntl.h> +#include <sys/mman.h> +#include <sys/ptrace.h> +#include <sys/resource.h> +#include <sys/stat.h> +#include <sys/syscall.h> +#include <sys/wait.h> +#include <unistd.h> +#include <linux/io_uring.h> + +#include "coredump_test.h" + +#ifndef PTRACE_SETSIGMASK +#define PTRACE_SETSIGMASK 0x420b +#endif + +/* The dump of the tiny child takes well under a second. */ +#define EXIT_TIMEOUT_MS 5000 + +FIXTURE_SETUP(coredump) +{ + FILE *file; + int ret; + + self->pid_coredump_server = -ESRCH; + self->fd_tmpfs_detached = -1; + file = fopen("/proc/sys/kernel/core_pattern", "r"); + ASSERT_NE(NULL, file); + + ret = fread(self->original_core_pattern, 1, sizeof(self->original_core_pattern), file); + ASSERT_TRUE(ret || feof(file)); + ASSERT_LT(ret, sizeof(self->original_core_pattern)); + + self->original_core_pattern[ret] = '\0'; + + ret = fclose(file); + ASSERT_EQ(0, ret); +} + +FIXTURE_TEARDOWN(coredump) +{ + const char *reason; + FILE *file; + int ret; + + file = fopen("/proc/sys/kernel/core_pattern", "w"); + if (!file) { + reason = "Unable to open core_pattern"; + goto fail; + } + + ret = fprintf(file, "%s", self->original_core_pattern); + if (ret < 0) { + reason = "Unable to write to core_pattern"; + goto fail; + } + + ret = fclose(file); + if (ret) { + reason = "Unable to close core_pattern"; + goto fail; + } + + return; +fail: + /* This should never happen */ + fprintf(stderr, "Failed to cleanup coredump test: %s\n", reason); +} + +/* A raw ring, no liburing. */ +struct uring { + int fd; + struct io_uring_params params; + void *sq; + size_t sq_len; + struct io_uring_sqe *sqes; + size_t sqes_len; + unsigned int *sq_tail, *sq_mask, *sq_array; + unsigned int *cq_head, *cq_tail, *cq_mask; + struct io_uring_cqe *cqes; +}; + +static int uring_setup(struct uring *r, unsigned int flags) +{ + size_t cq_len; + + memset(r, 0, sizeof(*r)); + r->params.flags = flags; + if (flags & IORING_SETUP_SQPOLL) + r->params.sq_thread_idle = 2000; + r->fd = syscall(__NR_io_uring_setup, 8, &r->params); + if (r->fd < 0) + return -1; + if (!(r->params.features & IORING_FEAT_SINGLE_MMAP)) + return -1; + + r->sq_len = r->params.sq_off.array + r->params.sq_entries * sizeof(unsigned int); + cq_len = r->params.cq_off.cqes + r->params.cq_entries * sizeof(struct io_uring_cqe); + if (cq_len > r->sq_len) + r->sq_len = cq_len; + r->sq = mmap(NULL, r->sq_len, PROT_READ | PROT_WRITE, + MAP_SHARED | MAP_POPULATE, r->fd, IORING_OFF_SQ_RING); + if (r->sq == MAP_FAILED) + return -1; + r->sqes_len = r->params.sq_entries * sizeof(struct io_uring_sqe); + r->sqes = mmap(NULL, r->sqes_len, PROT_READ | PROT_WRITE, + MAP_SHARED | MAP_POPULATE, r->fd, IORING_OFF_SQES); + if (r->sqes == MAP_FAILED) + return -1; + + r->sq_tail = r->sq + r->params.sq_off.tail; + r->sq_mask = r->sq + r->params.sq_off.ring_mask; + r->sq_array = r->sq + r->params.sq_off.array; + r->cq_head = r->sq + r->params.cq_off.head; + r->cq_tail = r->sq + r->params.cq_off.tail; + r->cq_mask = r->sq + r->params.cq_off.ring_mask; + r->cqes = r->sq + r->params.cq_off.cqes; + return 0; +} + +/* Submit one sqe, wait for its completion and return the result. */ +static int uring_submit_wait(struct uring *r, const struct io_uring_sqe *sqe) +{ + unsigned int tail = *r->sq_tail, idx = tail & *r->sq_mask; + unsigned int flags = IORING_ENTER_GETEVENTS; + int i; + + r->sqes[idx] = *sqe; + r->sq_array[idx] = idx; + __atomic_store_n(r->sq_tail, tail + 1, __ATOMIC_RELEASE); + + if (r->params.flags & IORING_SETUP_SQPOLL) + flags |= IORING_ENTER_SQ_WAKEUP; + + for (i = 0; i < 100; i++) { + if (syscall(__NR_io_uring_enter, r->fd, 1, 1, flags, NULL, 0) < 0 && + errno != EINTR) + return -1; + if (__atomic_load_n(r->cq_tail, __ATOMIC_ACQUIRE) != *r->cq_head) { + unsigned int head = *r->cq_head; + int res = r->cqes[head & *r->cq_mask].res; + + __atomic_store_n(r->cq_head, head + 1, __ATOMIC_RELEASE); + return res; + } + flags &= ~IORING_ENTER_SQ_WAKEUP; + usleep(10 * 1000); + } + return -1; +} + +static bool uring_available(unsigned int flags) +{ + struct io_uring_params params = { .flags = flags }; + int fd; + + fd = syscall(__NR_io_uring_setup, 2, ¶ms); + if (fd < 0) + return false; + close(fd); + return true; +} + +/* + * Keep a ring, an idle io-wq worker and with SQPOLL an SQPOLL thread + * alive. The last worker of a ring never exits on its idle timeout. + */ +static void worker_child(bool sqpoll, int fd_ipc) +{ + struct rlimit rl = { RLIM_INFINITY, RLIM_INFINITY }; + struct io_uring_sqe sqe = {}; + static char buf[64]; + struct uring ring; + int memfd; + + if (setrlimit(RLIMIT_CORE, &rl)) + _exit(EXIT_FAILURE); + + memfd = memfd_create("coredump_worker", 0); + if (memfd < 0 || write(memfd, "hello", 5) != 5) + _exit(EXIT_FAILURE); + + if (uring_setup(&ring, sqpoll ? IORING_SETUP_SQPOLL : 0)) + _exit(EXIT_FAILURE); + + /* IOSQE_ASYNC forces the read through io-wq so a worker appears. */ + sqe.opcode = IORING_OP_READ; + sqe.fd = memfd; + sqe.addr = (__u64)(uintptr_t)buf; + sqe.len = sizeof(buf); + sqe.flags = IOSQE_ASYNC; + if (uring_submit_wait(&ring, &sqe) != 5) + _exit(EXIT_FAILURE); + + if (write_nointr(fd_ipc, "1", 1) != 1) + _exit(EXIT_FAILURE); + close(fd_ipc); + + for (;;) + pause(); +} + +/* Find the thread of @pid whose comm starts with @prefix. */ +static pid_t find_thread(pid_t pid, const char *prefix) +{ + char path[64], comm[64]; + pid_t tid = -1; + struct dirent *de; + ssize_t bytes; + DIR *dir; + int fd; + + snprintf(path, sizeof(path), "/proc/%d/task", pid); + dir = opendir(path); + if (!dir) + return -1; + + while (tid < 0 && (de = readdir(dir))) { + if (!isdigit(de->d_name[0])) + continue; + snprintf(path, sizeof(path), "/proc/%d/task/%s/comm", pid, de->d_name); + fd = open(path, O_RDONLY | O_CLOEXEC); + if (fd < 0) + continue; + bytes = read(fd, comm, sizeof(comm) - 1); + close(fd); + if (bytes <= 0) + continue; + comm[bytes] = '\0'; + if (!strncmp(comm, prefix, strlen(prefix))) + tid = atoi(de->d_name); + } + closedir(dir); + return tid; +} + +/* + * Attach, stop the thread with SIGSTOP, drop the signal mask that + * copy_process() gave it and resume it with SIGSEGV. Returns 1 when the + * mask was changed, 0 when PTRACE_SETSIGMASK was refused (a user worker + * keeps its mask and the SIGSEGV stays pending), -1 on any other failure. + */ +static int inject_coredump_signal(pid_t pid, pid_t tid) +{ + __u64 mask = 0; + int status, ret = 1; + + if (ptrace(PTRACE_SEIZE, tid, NULL, NULL)) + return -1; + if (syscall(SYS_tgkill, pid, tid, SIGSTOP)) + return -1; + if (waitpid(tid, &status, __WALL) != tid) + return -1; + if (!WIFSTOPPED(status) || WSTOPSIG(status) != SIGSTOP) + return -1; + if (ptrace(PTRACE_SETSIGMASK, tid, sizeof(mask), &mask)) { + if (errno != EPERM) + return -1; + ret = 0; + } + if (ptrace(PTRACE_DETACH, tid, NULL, (void *)(long)SIGSEGV)) + return -1; + return ret; +} + +/* Reap @pid within @timeout_ms, -1 when it is still there. */ +static int wait_exit(pid_t pid, int *status, int timeout_ms) +{ + int i; + + for (i = 0; i < timeout_ms / 10; i++) { + pid_t ret = waitpid(pid, status, WNOHANG); + + if (ret == pid) + return 0; + if (ret < 0) + return -1; + usleep(10 * 1000); + } + return -1; +} + +static void log_threads(struct __test_metadata *const _metadata, pid_t pid) +{ + char path[64], line[256], comm[64] = {}; + struct dirent *de; + DIR *dir; + FILE *f; + + snprintf(path, sizeof(path), "/proc/%d/task", pid); + dir = opendir(path); + if (!dir) + return; + while ((de = readdir(dir))) { + if (!isdigit(de->d_name[0])) + continue; + snprintf(path, sizeof(path), "/proc/%d/task/%s/status", pid, de->d_name); + f = fopen(path, "r"); + if (!f) + continue; + while (fgets(line, sizeof(line), f)) { + line[strcspn(line, "\n")] = '\0'; + if (!strncmp(line, "Name:", 5)) + snprintf(comm, sizeof(comm), "%s", line + 6); + else if (!strncmp(line, "State:", 6)) + TH_LOG("tid %s (%s) %s", de->d_name, comm, line + 7); + } + fclose(f); + } + closedir(dir); +} + +enum dumper { + DUMPER_MAIN, + DUMPER_WORKER, + DUMPER_SQPOLL, +}; + +static void run_dumper(struct __test_metadata *const _metadata, bool sqpoll, + enum dumper dumper) +{ + bool killed = false; + char path[64], c; + int ipc[2], status, fd, ret; + pid_t pid, tid; + + ASSERT_TRUE(set_core_pattern("/tmp/coredump.file.%p")); + ASSERT_EQ(pipe2(ipc, O_CLOEXEC), 0); + + pid = fork(); + ASSERT_GE(pid, 0); + if (pid == 0) { + close(ipc[0]); + worker_child(sqpoll, ipc[1]); + } + close(ipc[1]); + ASSERT_EQ(read_nointr(ipc[0], &c, 1), 1); + close(ipc[0]); + + switch (dumper) { + case DUMPER_MAIN: + tid = pid; + break; + case DUMPER_WORKER: + tid = find_thread(pid, "iou-wrk-"); + break; + case DUMPER_SQPOLL: + tid = find_thread(pid, "iou-sqp-"); + break; + } + ASSERT_GT(tid, 0); + ret = inject_coredump_signal(pid, tid); + ASSERT_GE(ret, 0); + if (!ret) { + /* The signal sits on the worker, the group must be untouched. */ + ASSERT_NE(dumper, DUMPER_MAIN); + TH_LOG("PTRACE_SETSIGMASK refused for tid %d, the SIGSEGV stays pending", tid); + ASSERT_EQ(wait_exit(pid, &status, 1000), -1); + kill(pid, SIGKILL); + ASSERT_EQ(wait_exit(pid, &status, EXIT_TIMEOUT_MS), 0); + ASSERT_TRUE(WIFSIGNALED(status)); + ASSERT_EQ(WTERMSIG(status), SIGKILL); + return; + } + + if (wait_exit(pid, &status, EXIT_TIMEOUT_MS)) { + /* No dump. Whatever happened, SIGKILL must still work. */ + log_threads(_metadata, pid); + kill(pid, SIGKILL); + killed = true; + ASSERT_EQ(wait_exit(pid, &status, EXIT_TIMEOUT_MS), 0) { + TH_LOG("thread group %d is stuck after SIGSEGV to tid %d", + pid, tid); + } + } + + ASSERT_TRUE(WIFSIGNALED(status)); + if (killed) { + TH_LOG("tid %d did not dump, the group was killed instead", tid); + ASSERT_EQ(WTERMSIG(status), SIGKILL); + return; + } + ASSERT_EQ(WTERMSIG(status), SIGSEGV); + ASSERT_TRUE(WCOREDUMP(status)); + + snprintf(path, sizeof(path), "/tmp/coredump.file.%d", pid); + fd = open(path, O_RDONLY | O_CLOEXEC); + unlink(path); + ASSERT_GE(fd, 0); + ASSERT_TRUE(check_coredump_extent(fd)); + close(fd); +} + +/* The mechanics: an injected SIGSEGV into a normal thread dumps core. */ +TEST_F(coredump, main_thread_dumper) +{ + if (!uring_available(0)) + SKIP(return, "io_uring is not available"); + run_dumper(_metadata, false, DUMPER_MAIN); +} + +TEST_F(coredump, plain_worker_dumper) +{ + if (!uring_available(0)) + SKIP(return, "io_uring is not available"); + run_dumper(_metadata, false, DUMPER_WORKER); +} + +TEST_F(coredump, sqpoll_thread_dumper) +{ + if (!uring_available(IORING_SETUP_SQPOLL)) + SKIP(return, "io_uring SQPOLL is not available"); + run_dumper(_metadata, true, DUMPER_SQPOLL); +} + +/* + * The SQPOLL thread leaves its loop on the zap and waits for its io-wq + * workers to exit before it parks. The dumping worker never does. + */ +TEST_F(coredump, sqpoll_worker_dumper) +{ + if (!uring_available(IORING_SETUP_SQPOLL)) + SKIP(return, "io_uring SQPOLL is not available"); + run_dumper(_metadata, true, DUMPER_WORKER); +} + +TEST_HARNESS_MAIN diff --git a/tools/testing/selftests/exec/Makefile b/tools/testing/selftests/exec/Makefile index b640af8f02b5..2220ed345e92 100644 --- a/tools/testing/selftests/exec/Makefile +++ b/tools/testing/selftests/exec/Makefile @@ -45,6 +45,10 @@ TEST_GEN_FILES += binfmt_transparent_interp TEST_GEN_PROGS += binfmt_misc_loader TEST_GEN_FILES += binfmt_loader_payload binfmt_loader_payload_static +# Only ASCII punctuation delimits the fields of a register string, so a new +# flag character cannot change which strings register. No bpf toolchain. +TEST_GEN_PROGS += binfmt_misc_delim + # binfmt_misc bpf-backed ('B') handler test: a libbpf harness plus its # struct_ops objects and the test interpreter/app it routes between. Only # built when clang, bpftool, the vmlinux BTF and libbpf are all present diff --git a/tools/testing/selftests/exec/binfmt_misc_delim.c b/tools/testing/selftests/exec/binfmt_misc_delim.c new file mode 100644 index 000000000000..ffc17cb78545 --- /dev/null +++ b/tools/testing/selftests/exec/binfmt_misc_delim.c @@ -0,0 +1,127 @@ +// SPDX-License-Identifier: GPL-2.0 +/* + * Test which characters may delimit the fields of a register string. + */ +#define _GNU_SOURCE +#include <stdio.h> +#include <stdlib.h> + +#include "binfmt_misc_common.h" +#include "kselftest_harness.h" + +#define ENTRY "bmdelim" +/* Shares no character with the sets below, or a refusal proves nothing. */ +#define MAGIC "bmmagic" +#define INTERP "/bin/true" + +/* + * ASCII punctuation without '\' and '/'. The backslash is refused because + * it would cut a magic that uses \x to escape short. '/' is accepted + * but cannot delimit a rule that names an absolute interpreter. + */ +#define PUNCTUATION "!\"#$%&'()*+,-.:;<=>?@[]^_`{|}~" + +/* 'M', 'E' and 'B' name types, 'P' through 'D' are the flags. */ +#define LETTERS "MEBPOCFTLDqz" +#define DIGITS "0157" +#define WHITESPACE " \t\n" +#define CONTROL "\001\033\177" +#define NON_ASCII "\200\244\377" + +/* ':bmdelim:E::bmmagic::/bin/true:' with @del in place of every ':'. */ +static int register_with(char del) +{ + char rule[128]; + + snprintf(rule, sizeof(rule), "%c%s%cE%c%c%s%c%c%s%c", del, ENTRY, del, + del, del, MAGIC, del, del, INTERP, del); + return write_reg(rule); +} + +/* No character of @set may delimit a register string. */ +static void expect_refused(struct __test_metadata *_metadata, const char *set) +{ + const char *d; + + for (d = set; *d; d++) { + int rc = register_with(*d); + + EXPECT_EQ(rc, -1) + TH_LOG("%#x delimited a register string", + (unsigned char)*d); + if (rc == 0) { + unregister(ENTRY); + continue; + } + EXPECT_EQ(errno, EINVAL); + } +} + +FIXTURE(delim) { +}; + +FIXTURE_SETUP(delim) +{ + if (getuid() != 0) + SKIP(return, "test must be run as root"); + if (!binfmt_misc_available()) + SKIP(return, "no binfmt_misc"); + + /* A kernel without the allow-list takes any character but a flag. */ + if (register_with('q') == 0) { + unregister(ENTRY); + SKIP(return, "kernel without the delimiter allow-list"); + } +} + +FIXTURE_TEARDOWN(delim) +{ + unregister(ENTRY); +} + +/* Punctuation delimits, which is all anything deployed ever uses. */ +TEST_F(delim, punctuation_accepted) +{ + const char *d; + + for (d = PUNCTUATION; *d; d++) { + EXPECT_EQ(register_with(*d), 0) + TH_LOG("'%c' refused with errno %d", *d, errno); + unregister(ENTRY); + } +} + +/* Letters name the types and the flags, so none of them can delimit. */ +TEST_F(delim, letters_refused) +{ + expect_refused(_metadata, LETTERS); +} + +/* The offset field is written in digits. */ +TEST_F(delim, digits_refused) +{ + expect_refused(_metadata, DIGITS); +} + +TEST_F(delim, whitespace_refused) +{ + expect_refused(_metadata, WHITESPACE); +} + +TEST_F(delim, control_refused) +{ + expect_refused(_metadata, CONTROL); +} + +TEST_F(delim, non_ascii_refused) +{ + expect_refused(_metadata, NON_ASCII); +} + +/* The escape character would cut every magic that uses one short. */ +TEST_F(delim, backslash_refused) +{ + expect_refused(_metadata, "\\"); +} + +TEST_HARNESS_MAIN diff --git a/tools/testing/selftests/filesystems/.gitignore b/tools/testing/selftests/filesystems/.gitignore index 9eb185fb2f9d..57f5bbdbedff 100644 --- a/tools/testing/selftests/filesystems/.gitignore +++ b/tools/testing/selftests/filesystems/.gitignore @@ -2,7 +2,6 @@ dnotify_test devpts_pts fclog -file_stressor anon_inode_test kernfs_test idmapped_tmpfile diff --git a/tools/testing/selftests/filesystems/Makefile b/tools/testing/selftests/filesystems/Makefile index 03be337c1f35..bc4bfb677589 100644 --- a/tools/testing/selftests/filesystems/Makefile +++ b/tools/testing/selftests/filesystems/Makefile @@ -1,7 +1,7 @@ # SPDX-License-Identifier: GPL-2.0 CFLAGS += $(KHDR_INCLUDES) -TEST_GEN_PROGS := devpts_pts file_stressor anon_inode_test kernfs_test fclog ustat_test +TEST_GEN_PROGS := devpts_pts anon_inode_test kernfs_test fclog ustat_test TEST_GEN_PROGS += idmapped_tmpfile TEST_GEN_PROGS_EXTENDED := dnotify_test diff --git a/tools/testing/selftests/filesystems/config b/tools/testing/selftests/filesystems/config new file mode 100644 index 000000000000..9f45bc493a30 --- /dev/null +++ b/tools/testing/selftests/filesystems/config @@ -0,0 +1,8 @@ +CONFIG_CGROUPS=y +CONFIG_CGROUP_PIDS=y +CONFIG_FHANDLE=y +CONFIG_NAMESPACES=y +CONFIG_NET=y +CONFIG_NET_NS=y +CONFIG_INET=y +CONFIG_SYSFS=y diff --git a/tools/testing/selftests/filesystems/configfs/.gitignore b/tools/testing/selftests/filesystems/configfs/.gitignore new file mode 100644 index 000000000000..accfb6bb4826 --- /dev/null +++ b/tools/testing/selftests/filesystems/configfs/.gitignore @@ -0,0 +1,2 @@ +# SPDX-License-Identifier: GPL-2.0-only +configfs_test diff --git a/tools/testing/selftests/filesystems/configfs/Makefile b/tools/testing/selftests/filesystems/configfs/Makefile new file mode 100644 index 000000000000..40a91ed788ed --- /dev/null +++ b/tools/testing/selftests/filesystems/configfs/Makefile @@ -0,0 +1,8 @@ +# SPDX-License-Identifier: GPL-2.0 +# Copyright (c) 2026 Meta Platforms, Inc. and affiliates +# Copyright (c) 2026 Breno Leitao <leitao@debian.org> + +CFLAGS += -Wall -Werror -pthread +TEST_GEN_PROGS := configfs_test + +include ../../lib.mk diff --git a/tools/testing/selftests/filesystems/configfs/config b/tools/testing/selftests/filesystems/configfs/config new file mode 100644 index 000000000000..5ea17df535b3 --- /dev/null +++ b/tools/testing/selftests/filesystems/configfs/config @@ -0,0 +1,5 @@ +CONFIG_CONFIGFS_FS=y +CONFIG_MODULES=y +CONFIG_MODULE_UNLOAD=y +CONFIG_SAMPLES=y +CONFIG_SAMPLE_CONFIGFS=m diff --git a/tools/testing/selftests/filesystems/configfs/configfs_test.c b/tools/testing/selftests/filesystems/configfs/configfs_test.c new file mode 100644 index 000000000000..c6a1049e5852 --- /dev/null +++ b/tools/testing/selftests/filesystems/configfs/configfs_test.c @@ -0,0 +1,481 @@ +// SPDX-License-Identifier: GPL-2.0 +/* + * Exercise the configfs userspace interface through the subsystems + * registered by samples/configfs. + * + * Copyright (c) 2026 Meta Platforms, Inc. and affiliates + * Copyright (c) 2026 Breno Leitao <leitao@debian.org> + */ +#define _GNU_SOURCE + +#include <errno.h> +#include <fcntl.h> +#include <limits.h> +#include <pthread.h> +#include <sched.h> +#include <stdbool.h> +#include <stdio.h> +#include <stdlib.h> +#include <string.h> +#include <sys/mount.h> +#include <sys/stat.h> +#include <sys/syscall.h> +#include <sys/vfs.h> +#include <unistd.h> + +#include "kselftest_harness.h" + +/* Private to fs/configfs/mount.c. */ +#define CONFIGFS_MAGIC 0x62656570 + +#define SAMPLE_MODULE "configfs_sample" + +#define CHILDLESS "01-childless" +#define SIMPLE "02-simple-children" +#define GROUPS "03-group-children" +#define SYMLINKS "04-symlink-children" + +#define ITEM_A SIMPLE "/kselftest-a" +#define ITEM_B SIMPLE "/kselftest-b" +#define GROUP GROUPS "/kselftest-group" +#define GROUP_ITEM GROUP "/kselftest-a" +#define LINK_SRC SYMLINKS "/kselftest-src" +#define LINK LINK_SRC "/kselftest-link" + +#define RACE_ITERATIONS 20000 + +static const char * const test_links[] = { + LINK, +}; + +/* Deepest first, so one pass empties the tree. */ +static const char * const test_dirs[] = { + GROUP_ITEM, + GROUP, + LINK_SRC, + ITEM_A, + ITEM_B, +}; + +static void drop_test_dirs(void) +{ + size_t i; + + /* Links first: they hold both their source and their target. */ + for (i = 0; i < ARRAY_SIZE(test_links); i++) + unlink(test_links[i]); + + for (i = 0; i < ARRAY_SIZE(test_dirs); i++) + rmdir(test_dirs[i]); +} + +static ssize_t read_attr(const char *path, char *buf, size_t len) +{ + ssize_t ret; + int fd; + + fd = open(path, O_RDONLY); + if (fd < 0) + return -1; + + ret = read(fd, buf, len - 1); + close(fd); + if (ret < 0) + return -1; + + buf[ret] = '\0'; + return ret; +} + +static ssize_t write_attr(const char *path, const char *val) +{ + ssize_t ret; + int fd, err; + + fd = open(path, O_WRONLY); + if (fd < 0) + return -1; + + ret = write(fd, val, strlen(val)); + err = errno; + close(fd); + errno = err; + + return ret; +} + +FIXTURE(configfs) { + char mnt[sizeof(P_tmpdir "/configfs_XXXXXX")]; + bool mounted; +}; + +FIXTURE_SETUP(configfs) +{ + char tmpl[] = P_tmpdir "/configfs_XXXXXX"; + + if (geteuid()) + SKIP(return, "need root to load modules and mount configfs"); + + ASSERT_EQ(system("modprobe -q " SAMPLE_MODULE), 0) + TH_LOG(SAMPLE_MODULE " missing, is CONFIG_SAMPLE_CONFIGFS=m?"); + + ASSERT_EQ(unshare(CLONE_NEWNS), 0); + ASSERT_EQ(mount(NULL, "/", NULL, MS_REC | MS_PRIVATE, NULL), 0); + + ASSERT_NE(mkdtemp(tmpl), NULL); + strcpy(self->mnt, tmpl); + + ASSERT_EQ(mount("configfs", self->mnt, "configfs", 0, NULL), 0); + ASSERT_EQ(chdir(self->mnt), 0); + self->mounted = true; + + /* configfs items outlive the mount, so a killed run leaves some. */ + drop_test_dirs(); +} + +FIXTURE_TEARDOWN(configfs) +{ + if (self->mounted) { + drop_test_dirs(); + EXPECT_EQ(chdir("/"), 0); + EXPECT_EQ(umount2(self->mnt, MNT_DETACH), 0); + } + + if (self->mnt[0]) + EXPECT_EQ(rmdir(self->mnt), 0); +} + +TEST_F(configfs, mount_and_subsystems) +{ + const char * const subsys[] = { CHILDLESS, SIMPLE, GROUPS, SYMLINKS }; + struct statfs sfs; + struct stat st; + size_t i; + + ASSERT_EQ(statfs(".", &sfs), 0); + EXPECT_EQ(sfs.f_type, CONFIGFS_MAGIC); + + for (i = 0; i < ARRAY_SIZE(subsys); i++) { + ASSERT_EQ(stat(subsys[i], &st), 0) + TH_LOG("%s is missing", subsys[i]); + EXPECT_TRUE(S_ISDIR(st.st_mode)); + } +} + +TEST_F(configfs, mkdir_at_root) +{ + /* The root has no ->mkdir(); only subsystems register there. */ + ASSERT_EQ(mkdir("kselftest-root", 0755), -1); + EXPECT_EQ(errno, EPERM); +} + +TEST_F(configfs, rmdir_subsystem) +{ + ASSERT_EQ(rmdir(CHILDLESS), -1); + EXPECT_EQ(errno, EPERM); +} + +TEST_F(configfs, mkdir_without_group_ops) +{ + /* 01-childless has attributes but no ->make_item()/->make_group(). */ + ASSERT_EQ(mkdir(CHILDLESS "/kselftest-a", 0755), -1); + EXPECT_EQ(errno, EPERM); +} + +TEST_F(configfs, attr_store_and_show) +{ + char buf[64]; + + ASSERT_GT(write_attr(CHILDLESS "/storeme", "42"), 0); + ASSERT_GT(read_attr(CHILDLESS "/storeme", buf, sizeof(buf)), 0); + EXPECT_STREQ(buf, "42\n"); +} + +TEST_F(configfs, attr_store_rejects_garbage) +{ + ASSERT_EQ(write_attr(CHILDLESS "/storeme", "not-a-number"), -1); + EXPECT_EQ(errno, EINVAL); +} + +TEST_F(configfs, attr_show_runs_on_every_open) +{ + char first[64], second[64]; + + /* 01-childless/showme increments the value it just returned. */ + ASSERT_GT(read_attr(CHILDLESS "/showme", first, sizeof(first)), 0); + ASSERT_GT(read_attr(CHILDLESS "/showme", second, sizeof(second)), 0); + EXPECT_EQ(atoi(second), atoi(first) + 1); +} + +TEST_F(configfs, attr_read_only) +{ + ASSERT_EQ(open(CHILDLESS "/description", O_WRONLY), -1); + EXPECT_EQ(errno, EACCES); +} + +TEST_F(configfs, attr_unlink) +{ + /* ->unlink() only accepts the symlinks configfs itself created. */ + ASSERT_EQ(unlink(CHILDLESS "/storeme"), -1); + EXPECT_EQ(errno, EPERM); +} + +TEST_F(configfs, attr_read_length) +{ + char buf[8192]; + struct stat st; + ssize_t n; + int fd; + + fd = open(CHILDLESS "/description", O_RDONLY); + ASSERT_GE(fd, 0); + + /* Attributes report a page, whatever ->show() ends up producing. */ + ASSERT_EQ(fstat(fd, &st), 0); + EXPECT_EQ(st.st_size, sysconf(_SC_PAGESIZE)); + + n = read(fd, buf, sizeof(buf)); + ASSERT_GT(n, 0); + EXPECT_LT(n, st.st_size); + EXPECT_EQ(read(fd, buf, sizeof(buf)), 0); + + EXPECT_EQ(close(fd), 0); +} + +TEST_F(configfs, attr_write_is_not_incremental) +{ + char buf[64]; + int fd; + + /* + * Every write hands the whole buffer to ->store() and the file + * position is ignored, so the second write replaces the first. + */ + fd = open(CHILDLESS "/storeme", O_WRONLY); + ASSERT_GE(fd, 0); + ASSERT_EQ(write(fd, "1", 1), 1); + ASSERT_EQ(write(fd, "2", 1), 1); + EXPECT_EQ(close(fd), 0); + + ASSERT_GT(read_attr(CHILDLESS "/storeme", buf, sizeof(buf)), 0); + EXPECT_STREQ(buf, "2\n"); +} + +TEST_F(configfs, item_create_and_drop) +{ + struct stat st; + + ASSERT_EQ(mkdir(ITEM_A, 0755), 0); + EXPECT_EQ(stat(ITEM_A "/storeme", &st), 0); + + /* The item carries its own attributes, not the subsystem's. */ + ASSERT_EQ(stat(ITEM_A "/description", &st), -1); + EXPECT_EQ(errno, ENOENT); + + ASSERT_EQ(rmdir(ITEM_A), 0); + ASSERT_EQ(stat(ITEM_A, &st), -1); + EXPECT_EQ(errno, ENOENT); +} + +TEST_F(configfs, item_create_twice) +{ + ASSERT_EQ(mkdir(ITEM_A, 0755), 0); + ASSERT_EQ(mkdir(ITEM_A, 0755), -1); + EXPECT_EQ(errno, EEXIST); +} + +TEST_F(configfs, item_has_no_children) +{ + /* ->make_item() produces an item, so it cannot nest. */ + ASSERT_EQ(mkdir(ITEM_A, 0755), 0); + ASSERT_EQ(mkdir(ITEM_A "/kselftest-b", 0755), -1); + EXPECT_EQ(errno, EPERM); +} + +TEST_F(configfs, item_attrs_are_private) +{ + char buf[64]; + + ASSERT_EQ(mkdir(ITEM_A, 0755), 0); + ASSERT_EQ(mkdir(ITEM_B, 0755), 0); + + ASSERT_GT(write_attr(ITEM_A "/storeme", "11"), 0); + ASSERT_GT(write_attr(ITEM_B "/storeme", "22"), 0); + + ASSERT_GT(read_attr(ITEM_A "/storeme", buf, sizeof(buf)), 0); + EXPECT_STREQ(buf, "11\n"); + ASSERT_GT(read_attr(ITEM_B "/storeme", buf, sizeof(buf)), 0); + EXPECT_STREQ(buf, "22\n"); +} + +TEST_F(configfs, group_create_and_drop) +{ + struct stat st; + + /* 03-group-children hands out groups that take items of their own. */ + ASSERT_EQ(mkdir(GROUP, 0755), 0); + EXPECT_EQ(stat(GROUP "/description", &st), 0); + + ASSERT_EQ(mkdir(GROUP_ITEM, 0755), 0); + EXPECT_EQ(stat(GROUP_ITEM "/storeme", &st), 0); + + ASSERT_EQ(rmdir(GROUP), -1); + EXPECT_EQ(errno, ENOTEMPTY); + + ASSERT_EQ(rmdir(GROUP_ITEM), 0); + ASSERT_EQ(rmdir(GROUP), 0); +} + +TEST_F(configfs, rename_item) +{ + ASSERT_EQ(mkdir(ITEM_A, 0755), 0); + ASSERT_EQ(rename(ITEM_A, ITEM_B), -1); + EXPECT_EQ(errno, EPERM); +} + +TEST_F(configfs, symlink_without_allow_link) +{ + ASSERT_EQ(mkdir(ITEM_A, 0755), 0); + ASSERT_EQ(symlink(ITEM_A, SIMPLE "/kselftest-link"), -1); + EXPECT_EQ(errno, EPERM); +} + +TEST_F(configfs, symlink_and_unlink) +{ + char buf[PATH_MAX]; + struct stat st; + ssize_t n; + + ASSERT_EQ(mkdir(ITEM_A, 0755), 0); + ASSERT_EQ(mkdir(LINK_SRC, 0755), 0); + + ASSERT_EQ(symlink(ITEM_A, LINK), 0); + + /* configfs stores its own body, a path relative to the link. */ + n = readlink(LINK, buf, sizeof(buf) - 1); + ASSERT_GT(n, 0); + buf[n] = '\0'; + EXPECT_STREQ(buf, "../../" ITEM_A); + EXPECT_EQ(stat(LINK "/storeme", &st), 0); + + /* ->allow_link() ran on the source, not on the target. */ + ASSERT_GT(read_attr(LINK_SRC "/nlinks", buf, sizeof(buf)), 0); + EXPECT_STREQ(buf, "1\n"); + + ASSERT_EQ(unlink(LINK), 0); + ASSERT_GT(read_attr(LINK_SRC "/nlinks", buf, sizeof(buf)), 0); + EXPECT_STREQ(buf, "0\n"); +} + +TEST_F(configfs, symlink_pins_both_ends) +{ + ASSERT_EQ(mkdir(ITEM_A, 0755), 0); + ASSERT_EQ(mkdir(LINK_SRC, 0755), 0); + ASSERT_EQ(symlink(ITEM_A, LINK), 0); + + /* A linked item cannot go away under the link. */ + ASSERT_EQ(rmdir(ITEM_A), -1); + EXPECT_EQ(errno, EBUSY); + + /* The link counts as a child of its source. */ + ASSERT_EQ(rmdir(LINK_SRC), -1); + EXPECT_EQ(errno, ENOTEMPTY); + + ASSERT_EQ(unlink(LINK), 0); + EXPECT_EQ(rmdir(ITEM_A), 0); +} + +TEST_F(configfs, symlink_target_outside_configfs) +{ + ASSERT_EQ(mkdir(LINK_SRC, 0755), 0); + ASSERT_EQ(symlink("/", LINK), -1); + EXPECT_EQ(errno, EPERM); +} + +TEST_F(configfs, symlink_target_missing) +{ + ASSERT_EQ(mkdir(LINK_SRC, 0755), 0); + ASSERT_EQ(symlink(SIMPLE "/kselftest-gone", LINK), -1); + EXPECT_EQ(errno, ENOENT); +} + +TEST_F(configfs, symlink_target_is_an_attribute) +{ + ASSERT_EQ(mkdir(LINK_SRC, 0755), 0); + + /* The target is resolved with LOOKUP_DIRECTORY. */ + ASSERT_EQ(symlink(CHILDLESS "/storeme", LINK), -1); + EXPECT_EQ(errno, ENOTDIR); +} + +static volatile int race_stop; + +static void *rmdir_target(void *arg) +{ + while (!race_stop) { + if (mkdir(ITEM_A, 0755) == 0 || errno == EEXIST) + rmdir(ITEM_A); + } + + return NULL; +} + +/* -1 if the kernel does not export a warning counter. */ +static long warn_count(void) +{ + char buf[32]; + + if (read_attr("/sys/kernel/warn_count", buf, sizeof(buf)) < 0) + return -1; + + return strtol(buf, NULL, 10); +} + +TEST_F(configfs, symlink_races_with_target_rmdir) +{ + pthread_t thread; + long warns; + int i; + + warns = warn_count(); + if (warns < 0) + SKIP(return, "no /sys/kernel/warn_count to watch"); + + ASSERT_EQ(mkdir(LINK_SRC, 0755), 0); + ASSERT_EQ(pthread_create(&thread, NULL, rmdir_target, NULL), 0); + + /* + * configfs_rmdir() drops the last reference to the item while its + * dentry is still hashed, and get_target() takes a hashed dentry as + * proof that the item behind it is alive. The symlink then walks + * ->ci_dentry into a released dirent, which configfs_get() warns + * about. KASAN sees the freed item itself. + */ + for (i = 0; i < RACE_ITERATIONS; i++) { + if (symlink(ITEM_A, LINK) == 0) + unlink(LINK); + + /* Give up on the first splat rather than flood the log. */ + if (!(i % 128) && warn_count() != warns) + break; + } + + race_stop = 1; + ASSERT_EQ(pthread_join(thread, NULL), 0); + + EXPECT_EQ(warn_count(), warns) + TH_LOG("kernel warned after %d iterations", i); +} + +TEST_F(configfs, module_pinned_by_item) +{ + ASSERT_EQ(mkdir(ITEM_A, 0755), 0); + + /* mkdir() pins both the subsystem's module and the new item's. */ + ASSERT_EQ(syscall(__NR_delete_module, SAMPLE_MODULE, O_NONBLOCK), -1); + if (errno == ENOSYS) + SKIP(return, "kernel built without CONFIG_MODULE_UNLOAD"); + EXPECT_EQ(errno, EWOULDBLOCK); +} + +TEST_HARNESS_MAIN diff --git a/tools/testing/selftests/filesystems/file_stressor/.gitignore b/tools/testing/selftests/filesystems/file_stressor/.gitignore new file mode 100644 index 000000000000..1eb3f40077d3 --- /dev/null +++ b/tools/testing/selftests/filesystems/file_stressor/.gitignore @@ -0,0 +1,2 @@ +# SPDX-License-Identifier: GPL-2.0-only +file_stressor diff --git a/tools/testing/selftests/filesystems/file_stressor/Makefile b/tools/testing/selftests/filesystems/file_stressor/Makefile new file mode 100644 index 000000000000..88c8231ac144 --- /dev/null +++ b/tools/testing/selftests/filesystems/file_stressor/Makefile @@ -0,0 +1,6 @@ +# SPDX-License-Identifier: GPL-2.0 + +CFLAGS += $(KHDR_INCLUDES) +TEST_GEN_PROGS := file_stressor + +include ../../lib.mk diff --git a/tools/testing/selftests/filesystems/file_stressor.c b/tools/testing/selftests/filesystems/file_stressor/file_stressor.c index 141badd671a9..141badd671a9 100644 --- a/tools/testing/selftests/filesystems/file_stressor.c +++ b/tools/testing/selftests/filesystems/file_stressor/file_stressor.c diff --git a/tools/testing/selftests/filesystems/file_stressor/settings b/tools/testing/selftests/filesystems/file_stressor/settings new file mode 100644 index 000000000000..b675ca93f936 --- /dev/null +++ b/tools/testing/selftests/filesystems/file_stressor/settings @@ -0,0 +1,3 @@ +# Timeout for file_stressor test +# The test runs for 900 seconds (15 minutes) plus setup/teardown time +timeout=1800 diff --git a/tools/testing/selftests/filesystems/fscontext_ns/.gitignore b/tools/testing/selftests/filesystems/fscontext_ns/.gitignore new file mode 100644 index 000000000000..a632905257a3 --- /dev/null +++ b/tools/testing/selftests/filesystems/fscontext_ns/.gitignore @@ -0,0 +1,2 @@ +# SPDX-License-Identifier: GPL-2.0-only +fscontext_ns_test diff --git a/tools/testing/selftests/filesystems/fuse/.gitignore b/tools/testing/selftests/filesystems/fuse/.gitignore index fb51603fe419..b5b03db1118c 100644 --- a/tools/testing/selftests/filesystems/fuse/.gitignore +++ b/tools/testing/selftests/filesystems/fuse/.gitignore @@ -1,4 +1,5 @@ # SPDX-License-Identifier: GPL-2.0-only fuse_mnt fusectl_test +test_syncfs write_extend_eof_test diff --git a/tools/testing/selftests/filesystems/fuse/Makefile b/tools/testing/selftests/filesystems/fuse/Makefile index 95a1ee947ca7..c2de8d225447 100644 --- a/tools/testing/selftests/filesystems/fuse/Makefile +++ b/tools/testing/selftests/filesystems/fuse/Makefile @@ -2,7 +2,7 @@ CFLAGS += -Wall -O2 -g $(KHDR_INCLUDES) -TEST_GEN_PROGS := fusectl_test +TEST_GEN_PROGS := fusectl_test test_syncfs TEST_GEN_PROGS += write_extend_eof_test TEST_GEN_FILES := fuse_mnt diff --git a/tools/testing/selftests/filesystems/kernfs_test.c b/tools/testing/selftests/filesystems/kernfs_test.c index 84c2b910a60d..f4bcf3caf1d5 100644 --- a/tools/testing/selftests/filesystems/kernfs_test.c +++ b/tools/testing/selftests/filesystems/kernfs_test.c @@ -2,9 +2,23 @@ #define _GNU_SOURCE #define __SANE_USERSPACE_TYPES__ +#include <dirent.h> +#include <errno.h> #include <fcntl.h> +#include <limits.h> +#include <net/if.h> +#include <sched.h> +#include <signal.h> #include <stdio.h> +#include <stdlib.h> +#include <string.h> +#include <time.h> +#include <unistd.h> +#include <sys/ioctl.h> +#include <sys/mount.h> +#include <sys/socket.h> #include <sys/stat.h> +#include <sys/syscall.h> #include <sys/xattr.h> #include "kselftest_harness.h" @@ -12,12 +26,24 @@ TEST(kernfs_listxattr) { + ssize_t len; int fd; - /* Read-only file that can never have any extended attributes set. */ + /* Read-only file that can never have any extended attributes set. + * However, on systems with SELinux enabled, security.selinux xattr + * may be present. Skip the content check if any xattrs are found. + */ fd = open("/sys/kernel/warn_count", O_RDONLY | O_CLOEXEC); ASSERT_GE(fd, 0); - ASSERT_EQ(flistxattr(fd, NULL, 0), 0); + + len = flistxattr(fd, NULL, 0); + ASSERT_GE(len, 0); + + if (len > 0) { + close(fd); + SKIP(return, "xattrs present on /sys/kernel/warn_count, skipping xattr content check"); + } + EXPECT_EQ(close(fd), 0); } @@ -34,5 +60,1176 @@ TEST(kernfs_getxattr) EXPECT_EQ(close(fd), 0); } -TEST_HARNESS_MAIN +/* + * Exercise the kernfs dentry cache: lookup, revalidation of positive and + * negative dentries, readdir and namespace tagging. + * + * These drive kernfs from kernel context rather than VFS create/unlink, + * which is what ->d_revalidate() exists for: writing cgroup.subtree_control + * adds and removes files in every child cgroup with no VFS operation + * touching those names. + */ + +#define CG_SCRATCH "kernfs_selftest" +#define TEST_IFNAME "kfstest0" + +/* + * Controllers that add a file to each child cgroup when enabled. The probe + * file must be owned by the controller: cgroup_base_files[] entries such as + * cpu.stat exist in every cgroup regardless, and every file the cpu + * controller does own is behind a Kconfig symbol, so cpu is not usable here. + */ +static const struct { + const char *name; + const char *probe_file; +} controllers[] = { + { "memory", "memory.current" }, + { "pids", "pids.current" }, +}; + +static int find_cgroup2_root(char *buf, size_t len) +{ + char line[PATH_MAX * 2]; + FILE *f; + int ret = -1; + + f = fopen("/proc/self/mounts", "re"); + if (!f) + return -1; + + while (fgets(line, sizeof(line), f)) { + char mnt[PATH_MAX], type[64]; + + /* Octal escaping can expand a path fourfold; bound both %s. */ + if (sscanf(line, "%*s %4095s %63s", mnt, type) != 2) + continue; + if (strcmp(type, "cgroup2")) + continue; + if (strlen(mnt) >= len) + break; + strcpy(buf, mnt); + ret = 0; + break; + } + + fclose(f); + return ret; +} + +static int write_file(const char *path, const char *val) +{ + ssize_t len = strlen(val); + int fd, ret; + + fd = open(path, O_WRONLY | O_CLOEXEC); + if (fd < 0) + return -1; + ret = write(fd, val, len) == len ? 0 : -1; + close(fd); + return ret; +} + +static bool file_has_word(const char *path, const char *word) +{ + char buf[4096], *tok, *save; + bool found = false; + ssize_t n; + int fd; + + fd = open(path, O_RDONLY | O_CLOEXEC); + if (fd < 0) + return false; + n = read(fd, buf, sizeof(buf) - 1); + close(fd); + if (n < 0) + return false; + buf[n] = '\0'; + + for (tok = strtok_r(buf, "\n ", &save); tok; + tok = strtok_r(NULL, "\n ", &save)) { + if (!strcmp(tok, word)) { + found = true; + break; + } + } + return found; +} + +static bool path_is_mounted(const char *path) +{ + char line[PATH_MAX * 2]; + bool found = false; + FILE *f; + + f = fopen("/proc/self/mounts", "re"); + if (!f) + return false; + while (fgets(line, sizeof(line), f)) { + char mnt[PATH_MAX]; + + if (sscanf(line, "%*s %4095s", mnt) != 1) + continue; + if (!strcmp(mnt, path)) { + found = true; + break; + } + } + fclose(f); + return found; +} + +/* Shared by the stress tests below. */ +static bool stress_deadline(const struct timespec *end) +{ + struct timespec now; + + clock_gettime(CLOCK_MONOTONIC, &now); + return now.tv_sec > end->tv_sec || + (now.tv_sec == end->tv_sec && now.tv_nsec >= end->tv_nsec); +} + +FIXTURE(kernfs_cgroup) +{ + char scratch[PATH_MAX]; /* <cg2>/kernfs_selftest.<pid> */ + char child[PATH_MAX]; /* <scratch>/child */ + char probe[PATH_MAX]; /* child's controller file */ + char scratch_sc[PATH_MAX]; /* scratch's cgroup.subtree_control */ + char root_sc[PATH_MAX]; /* root's cgroup.subtree_control */ + char enable[32]; /* "+<controller>" */ + char disable[32]; /* "-<controller>" */ + char mnt[PATH_MAX]; /* our own mount, if we made one */ + bool mounted; + bool enabled_at_root; +}; + +/* A cgroup stays busy briefly after its last task exits. */ +static void rmdir_retry(const char *path) +{ + int i; + + for (i = 0; i < 500; i++) { + if (!rmdir(path) || errno != EBUSY) + return; + usleep(10000); + } +} + +/* + * Undo whatever SETUP managed to do. The harness skips TEARDOWN after a + * failed or skipped SETUP, so SETUP must call this before returning early. + */ +static void kernfs_cgroup_undo(FIXTURE_DATA(kernfs_cgroup) *self) +{ + rmdir_retry(self->child); + rmdir(self->scratch); + if (self->enabled_at_root) + write_file(self->root_sc, self->disable); + if (self->mounted) { + umount2(self->mnt, MNT_DETACH); + rmdir(self->mnt); + } + self->enabled_at_root = false; + self->mounted = false; +} + +FIXTURE_SETUP(kernfs_cgroup) +{ + char root[PATH_MAX], ctl[PATH_MAX]; + const char *probe_file = NULL; + size_t i; + + if (geteuid()) + SKIP(return, "test needs to run as root"); + + /* + * A private mount namespace stops our mounts leaking, but does not + * isolate the cgroup hierarchy: cgroup2 has one default hierarchy + * however many times it is mounted. The scratch cgroups live in the + * host's and must be removed, not discarded with the namespace. + */ + if (unshare(CLONE_NEWNS)) + SKIP(return, "unshare(CLONE_NEWNS): %s", strerror(errno)); + if (mount(NULL, "/", NULL, MS_REC | MS_PRIVATE, NULL)) + SKIP(return, "make / private: %s", strerror(errno)); + + /* Use an existing cgroup2 mount if there is one, else make our own. */ + if (find_cgroup2_root(root, sizeof(root))) { + strcpy(self->mnt, "/tmp/kernfs_selftest_cg2.XXXXXX"); + if (!mkdtemp(self->mnt)) + SKIP(return, "mkdtemp: %s", strerror(errno)); + if (mount("none", self->mnt, "cgroup2", 0, NULL)) { + rmdir(self->mnt); + SKIP(return, "mount cgroup2: %s", strerror(errno)); + } + self->mounted = true; + strcpy(root, self->mnt); + } + + snprintf(self->root_sc, sizeof(self->root_sc), + "%s/cgroup.subtree_control", root); + snprintf(ctl, sizeof(ctl), "%s/cgroup.controllers", root); + + /* Named after our pid so we cannot collide with anything else. */ + snprintf(self->scratch, sizeof(self->scratch), "%s/%s.%d", root, + CG_SCRATCH, getpid()); + snprintf(self->child, sizeof(self->child), "%s/child", self->scratch); + snprintf(self->scratch_sc, sizeof(self->scratch_sc), + "%s/cgroup.subtree_control", self->scratch); + + for (i = 0; i < ARRAY_SIZE(controllers); i++) { + if (!file_has_word(ctl, controllers[i].name)) + continue; + snprintf(self->enable, sizeof(self->enable), "+%s", + controllers[i].name); + snprintf(self->disable, sizeof(self->disable), "-%s", + controllers[i].name); + probe_file = controllers[i].probe_file; + + /* + * A controller must be in the root's subtree_control before + * it appears in our scratch cgroup. Note if we enabled it, + * so it can be put back. + */ + self->enabled_at_root = !file_has_word(self->root_sc, + controllers[i].name); + if (self->enabled_at_root && + write_file(self->root_sc, self->enable)) { + self->enabled_at_root = false; + probe_file = NULL; + continue; + } + break; + } + if (!probe_file) { + kernfs_cgroup_undo(self); + SKIP(return, "no usable cgroup2 controller"); + } + + snprintf(self->probe, sizeof(self->probe), "%s/%s", self->child, + probe_file); + + /* + * Only an unusable environment may skip. A scratch cgroup named + * after our own pid should always be creatable, so failing to make + * one is a result -- skipping would let a broken kernel look green. + */ + if (mkdir(self->scratch, 0755)) { + int err = errno; + + kernfs_cgroup_undo(self); + if (err == EROFS || err == EACCES || err == EPERM) + SKIP(return, "mkdir %s: %s", self->scratch, + strerror(err)); + ASSERT_EQ(err, 0) TH_LOG("mkdir %s: %s", self->scratch, + strerror(err)); + } + if (mkdir(self->child, 0755)) { + int err = errno; + + kernfs_cgroup_undo(self); + ASSERT_EQ(err, 0) TH_LOG("mkdir %s: %s", self->child, + strerror(err)); + } + + /* + * The tests below need the probe file to appear and disappear with + * the controller, so it must be absent now, before anything enables + * it in the scratch cgroup. Skip rather than fail: a probe file that + * is already there means the table names one the controller does not + * own, not that the kernel is broken. + */ + if (!access(self->probe, F_OK)) { + kernfs_cgroup_undo(self); + SKIP(return, "%s is not owned by the %s controller", + probe_file, self->enable + 1); + } +} + +FIXTURE_TEARDOWN(kernfs_cgroup) +{ + write_file(self->scratch_sc, self->disable); + kernfs_cgroup_undo(self); +} + +/* + * Walking already-cached dentries must not invalidate them. Spurious + * invalidation is not merely slow: d_invalidate() calls detach_mounts(), so + * an unrelated lookup would silently tear down any mount below. + */ +TEST_F(kernfs_cgroup, path_walk_does_not_invalidate) +{ + char src[] = "/tmp/kernfs_selftest_bind.XXXXXX"; + char sub[PATH_MAX], probe[PATH_MAX]; + int i; + + snprintf(sub, sizeof(sub), "%s/sub", self->child); + ASSERT_EQ(mkdir(sub, 0755), 0); + + if (!mkdtemp(src)) { + rmdir(sub); + SKIP(return, "mkdtemp: %s", strerror(errno)); + } + if (mount(src, sub, NULL, MS_BIND, NULL)) { + int err = errno; + + rmdir(sub); + rmdir(src); + SKIP(return, "bind mount onto a cgroup dir: %s", strerror(err)); + } + ASSERT_TRUE(path_is_mounted(sub)); + + /* Walk a sibling path through the same directory, repeatedly. */ + snprintf(probe, sizeof(probe), "%s/cgroup.procs", self->child); + for (i = 0; i < 8; i++) { + int fd = open(probe, O_RDONLY | O_CLOEXEC); + + if (fd >= 0) + close(fd); + } + + EXPECT_TRUE(path_is_mounted(sub)); + + umount2(sub, MNT_DETACH); + rmdir(sub); + rmdir(src); +} + +/* + * A cached negative dentry must be invalidated when the kernel creates the + * name behind the dcache's back. That is what kernfs_dir_changed() and + * kernfs_elem_dir::rev are for. + */ +TEST_F(kernfs_cgroup, negative_dentry_invalidated_by_kernel_create) +{ + struct stat st; + + /* Caches a negative dentry for the probe file. */ + ASSERT_EQ(stat(self->probe, &st), -1); + ASSERT_EQ(errno, ENOENT); + + /* The kernel now creates it, with no VFS operation on that name. */ + ASSERT_EQ(write_file(self->scratch_sc, self->enable), 0); + + EXPECT_EQ(stat(self->probe, &st), 0); +} + +/* The mirror image: a cached positive dentry must go when the node does. */ +TEST_F(kernfs_cgroup, positive_dentry_invalidated_by_kernel_remove) +{ + struct stat st; + + ASSERT_EQ(write_file(self->scratch_sc, self->enable), 0); + /* Caches a positive dentry. */ + ASSERT_EQ(stat(self->probe, &st), 0); + + ASSERT_EQ(write_file(self->scratch_sc, self->disable), 0); + + ASSERT_EQ(stat(self->probe, &st), -1); + EXPECT_EQ(errno, ENOENT); +} + +/* Opening a removed node fails; it never returns stale content. */ +TEST_F(kernfs_cgroup, open_after_rmdir_fails) +{ + char path[PATH_MAX]; + char buf[64]; + int fd; + + snprintf(path, sizeof(path), "%s/cgroup.procs", self->child); + + fd = open(path, O_RDONLY | O_CLOEXEC); + ASSERT_GE(fd, 0); + + ASSERT_EQ(rmdir(self->child), 0); + + /* Lookup by path must fail. */ + EXPECT_EQ(open(path, O_RDONLY | O_CLOEXEC), -1); + EXPECT_EQ(errno, ENOENT); + + /* + * An fd held across removal must fail rather than return stale + * content. rmdir() deactivates the node before it returns, so + * kernfs_seq_start() fails to get an active reference. + */ + EXPECT_LT(read(fd, buf, sizeof(buf)), 0); + EXPECT_EQ(errno, ENODEV); + EXPECT_EQ(close(fd), 0); + + ASSERT_EQ(mkdir(self->child, 0755), 0); +} + +/* readdir returns every entry exactly once. */ +TEST_F(kernfs_cgroup, readdir_no_duplicates) +{ + char names[512][NAME_MAX + 1]; + struct dirent *de; + int n = 0, i, j; + DIR *d; + + ASSERT_EQ(write_file(self->scratch_sc, self->enable), 0); + + d = opendir(self->child); + ASSERT_NE(d, NULL); + while ((de = readdir(d))) { + if (!strcmp(de->d_name, ".") || !strcmp(de->d_name, "..")) + continue; + ASSERT_LT(n, (int)ARRAY_SIZE(names)); + strncpy(names[n], de->d_name, NAME_MAX); + names[n][NAME_MAX] = '\0'; + n++; + } + closedir(d); + + ASSERT_GT(n, 0); + for (i = 0; i < n; i++) + for (j = i + 1; j < n; j++) + EXPECT_STRNE(names[i], names[j]); +} + +#define RESUME_DIRS 24 + +/* + * Resuming at an entry that has gone must carry on after it, never before. + * Take a cookie for every entry, then remove each one, seek to its cookie + * and read the rest; nothing already reported may come back. + */ +TEST_F(kernfs_cgroup, readdir_resume_at_removed_entry) +{ + /* The cgroup's own control files are listed alongside ours. */ + char names[128][NAME_MAX + 1]; + long pos[128]; + char path[PATH_MAX]; + struct dirent *de; + int n = 0, i, j; + DIR *d; + + for (i = 0; i < RESUME_DIRS; i++) { + snprintf(path, sizeof(path), "%s/e%02d", self->scratch, i); + ASSERT_EQ(mkdir(path, 0755), 0); + } + + /* Record the cookie before reading each entry, with its name. */ + d = opendir(self->scratch); + ASSERT_NE(d, NULL); + while (1) { + long here = telldir(d); + + de = readdir(d); + if (!de) + break; + if (!strcmp(de->d_name, ".") || !strcmp(de->d_name, "..")) + continue; + ASSERT_LT(n, (int)ARRAY_SIZE(pos)); + pos[n] = here; + strncpy(names[n], de->d_name, NAME_MAX); + names[n][NAME_MAX] = '\0'; + n++; + } + closedir(d); + ASSERT_GT(n, 1); + + for (i = 0; i < n; i++) { + /* Only the directories we made can be removed and put back. */ + if (strncmp(names[i], "e", 1)) + continue; + + snprintf(path, sizeof(path), "%s/%s", self->scratch, names[i]); + ASSERT_EQ(rmdir(path), 0); + + /* Reopen so the seek has to reach the kernel. */ + d = opendir(self->scratch); + ASSERT_NE(d, NULL); + seekdir(d, pos[i]); + while ((de = readdir(d))) { + if (!strcmp(de->d_name, ".") || !strcmp(de->d_name, "..")) + continue; + for (j = 0; j < i; j++) + ASSERT_STRNE(de->d_name, names[j]) + TH_LOG("resuming at %s (gone) went back to %s", + names[i], names[j]); + } + closedir(d); + + ASSERT_EQ(mkdir(path, 0755), 0); + } + + for (i = 0; i < RESUME_DIRS; i++) { + snprintf(path, sizeof(path), "%s/e%02d", self->scratch, i); + EXPECT_EQ(rmdir(path), 0); + } +} + +#define CHURN_ROUNDS 400 +#define CHURN_BUFSZ 512 /* small, so a listing takes several calls */ + +/* + * The files appear at the end of the enabling write and go at the start of + * the disabling one, so the window where they exist is the short one. + */ +#define CHURN_DWELL_ON 2000 +#define CHURN_DWELL_OFF 200 + +struct kernfs_dirent64 { + unsigned long long d_ino; + long long d_off; + unsigned short d_reclen; + unsigned char d_type; + char d_name[]; +}; + +/* + * The same resume, but inside one getdents(2) call. rmdir(2) cannot reach + * that window because iterate_dir() holds the listed directory's i_rwsem + * for the whole listing; cgroup.subtree_control can, having no VFS + * operation on the names it adds and removes. The files that are not the + * controller's stay throughout, so each must appear exactly once. + * + * A stress test: it has not been seen to catch the ordering bug, and is + * here to keep the unlocked window under load for lockdep and KASAN. + */ +TEST_F(kernfs_cgroup, readdir_resume_vs_internal_remove) +{ + char buf[CHURN_BUFSZ] __attribute__((aligned(8))); + char stable[128][NAME_MAX + 1]; + int nstable = 0, i, r; + int withctl = 0, without = 0; + int seen[128], status; + pid_t churner; + DIR *d; + + /* With the controller off, whatever is left is what must persist. */ + ASSERT_EQ(write_file(self->scratch_sc, self->disable), 0); + d = opendir(self->child); + ASSERT_NE(d, NULL); + for (;;) { + struct dirent *de = readdir(d); + + if (!de) + break; + if (!strcmp(de->d_name, ".") || !strcmp(de->d_name, "..")) + continue; + ASSERT_LT(nstable, (int)ARRAY_SIZE(stable)); + strncpy(stable[nstable], de->d_name, NAME_MAX); + stable[nstable][NAME_MAX] = '\0'; + nstable++; + } + closedir(d); + ASSERT_GT(nstable, 0); + + churner = fork(); + ASSERT_GE(churner, 0); + if (churner == 0) { + for (;;) { + if (write_file(self->scratch_sc, self->enable)) + _exit(10); + usleep(CHURN_DWELL_ON); + if (write_file(self->scratch_sc, self->disable)) + _exit(11); + usleep(CHURN_DWELL_OFF); + } + } + + for (r = 0; r < CHURN_ROUNDS; r++) { + int fd = open(self->child, O_RDONLY | O_DIRECTORY); + int extra = 0; + int n; + ASSERT_GE(fd, 0); + memset(seen, 0, sizeof(seen)); + + while ((n = syscall(SYS_getdents64, fd, buf, sizeof(buf))) > 0) { + int off = 0; + + while (off < n) { + struct kernfs_dirent64 *de = (void *)(buf + off); + bool known = false; + + off += de->d_reclen; + for (i = 0; i < nstable; i++) + if (!strcmp(de->d_name, stable[i])) { + seen[i]++; + known = true; + } + if (!known && strcmp(de->d_name, ".") && + strcmp(de->d_name, "..")) + extra++; + } + } + ASSERT_GE(n, 0); + EXPECT_EQ(close(fd), 0); + + if (extra) + withctl++; + else + without++; + + for (i = 0; i < nstable; i++) + ASSERT_EQ(seen[i], 1) + TH_LOG("round %d: %s seen %d times", + r, stable[i], seen[i]); + } + + /* The churn must have been running, or the listings prove nothing. */ + EXPECT_EQ(kill(churner, SIGKILL), 0); + ASSERT_EQ(waitpid(churner, &status, 0), churner); + ASSERT_TRUE(WIFSIGNALED(status) && WTERMSIG(status) == SIGKILL) + TH_LOG("churner exited on its own: status %d", status); + + /* + * They also have to have overlapped it. How much depends on the + * machine, so say the race could not be arranged rather than fail. + */ + if (!withctl || !without) + SKIP(return, "listings did not span the churn: %d with, %d without", + withctl, without); +} + +/* + * A telldir() cookie must resolve back to the same entry after seekdir(). + * kernfs encodes the cookie as the node's name hash, so this covers + * kernfs_dir_pos() as well as plain iteration. + */ +TEST_F(kernfs_cgroup, readdir_seekdir_roundtrip) +{ + char names[512][NAME_MAX + 1]; + struct dirent *de; + long pos[512]; + int n = 0, i; + DIR *d; + + ASSERT_EQ(write_file(self->scratch_sc, self->enable), 0); + + d = opendir(self->child); + ASSERT_NE(d, NULL); + + /* Record the cookie *before* reading each entry, with its name. */ + while (1) { + long here = telldir(d); + + de = readdir(d); + if (!de) + break; + if (!strcmp(de->d_name, ".") || !strcmp(de->d_name, "..")) + continue; + ASSERT_LT(n, (int)ARRAY_SIZE(pos)); + pos[n] = here; + strncpy(names[n], de->d_name, NAME_MAX); + names[n][NAME_MAX] = '\0'; + n++; + } + ASSERT_GT(n, 0); + + /* Seeking back to a cookie must land on the entry it was taken at. */ + for (i = 0; i < n; i++) { + seekdir(d, pos[i]); + de = readdir(d); + ASSERT_NE(de, NULL); + EXPECT_STREQ(de->d_name, names[i]); + } + + closedir(d); +} + +#define STRESS_SECS 2 +#define STRESS_DIRS 4 +#define STRESS_READERS 4 + +/* + * Hammer lookup against creation and removal. Revalidation holds no lock + * against the writers, so what makes it safe is that every answer it can + * give is one the caller already handles: a reader must only ever see + * success or an errno meaning "it went away", never garbage or a hang. + */ +TEST_F(kernfs_cgroup, lookup_vs_create_remove_stress) +{ + pid_t pids[STRESS_DIRS + STRESS_READERS]; + struct timespec end; + int i, status, n = 0; + + clock_gettime(CLOCK_MONOTONIC, &end); + end.tv_sec += STRESS_SECS; + + for (i = 0; i < STRESS_DIRS; i++) { + pid_t pid = fork(); + + ASSERT_GE(pid, 0); + if (pid == 0) { + char dir[PATH_MAX]; + + snprintf(dir, sizeof(dir), "%s/s%d", self->scratch, i); + while (!stress_deadline(&end)) { + if (mkdir(dir, 0755) && errno != EEXIST) + _exit(10); + if (rmdir(dir) && errno != ENOENT && + errno != EBUSY) + _exit(11); + } + _exit(0); + } + pids[n++] = pid; + } + + for (i = 0; i < STRESS_READERS; i++) { + pid_t pid = fork(); + + ASSERT_GE(pid, 0); + if (pid == 0) { + /* Start each reader on a different directory. */ + unsigned int seq = i; + + while (!stress_deadline(&end)) { + int which = seq++ % STRESS_DIRS; + char path[PATH_MAX]; + struct stat st; + int fd; + + snprintf(path, sizeof(path), + "%s/s%d/cgroup.procs", + self->scratch, which); + + if (stat(path, &st) && errno != ENOENT && + errno != ENODEV) + _exit(20); + + fd = open(path, O_RDONLY | O_CLOEXEC); + if (fd < 0) { + if (errno != ENOENT && errno != ENODEV) + _exit(21); + } else { + close(fd); + } + + if (access(path, F_OK) && errno != ENOENT && + errno != ENODEV) + _exit(22); + } + _exit(0); + } + pids[n++] = pid; + } + + for (i = 0; i < n; i++) { + ASSERT_EQ(waitpid(pids[i], &status, 0), pids[i]); + ASSERT_TRUE(WIFEXITED(status)); + EXPECT_EQ(WEXITSTATUS(status), 0); + } + + for (i = 0; i < STRESS_DIRS; i++) { + char dir[PATH_MAX]; + + snprintf(dir, sizeof(dir), "%s/s%d", self->scratch, i); + rmdir_retry(dir); + } +} + +struct kernfs_handle { + struct file_handle h; + unsigned char buf[MAX_HANDLE_SZ]; +}; + +static int kernfs_encode(const char *path, struct kernfs_handle *fh) +{ + int mount_id; + + memset(fh, 0, sizeof(*fh)); + fh->h.handle_bytes = sizeof(fh->buf); + return name_to_handle_at(AT_FDCWD, path, &fh->h, &mount_id, 0); +} + +/* + * Skip only where file handles do not work at all. ENOENT must still + * fail: mkdir leaves a negative dentry cached, so the name resolves only + * after ->d_revalidate() drops it. The encode tests revalidation too. + */ +static bool fh_unsupported(int err) +{ + return err == EOPNOTSUPP || err == EPERM || err == ENOSYS; +} + +/* + * Decoding a file needs CAP_DAC_READ_SEARCH in the initial user + * namespace. Probe once so the tests skip instead of fail. + */ +static bool fh_can_decode(int mfd, struct kernfs_handle *fh) +{ + int fd = open_by_handle_at(mfd, &fh->h, O_PATH); + + if (fd < 0) + return errno != EPERM; + close(fd); + return true; +} + +/* + * A file handle reaches a node without a lookup through its parent. A + * live node must decode. A removed one must not, because + * kernfs_find_and_get_node_by_id() refuses inactive nodes. + * + * Use O_PATH: opening a removed node fails with ENODEV, which would hide + * what is being tested. + */ +TEST_F(kernfs_cgroup, exportfs_decode_and_stale) +{ + char victim[PATH_MAX], procs[PATH_MAX]; + struct kernfs_handle fh; + struct stat st; + int mfd, fd; + + snprintf(victim, sizeof(victim), "%s/fh", self->scratch); + snprintf(procs, sizeof(procs), "%s/cgroup.procs", victim); + ASSERT_EQ(mkdir(victim, 0755), 0); + + /* Any fd on the filesystem identifies it to open_by_handle_at(). */ + mfd = open(self->scratch, O_RDONLY | O_DIRECTORY | O_CLOEXEC); + ASSERT_GE(mfd, 0); + + if (kernfs_encode(procs, &fh)) { + int err = errno; + + close(mfd); + rmdir(victim); + ASSERT_TRUE(fh_unsupported(err)) + TH_LOG("name_to_handle_at: %s", strerror(err)); + SKIP(return, "name_to_handle_at: %s", strerror(err)); + } + + if (!fh_can_decode(mfd, &fh)) { + close(mfd); + rmdir(victim); + SKIP(return, "open_by_handle_at: no CAP_DAC_READ_SEARCH"); + } + + fd = open_by_handle_at(mfd, &fh.h, O_PATH); + ASSERT_GE(fd, 0); + EXPECT_EQ(fstat(fd, &st), 0); + EXPECT_EQ(st.st_nlink, 1); + EXPECT_EQ(close(fd), 0); + + ASSERT_EQ(rmdir(victim), 0); + + fd = open_by_handle_at(mfd, &fh.h, O_PATH); + EXPECT_LT(fd, 0); + if (fd >= 0) + close(fd); + else + EXPECT_EQ(errno, ESTALE); + + EXPECT_EQ(close(mfd), 0); +} + +#define FH_STRESS_SECS 2 +#define FH_DECODE_CAP 10000 + +/* + * Decode file handles while the node is being removed. A decode must + * answer with a usable handle or ESTALE, never garbage and never a hang. + * + * The link count is checked too. An inode that reaches the inode hash + * after the removal cleared link counts keeps the 1 it was born with, so + * it never gets an IN_DELETE_SELF. This has not been seen to fire: it + * needs the decode to stall between the lookup by id and the hash insert, + * and nothing there blocks. It is kept because it is cheap and only + * looks once the directory is gone, so it cannot fail falsely. + */ +TEST_F(kernfs_cgroup, exportfs_decode_vs_rmdir_stress) +{ + int mfd, bad = 0, rounds = 0; + struct timespec end; + + mfd = open(self->scratch, O_RDONLY | O_DIRECTORY | O_CLOEXEC); + ASSERT_GE(mfd, 0); + + clock_gettime(CLOCK_MONOTONIC, &end); + end.tv_sec += FH_STRESS_SECS; + + while (!stress_deadline(&end)) { + char victim[PATH_MAX], procs[PATH_MAX]; + int last = -1, fd, i; + struct kernfs_handle fh; + struct stat st; + pid_t pid; + + snprintf(victim, sizeof(victim), "%s/fh%d", self->scratch, + rounds++); + snprintf(procs, sizeof(procs), "%s/cgroup.procs", victim); + if (mkdir(victim, 0755)) + break; + if (kernfs_encode(procs, &fh)) { + int err = errno; + + rmdir(victim); + ASSERT_TRUE(fh_unsupported(err)) + TH_LOG("name_to_handle_at: %s", strerror(err)); + SKIP(goto out, "name_to_handle_at: %s", strerror(err)); + } + if (rounds == 1 && !fh_can_decode(mfd, &fh)) { + rmdir(victim); + SKIP(goto out, + "open_by_handle_at: no CAP_DAC_READ_SEARCH"); + } + + pid = fork(); + ASSERT_GE(pid, 0); + if (pid == 0) { + rmdir_retry(victim); + _exit(0); + } + + /* + * Decode until the removal deactivates the node. Keep the + * last one that worked: it ran closest to the removal. + */ + for (i = 0; i < FH_DECODE_CAP; i++) { + fd = open_by_handle_at(mfd, &fh.h, O_PATH); + if (fd < 0) + break; + if (last >= 0) + close(last); + last = fd; + } + ASSERT_EQ(waitpid(pid, NULL, 0), pid); + + if (last >= 0) { + if (access(victim, F_OK) && errno == ENOENT && + !fstat(last, &st) && st.st_nlink != 0) + bad++; + close(last); + } + rmdir(victim); + } + + EXPECT_EQ(bad, 0) + TH_LOG("%d of %d rounds decoded a removed node whose inode kept its link count", + bad, rounds); +out: + close(mfd); +} + +/* + * sysfs is namespace tagged (KERNFS_NS) and supports rename; cgroup2 does + * neither. Run in a private netns with its own sysfs so the host is + * untouched. + */ +FIXTURE(kernfs_netns) +{ + char mnt[PATH_MAX]; + char net[PATH_MAX]; + bool mounted; +}; + +FIXTURE_SETUP(kernfs_netns) +{ + if (geteuid()) + SKIP(return, "test needs to run as root"); + + if (unshare(CLONE_NEWNS | CLONE_NEWNET)) + SKIP(return, "unshare(CLONE_NEWNS|CLONE_NEWNET): %s", + strerror(errno)); + + /* Don't let our sysfs mount escape into the parent namespace. */ + ASSERT_EQ(mount(NULL, "/", NULL, MS_REC | MS_PRIVATE, NULL), 0); + + strcpy(self->mnt, "/tmp/kernfs_selftest_sysfs.XXXXXX"); + if (!mkdtemp(self->mnt)) + SKIP(return, "mkdtemp: %s", strerror(errno)); + + if (mount("none", self->mnt, "sysfs", 0, NULL)) { + rmdir(self->mnt); + SKIP(return, "mount sysfs: %s", strerror(errno)); + } + self->mounted = true; + + snprintf(self->net, sizeof(self->net), "%s/class/net", self->mnt); +} + +FIXTURE_TEARDOWN(kernfs_netns) +{ + if (self->mounted) + umount2(self->mnt, MNT_DETACH); + rmdir(self->mnt); +} + +/* + * sysfs must show this network namespace's interfaces, not the parent's. + * + * Do not assume a fresh netns contains only "lo": fallback tunnel devices + * (tunl0, sit0, gre0, ...) are created in every namespace unless + * net.core.fb_tunnels_only_for_init_net is set, so which names appear + * depends on the modules the host has. Check the set instead -- + * if_nametoindex() resolves in the current netns, so every name sysfs shows + * must resolve there, and the counts must agree. + * + * Count only symlinks. Not every entry is a device: bonding adds a + * bonding_masters attribute to /sys/class/net in every namespace. + */ +TEST_F(kernfs_netns, ns_tag_isolates_class_net) +{ + struct if_nameindex *idx, *i; + bool found_lo = false; + int n = 0, want = 0; + struct dirent *de; + DIR *d; + + d = opendir(self->net); + ASSERT_NE(d, NULL); + while ((de = readdir(d))) { + if (de->d_type != DT_LNK) + continue; + EXPECT_NE(if_nametoindex(de->d_name), 0u) + TH_LOG("%s is not in this netns", de->d_name); + if (!strcmp(de->d_name, "lo")) + found_lo = true; + n++; + } + closedir(d); + + idx = if_nameindex(); + ASSERT_NE(idx, NULL); + for (i = idx; i->if_index; i++) + want++; + if_freenameindex(idx); + + EXPECT_TRUE(found_lo); + EXPECT_EQ(n, want); +} + +/* + * After a rename the old name must stop resolving and the new one must + * start, even though both dentries are already cached. + */ +TEST_F(kernfs_netns, rename_is_revalidated) +{ + char old_path[PATH_MAX], new_path[PATH_MAX]; + struct ifreq ifr = {}; + struct stat st; + int sk; + + snprintf(old_path, sizeof(old_path), "%s/lo", self->net); + snprintf(new_path, sizeof(new_path), "%s/%s", self->net, TEST_IFNAME); + + /* Warm both dentries: one positive, one negative. */ + ASSERT_EQ(stat(old_path, &st), 0); + ASSERT_EQ(stat(new_path, &st), -1); + ASSERT_EQ(errno, ENOENT); + + sk = socket(AF_INET, SOCK_DGRAM | SOCK_CLOEXEC, 0); + ASSERT_GE(sk, 0); + strcpy(ifr.ifr_name, "lo"); + strcpy(ifr.ifr_newname, TEST_IFNAME); + if (ioctl(sk, SIOCSIFNAME, &ifr)) { + close(sk); + SKIP(return, "SIOCSIFNAME: %s", strerror(errno)); + } + close(sk); + + EXPECT_EQ(stat(old_path, &st), -1); + EXPECT_EQ(errno, ENOENT); + EXPECT_EQ(stat(new_path, &st), 0); +} + +static int netdev_rename(const char *from, const char *to) +{ + struct ifreq ifr = {}; + int sk, ret; + + sk = socket(AF_INET, SOCK_DGRAM | SOCK_CLOEXEC, 0); + if (sk < 0) + return -1; + strncpy(ifr.ifr_name, from, IFNAMSIZ - 1); + strncpy(ifr.ifr_newname, to, IFNAMSIZ - 1); + ret = ioctl(sk, SIOCSIFNAME, &ifr); + close(sk); + return ret; +} + +/* + * Bounded by a count, not by time: every rename is logged and not rate + * limited, so a timed loop would flood the kernel log. + */ +#define RENAME_FLIPS 200 +#define RENAME_READERS 4 + +/* + * Rename an interface while other tasks look up the names it moves + * between. This renames its /sys/class/net entry through + * kernfs_rename_ns() with the parent unchanged. + * + * The renamer checks what is certain: SIOCSIFNAME returns once the rename + * is done and nothing else renames here, so the new name must resolve and + * the old must not. The readers cannot check that, because the name can + * move between their two lstat() calls. They only check that a lookup + * returns success or ENOENT, and keep the lock busy while renames run. + * + * lstat() not stat(): /sys/class/net/<dev> is a symlink and is renamed + * before the directory it points at, so the two are not atomic. + */ +TEST_F(kernfs_netns, rename_vs_lookup_stress) +{ + char old_path[PATH_MAX], new_path[PATH_MAX]; + pid_t pids[RENAME_READERS]; + int i, status, n = 0, bad = 0; + struct stat st; + int done[2]; + + snprintf(old_path, sizeof(old_path), "%s/lo", self->net); + snprintf(new_path, sizeof(new_path), "%s/%s", self->net, TEST_IFNAME); + + if (netdev_rename("lo", TEST_IFNAME)) + SKIP(return, "SIOCSIFNAME: %s", strerror(errno)); + if (netdev_rename(TEST_IFNAME, "lo")) + SKIP(return, "SIOCSIFNAME back: %s", strerror(errno)); + + /* Readers run until the renamer closes the write end. */ + ASSERT_EQ(pipe2(done, O_NONBLOCK | O_CLOEXEC), 0); + + for (i = 0; i < RENAME_READERS; i++) { + pid_t pid = fork(); + + ASSERT_GE(pid, 0); + if (pid == 0) { + struct stat rst; + char c; + + close(done[1]); + while (read(done[0], &c, 1) < 0 && errno == EAGAIN) { + if (lstat(old_path, &rst) && errno != ENOENT) + _exit(20); + if (lstat(new_path, &rst) && errno != ENOENT) + _exit(21); + } + _exit(0); + } + pids[n++] = pid; + } + close(done[0]); + + for (i = 0; i < RENAME_FLIPS; i++) { + if (netdev_rename("lo", TEST_IFNAME)) + break; + if (lstat(new_path, &st) || !lstat(old_path, &st)) { + bad++; + break; + } + if (netdev_rename(TEST_IFNAME, "lo")) + break; + if (lstat(old_path, &st) || !lstat(new_path, &st)) { + bad++; + break; + } + } + close(done[1]); + + for (i = 0; i < n; i++) { + ASSERT_EQ(waitpid(pids[i], &status, 0), pids[i]); + ASSERT_TRUE(WIFEXITED(status)); + EXPECT_EQ(WEXITSTATUS(status), 0); + } + + EXPECT_EQ(bad, 0) + TH_LOG("a completed rename left the wrong name resolving"); + + /* Leave the interface as the fixture found it. */ + netdev_rename(TEST_IFNAME, "lo"); +} + +TEST_HARNESS_MAIN diff --git a/tools/testing/selftests/filesystems/mntns_unbindable/Makefile b/tools/testing/selftests/filesystems/mntns_unbindable/Makefile new file mode 100644 index 000000000000..33a311c5bd72 --- /dev/null +++ b/tools/testing/selftests/filesystems/mntns_unbindable/Makefile @@ -0,0 +1,6 @@ +# SPDX-License-Identifier: GPL-2.0 +TEST_GEN_PROGS := mntns_unbindable_test + +CFLAGS += -Wall -O2 -g $(KHDR_INCLUDES) + +include ../../lib.mk diff --git a/tools/testing/selftests/filesystems/mntns_unbindable/mntns_unbindable_test.c b/tools/testing/selftests/filesystems/mntns_unbindable/mntns_unbindable_test.c new file mode 100644 index 000000000000..9aebc37cf74d --- /dev/null +++ b/tools/testing/selftests/filesystems/mntns_unbindable/mntns_unbindable_test.c @@ -0,0 +1,227 @@ +// SPDX-License-Identifier: GPL-2.0 +/* + * An unbindable mount stays unbindable in a cloned mount namespace. + */ +#define _GNU_SOURCE +#include <errno.h> +#include <fcntl.h> +#include <sched.h> +#include <stdio.h> +#include <stdlib.h> +#include <string.h> +#include <sys/mount.h> +#include <sys/stat.h> +#include <sys/syscall.h> +#include <sys/wait.h> +#include <unistd.h> + +#include "../../kselftest_harness.h" + +#ifndef OPEN_TREE_CLONE +#define OPEN_TREE_CLONE 1 +#endif +#ifndef OPEN_TREE_CLOEXEC +#define OPEN_TREE_CLOEXEC O_CLOEXEC +#endif + +static int sys_open_tree(int dfd, const char *filename, unsigned int flags) +{ + return syscall(__NR_open_tree, dfd, filename, flags); +} + +/* Child exit codes. */ +enum { + CHILD_OK, /* the operation failed with EINVAL as it must */ + CHILD_ALLOWED, /* the operation succeeded: the flag was lost */ + CHILD_UNSHARE, /* unshare(CLONE_NEWNS) failed */ + CHILD_ERRNO, /* the operation failed with some other errno */ + CHILD_MOUNTINFO, /* the mount was not found in mountinfo */ +}; + +FIXTURE(mntns_unbindable) +{ + char base[64]; + char src[80]; + char dst[80]; + bool mounted; +}; + +FIXTURE_SETUP(mntns_unbindable) +{ + self->mounted = false; + + if (geteuid() != 0) + SKIP(return, "test requires CAP_SYS_ADMIN"); + + ASSERT_EQ(unshare(CLONE_NEWNS), 0); + ASSERT_EQ(mount("", "/", NULL, MS_REC | MS_PRIVATE, NULL), 0); + + snprintf(self->base, sizeof(self->base), "/tmp/mntns_unbindable.XXXXXX"); + ASSERT_NE(mkdtemp(self->base), NULL); + ASSERT_EQ(mount("tmpfs", self->base, "tmpfs", 0, NULL), 0); + self->mounted = true; + + snprintf(self->src, sizeof(self->src), "%s/src", self->base); + snprintf(self->dst, sizeof(self->dst), "%s/dst", self->base); + ASSERT_EQ(mkdir(self->src, 0755), 0); + ASSERT_EQ(mkdir(self->dst, 0755), 0); + + ASSERT_EQ(mount("tmpfs", self->src, "tmpfs", 0, NULL), 0); + ASSERT_EQ(mount(NULL, self->src, NULL, MS_UNBINDABLE, NULL), 0); +} + +FIXTURE_TEARDOWN(mntns_unbindable) +{ + if (self->mounted) + umount2(self->base, MNT_DETACH); + rmdir(self->base); +} + +static int classify(int ret, int err) +{ + if (ret >= 0) + return CHILD_ALLOWED; + return err == EINVAL ? CHILD_OK : CHILD_ERRNO; +} + +/* Is the mount on @mountpoint marked unbindable in /proc/self/mountinfo? */ +static int mountinfo_unbindable(const char *mountpoint) +{ + char line[4096]; + FILE *f; + int ret = CHILD_MOUNTINFO; + + f = fopen("/proc/self/mountinfo", "re"); + if (!f) + return CHILD_ERRNO; + + while (fgets(line, sizeof(line), f)) { + char *fields[6], *p = line, *opt; + int i; + + for (i = 0; i < 6; i++) { + fields[i] = strsep(&p, " "); + if (!fields[i]) + break; + } + if (i < 6 || strcmp(fields[4], mountpoint)) + continue; + + /* the optional fields, up to the "-" separator */ + ret = CHILD_ALLOWED; + while ((opt = strsep(&p, " ")) && strcmp(opt, "-")) { + if (!strcmp(opt, "unbindable")) + ret = CHILD_OK; + } + break; + } + fclose(f); + return ret; +} + +static int run_in_child(int (*fn)(const char *src, const char *dst), + const char *src, const char *dst) +{ + int status; + pid_t pid; + + pid = fork(); + if (pid < 0) + return -1; + if (pid == 0) + _exit(fn(src, dst)); + if (waitpid(pid, &status, 0) != pid || !WIFEXITED(status)) + return -1; + return WEXITSTATUS(status); +} + +static int bind_after_clone(const char *src, const char *dst) +{ + int ret; + + if (unshare(CLONE_NEWNS)) + return CHILD_UNSHARE; + ret = mount(src, dst, NULL, MS_BIND, NULL); + return classify(ret, errno); +} + +static int rbind_after_clone(const char *src, const char *dst) +{ + int ret; + + if (unshare(CLONE_NEWNS)) + return CHILD_UNSHARE; + ret = mount(src, dst, NULL, MS_BIND | MS_REC, NULL); + return classify(ret, errno); +} + +static int open_tree_after_clone(const char *src, const char *dst) +{ + int ret; + + if (unshare(CLONE_NEWNS)) + return CHILD_UNSHARE; + ret = sys_open_tree(AT_FDCWD, src, OPEN_TREE_CLONE | OPEN_TREE_CLOEXEC); + return classify(ret, errno); +} + +static int mountinfo_after_clone(const char *src, const char *dst) +{ + if (unshare(CLONE_NEWNS)) + return CHILD_UNSHARE; + return mountinfo_unbindable(src); +} + +static int bind_after_two_clones(const char *src, const char *dst) +{ + int ret; + + if (unshare(CLONE_NEWNS)) + return CHILD_UNSHARE; + if (unshare(CLONE_NEWNS)) + return CHILD_UNSHARE; + ret = mount(src, dst, NULL, MS_BIND, NULL); + return classify(ret, errno); +} + +/* The namespace the mount was made unbindable in. */ +TEST_F(mntns_unbindable, refuses_bind) +{ + int ret = mount(self->src, self->dst, NULL, MS_BIND, NULL); + + ASSERT_EQ(classify(ret, errno), CHILD_OK); + ASSERT_EQ(mountinfo_unbindable(self->src), CHILD_OK); +} + +/* A copy of the namespace must not turn the mount bindable. */ +TEST_F(mntns_unbindable, refuses_bind_after_clone) +{ + ASSERT_EQ(run_in_child(bind_after_clone, self->src, self->dst), CHILD_OK) + TH_LOG("bind of an unbindable mount allowed in a copied mount namespace"); +} + +TEST_F(mntns_unbindable, refuses_rbind_after_clone) +{ + ASSERT_EQ(run_in_child(rbind_after_clone, self->src, self->dst), CHILD_OK) + TH_LOG("rbind of an unbindable mount allowed in a copied mount namespace"); +} + +TEST_F(mntns_unbindable, refuses_open_tree_after_clone) +{ + ASSERT_EQ(run_in_child(open_tree_after_clone, self->src, self->dst), CHILD_OK) + TH_LOG("OPEN_TREE_CLONE of an unbindable mount allowed in a copied mount namespace"); +} + +TEST_F(mntns_unbindable, mountinfo_after_clone) +{ + ASSERT_EQ(run_in_child(mountinfo_after_clone, self->src, self->dst), CHILD_OK) + TH_LOG("mountinfo does not show the mount as unbindable in a copied mount namespace"); +} + +TEST_F(mntns_unbindable, refuses_bind_after_two_clones) +{ + ASSERT_EQ(run_in_child(bind_after_two_clones, self->src, self->dst), CHILD_OK) + TH_LOG("bind of an unbindable mount allowed two mount namespace copies down"); +} + +TEST_HARNESS_MAIN diff --git a/tools/testing/selftests/filesystems/openat2/openat2_test.c b/tools/testing/selftests/filesystems/openat2/openat2_test.c index 6f5afbe2d8d3..e08c94ce0530 100644 --- a/tools/testing/selftests/filesystems/openat2/openat2_test.c +++ b/tools/testing/selftests/filesystems/openat2/openat2_test.c @@ -23,8 +23,16 @@ * XXX: This is wrong on {mips, parisc, powerpc, sparc}. */ #undef O_LARGEFILE -#ifdef __aarch64__ +#if defined(__aarch64__) || defined(__alpha__) #define O_LARGEFILE 0x20000 +#elif defined(__powerpc__) || defined(__ppc__) +#define O_LARGEFILE 0x10000 +#elif defined(__sparc__) +#define O_LARGEFILE 0x40000 +#elif defined(__mips__) +#define O_LARGEFILE 0x2000 +#elif defined(__parisc__) +#define O_LARGEFILE 0x800 #else #define O_LARGEFILE 0x8000 #endif diff --git a/tools/testing/selftests/filesystems/openat2/resolve_test.c b/tools/testing/selftests/filesystems/openat2/resolve_test.c index eacde59ce158..6216a8546388 100644 --- a/tools/testing/selftests/filesystems/openat2/resolve_test.c +++ b/tools/testing/selftests/filesystems/openat2/resolve_test.c @@ -140,9 +140,12 @@ FIXTURE_SETUP(openat2_resolve) if (!openat2_supported) SKIP(return, "openat2(2) not supported"); - /* Unshare and make /tmp a new directory. */ + /* Unshare and make the mount tree private. */ ASSERT_EQ(unshare(CLONE_NEWNS), 0); - ASSERT_EQ(mount("", "/tmp", "", MS_PRIVATE, ""), 0); + ASSERT_EQ(mount("", "/", "", MS_PRIVATE | MS_REC, ""), 0); + + /* Ensure /tmp is a mountpoint for RESOLVE_NO_XDEV test crossing into /tmp. */ + ASSERT_EQ(mount("/tmp", "/tmp", NULL, MS_BIND, NULL), 0); /* Make the top-level directory. */ ASSERT_NE(mkdtemp(dirname), NULL); diff --git a/tools/testing/selftests/filesystems/statmount/statmount_test.c b/tools/testing/selftests/filesystems/statmount/statmount_test.c index 60c2c544db6a..52544512edc7 100644 --- a/tools/testing/selftests/filesystems/statmount/statmount_test.c +++ b/tools/testing/selftests/filesystems/statmount/statmount_test.c @@ -17,7 +17,7 @@ static const char *const known_fs[] = { "9p", "adfs", "affs", "afs", "aio", "anon_inodefs", "apparmorfs", - "autofs", "bcachefs", "bdev", "befs", "bfs", "binder", "binfmt_misc", + "autofs", "bcachefs", "bdev", "befs", "binder", "binfmt_misc", "bpf", "btrfs", "btrfs_test_fs", "ceph", "cgroup", "cgroup2", "cifs", "coda", "configfs", "cpuset", "cramfs", "cxl", "dax", "debugfs", "devpts", "devtmpfs", "dmabuf", "drm", "ecryptfs", "efivarfs", "efs", diff --git a/tools/testing/selftests/filesystems/umount_propagation/Makefile b/tools/testing/selftests/filesystems/umount_propagation/Makefile new file mode 100644 index 000000000000..fc0a0783018b --- /dev/null +++ b/tools/testing/selftests/filesystems/umount_propagation/Makefile @@ -0,0 +1,6 @@ +# SPDX-License-Identifier: GPL-2.0 +TEST_GEN_PROGS := umount_propagation_test + +CFLAGS += -Wall -O2 -g $(KHDR_INCLUDES) + +include ../../lib.mk diff --git a/tools/testing/selftests/filesystems/umount_propagation/umount_propagation_test.c b/tools/testing/selftests/filesystems/umount_propagation/umount_propagation_test.c new file mode 100644 index 000000000000..9e18d54dfb32 --- /dev/null +++ b/tools/testing/selftests/filesystems/umount_propagation/umount_propagation_test.c @@ -0,0 +1,226 @@ +// SPDX-License-Identifier: GPL-2.0 +/* + * A synchronous umount fails with EBUSY when a mount it would pull out by + * propagation is still in use. + */ +#define _GNU_SOURCE +#include <errno.h> +#include <fcntl.h> +#include <sched.h> +#include <stdio.h> +#include <stdlib.h> +#include <string.h> +#include <sys/mount.h> +#include <sys/stat.h> +#include <sys/syscall.h> +#include <sys/wait.h> +#include <unistd.h> +#include <linux/mount.h> +#include <linux/stat.h> + +#include "../../kselftest_harness.h" + +#ifndef OPEN_TREE_CLONE +#define OPEN_TREE_CLONE 1 +#endif +#ifndef OPEN_TREE_CLOEXEC +#define OPEN_TREE_CLOEXEC O_CLOEXEC +#endif +#ifndef AT_RECURSIVE +#define AT_RECURSIVE 0x8000 +#endif +#ifndef MOVE_MOUNT_F_EMPTY_PATH +#define MOVE_MOUNT_F_EMPTY_PATH 0x00000004 +#endif +#ifndef MOVE_MOUNT_BENEATH +#define MOVE_MOUNT_BENEATH 0x00000200 +#endif +#ifndef STATX_MNT_ID +#define STATX_MNT_ID 0x00001000U +#endif + +static int sys_open_tree(int dfd, const char *filename, unsigned int flags) +{ + return syscall(__NR_open_tree, dfd, filename, flags); +} + +static int sys_move_mount(int from_dfd, const char *from_pathname, + int to_dfd, const char *to_pathname, + unsigned int flags) +{ + return syscall(__NR_move_mount, from_dfd, from_pathname, to_dfd, + to_pathname, flags); +} + +/* Child exit codes. */ +enum { + CHILD_OK, + CHILD_UNSHARE, /* could not set up the slave namespace */ + CHILD_OPEN_TREE, /* open_tree() failed */ + CHILD_MOVE_MOUNT, /* move_mount() failed */ + CHILD_STATX, /* statx() failed */ + CHILD_PIPE, /* the parent went away */ +}; + +/* Messages between parent and child. */ +enum { + MSG_READY = 'r', /* child: the copy is mounted and referenced */ + MSG_CHECK = 'c', /* parent: check that the copy is still attached */ + MSG_ATTACHED = 'a', /* child: it is */ + MSG_DETACHED = 'd', /* child: it is not */ + MSG_CLOSE = 'x', /* parent: drop the reference */ + MSG_CLOSED = 'y', /* child: dropped */ + MSG_EXIT = 'e', /* parent: done */ +}; + +FIXTURE(umount_propagation) +{ + char base[64]; + char victim[80]; + bool mounted; +}; + +FIXTURE_SETUP(umount_propagation) +{ + self->mounted = false; + + if (geteuid() != 0) + SKIP(return, "test requires CAP_SYS_ADMIN"); + + ASSERT_EQ(unshare(CLONE_NEWNS), 0); + ASSERT_EQ(mount("", "/", NULL, MS_REC | MS_PRIVATE, NULL), 0); + + snprintf(self->base, sizeof(self->base), "/tmp/umount_propagation.XXXXXX"); + ASSERT_NE(mkdtemp(self->base), NULL); + ASSERT_EQ(mount("tmpfs", self->base, "tmpfs", 0, NULL), 0); + self->mounted = true; + ASSERT_EQ(mount(NULL, self->base, NULL, MS_SHARED, NULL), 0); + + snprintf(self->victim, sizeof(self->victim), "%s/victim", self->base); + ASSERT_EQ(mkdir(self->victim, 0755), 0); + ASSERT_EQ(mount("tmpfs", self->victim, "tmpfs", 0, NULL), 0); +} + +FIXTURE_TEARDOWN(umount_propagation) +{ + if (self->mounted) + umount2(self->base, MNT_DETACH); + rmdir(self->base); +} + +static int send_msg(int fd, char msg) +{ + return write(fd, &msg, 1) == 1 ? 0 : -1; +} + +static char recv_msg(int fd) +{ + char msg; + + if (read(fd, &msg, 1) != 1) + return 0; + return msg; +} + +/* Is the mount with id @mnt_id attached in this mount namespace? */ +static bool mount_attached(__u64 mnt_id) +{ + char line[4096]; + bool found = false; + FILE *f; + + f = fopen("/proc/self/mountinfo", "re"); + if (!f) + return false; + + while (fgets(line, sizeof(line), f)) { + if (strtoull(line, NULL, 10) == mnt_id) { + found = true; + break; + } + } + fclose(f); + return found; +} + +/* + * The slave namespace: take a detached copy of the shared tree and move it + * beneath the propagated copy of the victim, keeping the open_tree() + * descriptor as a reference on it. + */ +static int slave_child(const char *base, const char *victim, int to_parent, + int from_parent) +{ + struct statx stx; + int fd; + + if (unshare(CLONE_NEWNS)) + return CHILD_UNSHARE; + if (mount("", "/", NULL, MS_REC | MS_SLAVE, NULL)) + return CHILD_UNSHARE; + + fd = sys_open_tree(AT_FDCWD, base, + OPEN_TREE_CLONE | OPEN_TREE_CLOEXEC | AT_RECURSIVE); + if (fd < 0) + return CHILD_OPEN_TREE; + if (sys_move_mount(fd, "", AT_FDCWD, victim, + MOVE_MOUNT_F_EMPTY_PATH | MOVE_MOUNT_BENEATH)) + return CHILD_MOVE_MOUNT; + if (statx(fd, "", AT_EMPTY_PATH, STATX_MNT_ID, &stx)) + return CHILD_STATX; + + if (send_msg(to_parent, MSG_READY) || recv_msg(from_parent) != MSG_CHECK) + return CHILD_PIPE; + if (send_msg(to_parent, mount_attached(stx.stx_mnt_id) ? + MSG_ATTACHED : MSG_DETACHED)) + return CHILD_PIPE; + + if (recv_msg(from_parent) != MSG_CLOSE) + return CHILD_PIPE; + close(fd); + if (send_msg(to_parent, MSG_CLOSED) || recv_msg(from_parent) != MSG_EXIT) + return CHILD_PIPE; + return CHILD_OK; +} + +TEST_F(umount_propagation, busy_copy_pulled_out) +{ + int to_child[2], to_parent[2]; + int status; + pid_t pid; + + ASSERT_EQ(pipe(to_child), 0); + ASSERT_EQ(pipe(to_parent), 0); + + pid = fork(); + ASSERT_GE(pid, 0); + if (pid == 0) { + close(to_child[1]); + close(to_parent[0]); + _exit(slave_child(self->base, self->victim, to_parent[1], + to_child[0])); + } + close(to_child[0]); + close(to_parent[1]); + + ASSERT_EQ(recv_msg(to_parent[0]), MSG_READY); + + /* the copy in the slave namespace is in use */ + ASSERT_EQ(umount2(self->victim, 0), -1); + ASSERT_EQ(errno, EBUSY); + + ASSERT_EQ(send_msg(to_child[1], MSG_CHECK), 0); + ASSERT_EQ(recv_msg(to_parent[0]), MSG_ATTACHED); + + /* and once it is not, the umount goes through */ + ASSERT_EQ(send_msg(to_child[1], MSG_CLOSE), 0); + ASSERT_EQ(recv_msg(to_parent[0]), MSG_CLOSED); + ASSERT_EQ(umount2(self->victim, 0), 0); + + ASSERT_EQ(send_msg(to_child[1], MSG_EXIT), 0); + ASSERT_EQ(waitpid(pid, &status, 0), pid); + ASSERT_TRUE(WIFEXITED(status)); + ASSERT_EQ(WEXITSTATUS(status), CHILD_OK); +} + +TEST_HARNESS_MAIN diff --git a/tools/testing/selftests/move_mount_set_group/move_mount_set_group_test.c b/tools/testing/selftests/move_mount_set_group/move_mount_set_group_test.c index 12434415ec36..9c8fc8c7f62c 100644 --- a/tools/testing/selftests/move_mount_set_group/move_mount_set_group_test.c +++ b/tools/testing/selftests/move_mount_set_group/move_mount_set_group_test.c @@ -146,17 +146,19 @@ static void null_endofword(char *word) *word = '\0'; } -static bool is_shared_mount(const char *path) +/* Does the mount on @path carry the optional field @field in mountinfo? */ +static bool mount_has_field(const char *path, const char *field) { size_t len = 0; char *line = NULL; FILE *f = NULL; + bool found = false; f = fopen("/proc/self/mountinfo", "re"); if (!f) return false; - while (getline(&line, &len, f) != -1) { + while (!found && getline(&line, &len, f) != -1) { char *opts, *target; target = get_field(line, 4); @@ -172,15 +174,29 @@ static bool is_shared_mount(const char *path) if (strcmp(target, path) != 0) continue; - null_endofword(opts); - if (strstr(opts, "shared:")) - return true; + /* the optional fields end at the "-" separator */ + while (opts && *opts != '-') { + char *next = strchr(opts, ' '); + + if (next) + *next++ = '\0'; + if (!strncmp(opts, field, strlen(field))) { + found = true; + break; + } + opts = next; + } } free(line); fclose(f); - return false; + return found; +} + +static bool is_shared_mount(const char *path) +{ + return mount_has_field(path, "shared:"); } /* Attempt to de-conflict with the selftests tree. */ @@ -372,4 +388,50 @@ TEST_F(move_mount_set_group, complex_sharing_copying) ASSERT_EQ(is_shared_mount(SET_GROUP_A), 1); } +#define SET_GROUP_B "/tmp/B" +#define SET_GROUP_C "/tmp/C" + +/* + * An unbindable mount is neither shared nor a slave, so it must not be + * accepted as the target: with a slave source it would end up unbindable + * and a slave at the same time. + */ +TEST_F(move_mount_set_group, unbindable_target) +{ + bool ret; + + ret = move_mount_set_group_supported(); + ASSERT_GE(ret, 0); + if (!ret) + SKIP(return, "move_mount(MOVE_MOUNT_SET_GROUP) is not supported"); + + ASSERT_EQ(mount(NULL, SET_GROUP_A, NULL, MS_SHARED, 0), 0); + + /* B: a slave of A's peer group */ + ASSERT_EQ(mkdir(SET_GROUP_B, 0777), 0); + ASSERT_EQ(mount(SET_GROUP_A, SET_GROUP_B, NULL, MS_BIND, NULL), 0); + ASSERT_EQ(mount(NULL, SET_GROUP_B, NULL, MS_SLAVE, 0), 0); + ASSERT_TRUE(mount_has_field(SET_GROUP_B, "master:")); + + /* C: unbindable */ + ASSERT_EQ(mkdir(SET_GROUP_C, 0777), 0); + ASSERT_EQ(mount(SET_GROUP_A, SET_GROUP_C, NULL, MS_BIND, NULL), 0); + ASSERT_EQ(mount(NULL, SET_GROUP_C, NULL, MS_UNBINDABLE, 0), 0); + ASSERT_TRUE(mount_has_field(SET_GROUP_C, "unbindable")); + + /* from a slave */ + ASSERT_EQ(syscall(__NR_move_mount, AT_FDCWD, SET_GROUP_B, + AT_FDCWD, SET_GROUP_C, MOVE_MOUNT_SET_GROUP), -1); + ASSERT_EQ(errno, EINVAL); + ASSERT_FALSE(mount_has_field(SET_GROUP_C, "master:")); + ASSERT_TRUE(mount_has_field(SET_GROUP_C, "unbindable")); + + /* from a shared mount */ + ASSERT_EQ(syscall(__NR_move_mount, AT_FDCWD, SET_GROUP_A, + AT_FDCWD, SET_GROUP_C, MOVE_MOUNT_SET_GROUP), -1); + ASSERT_EQ(errno, EINVAL); + ASSERT_FALSE(mount_has_field(SET_GROUP_C, "shared:")); + ASSERT_TRUE(mount_has_field(SET_GROUP_C, "unbindable")); +} + TEST_HARNESS_MAIN diff --git a/tools/testing/selftests/pidfd/pidfd_open_test.c b/tools/testing/selftests/pidfd/pidfd_open_test.c index 318e6f09c8e0..c6698b7cdb47 100644 --- a/tools/testing/selftests/pidfd/pidfd_open_test.c +++ b/tools/testing/selftests/pidfd/pidfd_open_test.c @@ -168,7 +168,7 @@ int main(int argc, char **argv) } if (info.ppid != getppid()) { ksft_print_msg("ppid %d does not match ppid from ioctl %d\n", - pid, info.pid); + getppid(), info.ppid); goto on_error; } if (info.ruid != getuid()) { diff --git a/virt/kvm/guest_memfd.c b/virt/kvm/guest_memfd.c index 625e62e1a031..5c084d5393a6 100644 --- a/virt/kvm/guest_memfd.c +++ b/virt/kvm/guest_memfd.c @@ -506,7 +506,7 @@ static const struct address_space_operations kvm_gmem_aops = { #endif }; -static int kvm_gmem_setattr(struct mnt_idmap *idmap, struct dentry *dentry, +static int kvm_gmem_setattr(const struct mnt_idmap *idmap, struct dentry *dentry, struct iattr *attr) { return -EINVAL; |
