Commit Graph

8949 Commits

Author SHA1 Message Date
Peng Hao d0887a5df5 ext4: introduce direct I/O write using iomap infrastructure
commit: 378f32bab3714f04c4e0c3aee4129f6703805550

This patch introduces a new direct I/O write path which makes use of
the iomap infrastructure.

All direct I/O writes are now passed from the ->write_iter() callback
through to the new direct I/O handler ext4_dio_write_iter(). This
function is responsible for calling into the iomap infrastructure via
iomap_dio_rw().

Code snippets from the existing direct I/O write code within
ext4_file_write_iter() such as, checking whether the I/O request is
unaligned asynchronous I/O, or whether the write will result in an
overwrite have effectively been moved out and into the new direct I/O
->write_iter() handler.
The block mapping flags that are eventually passed down to
ext4_map_blocks() from the *_get_block_*() suite of routines have been
taken out and introduced within ext4_iomap_alloc().

For inode extension cases, ext4_handle_inode_extension() is
effectively the function responsible for performing such metadata
updates. This is called after iomap_dio_rw() has returned so that we
can safely determine whether we need to potentially truncate any
allocated blocks that may have been prepared for this direct I/O
write. We don't perform the inode extension, or truncate operations
from the ->end_io() handler as we don't have the original I/O 'length'
available there. The ->end_io() however is responsible fo converting
allocated unwritten extents to written extents.

In the instance of a short write, we fallback and complete the
remainder of the I/O using buffered I/O via
ext4_buffered_write_iter().

The existing buffer_head direct I/O implementation has been removed as
it's now redundant.

[ Fix up ext4_dio_write_iter() per Jan's comments at
  https://lore.kernel.org/r/20191105135932.GN22379@quack2.suse.cz -- TYT ]

Signed-off-by: Matthew Bobrowski <mbobrowski@mbobrowski.org>
Reviewed-by: Jan Kara <jack@suse.cz>
Reviewed-by: Ritesh Harjani <riteshh@linux.ibm.com>
Signed-off-by: Theodore Ts'o <tytso@mit.edu>
2022-03-02 10:48:04 +08:00
Peng Hao e9c5b25f18 ext4: update ext4_sync_file() to not use __generic_file_fsync()
commit: 3eaf9cc62f447a742b26fa601993e94406aa1ea1

When the filesystem is created without a journal, we eventually call
into __generic_file_fsync() in order to write out all the modified
in-core data to the permanent storage device. This function happens to
try and obtain an inode_lock() while synchronizing the files buffer
and it's associated metadata.

Generally, this is fine, however it becomes a problem when there is
higher level code that has already obtained an inode_lock() as this
leads to a recursive lock situation. This case is especially true when
porting across direct I/O to iomap infrastructure as we obtain an
inode_lock() early on in the I/O within ext4_dio_write_iter() and hold
it until the I/O has been completed. Consequently, to not run into
this specific issue, we move away from calling into
__generic_file_fsync() and perform the necessary synchronization tasks
within ext4_sync_file().

Signed-off-by: Matthew Bobrowski <mbobrowski@mbobrowski.org>
Reviewed-by: Ritesh Harjani <riteshh@linux.ibm.com>
Reviewed-by: Jan Kara <jack@suse.cz>
Signed-off-by: Theodore Ts'o <tytso@mit.edu>
2022-03-02 10:48:04 +08:00
Peng Hao be0ab19492 ext4: move inode extension check out from ext4_iomap_alloc()
commit: 0b9f230b94dd7457802264dc4c16921b3527dcf1

Lift the inode extension/orphan list handling code out from
ext4_iomap_alloc() and apply it within the ext4_dax_write_iter().

Signed-off-by: Matthew Bobrowski <mbobrowski@mbobrowski.org>
Reviewed-by: Jan Kara <jack@suse.cz>
Reviewed-by: Ritesh Harjani <riteshh@linux.ibm.com>
Signed-off-by: Theodore Ts'o <tytso@mit.edu>
2022-03-02 10:48:03 +08:00
Peng Hao 3b69d0370b ext4: move inode extension/truncate code out from ->iomap_end() callback
commit: 569342dc2485392e95b6a626281708c25014ba37

In preparation for implementing the iomap direct I/O modifications,
the inode extension/truncate code needs to be moved out from the
ext4_iomap_end() callback. For direct I/O, if the current code
remained, it would behave incorrrectly. Updating the inode size prior
to converting unwritten extents would potentially allow a racing
direct I/O read to find unwritten extents before being converted
correctly.

The inode extension/truncate code now resides within a new helper
ext4_handle_inode_extension(). This function has been designed so that
it can accommodate for both DAX and direct I/O extension/truncate
operations.

[flyingpeng: compatibility modification]

Signed-off-by: Matthew Bobrowski <mbobrowski@mbobrowski.org>
Reviewed-by: Jan Kara <jack@suse.cz>
Reviewed-by: Ritesh Harjani <riteshh@linux.ibm.com>
Signed-off-by: Theodore Ts'o <tytso@mit.edu>
Signed-off-by: Peng Hao <flyingpeng@tencent.com>
2022-03-02 10:48:03 +08:00
Peng Hao c8b28618d5 ext4: introduce direct I/O read using iomap infrastructure
commit: b1b4705d54abedfd69dcdf42779c521aa1e0fbd3

This patch introduces a new direct I/O read path which makes use of
the iomap infrastructure.

The new function ext4_do_read_iter() is responsible for calling into
the iomap infrastructure via iomap_dio_rw(). If the read operation
performed on the inode is not supported, which is checked via
ext4_dio_supported(), then we simply fallback and complete the I/O
using buffered I/O.

Existing direct I/O read code path has been removed, as it is now
redundant.

[flyingpeng: compatibility modification]
Signed-off-by: Matthew Bobrowski <mbobrowski@mbobrowski.org>
Reviewed-by: Jan Kara <jack@suse.cz>
Reviewed-by: Ritesh Harjani <riteshh@linux.ibm.com>
Signed-off-by: Peng Hao <flyingpeng@tencent.com>
2022-03-02 10:48:03 +08:00
Peng Hao 67105d499b ext4: introduce new callback for IOMAP_REPORT
commit: 09edf4d381957b144440bac18a4769c53063b943

As part of the ext4_iomap_begin() cleanups that precede this patch, we
also split up the IOMAP_REPORT branch into a completely separate
->iomap_begin() callback named ext4_iomap_begin_report(). Again, the
raionale for this change is to reduce the overall clutter within
ext4_iomap_begin().

Signed-off-by: Matthew Bobrowski <mbobrowski@mbobrowski.org>
Reviewed-by: Jan Kara <jack@suse.cz>
Reviewed-by: Ritesh Harjani <riteshh@linux.ibm.com>
2022-03-02 10:48:03 +08:00
Peng Hao a9e8b1a59f ext4: split IOMAP_WRITE branch in ext4_iomap_begin() into helper
commit: f063db5ee989aafe2dc9d571b5538f2a1f1cbad2

In preparation for porting across the ext4 direct I/O path over to the
iomap infrastructure, split up the IOMAP_WRITE branch that's currently
within ext4_iomap_begin() into a separate helper
ext4_alloc_iomap(). This way, when we add in the necessary code for
direct I/O, we don't end up with ext4_iomap_begin() becoming a
monstrous twisty maze.

Signed-off-by: Matthew Bobrowski <mbobrowski@mbobrowski.org>
Reviewed-by: Jan Kara <jack@suse.cz>
Reviewed-by: Ritesh Harjani <riteshh@linux.ibm.com>
2022-03-02 10:48:03 +08:00
Peng Hao 5b59f6e470 ext4: move set iomap routines into a separate helper ext4_set_iomap()
commit: c8fdfe294187455b70e42a15df35a3e1882f332d

Separate the iomap field population code that is currently within
ext4_iomap_begin() into a separate helper ext4_set_iomap(). The intent
of this function is self explanatory, however the rationale behind
taking this step is to reeduce the overall clutter that we currently
have within the ext4_iomap_begin() callback.

Signed-off-by: Matthew Bobrowski <mbobrowski@mbobrowski.org>
Reviewed-by: Jan Kara <jack@suse.cz>
Reviewed-by: Ritesh Harjani <riteshh@linux.ibm.com>
2022-03-02 10:48:03 +08:00
Peng Hao 9d339cae4a ext4: reorder map.m_flags checks within ext4_iomap_begin()
commit: 53e5cca56795a301bbe8465781dab084f7ae8d54

For the direct I/O changes that follow in this patch series, we need
to accommodate for the case where the block mapping flags passed
through to ext4_map_blocks() result in m_flags having both
EXT4_MAP_MAPPED and EXT4_MAP_UNWRITTEN bits set. In order for any
allocated unwritten extents to be converted correctly in the
->end_io() handler, the iomap->type must be set to IOMAP_UNWRITTEN for
cases where the EXT4_MAP_UNWRITTEN bit has been set within
m_flags. Hence the reason why we need to reshuffle this conditional
statement around.

This change is a no-op for DAX as the block mapping flags passed
through to ext4_map_blocks() i.e. EXT4_GET_BLOCKS_CREATE_ZERO never
results in both EXT4_MAP_MAPPED and EXT4_MAP_UNWRITTEN being set at
once.

Signed-off-by: Matthew Bobrowski <mbobrowski@mbobrowski.org>
Reviewed-by: Jan Kara <jack@suse.cz>
Reviewed-by: Ritesh Harjani <riteshh@linux.ibm.com>
Link: https://lore.kernel.org/r/1309ad80d31a637b2deed55a85283d582a54a26a.1572949325.git.mbobrowski@mbobrowski.org
Signed-off-by: Theodore Ts'o <tytso@mit.edu>
2021-11-30 10:40:21 +08:00
Menglong Dong 3a4384db67 bpf: fix double free for btf in map_create()
'btf' is freed by btf_put() in map_create() if map_check_btf() fails.
However, it is freed again in free_map, which cause use-after-free
in the later process.

Fix this by remove redundant btf_put().

Fixes: c421e6e7b3 ("copy part code from commid id 85d33df357b634649ddbe0a20fd2d0fc5732c3cb to linux 5.4 to fix sockhash map creation fail problem")

Signed-off-by: Menglong Dong <imagedong@tencent.com>
2021-11-30 10:38:54 +08:00
katrinzhou 1265b88089 default/config: set CONFIG_TCM_USER2 as module 2021-11-02 16:19:31 +08:00
markwhwu 2f5128a765 tcp: export sysctl_tcp_pingpong_thresh
Signed-off-by: Mark Whwu <markwhwu@tencent.com>
2021-11-01 15:13:59 +08:00
caelli b80247552b cgroupfs: support for proc and sys
export files such as cpuinfo, meminfo, stat and so on, which can by used
by containers.

Signed-off-by: caelli <caelli@tencent.com>
Reviewed-by: Peng Hao <flyingpeng@tencent.com>
Reviewed-by: Bin Lai <robinlai@tencent.com>
2021-11-01 15:00:06 +08:00
caelli 5990ae24e7 cgroupfs: quota aware support
add quota aware support when showing container
cpuinfo and stat information.

Signed-off-by: caelli <caelli@tencent.com>
Reviewed-by: Peng Hao <flyingpeng@tencent.com>
Reviewed-by: Bin Lai <robinlai@tencent.com>
2021-11-01 15:00:06 +08:00
caelli 2d555b2cfe cgroupfs: refactor cgroup resource statistics for reuse.
Signed-off-by: caelli <caelli@tencent.com>
Reviewed-by: Peng Hao <flyingpeng@tencent.com>
Reviewed-by: Bin Lai <robinlai@tencent.com>
2021-11-01 15:00:06 +08:00
markwhwu a518984de2 virtio_net: disable cb aggressively
upstream commit:a7766ef18b33674fa164e2e2916cef16d4e17f43

    There are currently two cases where we poll TX vq not in response to a
    callback: start xmit and rx napi.  We currently do this with callbacks
    enabled which can cause extra interrupts from the card.  Used not to be
    a big issue as we run with interrupts disabled but that is no longer the
    case, and in some cases the rate of spurious interrupts is so high
    linux detects this and actually kills the interrupt.

    Fix up by disabling the callbacks before polling the tx vq.

    Signed-off-by: Michael S. Tsirkin <mst@redhat.com>
2021-10-29 08:50:28 +00:00
markwhwu eb4eba5e09 virtio_net: move txq wakeups under tx q lock
upstream commit:22bc63c58e876cc359d0b1566dee3db8ecc16722

    We currently check num_free outside tx q lock
    which is unsafe: new packets can arrive meanwhile
    and there won't be space in the queue.
    Thus a spurious queue wakeup causing overhead
    and even packet drops.

    Move the check under the lock to fix that.

    Signed-off-by: Michael S. Tsirkin <mst@redhat.com>
2021-10-29 08:50:28 +00:00
markwhwu 6ea1ce5294 virtio_net: move tx vq operation under tx queue lock
upstream commit: 5a2f966d0f3fa0ef6dada7ab9eda74cacee96b8a

    It's unsafe to operate a vq from multiple threads.
    Unfortunately this is exactly what we do when invoking
    clean tx poll from rx napi.
    Same happens with napi-tx even without the
    opportunistic cleaning from the receive interrupt: that races
    with processing the vq in start_xmit.

    As a fix move everything that deals with the vq to under tx lock.

    Fixes: b92f1e6751a6 ("virtio-net: transmit napi")
    Signed-off-by: Michael S. Tsirkin <mst@redhat.com>
2021-10-29 08:50:28 +00:00
markwhwu efa6c27cf1 virtio_net: Remove BUG() to avoid machine dead
upstream commit:85eb1389458d134bdb75dad502cc026c3753a619

    We should not directly BUG() when there is hdr error, it is
    better to output a print when such error happens. Currently,
    the caller of xmit_skb() already did it.

    Signed-off-by: Xianting Tian <xianting.tian@linux.alibaba.com>
    Reviewed-by: Leon Romanovsky <leonro@nvidia.com>
    Signed-off-by: David S. Miller <davem@davemloft.net>
2021-10-29 08:50:28 +00:00
markwhwu 109f6104d2 virtio_net: check virtqueue_add_sgs() return value
upstream commit: 222722bc6ebfabf5d54467070f05cf9c0a55ea8c

    As virtqueue_add_sgs() can fail, we should check the return value.

    Addresses-Coverity-ID: 1464439 ("Unchecked return value")
    Signed-off-by: Yunjian Wang <wangyunjian@huawei.com>
    Signed-off-by: David S. Miller <davem@davemloft.net>
2021-10-29 08:50:28 +00:00
markwhwu 0ff3e24928 virtio_net: disable napi_tx by default
NAPI-TX causes high hi and si cpu usage and a regression of
18% when perfoming wrk http benchmark.Also,NAPI-TX can be enabled by:
    ifconfig eth0 down
    ethtool -C eth0 tx-frames 1
    ifconfig eth0 up

Signed-off-by: Munger Jiang <mungerjiang@tencent.com>
2021-10-29 08:50:28 +00:00
markwhwu ad9aa4da1e tcp: Change the TCP_PINGPONG_THRESH macro to a variable
The TCP_PINGPONG_THRESH parameter affects nginx performance.
When the default value is set to 1, the requested QPS is 13% higher than when it is set to 3

Use sysctl -w net.core.tcp_pingpong_thresh=%d to change the default value

Signed-off-by: Mark Whwu <markwhwu@tencent.com>
2021-10-29 08:50:28 +00:00
markwhwu 4bba8351f9 mm: enable transparent hugepage by default
Signed-off-by: Mark whwu <markwhwu@tencent.com>
2021-10-29 08:50:28 +00:00
markwhwu 112b81aef0 sched: disable sched_auto_group by default
Autogroup is optimized for desktop interaction and is obviously a negative optimization for UnixBench

Signed-off-by: Mark Whwu <markwhwu@tencent.com>
2021-10-29 08:50:28 +00:00
Bin Lai f91939a5b3 sli/cpu: introduce sli max latency metrics
Sli max latency is lightweight latency monitor metrics, the monitor
tool could get the data with a little overhead. And the accuracy of
the metrics is controlled by the sampling frequency of the monitor
tool. When the performance jitter was occurred in the system, we can
get some help from these lateny metris.

Signed-off-by: Bin Lai <robinlai@tencent.com>
Reviewed-by: benbjiang<benbjiang@tencent.com>
Reviewed-by: Mengmeng Chen <bauerchen@tencent.com>
Reviewed-by: mungerjiang <mungerjiang@tencent.com>
2021-10-29 06:51:52 +00:00
mungerjiang 95a8c85b7f sli/cpu: introduce longsys check
When a process is running in the system space, it cann't be preempted until
it return to userspace, even if the process should be reschedule and other
process was ready to run(because the server system close the kernel preempt
by default). This schedule delay may impact the performance of waiting process,
therefor we introduce the longsys check to collect the schedule delay information
of process. The longsys indication could help us spot the possible performance
jitter in the system.

Signed-off-by: Munger jiang <mungerjiang@tencent.com>
Signed-off-by: Bin Lai <robinlai@tencent.com>
Reviewed-by: benbjiang<benbjiang@tencent.com>
Reviewed-by: Bauerchen <bauerchen@tencent.com>
2021-10-29 06:51:52 +00:00
mungerjiang a76d37e2f4 sli: enable sli in cgroup v1
Signed-off-by: Munger jiang <mungerjiang@tencent.com>
Reviewed-by: Bauerchen <bauerchen@tencent.com>
2021-10-29 06:51:52 +00:00
mungerjiang 674dd9febe sli: fix sli_memlat_stat_start bug
Fix the issue that sli_no_enabled is true in sli_memlat_stat_start,
but sli_no_enabled is false in sli_memlat_stat_end,if *start* is
not zero,sli_memlat_stat_end would get a bad stat.

Signed-off-by: Munger jiang <mungerjiang@tencent.com>
2021-10-29 06:51:52 +00:00
mungerjiang 3585e8f134 sli/cpu: Add sched latency account
Signed-off-by: Munger jiang <mungerjiang@tencent.com>
2021-10-29 06:51:52 +00:00
mungerjiang e8c34acd40 sli/memory: Add memory latency account
Signed-off-by: Munger jiang <mungerjiang@tencent.com>
2021-10-29 06:51:52 +00:00
Bin Lai b879b1b64b sli: Introduce memory and sched latency stat infrastructure
Signed-off-by: Munger jiang <mungerjiang@tencent.com>
2021-10-29 06:51:52 +00:00
Bauerchen c44986b358 tqos/cpu: load.r/load.d support cgroup V2
Signed-off-by: Bauerchen <bauerchen@tencent.com>
Reviewed-by: Bin Lai <robinlai@tencent.com>
2021-10-29 06:51:52 +00:00
Bauerchen 8003819c11 tqos/cpu: add cgroup load.r/load.d statistic
Now cgroup loadavg calcuted according processes in R state and D state,
sometimes we need more specific information, so just calcuting loadavg
separately.

Line2 is load.r and line3 is load.d. When most process
do IO operation and we may get the output like this:

[root@VM-130-27-centos test]# cat cpuset.loadavg
20.03 9.85 3.88 1/339 3852
6.56 3.96 1.62
13.46 5.89 2.27

Signed-off-by: Bauerchen <bauerchen@tencent.com>
Reviewed-by: Bin Lai <robinlai@tencent.com>
2021-10-29 06:51:52 +00:00
Bauerchen c07a089a0f tqos/mbuf: export mbuf interface to cpuacct subsys
In cgroup V1, mbuf only exist in cpuacct subsys, so we may need a help
function to store buffer just according process task_struct.

Signed-off-by: Bauerchen <bauerchen@tencent.com>
Reviewed-by: mungerjiang<mungerjiang@tencent.com>
Reviewed-by: benbjiang<benbjiang@tencent.com>
2021-10-29 06:51:52 +00:00
Bauerchen f813e65898 tqos/mbuf: alloc mbuf slot for cpuacct in cgroup V1
Mbuf support cgroup V1, to compatible with cgroup V2, we only export it
to cpuacct subsys.

Signed-off-by: Bauerchen <bauerchen@tencent.com>
Reviewed-by: mungerjiang<mungerjiang@tencent.com>
Reviewed-by: benbjiang<benbjiang@tencent.com>
2021-10-29 06:51:52 +00:00
Bauerchen 7795ccbf82 tqos/mbuf: write a help function to get cgroup struct from task_struct.
In order to support cgroup V1 with mbuf and sli, we need a special
cgroup structure, cpuacct subsys cgroup is nice.

We only prepare mbuf and sli for cpuacct cgroup in V1. so first find cpuacct
cgroup and if return NULl or root, just find df1_cgrp.

Signed-off-by: Bauerchen <bauerchen@tencent.com>
Reviewed-by: mungerjiang<mungerjiang@tencent.com>
Reviewed-by: mungerjiang<mungerjiang@tencent.com>
2021-10-29 06:51:52 +00:00
Bauerchen 7600f2ca44 tqos/rqm: Tencent Quality Monitor Buffer
Providing back up buffer for Quality Monitor, can be used to catch key
context when abnormal jitters occur. And application can also use it
to detect system env exception.

Signed-off-by: Bauerchen <bauerchen@tencent.com>
Reviewed-by: Jiang Biao <benbjiang@tencent.com>
Reviewed-by: Bin Lai <robinlai@tencent.com>
2021-10-29 06:51:52 +00:00
Hongbo Li 589e8a96bc getsockopt: add SO_MARK2 to get the MARK of flow
add SO_MARK2 to get the MARK of flow
backport from tkernel2:
commit c6f2e27f7ad0b97ae0c162f70b6e0375f2270f27

Signed-off-by: Hongbo Li <herberthbli@tencent.com>
Signed-off-by: brookxu <brookxu@tencent.com>
2021-10-29 06:51:37 +00:00
Peng Hao 3df720cae4 mm/lru: revise the comments of lru_lock
Use this new function to replace repeated same code, no func change.

When testing for relock we can avoid the need for RCU locking if we simply
compare the page pgdat and memcg pointers versus those that the lruvec is
holding. By doing this we can avoid the extra pointer walks and accesses of
the memory cgroup.

In addition we can avoid the checks entirely if lruvec is currently NULL.

Signed-off-by: Alexander Duyck <alexander.h.duyck@linux.intel.com>
Signed-off-by: Alex Shi <alex.shi@linux.alibaba.com>
2021-10-29 06:51:21 +00:00
Peng Hao 1b7478d8d0 mm/lru: replace pgdat lru_lock with lruvec lock
This patch moves per node lru_lock into lruvec, thus bring a lru_lock for
each of memcg per node. So on a large machine, each of memcg don't
have to suffer from per node pgdat->lru_lock competition. They could go
fast with their self lru_lock.

After move memcg charge before lru inserting, page isolation could
serialize page's memcg, then per memcg lruvec lock is stable and could
replace per node lru lock.

In func isolate_migratepages_block, compact_unlock_should_abort and
lock_page_lruvec_irqsave are open coded to work with compact_control.
Also add a debug func in locking which may give some clues if there are
sth out of hands.

Daniel Jordan's testing show 62% improvement on modified readtwice case
on his 2P * 10 core * 2 HT broadwell box.
https://lore.kernel.org/lkml/20200915165807.kpp7uhiw7l3loofu@ca-dmjordan1.us.oracle.com/

On a large machine with memcg enabled but not used, the page's lruvec
seeking pass a few pointers, that may lead to lru_lock holding time
increase and a bit regression.

Hugh Dickins helped on patch polish, thanks!
[flyingpeng: compatibility modification]

Signed-off-by: Alex Shi <alex.shi@linux.alibaba.com>
Signed-off-by: Peng Hao <flyingpeng@tencent.com>
2021-10-29 06:51:21 +00:00
Peng Hao 821adde85c mm: memcontrol: charge swapin pages on instantiation
Right now, users that are otherwise memory controlled can easily escape
their containment and allocate significant amounts of memory that they're
not being charged for.  That's because swap readahead pages are not being
charged until somebody actually faults them into their page table.  This
can be exploited with MADV_WILLNEED, which triggers arbitrary readahead
allocations without charging the pages.

There are additional problems with the delayed charging of swap pages:

1. To implement refault/workingset detection for anonymous pages, we
   need to have a target LRU available at swapin time, but the LRU is not
   determinable until the page has been charged.

2. To implement per-cgroup LRU locking, we need page->mem_cgroup to be
   stable when the page is isolated from the LRU; otherwise, the locks
   change under us.  But swapcache gets charged after it's already on the
   LRU, and even if we cannot isolate it ourselves (since charging is not
   exactly optional).

The previous patch ensured we always maintain cgroup ownership records for
swap pages.  This patch moves the swapcache charging point from the fault
handler to swapin time to fix all of the above problems.

v2: simplify swapin error checking (Joonsoo)

[hughd@google.com: fix livelock in __read_swap_cache_async()]
[flyingpeng: port New API]
  Link: http://lkml.kernel.org/r/alpine.LSU.2.11.2005212246080.8458@eggly.anvils
Signed-off-by: Johannes Weiner <hannes@cmpxchg.org>
Signed-off-by: Hugh Dickins <hughd@google.com>
Signed-off-by: Peng Hao <flyingpeng@tencent.com>
2021-10-29 06:51:21 +00:00
Peng Hao 4e9dc3310b mm/swap.c: serialize memcg changes in pagevec_lru_move_fn
Hugh Dickins' found a memcg change bug on original version:
If we want to change the pgdat->lru_lock to memcg's lruvec lock, we have
to serialize mem_cgroup_move_account during pagevec_lru_move_fn. The
possible bad scenario would like:

	cpu 0					cpu 1
lruvec = mem_cgroup_page_lruvec()
					if (!isolate_lru_page())
						mem_cgroup_move_account

spin_lock_irqsave(&lruvec->lru_lock <== wrong lock.

So we need TestClearPageLRU to block isolate_lru_page(), that serializes
the memcg change. and then removing the PageLRU check in move_fn callee
as the consequence.

__pagevec_lru_add_fn() is different from the others, because the pages
it deals with are, by definition, not yet on the lru.  TestClearPageLRU
is not needed and would not work, so __pagevec_lru_add() goes its own
way.

Reported-by: Hugh Dickins <hughd@google.com>
Signed-off-by: Alex Shi <alex.shi@linux.alibaba.com>
2021-10-29 06:51:21 +00:00
Peng Hao f784a238c7 mm/compaction: do page isolation first in compaction
Currently, compaction would get the lru_lock and then do page isolation
 which works fine with pgdat->lru_lock, since any page isoltion would
 compete for the lru_lock. If we want to change to memcg lru_lock, we
 have to isolate the page before getting lru_lock, thus isoltion would
 block page's memcg change which relay on page isoltion too. Then we
 could safely use per memcg lru_lock later.

 The new page isolation use previous introduced TestClearPageLRU() +
 pgdat lru locking which will be changed to memcg lru lock later.

 Hugh Dickins <hughd@google.com> fixed following bugs in this patch's
 early version:

 Fix lots of crashes under compaction load: isolate_migratepages_block()
 must clean up appropriately when rejecting a page, setting PageLRU again
 if it had been cleared; and a put_page() after get_page_unless_zero()
 cannot safely be done while holding locked_lruvec - it may turn out to
 be the final put_page(), which will take an lruvec lock when PageLRU.
 And move __isolate_lru_page_prepare back after get_page_unless_zero to
 make trylock_page() safe:
 trylock_page() is not safe to use at this time: its setting PG_locked
 can race with the page being freed or allocated ("Bad page"), and can
 also erase flags being set by one of those "sole owners" of a freshly
 allocated page who use non-atomic __SetPageFlag().

Suggested-by: Johannes Weiner <hannes@cmpxchg.org>
Signed-off-by: Alex Shi <alex.shi@linux.alibaba.com>
2021-10-29 06:51:21 +00:00
Peng Hao 4702fd8039 mm/lru: introduce TestClearPageLRU
Currently lru_lock still guards both lru list and page's lru bit, that's
ok. but if we want to use specific lruvec lock on the page, we need to
pin down the page's lruvec/memcg during locking. Just taking lruvec
lock first may be undermined by the page's memcg charge/migration. To
fix this problem, we could clear the lru bit out of locking and use
it as pin down action to block the page isolation in memcg changing.

So now a standard steps of page isolation is following:
	1, get_page(); 	       #pin the page avoid to be free
	2, TestClearPageLRU(); #block other isolation like memcg change
	3, spin_lock on lru_lock; #serialize lru list access
	4, delete page from lru list;
The step 2 could be optimzed/replaced in scenarios which page is
unlikely be accessed or be moved between memcgs.

This patch start with the first part: TestClearPageLRU, which combines
PageLRU check and ClearPageLRU into a macro func TestClearPageLRU. This
function will be used as page isolation precondition to prevent other
isolations some where else. Then there are may !PageLRU page on lru
list, need to remove BUG() checking accordingly.

There 2 rules for lru bit now:
1, the lru bit still indicate if a page on lru list, just in some
   temporary moment(isolating), the page may have no lru bit when
   it's on lru list.  but the page still must be on lru list when the
   lru bit set.
2, have to remove lru bit before delete it from lru list.

As Andrew Morton mentioned this change would dirty cacheline for page
isn't on LRU. But the lost would be acceptable in Rong Chen
<rong.a.chen@intel.com> report:
https://lore.kernel.org/lkml/20200304090301.GB5972@shao2-debian/

[flyingpeng: compatibility code porting]
Suggested-by: Johannes Weiner <hannes@cmpxchg.org>
Signed-off-by: Alex Shi <alex.shi@linux.alibaba.com>
Signed-off-by: Peng Hao <flyingpeng@tencent.com>
2021-10-29 06:51:21 +00:00
Peng Hao f646d14e7d mm/mlock: remove __munlock_isolate_lru_page
The func only has one caller, remove it to clean up code and simplify
code.

Signed-off-by: Alex Shi <alex.shi@linux.alibaba.com>
2021-10-29 06:51:21 +00:00
Peng Hao fe81043c65 mm/mlock: remove lru_lock on TestClearPageMlocked
In the func munlock_vma_page, comments mentained lru_lock needed for
serialization with split_huge_pages. But the page must be PageLocked
as well as pages in split_huge_page series funcs. Thus the PageLocked
is enough to serialize both funcs.

Further more, Hugh Dickins pointed: before splitting in
split_huge_page_to_list, the page was unmap_page() to remove pmd/ptes
which protect the page from munlock. Thus, no needs to guard
__split_huge_page_tail for mlock clean, just keep the lru_lock there for
isolation purpose.

LKP found a preempt issue on __mod_zone_page_state which need change
to mod_zone_page_state. Thanks!

Signed-off-by: Alex Shi <alex.shi@linux.alibaba.com>
2021-10-29 06:51:21 +00:00
Peng Hao b80a4794f6 mm/vmscan: remove lruvec reget in move_pages_to_lru
A isolated page shouldn't be recharged by memcg since the memcg
migration isn't possible at the time.
So remove unnecessary regetting.

Thanks to Alexander Duyck for pointing this out.

Signed-off-by: Alex Shi <alex.shi@linux.alibaba.com>
2021-10-29 06:51:21 +00:00
Peng Hao fb8db998fd mm/swap.c: fold vm event PGROTATED into pagevec_move_tail_fn
Fold the PGROTATED event collection into pagevec_move_tail_fn call back
func like other funcs does in pagevec_lru_move_fn. Thus we could save
func call pagevec_move_tail().
Now all usage of pagevec_lru_move_fn are same and no needs of its 3rd
parameter.

It's just simply the calling. No functional change.

[lkp@intel.com: found a build issue in the original patch, thanks]
Signed-off-by: Alex Shi <alex.shi@linux.alibaba.com>
2021-10-29 06:51:21 +00:00
Peng Hao a50cb5b4fd mm/memcg: add debug checking in lock_page_memcg
Add a debug checking in lock_page_memcg, then we could get alarm
if anything wrong here.

Suggested-by: Johannes Weiner <hannes@cmpxchg.org>
Signed-off-by: Alex Shi <alex.shi@linux.alibaba.com>
2021-10-29 06:51:21 +00:00
Peng Hao b85b3df216 mm: page_idle_get_page() does not need lru_lock
It is necessary for page_idle_get_page() to recheck PageLRU() after
get_page_unless_zero(), but holding lru_lock around that serves no
useful purpose, and adds to lru_lock contention: delete it.

See https://lore.kernel.org/lkml/20150504031722.GA2768@blaptop for the
discussion that led to lru_lock there; but __page_set_anon_rmap() now
uses WRITE_ONCE(), and I see no other risk in page_idle_clear_pte_refs()
using rmap_walk() (beyond the risk of racing PageAnon->PageKsm, mostly
but not entirely prevented by page_count() check in ksm.c's
write_protect_page(): that risk being shared with page_referenced() and
not helped by lru_lock).

Signed-off-by: Hugh Dickins <hughd@google.com>
Signed-off-by: Alex Shi <alex.shi@linux.alibaba.com>
2021-10-29 06:51:21 +00:00