Commit Graph

344 Commits

Author SHA1 Message Date
Kairui Song ef9529c002 x86: kdump: move reserve_crashkernel[_low]() into crash_core.c
From: Chen Zhou <chenzhou10@huawei.com>
Link: https://lkml.org/lkml/2021/1/30/53
Link: 8cb8686864

Make the functions reserve_crashkernel[_low]() as generic.
Arm64 will use these to reimplement crashkernel=X.

Signed-off-by: Chen Zhou <chenzhou10@huawei.com>
Tested-by: John Donnelly <John.p.donnelly@oracle.com>
Signed-off-by: Chen Zhou <chenzhou10@huawei.com>
Acked-by: Xie XiuQi <xiexiuqi@huawei.com>
Signed-off-by: Zheng Zengkai <zhengzengkai@huawei.com>
Signed-off-by: Kairui Song <kasong@tencent.com>
2022-06-06 13:50:12 +08:00
frankjpliu 7503de0908
Merge pull request #194 from chenmengc/dev/sli_merge_main
Dev/sli merge main
2022-04-20 15:17:37 +08:00
frankjpliu b54509a3af
Merge pull request #177 from lanxinyuchs/lanxinyu/tk4/sysctl-write-forbid
sysctl: restrict the writting for key parameters
2022-04-20 14:58:24 +08:00
Lan Xinyu 2d927e094e sysctl: restrict the writing for key parameters
Users may modify some sysctl parameters intentionally or unintentionally,
which may cause problems such as disconnection of the container network
connection.

Therefore, the modification of key sysctl parameters is
restricted on the host and in the container that shares the network
namespace with the host.
2022-04-19 19:57:10 +08:00
frankjpliu 0d557e88b4
Merge pull request #185 from Tencent/tk4/haisuwang
Support buffer I/O qos function in cgroup v1
2022-04-19 16:50:14 +08:00
Bin Lai 36ce5bdb96 sli/kabi: add the kabi reserve space for sli structure
In order to avoid breaking system kabi, we add the reserve space
for sli structure. Then we can use the reserved space to fix bug
by hotpatch.

Signed-off-by: Bin Lai <robinlai@tencent.com>
Reviewed-by: Bauerchen <bauerchen@tencent.com>
2022-04-18 21:10:23 +08:00
Bin Lai 8979f4ffb9 kabi: introduce more kabi-related macro
Currently, there is only KABI_RESERVE macro in the kabi.h. So we
introduce more kabi-related macros, then we could use the reserved
space more easily by the macros.

Signed-off-by: Bin Lai <robinlai@tencent.com>
Reviewed-by: Mengmeng Chen <bauerchen@tencent.com>
Reviewed-by: benbjiang <benbjiang@tencent.com>
2022-04-18 21:10:10 +08:00
bauerchen 9620012f40 sli: sli.monitor support cgroup v1
Signed-off-by: bauerchen <bauerchen@tencent.com>
Reviewed-by: Bin Lai <robinlai@tencent.com>
2022-04-18 21:09:31 +08:00
bauerchen 34961244de sli: add sli_monitor notify mechanism
Sli can collect much information such as cpu stall and memory stall, but
now this key info only exported to userspace by cgroupfs interface, so
if user app want to feel the resource load, it have to monitor the
cgroupfs file periodly  which have some latency and more resource
consumpetion.

Sli montior notify mechanism can solve this issue. It prides interface
which can be monitored by use app using poll or epoll. So User app can  get
informed at the first time the resource load reaches threshold.

Signed-off-by: bauerchen <bauerchen@tencent.com>
Reviewed-by: Bin Lai <robinlai@tencent.com>
2022-04-18 21:08:43 +08:00
Bin Lai 549d03b9bf sli: remove the page_alloc trace item
Page_alloc is may be a hot spot in the system, the overhead of
trace will increase with the number of page alloc operations.
And we only focus on the page alloc latency caused by memory
compaction or reclaim. So we remove the useless page alloc item
to reduce the overhead of trace.

Signed-off-by: Bin Lai <robinlai@tencent.com>
Reviewed-by: Bauerchen <bauerchen@tencent.com>
2022-04-18 21:07:19 +08:00
Bin Lai 9a47b6193b sli: check cgroup's proactive monitoring statistics and report
the event if necessary

If there is not any task of cgroup running on the cpu, there is
no necessary to do the monitoring event check for cgroup. So we
put the checkpoint in the tick process to reduce the overhead
of system(if tasks of cgroup are not running, the tick interrupt
will not hit the task).

Signed-off-by: Bin Lai <robinlai@tencent.com>
Reviewed-by: Bauerchen <bauerchen@tencent.com>
2022-04-18 21:06:46 +08:00
Bin Lai a3b305afa3 sli: add sli event monitoring for cgroup
Sli proactive monitoring framework provid the separated event
monitoring control parameters for each cgroup. So we add the
sli event monitoring for each cgroup, then We can tune the
parameters for specified cgroup by related interface.

Signed-off-by: Bin Lai <robinlai@tencent.com>
Reviewed-by: Bauerchen <bauerchen@tencent.com>
2022-04-18 21:06:25 +08:00
Bin Lai 7b5a1444a3 sli: introduce sli event monitoring framework
In order to support proactive event monitoring, we introduce the
new sli_event_monitor framework and combine mbuf threshold tracing
with the proactive event monitoring framework. Then we can handle
these tracing event in the same framework, and it could help workloads
easily use these features.

Through the proactive event monitoring, workloads don't need to do
the periodic sampling and calculation. It only need to set the
monitoring rules according to defined format(for a specified cgroup
or all cgroups), then the sli event monitoring framework will trace
the related latency event and report the event when needed. It could
help system reduce the cost of sampling and make the interference
detection to be faster.

Signed-off-by: Bin Lai <robinlai@tencent.com>
Signed-off-by: bauerchen <bauerchen@tencent.com>
2022-04-18 21:05:16 +08:00
Bin Lai 08f6af622c sli: introuce the interrupt time metric for cgroup
The original sli_max is used for showing max value during the sampling.
But the workloads had not used it, so we intend to reuse this interface.
The sli_max will be applied to store the accumulated value for every
metric, then we can easily get the all interference information
during the samping.

Some workloads(such as ngix) are io-bound tasks(network), and the handler
may run in the softirq(take a lot of time). The VM is usually rps enabled,
so the network soft-interrupt will be fairly dispatched to the CPUs in the
system. The soft-interrupt is a high-priortity handler, it will interrupt
other cgroup's task and make some performance jitter. So we need the
interrupt time metric to indicate the interrupt interference from other
cgroups.

Signed-off-by: Bin Lai <robinlai@tencent.com>
Reviewed-by: Bauerchen <bauerchen@tencent.com>
2022-04-18 21:03:48 +08:00
Jan Kara 91d1fab653 rq-qos: fix missed wake-ups in rq_qos_throttle try two
Commit 545fbd0775ba ("rq-qos: fix missed wake-ups in rq_qos_throttle")
tried to fix a problem that a process could be sleeping in rq_qos_wait()
without anyone to wake it up. However the fix is not complete and the
following can still happen:

CPU1 (waiter1)		CPU2 (waiter2)		CPU3 (waker)
rq_qos_wait()		rq_qos_wait()
  acquire_inflight_cb() -> fails
			  acquire_inflight_cb() -> fails

						completes IOs, inflight
						  decreased
  prepare_to_wait_exclusive()
			  prepare_to_wait_exclusive()
  has_sleeper = !wq_has_single_sleeper() -> true as there are two sleepers
			  has_sleeper = !wq_has_single_sleeper() -> true
  io_schedule()		  io_schedule()

Deadlock as now there's nobody to wakeup the two waiters. The logic
automatically blocking when there are already sleepers is really subtle
and the only way to make it work reliably is that we check whether there
are some waiters in the queue when adding ourselves there. That way, we
are guaranteed that at least the first process to enter the wait queue
will recheck the waiting condition before going to sleep and thus
guarantee forward progress.

UPSTREAM COMMIT:
11c7aa0ddea8611007768d3e6b58d45dc60a19e1

Fixes: 545fbd0775ba ("rq-qos: fix missed wake-ups in rq_qos_throttle")
CC: stable@vger.kernel.org
Signed-off-by: Jan Kara <jack@suse.cz>
Link: https://lore.kernel.org/r/20210607112613.25344-1-jack@suse.cz
Signed-off-by: Jens Axboe <axboe@kernel.dk>
2022-04-12 16:54:12 +08:00
samuelliao bc99a5f248 block: fix zero ioutil if io stall or disk enter blocked state
Since commit 5b18b5a737
    block: delete part_round_stats and switch to less precise counting,
io_ticks don't advance if io stall, iostat will show 0% io util.
This patch add back the logical, advance io_ticks if inflight > 0.
show 100% ioutil if io stall. This patch also show 100% ioutil if
queue quiesced, eg: scsi device blocked.

Signed-off-by: samuelliao <samuelliao@tencent.com>
2022-04-01 10:13:10 +08:00
Haisu Wang 9d953f7b95 backport blkcg: add bufio isolation based on v2 infrastructure to v1
Add buffer IO isolation based on v2 infrastructure to v1, so we
can unify the interface for dio and bufio.

backport tk4 commit 9d1a9b49ad00e3dd9d6246a3c7e291a8d5a93622

Signed-off-by: Haisu Wang <haisuwang@tencent.com>
2022-03-31 10:17:14 +08:00
Haisu Wang 4a7c3a1caf backport tqos/io: add io_qos switch in sysctrl
Add io_qos switch for buffer I/O QoS.

backport tk4 commit df7c944c269e4765fe06f5358aebefb61d1108ec

Signed-off-by: Haisu Wang <haisuwang@tencent.com>
2022-03-31 10:17:14 +08:00
Lan Xinyu ec7b1baac2 Revert "mm/page_alloc: allow high-order pages on the per-cpu lists"
Due to unstable results on some machines, revert the backport.
This reverts commit 5312a9ff49.
2022-03-30 12:26:29 +08:00
linuszeng d955f2cc41 tqos/mm: set the default value of pagecache_limit_global to zero.
Signed-off-by: Zeng Jingxiang <linuszeng@tencent.com>
2022-03-16 12:44:28 +08:00
lanxinyuchs 5312a9ff49 mm/page_alloc: allow high-order pages to be stored on the per-cpu lists
commit:44042b4498728f4376e84bae1ac8016d146d850b

The per-cpu page allocator (PCP) only stores order-0 pages.  This means
that all THP and "cheap" high-order allocations including SLUB contends on
the zone->lock.  This patch extends the PCP allocator to store THP and
"cheap" high-order pages.  Note that struct per_cpu_pages increases in
size to 256 bytes (4 cache lines) on x86-64.

Note that this is not necessarily a universal performance win because of
how it is implemented.  High-order pages can cause pcp->high to be
exceeded prematurely for lower-orders so for example, a large number of
THP pages being freed could release order-0 pages from the PCP lists.
Hence, much depends on the allocation/free pattern as observed by a single
CPU to determine if caching helps or hurts a particular workload.

That said, basic performance testing passed.  The following is a netperf
UDP_STREAM test which hits the relevant patches as some of the network
allocations are high-order.

netperf-udp
                                 5.13.0-rc2             5.13.0-rc2
                           mm-pcpburst-v3r4   mm-pcphighorder-v1r7
Hmean     send-64         261.46 (   0.00%)      266.30 *   1.85%*
Hmean     send-128        516.35 (   0.00%)      536.78 *   3.96%*
Hmean     send-256       1014.13 (   0.00%)     1034.63 *   2.02%*
Hmean     send-1024      3907.65 (   0.00%)     4046.11 *   3.54%*
Hmean     send-2048      7492.93 (   0.00%)     7754.85 *   3.50%*
Hmean     send-3312     11410.04 (   0.00%)    11772.32 *   3.18%*
Hmean     send-4096     13521.95 (   0.00%)    13912.34 *   2.89%*
Hmean     send-8192     21660.50 (   0.00%)    22730.72 *   4.94%*
Hmean     send-16384    31902.32 (   0.00%)    32637.50 *   2.30%*

Functionally, a patch like this is necessary to make bulk allocation of
high-order pages work with similar performance to order-0 bulk
allocations.  The bulk allocator is not updated in this series as it would
have to be determined by bulk allocation users how they want to track the
order of pages allocated with the bulk allocator.

Link: https://lkml.kernel.org/r/20210611135753.GC30378@techsingularity.net
Signed-off-by: Mel Gorman <mgorman@techsingularity.net>
Acked-by: Vlastimil Babka <vbabka@suse.cz>
Cc: Zi Yan <ziy@nvidia.com>
Cc: Dave Hansen <dave.hansen@linux.intel.com>
Cc: Michal Hocko <mhocko@kernel.org>
Cc: Jesper Dangaard Brouer <brouer@redhat.com>
Signed-off-by: Andrew Morton <akpm@linux-foundation.org>
Signed-off-by: Linus Torvalds <torvalds@linux-foundation.org>
2022-03-14 15:17:27 +08:00
Menglong Dong 732ef51851 bpf: fix compile error when CGROUP_BPF disabled
Following functions are missed when CGROUP_BPF disabled, and compile
will fail:

BPF_CGROUP_RUN_PROG_TW_CLOSE
BPF_CGROUP_RUN_PROG_UDP_UNHASH
BPF_CGROUP_RUN_PROG_INET_POST_AUTOBIND

Fix it by define them while CGROUP_BPF disabled.

Signed-off-by: Menglong Dong <imagedong@tencent.com>
2022-02-21 17:22:51 +08:00
Menglong Dong 4c22274eac net: bpf: introduce BPF_CGROUP_UDP_UNHASH eBPF hook
Add new cgroup based eBPF hook 'BPF_CGROUP_UDP_UNHASH' which is called
when UDP sock unhashed. This is used to monitor the release of UDP
sock.

Signed-off-by: Menglong Dong <imagedong@tencent.com>
2022-02-17 12:58:28 +08:00
Menglong Dong e22ef12b56 net: bpf: add BPF_CGROUP_TWSK_CLOSE for tcp timewait sock close
For now, 'BPF_SOCK_OPS_STATE_CB' of sock_ops can be used to monitor
the state change of TCP sock. However, once tcp sock change to timewait
sock, 'TCP_CLOSE' event will be passed to the eBPF program, and it's
hard to capture the finish of a TCP connect.

Add 'BPF_CGROUP_TWSK_CLOSE', which will be called when timewait sock
close.

Signed-off-by: Menglong Dong <imagedong@tencent.com>
2022-02-17 12:58:28 +08:00
Menglong Dong 4c07420add net: bpf: introduce BPF_CGROUP_INET_POST_AUTOBIND attach type
Add new cgroup based eBPF type 'BPF_CGROUP_INET4_POST_AUTOBIND', which
is called after the success of port binding in inet_autobind().

The return value is used to determine if this port is usable, therefore
users have the chance to reject the autobind operation.

Signed-off-by: Menglong Dong <imagedong@tencent.com>
2022-02-17 12:58:28 +08:00
zgpeng fd57c9e500 Make CFS bandwidth controller burstable
Accumulate unused quota from previous periods, thus accumulated
bandwidth runtime can be used in the following periods. During
accumulation, take care of runtime overflow. Previous non-burstable
CFS bandwidth controller only assign quota to runtime, that saves a lot.

A sysctl parameter cpu_qos_cfs_bw_burst_enabled is introduced as a
switch for burst. It is disabled by default.

Signed-off-by: Huaixin Chang <changhuaixin@linux.alibaba.com>
Signed-off-by: Shanpei Chen <shanpeic@linux.alibaba.com>
2021-12-30 14:50:28 +08:00
Daniel Borkmann 3945c7126a bpf, net: Rework cookie generator as per-cpu one
With its use in BPF, the cookie generator can be called very frequently
in particular when used out of cgroup v2 hooks (e.g. connect / sendmsg)
and attached to the root cgroup, for example, when used in v1/v2 mixed
environments. In particular, when there's a high churn on sockets in the
system there can be many parallel requests to the bpf_get_socket_cookie()
and bpf_get_netns_cookie() helpers which then cause contention on the
atomic counter.

As similarly done in f991bd2e1421 ("fs: introduce a per-cpu last_ino
allocator"), add a small helper library that both can use for the 64 bit
counters. Given this can be called from different contexts, we also need
to deal with potential nested calls even though in practice they are
considered extremely rare. One idea as suggested by Eric Dumazet was
to use a reverse counter for this situation since we don't expect 64 bit
overflows anyways; that way, we can avoid bigger gaps in the 64 bit
counter space compared to just batch-wise increase. Even on machines
with small number of cores (e.g. 4) the cookie generation shrinks from
min/max/med/avg (ns) of 22/50/40/38.9 down to 10/35/14/17.3 when run
in parallel from multiple CPUs.

Signed-off-by: Daniel Borkmann <daniel@iogearbox.net>
Signed-off-by: Alexei Starovoitov <ast@kernel.org>
Reviewed-by: Eric Dumazet <edumazet@google.com>
Acked-by: Martin KaFai Lau <kafai@fb.com>
Cc: Eric Dumazet <eric.dumazet@gmail.com>
Link: https://lore.kernel.org/bpf/8a80b8d27d3c49f9a14e1d5213c19d8be87d1dc8.1601477936.git.daniel@iogearbox.net
2021-12-30 14:50:11 +08:00
Peng Hao 146ec6c502 ext4: introduce direct I/O write using iomap infrastructure
commit: 378f32bab3714f04c4e0c3aee4129f6703805550

This patch introduces a new direct I/O write path which makes use of
the iomap infrastructure.

All direct I/O writes are now passed from the ->write_iter() callback
through to the new direct I/O handler ext4_dio_write_iter(). This
function is responsible for calling into the iomap infrastructure via
iomap_dio_rw().

Code snippets from the existing direct I/O write code within
ext4_file_write_iter() such as, checking whether the I/O request is
unaligned asynchronous I/O, or whether the write will result in an
overwrite have effectively been moved out and into the new direct I/O
->write_iter() handler.
The block mapping flags that are eventually passed down to
ext4_map_blocks() from the *_get_block_*() suite of routines have been
taken out and introduced within ext4_iomap_alloc().

For inode extension cases, ext4_handle_inode_extension() is
effectively the function responsible for performing such metadata
updates. This is called after iomap_dio_rw() has returned so that we
can safely determine whether we need to potentially truncate any
allocated blocks that may have been prepared for this direct I/O
write. We don't perform the inode extension, or truncate operations
from the ->end_io() handler as we don't have the original I/O 'length'
available there. The ->end_io() however is responsible fo converting
allocated unwritten extents to written extents.

In the instance of a short write, we fallback and complete the
remainder of the I/O using buffered I/O via
ext4_buffered_write_iter().

The existing buffer_head direct I/O implementation has been removed as
it's now redundant.

[ Fix up ext4_dio_write_iter() per Jan's comments at
  https://lore.kernel.org/r/20191105135932.GN22379@quack2.suse.cz -- TYT ]

Signed-off-by: Matthew Bobrowski <mbobrowski@mbobrowski.org>
Reviewed-by: Jan Kara <jack@suse.cz>
Reviewed-by: Ritesh Harjani <riteshh@linux.ibm.com>
Signed-off-by: Theodore Ts'o <tytso@mit.edu>
2021-11-30 10:40:21 +08:00
Bin Lai f91939a5b3 sli/cpu: introduce sli max latency metrics
Sli max latency is lightweight latency monitor metrics, the monitor
tool could get the data with a little overhead. And the accuracy of
the metrics is controlled by the sampling frequency of the monitor
tool. When the performance jitter was occurred in the system, we can
get some help from these lateny metris.

Signed-off-by: Bin Lai <robinlai@tencent.com>
Reviewed-by: benbjiang<benbjiang@tencent.com>
Reviewed-by: Mengmeng Chen <bauerchen@tencent.com>
Reviewed-by: mungerjiang <mungerjiang@tencent.com>
2021-10-29 06:51:52 +00:00
mungerjiang 95a8c85b7f sli/cpu: introduce longsys check
When a process is running in the system space, it cann't be preempted until
it return to userspace, even if the process should be reschedule and other
process was ready to run(because the server system close the kernel preempt
by default). This schedule delay may impact the performance of waiting process,
therefor we introduce the longsys check to collect the schedule delay information
of process. The longsys indication could help us spot the possible performance
jitter in the system.

Signed-off-by: Munger jiang <mungerjiang@tencent.com>
Signed-off-by: Bin Lai <robinlai@tencent.com>
Reviewed-by: benbjiang<benbjiang@tencent.com>
Reviewed-by: Bauerchen <bauerchen@tencent.com>
2021-10-29 06:51:52 +00:00
mungerjiang a76d37e2f4 sli: enable sli in cgroup v1
Signed-off-by: Munger jiang <mungerjiang@tencent.com>
Reviewed-by: Bauerchen <bauerchen@tencent.com>
2021-10-29 06:51:52 +00:00
Bin Lai b879b1b64b sli: Introduce memory and sched latency stat infrastructure
Signed-off-by: Munger jiang <mungerjiang@tencent.com>
2021-10-29 06:51:52 +00:00
Bauerchen c07a089a0f tqos/mbuf: export mbuf interface to cpuacct subsys
In cgroup V1, mbuf only exist in cpuacct subsys, so we may need a help
function to store buffer just according process task_struct.

Signed-off-by: Bauerchen <bauerchen@tencent.com>
Reviewed-by: mungerjiang<mungerjiang@tencent.com>
Reviewed-by: benbjiang<benbjiang@tencent.com>
2021-10-29 06:51:52 +00:00
Bauerchen 7795ccbf82 tqos/mbuf: write a help function to get cgroup struct from task_struct.
In order to support cgroup V1 with mbuf and sli, we need a special
cgroup structure, cpuacct subsys cgroup is nice.

We only prepare mbuf and sli for cpuacct cgroup in V1. so first find cpuacct
cgroup and if return NULl or root, just find df1_cgrp.

Signed-off-by: Bauerchen <bauerchen@tencent.com>
Reviewed-by: mungerjiang<mungerjiang@tencent.com>
Reviewed-by: mungerjiang<mungerjiang@tencent.com>
2021-10-29 06:51:52 +00:00
Bauerchen 7600f2ca44 tqos/rqm: Tencent Quality Monitor Buffer
Providing back up buffer for Quality Monitor, can be used to catch key
context when abnormal jitters occur. And application can also use it
to detect system env exception.

Signed-off-by: Bauerchen <bauerchen@tencent.com>
Reviewed-by: Jiang Biao <benbjiang@tencent.com>
Reviewed-by: Bin Lai <robinlai@tencent.com>
2021-10-29 06:51:52 +00:00
Peng Hao 3df720cae4 mm/lru: revise the comments of lru_lock
Use this new function to replace repeated same code, no func change.

When testing for relock we can avoid the need for RCU locking if we simply
compare the page pgdat and memcg pointers versus those that the lruvec is
holding. By doing this we can avoid the extra pointer walks and accesses of
the memory cgroup.

In addition we can avoid the checks entirely if lruvec is currently NULL.

Signed-off-by: Alexander Duyck <alexander.h.duyck@linux.intel.com>
Signed-off-by: Alex Shi <alex.shi@linux.alibaba.com>
2021-10-29 06:51:21 +00:00
Peng Hao 1b7478d8d0 mm/lru: replace pgdat lru_lock with lruvec lock
This patch moves per node lru_lock into lruvec, thus bring a lru_lock for
each of memcg per node. So on a large machine, each of memcg don't
have to suffer from per node pgdat->lru_lock competition. They could go
fast with their self lru_lock.

After move memcg charge before lru inserting, page isolation could
serialize page's memcg, then per memcg lruvec lock is stable and could
replace per node lru lock.

In func isolate_migratepages_block, compact_unlock_should_abort and
lock_page_lruvec_irqsave are open coded to work with compact_control.
Also add a debug func in locking which may give some clues if there are
sth out of hands.

Daniel Jordan's testing show 62% improvement on modified readtwice case
on his 2P * 10 core * 2 HT broadwell box.
https://lore.kernel.org/lkml/20200915165807.kpp7uhiw7l3loofu@ca-dmjordan1.us.oracle.com/

On a large machine with memcg enabled but not used, the page's lruvec
seeking pass a few pointers, that may lead to lru_lock holding time
increase and a bit regression.

Hugh Dickins helped on patch polish, thanks!
[flyingpeng: compatibility modification]

Signed-off-by: Alex Shi <alex.shi@linux.alibaba.com>
Signed-off-by: Peng Hao <flyingpeng@tencent.com>
2021-10-29 06:51:21 +00:00
Peng Hao 821adde85c mm: memcontrol: charge swapin pages on instantiation
Right now, users that are otherwise memory controlled can easily escape
their containment and allocate significant amounts of memory that they're
not being charged for.  That's because swap readahead pages are not being
charged until somebody actually faults them into their page table.  This
can be exploited with MADV_WILLNEED, which triggers arbitrary readahead
allocations without charging the pages.

There are additional problems with the delayed charging of swap pages:

1. To implement refault/workingset detection for anonymous pages, we
   need to have a target LRU available at swapin time, but the LRU is not
   determinable until the page has been charged.

2. To implement per-cgroup LRU locking, we need page->mem_cgroup to be
   stable when the page is isolated from the LRU; otherwise, the locks
   change under us.  But swapcache gets charged after it's already on the
   LRU, and even if we cannot isolate it ourselves (since charging is not
   exactly optional).

The previous patch ensured we always maintain cgroup ownership records for
swap pages.  This patch moves the swapcache charging point from the fault
handler to swapin time to fix all of the above problems.

v2: simplify swapin error checking (Joonsoo)

[hughd@google.com: fix livelock in __read_swap_cache_async()]
[flyingpeng: port New API]
  Link: http://lkml.kernel.org/r/alpine.LSU.2.11.2005212246080.8458@eggly.anvils
Signed-off-by: Johannes Weiner <hannes@cmpxchg.org>
Signed-off-by: Hugh Dickins <hughd@google.com>
Signed-off-by: Peng Hao <flyingpeng@tencent.com>
2021-10-29 06:51:21 +00:00
Peng Hao f784a238c7 mm/compaction: do page isolation first in compaction
Currently, compaction would get the lru_lock and then do page isolation
 which works fine with pgdat->lru_lock, since any page isoltion would
 compete for the lru_lock. If we want to change to memcg lru_lock, we
 have to isolate the page before getting lru_lock, thus isoltion would
 block page's memcg change which relay on page isoltion too. Then we
 could safely use per memcg lru_lock later.

 The new page isolation use previous introduced TestClearPageLRU() +
 pgdat lru locking which will be changed to memcg lru lock later.

 Hugh Dickins <hughd@google.com> fixed following bugs in this patch's
 early version:

 Fix lots of crashes under compaction load: isolate_migratepages_block()
 must clean up appropriately when rejecting a page, setting PageLRU again
 if it had been cleared; and a put_page() after get_page_unless_zero()
 cannot safely be done while holding locked_lruvec - it may turn out to
 be the final put_page(), which will take an lruvec lock when PageLRU.
 And move __isolate_lru_page_prepare back after get_page_unless_zero to
 make trylock_page() safe:
 trylock_page() is not safe to use at this time: its setting PG_locked
 can race with the page being freed or allocated ("Bad page"), and can
 also erase flags being set by one of those "sole owners" of a freshly
 allocated page who use non-atomic __SetPageFlag().

Suggested-by: Johannes Weiner <hannes@cmpxchg.org>
Signed-off-by: Alex Shi <alex.shi@linux.alibaba.com>
2021-10-29 06:51:21 +00:00
Peng Hao 4702fd8039 mm/lru: introduce TestClearPageLRU
Currently lru_lock still guards both lru list and page's lru bit, that's
ok. but if we want to use specific lruvec lock on the page, we need to
pin down the page's lruvec/memcg during locking. Just taking lruvec
lock first may be undermined by the page's memcg charge/migration. To
fix this problem, we could clear the lru bit out of locking and use
it as pin down action to block the page isolation in memcg changing.

So now a standard steps of page isolation is following:
	1, get_page(); 	       #pin the page avoid to be free
	2, TestClearPageLRU(); #block other isolation like memcg change
	3, spin_lock on lru_lock; #serialize lru list access
	4, delete page from lru list;
The step 2 could be optimzed/replaced in scenarios which page is
unlikely be accessed or be moved between memcgs.

This patch start with the first part: TestClearPageLRU, which combines
PageLRU check and ClearPageLRU into a macro func TestClearPageLRU. This
function will be used as page isolation precondition to prevent other
isolations some where else. Then there are may !PageLRU page on lru
list, need to remove BUG() checking accordingly.

There 2 rules for lru bit now:
1, the lru bit still indicate if a page on lru list, just in some
   temporary moment(isolating), the page may have no lru bit when
   it's on lru list.  but the page still must be on lru list when the
   lru bit set.
2, have to remove lru bit before delete it from lru list.

As Andrew Morton mentioned this change would dirty cacheline for page
isn't on LRU. But the lost would be acceptable in Rong Chen
<rong.a.chen@intel.com> report:
https://lore.kernel.org/lkml/20200304090301.GB5972@shao2-debian/

[flyingpeng: compatibility code porting]
Suggested-by: Johannes Weiner <hannes@cmpxchg.org>
Signed-off-by: Alex Shi <alex.shi@linux.alibaba.com>
Signed-off-by: Peng Hao <flyingpeng@tencent.com>
2021-10-29 06:51:21 +00:00
Peng Hao b18f5a48c9 Add VM_WARN_ON_ONCE_PAGE() macro.
Since readahead page is charged on memcg too, in theory we don't have to
check this exception now. Before safely remove them all, add a warning
for the unexpected !memcg.

Signed-off-by: Alex Shi <alex.shi@linux.alibaba.com>
2021-10-29 06:51:21 +00:00
linuszeng 06d37efeef mm: pagecache limit per cgroup support
Signed-off-by: Chen Xiaoguang <xiaoggchen@tencent.com>
Signed-off-by: Zeng Jingxiang <linuszeng@tencent.com>
Reviewed-by: Bin Lai <robinlai@tencent.com>
Reviewed-by: bauerchen <bauerchen@tencent.com>
2021-10-11 14:11:13 +08:00
caelli 298d9fcc44 ovl: check return value before using lookup_one_len_unlocked
[commit]: 1434a65ea625c51317ccdf06dabf4bd27d20fa10

After calling lookup_one_len_unlocked inside ovl_lookup_positive_unlocked,
the return dentry pointer is used before checking validity, which may represent
error code.

Signed-off-by: caelli <caelli@tencent.com>
Reviewed-by: benbjiang <benbjiang@tencent.com>
Reviewed-by: mengensun <mengensun@tencent.com>
2021-08-18 07:27:21 +00:00
Ingo Molnar 6d928aeddf compiler.h: Move instrumentation_begin()/end() to new <linux/instrumentation.h> header
commit d19e789f068b3d633cbac430764962f404198022 upstream.
Backport summary: for 5.4 kernel fsgsbase support.

Linus pointed out that compiler.h - which is a key header that gets included in every
single one of the 28,000+ kernel files during a kernel build - was bloated in:

  655389666643: ("vmlinux.lds.h: Create section for protection against instrumentation")

Linus noted:

 > I have pulled this, but do we really want to add this to a header file
 > that is _so_ core that it gets included for basically every single
 > file built?
 >
 > I don't even see those instrumentation_begin/end() things used
 > anywhere right now.
 >
 > It seems excessive. That 53 lines is maybe not a lot, but it pushed
 > that header file to over 12kB, and while it's mostly comments, it's
 > extra IO and parsing basically for _every_ single file compiled in the
 > kernel.
 >
 > For what appears to be absolutely zero upside right now, and I really
 > don't see why this should be in such a core header file!

Move these primitives into a new header: <linux/instrumentation.h>, and include that
header in the headers that make use of it.

Unfortunately one of these headers is asm-generic/bug.h, which does get included
in a lot of places, similarly to compiler.h. So the de-bloating effect isn't as
good as we'd like it to be - but at least the interfaces are defined separately.

No change to functionality intended.

Reported-by: Linus Torvalds <torvalds@linux-foundation.org>
Signed-off-by: Ingo Molnar <mingo@kernel.org>
Link: https://lore.kernel.org/r/20200604071921.GA1361070@gmail.com
Cc: Thomas Gleixner <tglx@linutronix.de>
Cc: Borislav Petkov <bp@alien8.de>
Cc: Peter Zijlstra <peterz@infradead.org>
(cherry picked from commit d19e789f068b3d633cbac430764962f404198022)
Signed-off-by: Ethan Zhao <Haifeng.Zhao@intel.com>
2021-08-17 06:29:11 +00:00
Peter Zijlstra 2be3d3ef5b x86, kcsan: Add __no_kcsan to noinstr
commit 5ddbc4082e1072eeeae52ff561a88620a05be08f upstream.
Backport summary: for 5.4 kernel fsgsbase support.

The 'noinstr' function attribute means no-instrumentation, this should
very much include *SAN. Because lots of that is broken at present,
only include KCSAN for now, as that is limited to clang11, which has
sane function attribute behaviour.

Signed-off-by: Peter Zijlstra (Intel) <peterz@infradead.org>
(cherry picked from commit 5ddbc4082e1072eeeae52ff561a88620a05be08f)
Signed-off-by: Ethan Zhao <Haifeng.Zhao@intel.com>
2021-08-17 06:29:11 +00:00
Denise Cheng ed00a88eb6 kabi: fix kabi reserve place bug
change place of some reserve words because they are placed under dynamic array

Signed-off-by: Denise Cheng <denisecheng@tencent.com>
2021-08-05 09:25:30 +00:00
Kaixu Xia 38e331e321 locking/percpu-rwsem: Use this_cpu_{inc,dec}() for read_count
The __this_cpu*() accessors are (in general) IRQ-unsafe which, given
that percpu-rwsem is a blocking primitive, should be just fine.

However, file_end_write() is used from IRQ context and will cause
load-store issues on architectures where the per-cpu accessors are not
natively irq-safe.

Fix it by using the IRQ-safe this_cpu_*() for operations on
read_count. This will generate more expensive code on a number of
platforms, which might cause a performance regression for some of the
other percpu-rwsem users.

If any such is reported, we can consider alternative solutions.

Fixes: 70fe2f48152e ("aio: fix freeze protection of aio writes")
Signed-off-by: Hou Tao <houtao1@huawei.com>
Signed-off-by: Peter Zijlstra (Intel) <peterz@infradead.org>
Acked-by: Will Deacon <will@kernel.org>
Acked-by: Oleg Nesterov <oleg@redhat.com>
Link: https://lkml.kernel.org/r/20200915140750.137881-1-houtao1@huawei.com
2021-07-15 15:53:01 +08:00
Fan Du 15675d9235 mm: Add 'mprotect' hook to struct vm_operations_struct
Upstream commit id:
95bb7c42ac8a94ce3d0eb059ad64430390351ccb

Background
==========

1. SGX enclave pages are populated with data by copying from normal memory
   via ioctl() (SGX_IOC_ENCLAVE_ADD_PAGES), which will be added later in
   this series.
2. It is desirable to be able to restrict those normal memory data sources.
   For instance, to ensure that the source data is executable before
   copying data to an executable enclave page.
3. Enclave page permissions are dynamic (just like normal permissions) and
   can be adjusted at runtime with mprotect().

This creates a problem because the original data source may have long since
vanished at the time when enclave page permissions are established (mmap()
or mprotect()).

The solution (elsewhere in this series) is to force enclave creators to
declare their paging permission *intent* up front to the ioctl().  This
intent can be immediately compared to the source data’s mapping and
rejected if necessary.

The “intent” is also stashed off for later comparison with enclave
PTEs. This ensures that any future mmap()/mprotect() operations
performed by the enclave creator or done on behalf of the enclave
can be compared with the earlier declared permissions.

Problem
=======

There is an existing mmap() hook which allows SGX to perform this
permission comparison at mmap() time.  However, there is no corresponding
->mprotect() hook.

Solution
========

Add a vm_ops->mprotect() hook so that mprotect() operations which are
inconsistent with any page's stashed intent can be rejected by the driver.

Signed-off-by: Sean Christopherson <sean.j.christopherson@intel.com>
Co-developed-by: Jarkko Sakkinen <jarkko@kernel.org>
Signed-off-by: Jarkko Sakkinen <jarkko@kernel.org>
Signed-off-by: Borislav Petkov <bp@suse.de>
Acked-by: Jethro Beekman <jethro@fortanix.com>
Acked-by: Dave Hansen <dave.hansen@intel.com>
Acked-by: Mel Gorman <mgorman@techsingularity.net>
Acked-by: Hillf Danton <hdanton@sina.com>
Cc: linux-mm@kvack.org
Link: https://lkml.kernel.org/r/20201112220135.165028-11-jarkko@kernel.org
2021-07-12 08:56:46 +00:00
denisecheng 320a3efd01 kabi: reserve space for kabi
Signed-off-by: denisecheng <denisecheng@tencent.com>
Signed-off-by: denisecheng <denisecheng@tencent.com>
2021-06-23 07:24:18 +00:00
Kaixu Xia d253516e87 Revert "modules: mark ref_module static"
This reverts commit af16ca3bc7.
2021-05-21 10:00:22 +08:00