comment "echo "%__os_install_post %{nil}" >> ~/.rpmmacros"
to make compile success in chroot envrionment.
Signed-off-by: Hongbo Li <herberthbli@tencent.com>
offline_class was defined as an early_param. But the callback
function return 1 which lead to a warning message
"[ 0.000000] Malformed early option 'offline_class'"
when booting.
Signed-off-by: Xiaoguang Chen <xiaoggchen@tencent.com>
cat/proc/net/nf_conntrack may cause long latency.
Add cond_resched_rcu() when reading this proc file.
Signed-off-by: Hongbo Li <herberthbli@tencent.com>
upstream: 5424ea27390f ("netns: get more entropy from net_hash_mix()")
struct net are effectively allocated from order-1 pages on x86,
with one object per slab, meaning that the 13 low order bits
of their addresses are zero.
Once shifted by L1_CACHE_SHIFT, this leaves 7 zero-bits,
meaning that net_hash_mix() does not help spreading
objects on various hash tables.
For example, TCP listen table has 32 buckets, meaning that
all netns use the same bucket for port 80 or port 443.
Signed-off-by: Hongbo Li <herberthbli@tencent.com>
[tkernel2 commit c6f2e27f7ad]
add SO_MARK2 to get the MARK of flow
Signed-off-by: brookxu <brookxu@tencent.com>
Signed-off-by: Zhiping Du <zhipingdu@tencent.com>
currently sas disk detected order is decided by the first interrupt
generated order, not by slot number, fix it by slot number.
Signed-off-by: Xiaoming Gao <newtongao@tencent.com>
previously, for clusterIP type service, no SNAT is done.
However, in some case, user may add rs ip outside vpc which may require
SNAT.
To address this issue, a non snat ip table of max 64 entries are added.
Usage:
1. echo -n "0.0.0.0/0" > /proc/net/ip_vs_non_masq_cidrs will make all ip
bypass snat.
2. echo -n ":" > /proc/net/ip_vs_non_masq_cidrs will make all ip do snat.
3. echo "a.b.c.d/24:a.b.c.e/24" > /proc/net/ip_vs_non_masq_cidrs
Test case:
create a cluster with 9 PODS
172.19.0.175 172.19.0.176 172.19.0.177 172.19.0.241 172.19.0.242
172.19.0.243 172.19.0.244 172.19.0.100 172.19.0.101
0)
/proc/net/ip_vs_non_masq_cidrs is empty, curl cluster ip shall do SNAT
result: pass
1)
Three ip/32 in /proc/net/ip_vs_non_masq_cidrs
echo "172.19.0.175/32:172.19.0.176/32:172.19.0.177/32" > /proc/net/ip_vs_non_masq_cidrs
Test: curl the clusterip, and watch the tcpdump log
Expected result:
access to the ip in list no SNAT; access to the ip not in list do SNAT.
Result: pass, no leak.
2) stress test
wrk the clusterip , at the same time, run a program to modify the ip_vs_non_masq_cidrs in a loop
Expected result: curl ok. lo leak
while [ 1 ]
do
echo "172.19.0.175/32:172.19.0.176/32:172.19.0.177/32" > /proc/net/ip_vs_non_masq_cidrs
sleep 1
cat /proc/net/ip_vs_non_masq_cidrs
echo ""
echo "172.19.0.175/32" > /proc/net/ip_vs_non_masq_cidrs
sleep 1
cat /proc/net/ip_vs_non_masq_cidrs
echo ""
done
result: no leak;
3) corner test
write "0.0.0.0/0" to it.
expected result: shall not do SNAT.
result: ok
echo -n ":
expected result: do SNAT
result: ok
Signed-off-by: jianmingfan <jianmingfan@tencent.com>
[upstream commit 71f2cc64d027d7]
MIME-Version: 1.0
Content-Type: text/plain; charset=UTF-8
Content-Transfer-Encoding: 8bit
This patch fixes a possible null pointer dereference in
check_quota_exceeded, detected by the static checker smatch, with the
following warning:
fs/ceph/quota.c:240 check_quota_exceeded()
error: we previously assumed 'realm' could be null (see line 188)
Fixes: b7a2921765cf ("ceph: quota: support for ceph.quota.max_files")
Reported-by: Dan Carpenter <dan.carpenter@oracle.com>
Signed-off-by: Luis Henriques <lhenriques@suse.com>
Reviewed-by: "Yan, Zheng" <zyan@redhat.com>
Signed-off-by: Ilya Dryomov <idryomov@gmail.com>
Signed-off-by: Zhiping Du <zhipingdu@tencent.com>
This commit add cpu.offline to cpu cgroup, echo 1 > cpu.offline would
convert all tasks under this cgroup to offline task. Beside, a new
sysctl sysctl_sched_bt_ignore_cpubind is added, which makes offline
tasks ignoring CPU binding and running on any CPU.
Signed-off-by: Xiaoming Gao <newtongao@tencent.com>
Signed-off-by: Hua Liu <shookliu@tencent.com>
Signed-off-by: Xiaogguang Chen <xiaoggchen@tencent.com>
Signed-off-by: Zhiguang Peng <zgpeng@tencent.com>
Signed-off-by: Bin Fan <tombinfan@tencent.com>
Signed-off-by: He Chen <heddchen@tencent.com>
Since 'commit f719e3754ee2 ("ipvs: drop first packet to
redirect conntrack")', when a new TCP connection meet
the conditions that need reschedule, the first syn packet
is dropped, this cause one second latency for the new
connection, more discussion about this problem can easy
search from google, such as:
1)One second connection delay in masque
https://marc.info/?t=151683118100004&r=1&w=2
2)IPVS low throughput #70747
https://github.com/kubernetes/kubernetes/issues/70747
3)Apache Bench can fill up ipvs service proxy in seconds #544https://github.com/cloudnativelabs/kube-router/issues/544
4)Additional 1s latency in `host -> service IP -> pod`
https://github.com/kubernetes/kubernetes/issues/90854
5)kube-proxy ipvs conn_reuse_mode setting causes errors
with high load from single client
https://github.com/kubernetes/kubernetes/issues/81775
The root cause is when the old session is expired, the
conntrack related to the session is dropped by
ip_vs_conn_drop_conntrack. The code is as follows:
```
static void ip_vs_conn_expire(struct timer_list *t)
{
...
if ((cp->flags & IP_VS_CONN_F_NFCT) &&
!(cp->flags & IP_VS_CONN_F_ONE_PACKET)) {
/* Do not access conntracks during subsys cleanup
* because nf_conntrack_find_get can not be used after
* conntrack cleanup for the net.
*/
smp_rmb();
if (ipvs->enable)
ip_vs_conn_drop_conntrack(cp);
}
...
}
```
As shown in the code, only when condition (cp->flags & IP_VS_CONN_F_NFCT)
is true, the function ip_vs_conn_drop_conntrack will be called.
So we optimize this by following steps (Administrators
can choose the following optimization by setting
net.ipv4.vs.conn_reuse_old_conntrack=1):
1) erase the IP_VS_CONN_F_NFCT flag (it is safely because
no packets will use the old session)
2) call ip_vs_conn_expire_now to release the old session,
then the related conntrack will not be dropped
3) then ipvs unnecessary to drop the first syn packet, it
just continue to pass the syn packet to the next process,
create a new ipvs session, and the new session will related
to the old conntrack(which is reopened by conntrack as a new
one), the next whole things is just as normal as that the old
session isn't used to exist.
The above processing has no problems except for passive FTP,
for passive FTP situation, ipvs can judging from
condition (atomic_read(&cp->n_control)) and condition (cp->control).
So, for other conditions(means not FTP), ipvs should give users
the right to choose,they can choose a high performance one processing
logical by setting net.ipv4.vs.conn_reuse_old_conntrack=1. It is necessary
because most business scenarios (such as kubernetes) are very sensitive
to TCP short connection latency.
This patch has been verified on our thousands of kubernets
node servers on Tencent Inc.
Signed-off-by: YangYuxi <yx.atom1@gmail.com>
year 2015 d752c364571743d696c2a54a449ce77550c35ac5
year 2016 f719e3754ee2f7275437e61a6afd520181fdd43b
current only fix it in bpf mode. It will later be promoted to ipvs mode.
The key is that add a ref count in ct so that old/new ip_vs_conn can share it
without packet loss.
Test case
./wrk http 1.0 test, the cps increases from 1.5K to 30K.
2) solve no route to host bug
when conn_reuse_mode = 0, new connection may be redirect to rs with weight=0
if client port reuse. This cause icmp no route to host if the rs is
terminating
Test case:
1. wrk http1.0 from client
2. set rs to zero on lb, then kill the rs
in ipvs mode, you can see icmp error like
14:17:28.509454 IP 10.0.0.4 > 10.0.0.17: ICMP host 172.16.0.16 unreachable, length 68
in bpf mode, this is fixed.
Signed-off-by: jianmingfan <jianmingfan@tencent.com>
Also a switch had been added to control whether show the
subtree or not.
Signed-off-by: Chen He <heddchen@tencent.com>
Signed-off-by: Peng Zhiguang <zgpeng@tencent.com>
Signed-off-by: Chen Xiaoguang <xiaoggchen@tencent.com>
Offline tasks (BT tasks) may have some performance impact to online
tasks.
In this commit, we introduce Intel RDT features to limit offline tasks
L3 cache usage to avoid the influence caused by offline tasks.
Signed-off-by: Xiaoguang Chen <xiaoggchen@tencent.com>
Signed-off-by: He Chen <heddchen@tencent.com>
since we disable clocksource switch when tsc not stable, we need a
counter to track tsc unstable did happend.
Signed-off-by: Xiaoming Gao <newtongao@tencent.com>
commit <0aa69fd32a5f766e997ca8ab4723c5a1146efa8b>
commit <17d51b10d7773e4618bcac64648f30f12d4078fb>
bio_iov_iter_get_pages() currently only adds pages for the next non-zero
segment from the iov_iter to the bio. That's suboptimal for callers,
which typically try to pin as many pages as fit into the bio. This patch
converts the current bio_iov_iter_get_pages() into a static helper, and
introduces a new helper that allocates as many pages as
1) fit into the bio,
2) are present in the iov_iter,
3) and can be pinned by MM.
Error is returned only if zero pages could be pinned. Because of 3), a
zero return value doesn't necessarily mean all pages have been pinned.
Callers that have to pin every page in the iov_iter must still call this
function in a loop (this is currently the case).
This change matters most for __blkdev_direct_IO_simple(), which calls
bio_iov_iter_get_pages() only once. If it obtains less pages than
requested, it returns a "short write" or "short read", and
__generic_file_write_iter() falls back to buffered writes, which may
lead to data corruption.
Signed-off-by: Chunguang Xu <brookxu@tencent.com>
output more information when clocksource got unstable,
and add kernel.clocksource_switch_unstable_cs to control
whether to switch off unstable clocksource.
Signed-off-by: Xiaoming Gao <newtongao@tencent.com>
Catch the inode allocation state mismatch corruption, and then return
an corruption error when creating a file. Retry to allocate inode and
find a fine inode no..
Signed-off-by: Kaixu Xia <kaixuxia@tencent.com>