It seems that after the 6.12.40 yocto kernel release, the cyclic test performance, observed when executed on isolated cores with the realtime kernel enabled, suffers a significant degradation especially up to version 6.12.57 (around 3-5%). --- Mainline kernel results: linux-mainline: 6.12.40 SMP PREEMPT_RT x86_64 Average:1882.00 Maximum Latency:6592 Six nines percentile:3904.17 Number of overflows:0 Median:1874.17 linux-mainline: 6.12.57 SMP PREEMPT_RT x86_64 Average:1864.67 Maximum Latency:6077 Six nines percentile:3791.00 Number of overflows:0 Median:1853.17 Yocto kernel results: linux-yocto: 6.12.40-rt6+ SMP PREEMPT_RT x86_64 Average:1863.67 Maximum Latency:20738 Six nines percentile:5551.92 Number of overflows:0 Median:1847.17 linux-yocto: 6.12.57-rt6+ SMP PREEMPT_RT x86_64 Average:1933.00 Maximum Latency:6039 Six nines percentile:4753.00 Number of overflows:0 Median:1921.33 linux-yocto: 6.12.77-rt6+ SMP PREEMPT_RT x86_64 Average:1875.92 Maximum Latency:16850 Six nines percentile:4321.75 Number of overflows:0 Median:1867.33 --- It seems that after the 6.12.57 release, the performance degradation was reduced significantly even if there isn't yet to the mainline or 6.12.40 levels yet (as we can see it with the 6.12.77 release). Tested environment: HARDWARE: Intel(R) Xeon(R) Gold 6433N (32 cores with SMP enabled) OS: Debian Bullseye GRUB_CMDLINE_LINUX_DEFAULT="iommu=pt nmi_watchdog=0 softlockup_panic=0 intel_iommu=on selinux=0 enforcing=0 softdog.soft_panic=1 systemd.unified_cgroup_hierarchy=0 user_namespace.enable=1 biosdevname=0 skew_tick=1 rcutree.kthread_prio=21 nopti nospectre_v2 nospectre_v1 nohz_full=1-15,33-47 isolcpus=nohz,domain,managed_irq,1-15,33-47 rcu_nocbs=1-31,33-63 kthread_cpus=0,32 irqaffinity=16-31,48-63 audit=0 audit_backlog_limit=8192 intel_pstate=none hugepagesz=1G hugepages=10 hugepagesz=2M hugepages=0 default_hugepagesz=1G crashkernel=2048M apparmor=0 security=apparmor TEST: ./cyclictest --priority 95 --nsecs --histofall 40000 --smi --histfile /home/user/cyclictest/hist-file-60-mins.txt --duration 1800 --affinity 1-15,33-47 --threads 30 --mainaffinity 0 SYSCTL OPTIMIZATIONS: kernel.sched_rt_runtime_us = -1 kernel.timer_migration = 0 With that in mind, for the 6.12.57 version, removing the following scheduler changes we are applying on top of mainline does indeed restore the performance up to the 6.12.40 level (as measured by cyclictest): Commit ID Title 1 b4daf14cca7fc60c27c30e802eb545a8fc7a715b tracing: Record task flag NEED_RESCHED_LAZY. 2 73f315ca082365d5e7faded40947341e2178b013 sched/isolation: really align nohz_full with rcu_nocbs 3 3197eeac91bd298e48f6acc564bc5c9598428b1d sched: Add TIF_NEED_RESCHED_LAZY infrastructure 4 b33e6b9e378d8ebe70401d1b2f7e42aa544f8177 sched: Add Lazy preemption model 5 9cef2865c9d7aab8b983f84c9fcb9ea5fa9eaf34 sched: Enable PREEMPT_DYNAMIC for PREEMPT_RT 6 27522cbdb904f34a55211b31dc85516ca96c39a2 sched, x86: Enable Lazy preemption 7 bb8f3750d225b33e8d1cdb6acc04c1eda07917b9 sched: Add laziest preempt model 8 6fb8608877a083454560b3eaa806e8f399eb1492 sched: Fixup the IS_ENABLED check for PREEMPT_LAZY
Using the latest 6.12.79 build, we can observe an increase in the performance degradation as compared to 6.12.77. The results are reproducible. linux-yocto: 6.12.79-rt6+ SMP PREEMPT_RT x86_64 Average:1892.83 Maximum Latency:26303 Six nines percentile:5102.58 Number of overflows:0 Median:1878.1 With that in mind, all results were obtained as an average of at least 8 cyclic test runs, each running for 60 minutes (1 hour). Regarding the BIOS settings for the tested machine: C1E: Disabled C-States: Disabled Monitor/Mwait: Enabled CPU C1 Auto Demotion: Disabled CPU C1 Auto UnDemotion: Disabled Package C-States: Disabled
Catalin, This is on whinlatter, right? Can you reproduce the problem on 6.18 (master) since whinlatter is almost at the end of support.
> Catalin, > This is on whinlatter, right? Not exactly. This is the latest version of the 6.12 yocto kernel branch put on top of Debian 11. The problem is the performance degradation observed between different Yocto kernel versions starting 6.12.40 up to 6.12.79. > Can you reproduce the problem on 6.18 (master) since whinlatter is almost at the end of support. Based on the latest 6.18 Yocto kernel vs mainline branches, the observed degradation is around 1% on Sapphire Rapids machines using the Yocto RT branch. linux-yocto: 6.18.20-rt3+ SMP PREEMPT_RT x86_64 Average:1877.72 Maximum Latency:24463 Six nines percentile:4969.89 Number of overflows:0 Median:1872.5 linux-mainline: 6.18.20 SMP PREEMPT_RT x86_64 Average:1861.08 Maximum Latency:27431 Six nines percentile:5196.33 Number of overflows:0 Median:1853.83
Paul Barker asks: What do you mean by "on top of Debian 11" ? Are you using linux-yocto-rt? How was that configure and built? Can you provide more info about the linux-mainline kernel config as well.
> What do you mean by "on top of Debian 11" ? Only the yocto kernel was compiled and used in a Debian Bullseye OS deployment. > Are you using linux-yocto-rt? Yes. The exact variant that I was using for Yocto 6.12.57 is: - VARIANT: Yocto 6.12.57: URL: https://git.yoctoproject.org/linux-yocto/snapshot/linux-yocto-af2d3ab81402c14f81072715d771097a0dfcb427.tar.gz SHA256: 2e685328b68c2659326580f5dd247db27d40e55ea62f857b90267ab8d4931514 > How was that configure and built? I've attached the kernel config to this bug. Other than that, the build didn't contained any additional patches or modifications. Just a regular kernel build. > Can you provide more info about the linux-mainline kernel config as well. For the mainline Linux kernel, I just used Linus' tree using the tag for v6.12.57 when built. The kernel config was the same as the one used in the previous build. The defaults were used for any flag what was not identical.
Created attachment 5220 [details] kernel config
Yocto project probably does not have this specific configuration in the YP AB. We're not experts in RT performance so you may be better off talking with the linux-rt community. Sorry, we can't help more, eh.