Bug 14871

Summary: Run the SSD stress test by FIO will causes system auto reboot
Product: [Runtime] General Runtime Reporter: Jacky Lee <s5814108>
Component: General RuntimeAssignee: Unassigned <unassigned>
Status: RESOLVED INVALID QA Contact:
Severity: critical    
Priority: Undecided CC: randy.macleod, s5814108
Version: unspecified   
Target Milestone: ---   
Hardware: x86   
OS: x86_64   
Whiteboard:
OS type for building Yocto: --- Type of Regression: ---
Verified: Documentation change: No (bug/feature does not impact docs)
Attachments:
Description Flags
FIO debug enabled log none

Description Jacky Lee 2022-08-05 02:29:20 UTC
I met a problem when executing FIO stress test on a RAID0 which built from 6x SSDs thru mdadm under Yocto OS(Linux version 5.4.143-rt63-alm-64-abl), below is the information:

1. 6x PCIe NVMe SSD are the same vendor and model which is with 1.02TB automotive grade.
2. FIO parameter used for the test: fio --filename=/dev/md127 --direct=1 --rw=randrw --bs=64k --ioengine=libaio --iodepth=64 --runtime=43200 --numjobs=16 --time_based --group_reporting --name=randomrw --eta-newline=1
3. The system auto restart after ~30 minutes run.
4. The question is that I'd want to know why it would cause the system auto restart randomly, is that a software issue or software limitation, or a hardware issue? Would you suggest on how to isolate the issue?

I found that:

1. Both RAID0 and non-RAID mode are failed with same FIO parameter(only the --filename is with different target).
2. When issue occurs, re-run the test by same FIO parameter will encounter the issue again immediately, except that you format the SSD, but will fail again after ~30mins run.
3. Did not enounter this issue with given FIO --size parameter.
4. Did not enounter this issue with given FIO --debug=all parameter. (log attached)
4. When issue occurs, the SSD encounters over current issue. (accept: under 2A, over current: 5.5A)

I've raised the same issue to FIO GitHub and the developer replied me that this could be a kernel bug or it could be a hardware issue. It's certainly not a fio issue.

Please confirm if it is a kernel issue or not, thanks.

Jacky
Comment 1 Jacky Lee 2022-08-05 02:31:03 UTC
Created attachment 4891 [details]
FIO debug enabled log
Comment 2 Jacky Lee 2022-08-05 02:33:58 UTC
The OS we used is with Poky.
Comment 3 Randy MacLeod 2022-08-11 14:35:30 UTC
fio is part of meta-openembedded which doesn't track bugs in this bugzilla.
Please discuss on openembedded-devel@lists.openembedded.org and if there's an issue with the kernel or some package in oe-core/yocto then re-open this issue.
Comment 4 Randy MacLeod 2022-08-11 15:29:53 UTC
Jacky, my previous response was written during the YP bug triage meeting so I was a bit terse. In general, the Yocto team doesn't help with general debugging but I do have a few questions that might help move the bug along.

1. Are you able to reproduce the issue using just poky + meta-oe to add fio?
2. What branch/commit ID do you see that with: master, kirkstone, dunfell - provide exact commit since the branch content changes.
3. What is your init system: sysvinit, systemd?
4. Is there no log on the system console or in /var/log that is useful?
5. Finally and probably more importantly, are you patching the kernel at all?