Bug 14871 - Run the SSD stress test by FIO will causes system auto reboot
Summary: Run the SSD stress test by FIO will causes system auto reboot
Status: RESOLVED INVALID
Alias: None
Product: General Runtime
Classification: Runtime
Component: General Runtime (show other bugs)
Version: unspecified
Hardware: x86 x86_64
: Undecided critical
Target Milestone: ---
Assignee: Unassigned
QA Contact:
URL:
Whiteboard:
Depends on:
Blocks:
 
Reported: 2022-08-05 02:29 UTC by Jacky Lee
Modified: 2022-08-11 15:29 UTC (History)
2 users (show)

See Also:
OS type for building Yocto: ---
Type of Regression: ---
Verified:
Documentation change: No (bug/feature does not impact docs)


Attachments
FIO debug enabled log (49.99 MB, application/x-zip-compressed)
2022-08-05 02:31 UTC, Jacky Lee
no flags Details

Note You need to log in before you can comment on or make changes to this bug.
Description Jacky Lee 2022-08-05 02:29:20 UTC
I met a problem when executing FIO stress test on a RAID0 which built from 6x SSDs thru mdadm under Yocto OS(Linux version 5.4.143-rt63-alm-64-abl), below is the information:

1. 6x PCIe NVMe SSD are the same vendor and model which is with 1.02TB automotive grade.
2. FIO parameter used for the test: fio --filename=/dev/md127 --direct=1 --rw=randrw --bs=64k --ioengine=libaio --iodepth=64 --runtime=43200 --numjobs=16 --time_based --group_reporting --name=randomrw --eta-newline=1
3. The system auto restart after ~30 minutes run.
4. The question is that I'd want to know why it would cause the system auto restart randomly, is that a software issue or software limitation, or a hardware issue? Would you suggest on how to isolate the issue?

I found that:

1. Both RAID0 and non-RAID mode are failed with same FIO parameter(only the --filename is with different target).
2. When issue occurs, re-run the test by same FIO parameter will encounter the issue again immediately, except that you format the SSD, but will fail again after ~30mins run.
3. Did not enounter this issue with given FIO --size parameter.
4. Did not enounter this issue with given FIO --debug=all parameter. (log attached)
4. When issue occurs, the SSD encounters over current issue. (accept: under 2A, over current: 5.5A)

I've raised the same issue to FIO GitHub and the developer replied me that this could be a kernel bug or it could be a hardware issue. It's certainly not a fio issue.

Please confirm if it is a kernel issue or not, thanks.

Jacky
Comment 1 Jacky Lee 2022-08-05 02:31:03 UTC
Created attachment 4891 [details]
FIO debug enabled log
Comment 2 Jacky Lee 2022-08-05 02:33:58 UTC
The OS we used is with Poky.
Comment 3 Randy MacLeod 2022-08-11 14:35:30 UTC
fio is part of meta-openembedded which doesn't track bugs in this bugzilla.
Please discuss on openembedded-devel@lists.openembedded.org and if there's an issue with the kernel or some package in oe-core/yocto then re-open this issue.
Comment 4 Randy MacLeod 2022-08-11 15:29:53 UTC
Jacky, my previous response was written during the YP bug triage meeting so I was a bit terse. In general, the Yocto team doesn't help with general debugging but I do have a few questions that might help move the bug along.

1. Are you able to reproduce the issue using just poky + meta-oe to add fio?
2. What branch/commit ID do you see that with: master, kirkstone, dunfell - provide exact commit since the branch content changes.
3. What is your init system: sysvinit, systemd?
4. Is there no log on the system console or in /var/log that is useful?
5. Finally and probably more importantly, are you patching the kernel at all?