I met a problem when executing FIO stress test on a RAID0 which built from 6x SSDs thru mdadm under Yocto OS(Linux version 5.4.143-rt63-alm-64-abl), below is the information: 1. 6x PCIe NVMe SSD are the same vendor and model which is with 1.02TB automotive grade. 2. FIO parameter used for the test: fio --filename=/dev/md127 --direct=1 --rw=randrw --bs=64k --ioengine=libaio --iodepth=64 --runtime=43200 --numjobs=16 --time_based --group_reporting --name=randomrw --eta-newline=1 3. The system auto restart after ~30 minutes run. 4. The question is that I'd want to know why it would cause the system auto restart randomly, is that a software issue or software limitation, or a hardware issue? Would you suggest on how to isolate the issue? I found that: 1. Both RAID0 and non-RAID mode are failed with same FIO parameter(only the --filename is with different target). 2. When issue occurs, re-run the test by same FIO parameter will encounter the issue again immediately, except that you format the SSD, but will fail again after ~30mins run. 3. Did not enounter this issue with given FIO --size parameter. 4. Did not enounter this issue with given FIO --debug=all parameter. (log attached) 4. When issue occurs, the SSD encounters over current issue. (accept: under 2A, over current: 5.5A) I've raised the same issue to FIO GitHub and the developer replied me that this could be a kernel bug or it could be a hardware issue. It's certainly not a fio issue. Please confirm if it is a kernel issue or not, thanks. Jacky
Created attachment 4891 [details] FIO debug enabled log
The OS we used is with Poky.
fio is part of meta-openembedded which doesn't track bugs in this bugzilla. Please discuss on openembedded-devel@lists.openembedded.org and if there's an issue with the kernel or some package in oe-core/yocto then re-open this issue.
Jacky, my previous response was written during the YP bug triage meeting so I was a bit terse. In general, the Yocto team doesn't help with general debugging but I do have a few questions that might help move the bug along. 1. Are you able to reproduce the issue using just poky + meta-oe to add fio? 2. What branch/commit ID do you see that with: master, kirkstone, dunfell - provide exact commit since the branch content changes. 3. What is your init system: sysvinit, systemd? 4. Is there no log on the system console or in /var/log that is useful? 5. Finally and probably more importantly, are you patching the kernel at all?