Bug 15391

Summary: [layerindex-web] Statistics page timeout
Product: [Yocto Project Subprojects] Layer Index Reporter: Tim Orling <tim.orling>
Component: Layer IndexAssignee: Tim Orling <tim.orling>
Status: RESOLVED FIXED QA Contact: apoorv sangal <apoorvsangal>
Severity: normal    
Priority: Medium+ CC: randy.macleod, ross.burton
Version: 5.0   
Target Milestone: 6.1 M1   
Hardware: x86   
OS: Multiple   
Whiteboard:
OS type for building Yocto: --- Type of Regression: ---
Verified: Documentation change: Don't know

Description Tim Orling 2024-02-08 23:31:17 UTC
https://layers.openembedded.org/layerindex/stats/ (also available to logged in users from the Tools -> Statistics menu) times out with 504 Gateway Time-out. This means that nginx gave up waiting for Django to respond to the request.

On the server side, we can see:
[2024-02-07 18:50:59 +0000] [1] [CRITICAL] WORKER TIMEOUT (pid:28)
[2024-02-07 18:51:00 +0000] [1] [ERROR] Worker (pid:28) was sent SIGKILL! Perhaps out of memory?

In this context, WORKER is most likely a gunicorn worker.

The code behind this view is:
https://git.yoctoproject.org/layerindex-web/tree/layerindex/views.py#n1512

And the template is:
https://git.yoctoproject.org/layerindex-web/tree/templates/layerindex/stats.html

If you comment out the context['perbranch'] annotate code in StatsView class, the code succeeds:

Statistics
Overall
Layers 541
Recipes 25843 (distinct names)
Machines 1468 (distinct names)
Classes 1620 (distinct names)
Distros 194 (distinct names)

So the problem is too much data (in memory presumably) in this code:

context['perbranch'] = Branch.objects.filter(hidden=False).order_by('sort_priority').annotate(
                layer_count=Count('layerbranch', distinct=True),
                recipe_count=Count('layerbranch__recipe', distinct=True),
                class_count=Count('layerbranch__bbclass', distinct=True),
                machine_count=Count('layerbranch__machine', distinct=True),
                distro_count=Count('layerbranch__distro', distinct=True))

We should probably instead loop over each branch individually? This would hopefully allow caching of each branch and be friendlier to the gunicorn worker.