Repository navigation
[BUG] High CPU usage at random moments #539
Description
Activity
Thanks for opening your first issue here! Be sure to follow the relevant issue templates, or risk having this issue marked as invalid.
Just came here to support this. I run almost an identical setup, my OS is Debian 13 + same Docker compose.
I have the exact same issue. Daily 1x high CPU usage for 1h and no apparent reason other than php process.
If I go back in Proxmox logs it's been going on for a while. Can't narrow down when it started, but at least a couple of months if not longer. I usually keep to latest nextcloud image by linuxserver.Reacted by mumbo2030 and Sergey Kibishupstream nextcloud-repo has a similar issue ticket
for me it renders my nextcloud completely unusable and i have to restart it..
I think its better we go to the upstream repo issues and request support there.
Or am i wrong annd they are different issues ?
Reacted by mumbo2030 and Sergey Kibishupstream nextcloud-repo has a similar issue ticket
for me it renders my nextcloud completely unusable and i have to restart it..
I think its better we go to the upstream repo issues and request support there.
Or am i wrong annd they are different issues ?
I am unsure, didn't the upstream Nextcloud devs resolve it or something?
your linuxserver-nextcloud image version is oudated.
we are at lscr.io/linuxserver/nextcloud:version-33.0.4
which is Linuxserver.io version:- 33.0.4-ls433 Build-date:- 2026-05-28T18:26:57+00:00
ID | sha256:a1208ad00ce8228de38d60fcbeacae3fe895c21a49192c927f1294d4c4076abfupstream nextcloud-repo has a similar issue ticket
nextcloud/server#59036
for me it renders my nextcloud completely unusable and i have to restart it..
I think its better we go to the upstream repo issues and request support there.
Or am i wrong annd they are different issues ?I am unsure, didn't the upstream Nextcloud devs resolve it or something?
i am not aware of any resolution of this bug at the upstream nextcloud repo.
As far as i know is neither the cause nor a workaround known.
Its still possible these are different issues.
nextcloud/server#59036your linuxserver-nextcloud image version is oudated. we are at lscr.io/linuxserver/nextcloud:version-33.0.4 which is Linuxserver.io version:- 33.0.4-ls433 Build-date:- 2026-05-28T18:26:57+00:00 ID | sha256:a1208ad00ce8228de38d60fcbeacae3fe895c21a49192c927f1294d4c4076abf
You're right, I copy/pasted the wrong image. I was already using the
33.0.4:ls433. But I updated it in the running config rather than in the k3s manifest I copied from into the opening post.Thanks for pointing that out though!
Reacted by perahokyI keep having this issue from time to time with podman under fedora 44 and daily auto updates of the image. I dunno what's causing it, as usually the container seems unable to respond. Usually it happens randomly and a manual restart of the container fixes it, but it keeps coming up multiple times a month, so it is annoying.
I'm gonna compare config.php files at work tomorrow, as for some reason it only happens to one of my two instances
Compared config files today. Except that loglevel was set to 2 instead of 1 they seemed pretty much identical. I dunno if this might be related, but I use the preview generator addon, to get preview images of video files etc through ffmpeg.
Hi All. Same issue with docker running on Debian, service stuck and huge CPU usage. Restarting the container fixes the issue.
Unfortunately I don't have any additional detail to help to debug the issue. I think this looks like the issue is the mainstream repository, since when the CPU usage spikes, the process consuming it is the internal php-fpm binary from nextcloud.
I will try to get more details the next time if I have the chance.
might explain why it doesn't happen on my other instance, as that one is running on arm, while the faulty one is x86. So it could be a bug with the php-fpm version.
Hi all, I'm running
34.0.0-ls438on unraid experiencing the same issue. Nothing on the logs other than high cpu usage up to 50 - 70%. Oddly enough, i noticed cospend (although is not enabled in settings) was out of date, and updating it seemed to have calmed it down. Unsure if related but thought worth mentioning. Will report back if the CPU load issue resurfaces and I get more data to share.inspired by you @jaychu I disabled an app (epubviewer that has not been updated for several months) - i am running 33.0.5 - and the CPU usages calmed down. Let's see if it is related, but generally many Nextcloud problems are related to apps
I've finally set up a redis with valkey after noticing that without the preview generator app working under 34 my instance got super unusable in general as it had to regenerate all previews for some reason. Maybe this might give some benefits for this as well?
17 remaining items
Three follow-ups on my analysis above — one correction of my own, and two things already in this
thread that I think the JIT explanation now accounts for.Correction to my post
I wrote that "PHP upstream default is
opcache.jit_buffer_size=0, which means the JIT is off by
default. Setting a non-zero buffer is what turns it on." That was true up to PHP 8.3 but not for
PHP 8.4, which is what this image runs. From the
PHP 8.4 UPGRADING notes:The JIT config defaults changed from
opcache.jit=tracingandopcache.jit_buffer_size=0to
opcache.jit=disableandopcache.jit_buffer_size=64M. This does not change behaviour — JIT
remains disabled by default.So the conclusion is unchanged — upstream PHP 8.4 still ships with the JIT off — but the
mechanism is different: it is nowopcache.jit=disablethat keeps it off, not a zero buffer. Two
practical consequences:- My suggestion "remove both lines from
00_opcache.ini" still gives the right result on PHP 8.4:
the defaultopcache.jit=disabletakes over and the JIT stays off. - But anyone applying the workaround should set
opcache.jit=disable, not just
opcache.jit_buffer_size=0, because on 8.4 the buffer size alone no longer controls it. The
snippet I posted sets both, so it is safe either way — I verified afterwards that the executable
r-xsJIT mapping disappears from the workers entirely.
@jaychu's Sunday 2am observation is, I think, exactly this
Started monitoring my uptime kuma and realized the first time this happened, it crashed around
2:13am EST on Sunday. Now a week later (today) it crashed at 2:17am (EST). [...] It's as if there's
a job that runs around 2am on SundaysThere is.
/etc/crontabs/rootruns logrotate daily at 02:00, and/etc/logrotate.d/php-fpmis a
weeklystanza whosepostrotateiss6-svc -t /run/service/svc-php-fpm— SIGTERM, i.e. a full
php-fpm restart, not a log reopen. Once a week, at ~02:00, the whole pool is restarted and the
opcache/JIT shared segment is recreated from scratch. Nothing appears in the Docker logs.That is the same event as a Watchtower update as far as the JIT is concerned, which is why some
people here correlate the failure with updates (#536) and others see it "at random" — the weekly one
is invisible unless you look at/config/log/php/error.log, where it shows up as:[16-Aug-2026 02:00:00] NOTICE: Terminating ... / exiting, bye-bye! [16-Aug-2026 02:00:02] NOTICE: fpm is running, pid 1559On my instances 5 out of 5 incidents were preceded within hours by a php-fpm restart, and the two we
had filed as "cause unknown" were the two where the restart came from logrotate.If anyone wants a quick check on their own instance: the weekly restart times are readable from the
mtimes of/config/log/nginx/access.log.Nand/config/log/php/error.log.N. Compare them with when
your instance died.@jerheij's log line is consistent, and so is the ARM observation
Maximum execution time of 3600 seconds exceeded at [...] CompressionMiddleware.php#66Different file from the one my workers were frozen in (
PresetManager.php:57), same signature: a
worker that burned 3600 s of CPU and got killed bymax_execution_time. A miscompiled trace can
root at whatever happened to be hot at the time, so the reported location moving between instances is
expected — and it is why grepping the Nextcloud log for a single file name never converged on
anything.@4liceD's "faulty on x86, fine on arm" is suggestive too, though I want to be careful with it: PHP's
JIT does support arm64 (since 8.1), so this is not an "arm has no JIT" situation. But the x86_64 and
arm64 JIT backends are separate code, and theC=1digit in1255is an x86-specific AVX flag, so a
backend-specific miscompilation would look exactly like that. Worth noting, not worth concluding
from.On the
request_terminate_timeoutPR@peterge-misoft's finding and @4liceD's PR are worth merging on their own merits, independently of
the JIT question.request_terminate_timeout=0is why a single runaway worker stays runaway forever
instead of being reaped — it is what turns "one bad request" into "the pool is gone and the host is
at load 117". It will not prevent the miscompilation, but it turns a total outage into a degradation.- My suggestion "remove both lines from
No need to repeat what everyone else is saying here - seeing this same issue with container on my Unraid server, too.
No need to repeat what everyone else is saying here - seeing this same issue with container on my Unraid server, too.
I am sorry but your reply refers to my previous posts ? If yes, i don't really get what i am repeating exactly, i am suggesting (AI but still) a fix that seems to be the root cause of the issue, i haven't seen any suggestion for similar fixes in the previous posts, unless i missed it (i am talking about the bullet "Suggested changes to the image")
By the way, i posted my replies more than a week ago, i can confirm that on the instances where i applied the fix, i did not have the issue anymore, on a single instance i did not patch, i got the issue this night. That suggests for now, that this fix is valid. Of course, if the team gives it a bless. I will continue to monitor it that way, and will report.I am sorry but your reply refers to my previous posts ? If yes, i don't really get what i am repeating exactly, i am suggesting (AI but still) a fix that seems to be the root cause of the issue, i haven't seen any suggestion for similar fixes in the previous posts, unless i missed it (i am talking about the bullet "Suggested changes to the image")
I read it as: "There's no need for me to give a big explanation, which just repeats what everyone here has been saying". :)
By the way, i posted my replies more than a week ago, i can confirm that on the instances where i applied the fix, i did not have the issue anymore, on a single instance i did not patch, i got the issue this night. That suggests for now, that this fix is valid.
Maybe it's time to open up a PR, then? :)
Reacted by blaine07Just to note, there is an open PR for this that was provided by @4liceD ( #542 )
PR's used to get auto built so the submitter and/or we could test what was submitted, unfortunately we changed that recently due to some security concerns.
Originally it did seem that this issue was related to a nextcloud bug but that might be the case since there have been more recent which had supposedly fixed the bug but it seems to still exist.
I do need to get around to upgrading my personal instance to latest and see if I experience the same issues.
I am sorry but your reply refers to my previous posts ? If yes, i don't really get what i am repeating exactly, i am suggesting (AI but still) a fix that seems to be the root cause of the issue, i haven't seen any suggestion for similar fixes in the previous posts, unless i missed it (i am talking about the bullet "Suggested changes to the image")
I read it as: "There's no need for me to give a big explanation, which just repeats what everyone here has been saying". :)
By the way, i posted my replies more than a week ago, i can confirm that on the instances where i applied the fix, i did not have the issue anymore, on a single instance i did not patch, i got the issue this night. That suggests for now, that this fix is valid.
Maybe it's time to open up a PR, then? :)
I am not against opening the PR, but i would suggest something i did not really made myself, but AI. I would just have checked it, is that tolerable ?
I have had this happening too, I edited
/php/www2.conftopm.max_children = 4 pm.start_servers = 4 pm.min_spare_servers = 2 pm.max_spare_servers = 4 pm.max_requests = 500And I haven't had it happening since.
Mine is running on an old i7-4790K with 6 out of 8 threads available to the container.
- added a commit that references this issue
on Sep 2, 2026 I have had this happening too, I edited
/php/www2.conftopm.max_children = 4 pm.start_servers = 4 pm.min_spare_servers = 2 pm.max_spare_servers = 4 pm.max_requests = 500And I haven't had it happening since.
Mine is running on an old i7-4790K with 6 out of 8 threads available to the container.
I have had this happening too, I edited
/php/www2.conftopm.max_children = 4 pm.start_servers = 4 pm.min_spare_servers = 2 pm.max_spare_servers = 4 pm.max_requests = 500And I haven't had it happening since.
Mine is running on an old i7-4790K with 6 out of 8 threads available to the container.
Mine says this? And it’s still doing it
pm = dynamic
pm.max_children = 660
pm.start_servers = 355
pm.min_spare_servers = 305
pm.max_spare_servers = 455
pm.max_requests = 500Better data than I had last week. I run four of these instances. Three have had opcache.jit disabled since 16 Aug and haven't had a single incident since. The fourth one I simply forgot to patch, and it went to 100% CPU again yesterday and needed the usual restart. Not a deliberate control group, but it's the same split @ErikDB87 described: patched fine, unpatched dies.
On the patched one that used to fail most often, php-fpm has restarted 8 times since, including the weekly logrotate restarts on 23 and 30 Aug, which is exactly the event that used to set it off. Zero upstream timed out in nginx. Before the fix it was roughly one incident every two or three restarts.
On the pool settings, I think the thread has answered that one by itself. @Th3M1k3y is at max_children = 4 and hasn't seen it since; @blaine07 is at 660 and still gets it. I've had it at 5, at 20 and at 120, and I already had pm.max_requests = 500 set when it last blew up. That's the whole range from 4 to 660 with the same failure, so I don't think worker count is the variable.
It fits mechanically too: a worker stuck in an infinite loop never finishes its request, so max_requests never recycles it, and with request_terminate_timeout at 0 fpm never kills it either. What the worker count changes is where the damage lands. With 120 my whole host went to load 117; with 4 on a smaller box Nextcloud still dies but the machine stays usable, so it looks better than it is. And a few days without a crash is unfortunately inside the normal gap between incidents, which is what has made this so hard to pin down. Several of us (me included) have "fixed" it more than once already.
@blaine07 unrelated to the bug, but pm.start_servers = 355 will spawn 355 workers at startup and hold at least 305 idle. That's a lot of RAM for no gain, and it hands this bug 355 processes to chew through. Worth bringing back down whatever else you do.
@ErikDB87 you nudged me to open a PR, so I did: #545, for the JIT change. It's complementary to #542 (request_terminate_timeout), which is worth merging too, but that one is a seatbelt rather than a fix: it turns a total outage into a slow instance, it doesn't prevent the miscompilation.
Reacted by ErikDB87Better data than I had last week. I run four of these instances. Three have had opcache.jit disabled since 16 Aug and haven't had a single incident since. The fourth one I simply forgot to patch, and it went to 100% CPU again yesterday and needed the usual restart. Not a deliberate control group, but it's the same split @ErikDB87 described: patched fine, unpatched dies.
On the patched one that used to fail most often, php-fpm has restarted 8 times since, including the weekly logrotate restarts on 23 and 30 Aug, which is exactly the event that used to set it off. Zero upstream timed out in nginx. Before the fix it was roughly one incident every two or three restarts.
On the pool settings, I think the thread has answered that one by itself. @Th3M1k3y is at max_children = 4 and hasn't seen it since; @blaine07 is at 660 and still gets it. I've had it at 5, at 20 and at 120, and I already had pm.max_requests = 500 set when it last blew up. That's the whole range from 4 to 660 with the same failure, so I don't think worker count is the variable.
It fits mechanically too: a worker stuck in an infinite loop never finishes its request, so max_requests never recycles it, and with request_terminate_timeout at 0 fpm never kills it either. What the worker count changes is where the damage lands. With 120 my whole host went to load 117; with 4 on a smaller box Nextcloud still dies but the machine stays usable, so it looks better than it is. And a few days without a crash is unfortunately inside the normal gap between incidents, which is what has made this so hard to pin down. Several of us (me included) have "fixed" it more than once already.
@blaine07 unrelated to the bug, but pm.start_servers = 355 will spawn 355 workers at startup and hold at least 305 idle. That's a lot of RAM for no gain, and it hands this bug 355 processes to chew through. Worth bringing back down whatever else you do.
@ErikDB87 you nudged me to open a PR, so I did: #545, for the JIT change. It's complementary to #542 (request_terminate_timeout), which is worth merging too, but that one is a seatbelt rather than a fix: it turns a total outage into a slow instance, it doesn't prevent the miscompilation.
@mauro2306 sorry a little OT.
What would you recommend I go to here?
pm = dynamic
pm.max_children = 660
pm.start_servers = 355
pm.min_spare_servers = 305
pm.max_spare_servers = 455
pm.max_requests = 500@blaine07 This issue affected my instance constantly for quite a while. Based on an earlier comment, I added only the following lines to
/config/php/php-local.inia couple of weeks ago and have not seen the issue since:opcache.jit=disable opcache.jit_buffer_size=0I'm currently running 34.0.3
php/php-src#21243
This seems to be an upstream php issue. i put the link above. Probably we can upgrade the php version or downgrade it to 8.3 to fix it.I can confirm this issue, which occurs randomly after running add-on updates. In my case, all workers and the CPU are at full capacity, causing the website to display the message “No server available.”
Update: in my case updating the Mailapp as a part of an Bulkupdate seems to exeed workers
Metadata
Metadata
Assignees
Labels
Type
Projects
- StatusShow more project fieldsIssues
Is there an existing issue for this?
Current Behavior
Hello, this is a recreation of the very similar bug report link.
I have not seen the errors in the logfiles and unfortunately I can't recheck them. But the symptoms are the same as in the before mentioned tickets just at random times.
I've had it occur twice over the past 5 days, the container just randomly starts using a lot of CPU resources and Nextcloud becomes unavailable. The only thing that seems to resolve it is restarting the container.
Example of the CPU spike on the server in question:

Expected Behavior
No CPU usage spike and application remaining available.
Steps To Reproduce
It seems to be happening randomly.
Environment
CPU architecture
x86-64
Docker creation
Container logs