Skip to content

[BUG] High CPU usage at random moments #539

Description

@jerheij

Is there an existing issue for this?

  • I have searched the existing issues

Current Behavior

Hello, this is a recreation of the very similar bug report link.

I have not seen the errors in the logfiles and unfortunately I can't recheck them. But the symptoms are the same as in the before mentioned tickets just at random times.

I've had it occur twice over the past 5 days, the container just randomly starts using a lot of CPU resources and Nextcloud becomes unavailable. The only thing that seems to resolve it is restarting the container.

Example of the CPU spike on the server in question:
Image

Expected Behavior

No CPU usage spike and application remaining available.

Steps To Reproduce

It seems to be happening randomly.

Environment

OS: 
Ubuntu 24.04.4 LTS

Image: 
lscr.io/linuxserver/nextcloud:33.0.3-ls428

Environment:
k3s (previously experienced the exact configuration in docker as well)

CPU architecture

x86-64

Docker creation

k3s manifest:
: apps/v1
kind: Deployment
metadata:
  name: nextcloud
  namespace: cloud
spec:
  replicas: 1
  selector:
    matchLabels:
      app: nextcloud
  template:
    metadata:
      labels:
        app: nextcloud
    spec:
      containers:
        - name: nextcloud
          image: lscr.io/linuxserver/nextcloud:33.0.3-ls428
          ports:
            - containerPort: 80
          envFrom:
            - secretRef:
                name: cloud-nextcloud
          volumeMounts:
            - name: config
              mountPath: /config
            - name: custom-apps
              mountPath: /app/www/public/custom_apps
            - name: documents
              mountPath: /documents
            - name: pictures
              mountPath: /pictures
            - name: games
              mountPath: /games
            - name: it
              mountPath: /IT
            - name: ebooks
              mountPath: /ebooks
            - name: data
              mountPath: /data


Which is the k3s equivalent of this docker-compose snippet:

  nextcloud:
    image: lscr.io/linuxserver/nextcloud:33.0.3-ls428
    healthcheck:
      test: ["CMD", "curl", "https://wolkje.heijlond.com/index.php/login"]
      timeout: 2s
      interval: 30s
      retries: 1
    environment:
      - PUID=2008
      - PGID=2008
      - TZ=Europe/London
    depends_on:
      - redis
    volumes:
<same mounts>

Container logs

I am working on a way of saving logs to be able to check/paste them. But the previous bug report linked contains logs etc.

Activity

  1. github-actions commented on Jun 1, 2026

    @github-actions

    Thanks for opening your first issue here! Be sure to follow the relevant issue templates, or risk having this issue marked as invalid.

  2. arktisk-varg commented on Jun 1, 2026

    @arktisk-varg

    Just came here to support this. I run almost an identical setup, my OS is Debian 13 + same Docker compose.

    I have the exact same issue. Daily 1x high CPU usage for 1h and no apparent reason other than php process.
    If I go back in Proxmox logs it's been going on for a while. Can't narrow down when it started, but at least a couple of months if not longer. I usually keep to latest nextcloud image by linuxserver.

  3. perahoky commented on Jun 1, 2026

    @perahoky

    upstream nextcloud-repo has a similar issue ticket

    nextcloud/server#59036

    for me it renders my nextcloud completely unusable and i have to restart it..

    I think its better we go to the upstream repo issues and request support there.

    Or am i wrong annd they are different issues ?

  4. jerheij commented on Jun 1, 2026

    @jerheij
    Author

    upstream nextcloud-repo has a similar issue ticket

    nextcloud/server#59036

    for me it renders my nextcloud completely unusable and i have to restart it..

    I think its better we go to the upstream repo issues and request support there.

    Or am i wrong annd they are different issues ?

    I am unsure, didn't the upstream Nextcloud devs resolve it or something?

  5. perahoky commented on Jun 1, 2026

    @perahoky

    your linuxserver-nextcloud image version is oudated.
    we are at lscr.io/linuxserver/nextcloud:version-33.0.4
    which is Linuxserver.io version:- 33.0.4-ls433 Build-date:- 2026-05-28T18:26:57+00:00
    ID | sha256:a1208ad00ce8228de38d60fcbeacae3fe895c21a49192c927f1294d4c4076abf

    upstream nextcloud-repo has a similar issue ticket
    nextcloud/server#59036
    for me it renders my nextcloud completely unusable and i have to restart it..
    I think its better we go to the upstream repo issues and request support there.
    Or am i wrong annd they are different issues ?

    I am unsure, didn't the upstream Nextcloud devs resolve it or something?

    i am not aware of any resolution of this bug at the upstream nextcloud repo.
    As far as i know is neither the cause nor a workaround known.
    Its still possible these are different issues.
    nextcloud/server#59036

  6. jerheij commented on Jun 1, 2026

    @jerheij
    Author

    your linuxserver-nextcloud image version is oudated. we are at lscr.io/linuxserver/nextcloud:version-33.0.4 which is Linuxserver.io version:- 33.0.4-ls433 Build-date:- 2026-05-28T18:26:57+00:00 ID | sha256:a1208ad00ce8228de38d60fcbeacae3fe895c21a49192c927f1294d4c4076abf

    You're right, I copy/pasted the wrong image. I was already using the 33.0.4:ls433. But I updated it in the running config rather than in the k3s manifest I copied from into the opening post.

    Thanks for pointing that out though!

  7. 4liceD commented on Jun 1, 2026

    @4liceD

    I keep having this issue from time to time with podman under fedora 44 and daily auto updates of the image. I dunno what's causing it, as usually the container seems unable to respond. Usually it happens randomly and a manual restart of the container fixes it, but it keeps coming up multiple times a month, so it is annoying.

    I'm gonna compare config.php files at work tomorrow, as for some reason it only happens to one of my two instances

  8. 4liceD commented on Jun 2, 2026

    @4liceD

    Compared config files today. Except that loglevel was set to 2 instead of 1 they seemed pretty much identical. I dunno if this might be related, but I use the preview generator addon, to get preview images of video files etc through ffmpeg.

  9. shaggyz commented on Jun 7, 2026

    @shaggyz

    Hi All. Same issue with docker running on Debian, service stuck and huge CPU usage. Restarting the container fixes the issue.

    Unfortunately I don't have any additional detail to help to debug the issue. I think this looks like the issue is the mainstream repository, since when the CPU usage spikes, the process consuming it is the internal php-fpm binary from nextcloud.

    I will try to get more details the next time if I have the chance.

  10. 4liceD commented on Jun 7, 2026

    @4liceD

    might explain why it doesn't happen on my other instance, as that one is running on arm, while the faulty one is x86. So it could be a bug with the php-fpm version.

  11. jaychu commented on Jun 21, 2026

    @jaychu

    Hi all, I'm running 34.0.0-ls438 on unraid experiencing the same issue. Nothing on the logs other than high cpu usage up to 50 - 70%. Oddly enough, i noticed cospend (although is not enabled in settings) was out of date, and updating it seemed to have calmed it down. Unsure if related but thought worth mentioning. Will report back if the CPU load issue resurfaces and I get more data to share.

  12. TheNomad11 commented on Jun 21, 2026

    @TheNomad11

    inspired by you @jaychu I disabled an app (epubviewer that has not been updated for several months) - i am running 33.0.5 - and the CPU usages calmed down. Let's see if it is related, but generally many Nextcloud problems are related to apps

  13. 4liceD commented on Jun 21, 2026

    @4liceD

    I've finally set up a redis with valkey after noticing that without the preview generator app working under 34 my instance got super unusable in general as it had to regenerate all previews for some reason. Maybe this might give some benefits for this as well?

  14. 17 remaining items

  15. mauro2306 commented on Aug 20, 2026

    @mauro2306

    Three follow-ups on my analysis above — one correction of my own, and two things already in this
    thread that I think the JIT explanation now accounts for.

    Correction to my post

    I wrote that "PHP upstream default is opcache.jit_buffer_size=0, which means the JIT is off by
    default. Setting a non-zero buffer is what turns it on." That was true up to PHP 8.3 but not for
    PHP 8.4
    , which is what this image runs. From the
    PHP 8.4 UPGRADING notes:

    The JIT config defaults changed from opcache.jit=tracing and opcache.jit_buffer_size=0 to
    opcache.jit=disable and opcache.jit_buffer_size=64M. This does not change behaviour — JIT
    remains disabled by default.

    So the conclusion is unchanged — upstream PHP 8.4 still ships with the JIT off — but the
    mechanism is different: it is now opcache.jit=disable that keeps it off, not a zero buffer. Two
    practical consequences:

    • My suggestion "remove both lines from 00_opcache.ini" still gives the right result on PHP 8.4:
      the default opcache.jit=disable takes over and the JIT stays off.
    • But anyone applying the workaround should set opcache.jit=disable, not just
      opcache.jit_buffer_size=0, because on 8.4 the buffer size alone no longer controls it. The
      snippet I posted sets both, so it is safe either way — I verified afterwards that the executable
      r-xs JIT mapping disappears from the workers entirely.

    @jaychu's Sunday 2am observation is, I think, exactly this

    Started monitoring my uptime kuma and realized the first time this happened, it crashed around
    2:13am EST on Sunday. Now a week later (today) it crashed at 2:17am (EST). [...] It's as if there's
    a job that runs around 2am on Sundays

    There is. /etc/crontabs/root runs logrotate daily at 02:00, and /etc/logrotate.d/php-fpm is a
    weekly stanza whose postrotate is s6-svc -t /run/service/svc-php-fpm — SIGTERM, i.e. a full
    php-fpm restart, not a log reopen. Once a week, at ~02:00, the whole pool is restarted and the
    opcache/JIT shared segment is recreated from scratch. Nothing appears in the Docker logs.

    That is the same event as a Watchtower update as far as the JIT is concerned, which is why some
    people here correlate the failure with updates (#536) and others see it "at random" — the weekly one
    is invisible unless you look at /config/log/php/error.log, where it shows up as:

    [16-Aug-2026 02:00:00] NOTICE: Terminating ... / exiting, bye-bye!
    [16-Aug-2026 02:00:02] NOTICE: fpm is running, pid 1559
    

    On my instances 5 out of 5 incidents were preceded within hours by a php-fpm restart, and the two we
    had filed as "cause unknown" were the two where the restart came from logrotate.

    If anyone wants a quick check on their own instance: the weekly restart times are readable from the
    mtimes of /config/log/nginx/access.log.N and /config/log/php/error.log.N. Compare them with when
    your instance died.

    @jerheij's log line is consistent, and so is the ARM observation

    Maximum execution time of 3600 seconds exceeded at [...] CompressionMiddleware.php#66

    Different file from the one my workers were frozen in (PresetManager.php:57), same signature: a
    worker that burned 3600 s of CPU and got killed by max_execution_time. A miscompiled trace can
    root at whatever happened to be hot at the time, so the reported location moving between instances is
    expected — and it is why grepping the Nextcloud log for a single file name never converged on
    anything.

    @4liceD's "faulty on x86, fine on arm" is suggestive too, though I want to be careful with it: PHP's
    JIT does support arm64 (since 8.1), so this is not an "arm has no JIT" situation. But the x86_64 and
    arm64 JIT backends are separate code, and the C=1 digit in 1255 is an x86-specific AVX flag, so a
    backend-specific miscompilation would look exactly like that. Worth noting, not worth concluding
    from.

    On the request_terminate_timeout PR

    @peterge-misoft's finding and @4liceD's PR are worth merging on their own merits, independently of
    the JIT question. request_terminate_timeout=0 is why a single runaway worker stays runaway forever
    instead of being reaped — it is what turns "one bad request" into "the pool is gone and the host is
    at load 117". It will not prevent the miscompilation, but it turns a total outage into a degradation.

  16. blaine07 commented on Aug 27, 2026

    @blaine07

    No need to repeat what everyone else is saying here - seeing this same issue with container on my Unraid server, too.

  17. mauro2306 commented on Aug 27, 2026

    @mauro2306

    No need to repeat what everyone else is saying here - seeing this same issue with container on my Unraid server, too.

    I am sorry but your reply refers to my previous posts ? If yes, i don't really get what i am repeating exactly, i am suggesting (AI but still) a fix that seems to be the root cause of the issue, i haven't seen any suggestion for similar fixes in the previous posts, unless i missed it (i am talking about the bullet "Suggested changes to the image")
    By the way, i posted my replies more than a week ago, i can confirm that on the instances where i applied the fix, i did not have the issue anymore, on a single instance i did not patch, i got the issue this night. That suggests for now, that this fix is valid. Of course, if the team gives it a bless. I will continue to monitor it that way, and will report.

  18. ErikDB87 commented on Aug 27, 2026

    @ErikDB87

    I am sorry but your reply refers to my previous posts ? If yes, i don't really get what i am repeating exactly, i am suggesting (AI but still) a fix that seems to be the root cause of the issue, i haven't seen any suggestion for similar fixes in the previous posts, unless i missed it (i am talking about the bullet "Suggested changes to the image")

    I read it as: "There's no need for me to give a big explanation, which just repeats what everyone here has been saying". :)

    By the way, i posted my replies more than a week ago, i can confirm that on the instances where i applied the fix, i did not have the issue anymore, on a single instance i did not patch, i got the issue this night. That suggests for now, that this fix is valid.

    Maybe it's time to open up a PR, then? :)

  19. j0nnymoe commented on Aug 27, 2026

    @j0nnymoe
    Member

    Just to note, there is an open PR for this that was provided by @4liceD ( #542 )

    PR's used to get auto built so the submitter and/or we could test what was submitted, unfortunately we changed that recently due to some security concerns.

    Originally it did seem that this issue was related to a nextcloud bug but that might be the case since there have been more recent which had supposedly fixed the bug but it seems to still exist.

    I do need to get around to upgrading my personal instance to latest and see if I experience the same issues.

  20. mauro2306 commented on Aug 27, 2026

    @mauro2306

    I am sorry but your reply refers to my previous posts ? If yes, i don't really get what i am repeating exactly, i am suggesting (AI but still) a fix that seems to be the root cause of the issue, i haven't seen any suggestion for similar fixes in the previous posts, unless i missed it (i am talking about the bullet "Suggested changes to the image")

    I read it as: "There's no need for me to give a big explanation, which just repeats what everyone here has been saying". :)

    By the way, i posted my replies more than a week ago, i can confirm that on the instances where i applied the fix, i did not have the issue anymore, on a single instance i did not patch, i got the issue this night. That suggests for now, that this fix is valid.

    Maybe it's time to open up a PR, then? :)

    I am not against opening the PR, but i would suggest something i did not really made myself, but AI. I would just have checked it, is that tolerable ?

  21. Th3M1k3y commented on Aug 28, 2026

    @Th3M1k3y

    I have had this happening too, I edited /php/www2.conf to

    pm.max_children = 4
    pm.start_servers = 4
    pm.min_spare_servers = 2
    pm.max_spare_servers = 4
    pm.max_requests = 500
    

    And I haven't had it happening since.

    Mine is running on an old i7-4790K with 6 out of 8 threads available to the container.

  22. blaine07 commented on Sep 2, 2026

    @blaine07

    I have had this happening too, I edited /php/www2.conf to

    pm.max_children = 4
    pm.start_servers = 4
    pm.min_spare_servers = 2
    pm.max_spare_servers = 4
    pm.max_requests = 500
    

    And I haven't had it happening since.

    Mine is running on an old i7-4790K with 6 out of 8 threads available to the container.

    I have had this happening too, I edited /php/www2.conf to

    pm.max_children = 4
    pm.start_servers = 4
    pm.min_spare_servers = 2
    pm.max_spare_servers = 4
    pm.max_requests = 500
    

    And I haven't had it happening since.

    Mine is running on an old i7-4790K with 6 out of 8 threads available to the container.

    Mine says this? And it’s still doing it

    pm = dynamic
    pm.max_children = 660
    pm.start_servers = 355
    pm.min_spare_servers = 305
    pm.max_spare_servers = 455
    pm.max_requests = 500

  23. mauro2306 commented on Sep 3, 2026

    @mauro2306

    Better data than I had last week. I run four of these instances. Three have had opcache.jit disabled since 16 Aug and haven't had a single incident since. The fourth one I simply forgot to patch, and it went to 100% CPU again yesterday and needed the usual restart. Not a deliberate control group, but it's the same split @ErikDB87 described: patched fine, unpatched dies.

    On the patched one that used to fail most often, php-fpm has restarted 8 times since, including the weekly logrotate restarts on 23 and 30 Aug, which is exactly the event that used to set it off. Zero upstream timed out in nginx. Before the fix it was roughly one incident every two or three restarts.

    On the pool settings, I think the thread has answered that one by itself. @Th3M1k3y is at max_children = 4 and hasn't seen it since; @blaine07 is at 660 and still gets it. I've had it at 5, at 20 and at 120, and I already had pm.max_requests = 500 set when it last blew up. That's the whole range from 4 to 660 with the same failure, so I don't think worker count is the variable.

    It fits mechanically too: a worker stuck in an infinite loop never finishes its request, so max_requests never recycles it, and with request_terminate_timeout at 0 fpm never kills it either. What the worker count changes is where the damage lands. With 120 my whole host went to load 117; with 4 on a smaller box Nextcloud still dies but the machine stays usable, so it looks better than it is. And a few days without a crash is unfortunately inside the normal gap between incidents, which is what has made this so hard to pin down. Several of us (me included) have "fixed" it more than once already.

    @blaine07 unrelated to the bug, but pm.start_servers = 355 will spawn 355 workers at startup and hold at least 305 idle. That's a lot of RAM for no gain, and it hands this bug 355 processes to chew through. Worth bringing back down whatever else you do.

    @ErikDB87 you nudged me to open a PR, so I did: #545, for the JIT change. It's complementary to #542 (request_terminate_timeout), which is worth merging too, but that one is a seatbelt rather than a fix: it turns a total outage into a slow instance, it doesn't prevent the miscompilation.

  24. blaine07 commented on Sep 3, 2026

    @blaine07

    Better data than I had last week. I run four of these instances. Three have had opcache.jit disabled since 16 Aug and haven't had a single incident since. The fourth one I simply forgot to patch, and it went to 100% CPU again yesterday and needed the usual restart. Not a deliberate control group, but it's the same split @ErikDB87 described: patched fine, unpatched dies.

    On the patched one that used to fail most often, php-fpm has restarted 8 times since, including the weekly logrotate restarts on 23 and 30 Aug, which is exactly the event that used to set it off. Zero upstream timed out in nginx. Before the fix it was roughly one incident every two or three restarts.

    On the pool settings, I think the thread has answered that one by itself. @Th3M1k3y is at max_children = 4 and hasn't seen it since; @blaine07 is at 660 and still gets it. I've had it at 5, at 20 and at 120, and I already had pm.max_requests = 500 set when it last blew up. That's the whole range from 4 to 660 with the same failure, so I don't think worker count is the variable.

    It fits mechanically too: a worker stuck in an infinite loop never finishes its request, so max_requests never recycles it, and with request_terminate_timeout at 0 fpm never kills it either. What the worker count changes is where the damage lands. With 120 my whole host went to load 117; with 4 on a smaller box Nextcloud still dies but the machine stays usable, so it looks better than it is. And a few days without a crash is unfortunately inside the normal gap between incidents, which is what has made this so hard to pin down. Several of us (me included) have "fixed" it more than once already.

    @blaine07 unrelated to the bug, but pm.start_servers = 355 will spawn 355 workers at startup and hold at least 305 idle. That's a lot of RAM for no gain, and it hands this bug 355 processes to chew through. Worth bringing back down whatever else you do.

    @ErikDB87 you nudged me to open a PR, so I did: #545, for the JIT change. It's complementary to #542 (request_terminate_timeout), which is worth merging too, but that one is a seatbelt rather than a fix: it turns a total outage into a slow instance, it doesn't prevent the miscompilation.

    @mauro2306 sorry a little OT.

    What would you recommend I go to here?

    pm = dynamic
    pm.max_children = 660
    pm.start_servers = 355
    pm.min_spare_servers = 305
    pm.max_spare_servers = 455
    pm.max_requests = 500

  25. mumbo2030 commented on Sep 4, 2026

    @mumbo2030

    @blaine07 This issue affected my instance constantly for quite a while. Based on an earlier comment, I added only the following lines to /config/php/php-local.ini a couple of weeks ago and have not seen the issue since:

    opcache.jit=disable
    opcache.jit_buffer_size=0
    

    I'm currently running 34.0.3

  26. Himyth commented on Sep 11, 2026

    @Himyth

    php/php-src#21243
    This seems to be an upstream php issue. i put the link above. Probably we can upgrade the php version or downgrade it to 8.3 to fix it.

  27. SchwarzeLanze commented on Sep 14, 2026

    @SchwarzeLanze

    I can confirm this issue, which occurs randomly after running add-on updates. In my case, all workers and the CPU are at full capacity, causing the website to display the message “No server available.”

  28. SchwarzeLanze commented on Sep 16, 2026

    @SchwarzeLanze

    Update: in my case updating the Mailapp as a part of an Bulkupdate seems to exeed workers

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    No labels
    No labels

    Type

    No type

    Projects

    Milestone

    No milestone

    Relationships

    None yet

    Development

    No branches or pull requests

    Issue actions