Skip to content

Delayed response in creating and merging PRs using the web UI #39410

Description

@TheFriendlyCoder

Gitea Version

1.27.3

What happened?

Description

On our self-hosted instance, creating a pull request and merging a pull request both take much longer than expected (5–11+ seconds of otherwise-unaccounted-for delay), even on a very small repository (13MB,
~700 git objects, single-digit-second operations everywhere else). We traced the delay with strace and found it's caused by Gitea's own long-lived git cat-file --batch-command helper subprocess hanging mid-request. This occurs even on change sets that have very small modifications (ie: single text file with a couple of modified lines)

Gitea Version

v1.27.3 (confirmed this is the current latest stable release as of this report)

Environment

  • OS: Debian 12 (bookworm)
  • Database: PostgreSQL 15, sub-millisecond query latency confirmed via psql \timing (ruled out as a factor)
  • Git signing enabled (REQUIRE_SIGNING = true), confirmed GPG signing itself takes ~0.26s (ruled out as a factor)

What we found

Using strace -f -T -tt -p <gitea-pid>, we captured two separate occurrences of the same hang, both traced to the exact same subprocess type:

execve("/usr/bin/git", ["/usr/bin/git", "cat-file", "--batch-command"], ...)

Occurrence 1 — creating a pull request:
<... read resumed>"info refs/heads/main\n", 4096) = 21 <6.189805>
A read() on the subprocess's stdout pipe blocked for 6.19 seconds waiting for a response to info refs/heads/main that should return near-instantly. Shortly after, Gitea's own internal timeout fired
and killed the child:
<... waitid resumed>{si_signo=SIGCHLD, si_code=CLD_KILLED, si_pid=, si_status=SIGKILL, ...}
The HTTP request then completed successfully (200 OK) immediately after the kill — total request time 6920ms for repo.CompareAndPullRequestPost, essentially all of it spent in this one blocked read.

Occurrence 2 — merging that same pull request:
Identical pattern, different ref, longer hang:
<... read resumed>"info refs/pull/457/head\n", 4096) = 24 <11.222310>
11.22 seconds blocked this time, again resolved by Gitea killing the subprocess, again followed by a successful 200 OK.

Analysis

This looks like the same class of bug as several historical issues (#17096, #17991, #17992, #19448, #19454 — all describing hangs from improper pipe/goroutine handling around git cat-file, fixed years
ago) and possibly related in spirit to #29402 (a similar-shaped regression in v1.21.6). Since v1.27.3 already includes all of those historical fixes and the changelog for 1.27.0–1.27.3 has no entries
mentioning cat-file/hang/pipe/deadlock, this appears to be either a new regression or an edge case not covered by the earlier fixes.

Our working theory: Gitea keeps a persistent cat-file --batch-command process per repository for efficient object lookups, reused across concurrent goroutines. Both operations we captured (PR creation, PR
merge) involve multiple near-simultaneous git-object lookups on the same repository (diff computation, Actions workflow-trigger evaluation reading .gitea/workflows/*.yml, and — for merge — the additional
ref lookups from the merge/push/branch-delete sequence). If two goroutines write/read this shared pipe concurrently without proper synchronization, one caller could end up waiting on a response consumed by
the other — a protocol desync that only resolves when Gitea's own timeout kills the stuck process.

How are you running Gitea?

No response

Activity

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    issue/needs-feedbackFor bugs, we need more details. For features, the feature must be described in more detailtype/bug

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions