Repository navigation
perf(client): read prime-rl's packed completion_logprobs - #174
Open
faresobeid wants to merge 3 commits into
Open
faresobeid wants to merge 3 commits into
faresobeid wants to merge 3 commits into
Conversation
This was referenced Oct 6, 2026
faresobeid
marked this pull request as ready for review
October 7, 2026 14:31
There was a problem hiding this comment.
Cursor Bugbot has reviewed your changes and found 1 potential issue.
❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, enable autofix in the Cursor dashboard.
Reviewed by Cursor Bugbot for commit 8a9619d. Configure here.
ApprovabilityVerdict: Approved at Macroscope's review found this PR approvable — This is a localized, backward-compatible performance path for decoding packed completion logprobs; existing response handling remains unchanged and NumPy was already a required dependency. The new behavior is covered by focused tests for valid and invalid payloads. Notes:
You can add or adjust custom eligibility rules. Learn more. |
This branch has not been deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.

Lets
generate()read prime-rl's packed per-token logprobs.prime-rl's
/inference/v1/generateserver now returns the sampled-token logprobs as one base64 float32 array,choice.completion_logprobs = {data, shape, dtype}, instead ofchoice.logprobs.content(one JSON object per token). The per-token objects were a large part of the vLLM API server's GC load under RL traffic (see the prime-rl PR)._parse_completion_logprobsreadscompletion_logprobswhen present and applies the same checks as the per-token path: count matchescompletion_ids, values finite, no-9999.0sentinel. Returns the samelist[float].sampling_maskis passed through as before. On prime-rl's server it is now a packed{ids, counts}dict; verifiers decodes it.Companions: prime-rl PrimeIntellect-ai/prime-rl#3900 / verifiers PrimeIntellect-ai/verifiers#2778. Land renderers -> verifiers -> prime-rl; the prime-rl submodule pins must move to the merged commits.
Validation
tests/test_client.py: 32 passed (one new test for the packed path, incl. the sentinel check).ruff check/ruff format --check(0.15.12) on the changed files: clean.Note
Medium Risk
Changes how sampled-token logprobs are parsed from engine responses; bugs could corrupt RL training signals, but behavior is backward compatible and guarded by existing validation plus new tests.
Overview
Adds support for prime-rl’s packed per-token completion logprobs on
/inference/v1/generateresponses, sogenerate()can avoid parsing the heavy OpenAI-stylechoice.logprobs.contentlist under RL traffic.When
choice.completion_logprobsis present ({data, shape, dtype}with base64 float32),_parse_completion_logprobsdecodes it via a new_parse_packed_completion_logprobshelper and still returns the samelist[float]as before. Validation mirrors the legacy path: length must matchcompletion_ids, values must be finite, and vLLM’s-9999missing-logprob sentinel is rejected. Stock vLLM responses without the field keep usinglogprobs.contentunchanged.A comment notes that prime-rl may also return
sampling_maskas packed CSR{ids, counts}(still passed through; decoding stays downstream). Tests cover successful packed decoding plus sentinel and non-finite failures.Reviewed by Cursor Bugbot for commit e134d56. Bugbot is set up for automated code reviews on this repo. Configure here.
Note
Add packed
completion_logprobsdecoding to client renderer_parse_packed_completion_logprobsin renderers/client.py, which decodes base64 packed logprob payloads into a NumPy array using the supplied dtype and returns per-token floatsgeneratenow preferschoice.completion_logprobswhen present, falling back to the existingchoice.logprobs.contentparsingMalformedGenerateResponseErrorrenderers.clientnow requires NumPy as a dependencyMacroscope summarized e134d56.