Skip to content

feat(aggregation): Add PCD - #786

Open
DaraVaram wants to merge 5 commits into
SimplexLab:mainfrom
DaraVaram:feat/pcd-aggregator
Open

DaraVaram wants to merge 5 commits into
SimplexLab:mainfrom
DaraVaram:feat/pcd-aggregator

Conversation

@DaraVaram

Copy link
Copy Markdown

Following the discussion in #665, this adds PCD and PCDWeighting, from our TMLR 2026 paper Not All Objectives Are Born Equal: Priority-Constrained Descent for Hierarchical Multi-Objective Optimization (arXiv:2606.29521, reference code: DaraVaram/priority-constrained-descent).

PCD treats the first row $g_1$ of the Jacobian as the primary objective. It follows $g_1$ as closely as possible, while each secondary objective gets at least a fraction $\tau$ of the first-order progress that a step along its own normalized gradient would give:

$$\tilde d = \arg\min_d \frac{1}{2}|d - \tilde g_1|^2 \quad \text{s.t.} \quad \tilde g_j^\top d \geq \tau |\tilde g_j|^2, \quad j = 2, \dots, m$$

Each $\tilde g_i$ is $g_i$ divided by the square root of a bias-corrected EMA of $|g_i|^2$, which makes $\tau$ scale-free. The output is $\tilde d$ rescaled to $|g_1|$. As agreed in #665, the normalization is done inside PCD for now.

Implementation

  • _GramianWeighting + GramianWeightedAggregator, stateful and non-differentiable, following GradVac. The EMA is computed on cpu in float64, reset() clears it, and it resets automatically when $m$ changes.
  • Source: the paper (Def. 4.5 and Algorithm 1), which our reference package also follows. Our original research code differed slightly: it normalized with $1/(\sqrt{\hat v} + \epsilon)$ instead of $1/\sqrt{\hat v + \epsilon}$.
  • Solver: the paper enumerates working sets (App. B.2), which is exponential in $m$ (10 to 17 s per call on the 20-row matrices of _inputs.py). Here the QP is solved with the dual active-set method of Goldfarb and Idnani, using only entries of the Gramian, in plain torch, with no new dependency (a few ms on the same matrices).
  • Special cases: $g_1 = 0$ gives $0$; mutually incompatible constraints (only possible for $m \geq 3$) give $g_1$; nan or inf in the input give nan and leave the EMA unchanged.
  • tau is a float or a vector with one value per secondary objective, validated to be in $[0, 1]$ as in the paper. Defaults: tau=0.02 (the paper's representative setting), beta=0.999, eps=1e-8.
  • No third-party code is adapted (the reference package is ours), so there is no NOTICES entry.

Verification

  • Against the reference package, over 50-step sequences ($m = 2$ to $5$, scalar and vector $\tau$), the output matches pcd.PCD.apply_gradients to $1.3 \times 10^{-10}$ relative error in float64. On about 90k random QPs, the solver gives the same direction as the reference solve_qp and detects infeasibility identically, except in 2 cases with secondary gradients within about $10^{-4}$ rad of anti-parallel.
  • pytest tests/unit -W error passes in float32 and float64, ruff check, ruff format and pre-commit pass, and the docs build with -W -n. I don't have a CUDA device, so I have not run the GPU tests. The docs doctest has one failure on my machine, in the autogram engine example, which this PR does not touch.

Properties

As discussed in #665, PCD is not permutation-invariant and not non-conflicting. With a fresh state and $\epsilon = 0$, scaling a secondary row by a positive constant leaves the output unchanged, and scaling the primary row scales the output (tested). The output has norm $|g_1|$ (tested).

The primary is always the first row. I can add a parameter to choose it if you prefer.

DaraVaram and others added 3 commits September 29, 2026 02:17
* Add PCD and PCDWeighting (Priority-Constrained Descent, TMLR 2026)
* Solve the QP with a Gramian-only Goldfarb-Idnani dual active-set method
* Add tests, docs page, README entry and changelog entry

Co-Authored-By: Claude Opus 5.5 <[email protected]>
…e key pattern

- nan or inf in the Gramian now gives nan weights and leaves the moving average unchanged.
  Before, it returned zeros and corrupted the moving average, so every later call returned
  zero until reset().
- Track the state with _state_key and _ensure_state, as in GradVac.
- Define the symbols of the docstring in a bullet list.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
- Cast the moving-average buffer after _ensure_state, as in GradVac, so that ty can type it.
- Remove the README table row: the docs page does not exist on torchjd.org/stable until the
  next release, which is when the table gets updated.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
@DaraVaram
DaraVaram marked this pull request as ready for review September 29, 2026 09:12

@ValerianRey ValerianRey left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Very good already, should be easy to merge!

I have a few comments / questions.

Also, could you add PCD to the interactive plotter (tests/plots/interactive_plotter.py) and double-check that it behaves as you would expect?
To run it, use:

export UV_NO_SYNC=1
export PYTHONPATH="$PYTHONPATH:<path_to_TorchJD>/tests"
uv run python tests/plots/interactive_plotter.py

Comment thread tests/unit/aggregation/test_pcd.py Outdated
Comment thread tests/unit/aggregation/test_pcd.py Outdated
Comment thread src/torchjd/aggregation/_pcd.py Outdated
Comment thread src/torchjd/aggregation/_pcd.py Outdated
return f"{self.__class__.__name__}(tau={self.tau!r}, beta={self.beta!r}, eps={self.eps!r})"


def _solve_qp(G: Tensor, taus: Tensor, tol: float = 1e-9) -> Tensor:

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

FYI I didn't review _solve_qp in details (it's not my area of expertise). I leave it to you to be sure that this is correct and it matches your official implementation / paper.

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Makes sense. Here is what I checked:

  • test_solve_qp_satisfies_kkt_conditions checks the KKT conditions of the returned multipliers. For this convex QP they certify optimality, independently of how the solver gets there.
  • Against our reference implementation, which enumerates the working sets as in the paper, it gives the same direction on about 90k random problems ($m = 2$ to $5$, including infeasible and degenerate ones), and the same outputs over 50-step sequences to $3 \times 10^{-10}$ in float64.
  • In the interactive plotter, it gives the closed-form answers (see my main reply).

@ValerianRey ValerianRey added cc: feat Conventional commit type for new features. package: aggregation labels Sep 29, 2026
@ValerianRey

Copy link
Copy Markdown
Member

@PierreQuinton do you agree to add this + release v0.18.0 with it?

- Cap the iterations of the active-set solver at 10 m (it needs about m in practice, and the cap
  never binds on 53,640 random and ill-conditioned problems with up to 100 objectives). When it is
  reached, the current iterate is returned.
- Raise the tolerance of the linear-dependence test to the precision of the Gramian (10 machine
  epsilons of its dtype, at least 1e-9). With float32 Gramians, rounding errors made dependent
  gradients look independent, which gave huge multipliers and a wrong direction.
- Return zero when the direction is zero up to the rounding errors of the Gramian, instead of
  rescaling rounding noise to the norm of the primary gradient.
- Give a zero scale to objectives whose gradient has always been zero, which avoids 1 / 0 when
  eps = 0.
- Parametrize the norm and closed-form tests on typical and scaled matrices.
- Add PCD to the interactive plotter.

Co-Authored-By: Claude Opus 5.5 <[email protected]>
@DaraVaram

Copy link
Copy Markdown
Author

Thanks for the review! I added PCD to the interactive plotter (76e2a61) and checked it against the answers worked out by hand:

Configuration PCD output Expected
Default ($g_1 = (0, 1)$, $g_2 = (1, -1)$, $g_3 = (1, 0)$) $(0.727, 0.687)$ the projection of $\tilde g_1$ onto the two constraints, rescaled to $\lVert g_1 \rVert$
Same, with $g_2$ scaled by 8.5 and $g_3$ by 0.05 $(0.727, 0.687)$ unchanged, since each gradient is normalized
Same, with $g_1$ scaled by 5 $(3.634, 3.434)$ scaled by 5
Both secondaries at 30° from $g_1$ $g_1$ $g_1$, since the constraints already hold
Opposite secondaries $g_1$ $g_1$, since the constraints are infeasible
$g_2 = -g_1 / 2$, $g_3 \perp g_1$ $(-1.414, 1.414)$ with $g_1 = (2, 0)$ $\tilde d = (-0.02, 0.02)$ rescaled to $\lVert g_1 \rVert$

The last row is expected: when a secondary directly opposes the primary, PCD gives it its $\tau$-share of progress, and the resulting short direction is rescaled to the norm of $g_1$.

Parametrizing the tests on typical_matrices and scaled_matrices also found two real issues, now fixed:

  • With float32 Gramians, the solver's linear-dependence test (tolerance 1e-9) was below the rounding error, so dependent gradients could look independent, giving huge multipliers and a wrong direction. Its tolerance now follows the precision of the Gramian (10 machine epsilons, at least 1e-9). Float64 behavior is unchanged, and float32 outputs are unchanged whenever the secondary gradients are well conditioned.
  • A direction that is zero up to rounding errors is now treated as zero, instead of rescaling the noise to $\lVert g_1 \rVert$.

I also fixed a division by zero with eps=0 when a secondary gradient has always been zero.

@DaraVaram DaraVaram closed this Sep 29, 2026
@DaraVaram DaraVaram reopened this Sep 29, 2026
Comment thread tests/unit/aggregation/test_pcd.py Outdated

@ValerianRey ValerianRey left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM.

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

cc: feat Conventional commit type for new features. package: aggregation

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants