Conversation
* Add PCD and PCDWeighting (Priority-Constrained Descent, TMLR 2026) * Solve the QP with a Gramian-only Goldfarb-Idnani dual active-set method * Add tests, docs page, README entry and changelog entry Co-Authored-By: Claude Opus 5.5 <[email protected]>
…e key pattern - nan or inf in the Gramian now gives nan weights and leaves the moving average unchanged. Before, it returned zeros and corrupted the moving average, so every later call returned zero until reset(). - Track the state with _state_key and _ensure_state, as in GradVac. - Define the symbols of the docstring in a bullet list. Co-Authored-By: Claude Opus 5.5 <[email protected]>
- Cast the moving-average buffer after _ensure_state, as in GradVac, so that ty can type it. - Remove the README table row: the docs page does not exist on torchjd.org/stable until the next release, which is when the table gets updated. Co-Authored-By: Claude Opus 5.5 <[email protected]>
ValerianRey
left a comment
There was a problem hiding this comment.
Very good already, should be easy to merge!
I have a few comments / questions.
Also, could you add PCD to the interactive plotter (tests/plots/interactive_plotter.py) and double-check that it behaves as you would expect?
To run it, use:
export UV_NO_SYNC=1
export PYTHONPATH="$PYTHONPATH:<path_to_TorchJD>/tests"
uv run python tests/plots/interactive_plotter.py
| return f"{self.__class__.__name__}(tau={self.tau!r}, beta={self.beta!r}, eps={self.eps!r})" | ||
|
|
||
|
|
||
| def _solve_qp(G: Tensor, taus: Tensor, tol: float = 1e-9) -> Tensor: |
There was a problem hiding this comment.
FYI I didn't review _solve_qp in details (it's not my area of expertise). I leave it to you to be sure that this is correct and it matches your official implementation / paper.
There was a problem hiding this comment.
Makes sense. Here is what I checked:
-
test_solve_qp_satisfies_kkt_conditionschecks the KKT conditions of the returned multipliers. For this convex QP they certify optimality, independently of how the solver gets there. - Against our reference implementation, which enumerates the working sets as in the paper, it gives the same direction on about 90k random problems (
$m = 2$ to$5$ , including infeasible and degenerate ones), and the same outputs over 50-step sequences to$3 \times 10^{-10}$ in float64. - In the interactive plotter, it gives the closed-form answers (see my main reply).
|
@PierreQuinton do you agree to add this + release v0.18.0 with it? |
- Cap the iterations of the active-set solver at 10 m (it needs about m in practice, and the cap never binds on 53,640 random and ill-conditioned problems with up to 100 objectives). When it is reached, the current iterate is returned. - Raise the tolerance of the linear-dependence test to the precision of the Gramian (10 machine epsilons of its dtype, at least 1e-9). With float32 Gramians, rounding errors made dependent gradients look independent, which gave huge multipliers and a wrong direction. - Return zero when the direction is zero up to the rounding errors of the Gramian, instead of rescaling rounding noise to the norm of the primary gradient. - Give a zero scale to objectives whose gradient has always been zero, which avoids 1 / 0 when eps = 0. - Parametrize the norm and closed-form tests on typical and scaled matrices. - Add PCD to the interactive plotter. Co-Authored-By: Claude Opus 5.5 <[email protected]>
|
Thanks for the review! I added PCD to the interactive plotter (76e2a61) and checked it against the answers worked out by hand:
The last row is expected: when a secondary directly opposes the primary, PCD gives it its Parametrizing the tests on
I also fixed a division by zero with |
Co-authored-by: Valérian Rey <[email protected]>
Following the discussion in #665, this adds
PCDandPCDWeighting, from our TMLR 2026 paper Not All Objectives Are Born Equal: Priority-Constrained Descent for Hierarchical Multi-Objective Optimization (arXiv:2606.29521, reference code: DaraVaram/priority-constrained-descent).PCD treats the first row$g_1$ of the Jacobian as the primary objective. It follows $g_1$ as closely as possible, while each secondary objective gets at least a fraction $\tau$ of the first-order progress that a step along its own normalized gradient would give:
Each$\tilde g_i$ is $g_i$ divided by the square root of a bias-corrected EMA of $|g_i|^2$ , which makes $\tau$ scale-free. The output is $\tilde d$ rescaled to $|g_1|$ . As agreed in #665, the normalization is done inside PCD for now.
Implementation
_GramianWeighting+GramianWeightedAggregator, stateful and non-differentiable, followingGradVac. The EMA is computed on cpu in float64,reset()clears it, and it resets automatically when_inputs.py). Here the QP is solved with the dual active-set method of Goldfarb and Idnani, using only entries of the Gramian, in plain torch, with no new dependency (a few ms on the same matrices).nanorinfin the input givenanand leave the EMA unchanged.tauis a float or a vector with one value per secondary objective, validated to be intau=0.02(the paper's representative setting),beta=0.999,eps=1e-8.NOTICESentry.Verification
pcd.PCD.apply_gradientstosolve_qpand detects infeasibility identically, except in 2 cases with secondary gradients within aboutpytest tests/unit -W errorpasses in float32 and float64,ruff check,ruff formatandpre-commitpass, and the docs build with-W -n. I don't have a CUDA device, so I have not run the GPU tests. The docs doctest has one failure on my machine, in theautogramengine example, which this PR does not touch.Properties
As discussed in #665, PCD is not permutation-invariant and not non-conflicting. With a fresh state and$\epsilon = 0$ , scaling a secondary row by a positive constant leaves the output unchanged, and scaling the primary row scales the output (tested). The output has norm $|g_1|$ (tested).
The primary is always the first row. I can add a parameter to choose it if you prefer.