Skip to content

Add PQC support - #141

Open
MarkNSweep wants to merge 3 commits into
kroxylicious:mainfrom
MarkNSweep:pqc-support
Open

MarkNSweep wants to merge 3 commits into
kroxylicious:mainfrom
MarkNSweep:pqc-support

Conversation

@MarkNSweep

@MarkNSweep MarkNSweep commented Sep 27, 2026 •

Copy link
Copy Markdown

Proposal : Post-Quantum Cryptography (PQC) Support

This proposal adds post-quantum cryptography support to Kroxylicious, enabling protection against quantum computer attacks on data in transit.

Summary

Enables ML-KEM (FIPS 203) for TLS key exchange and ML-DSA (FIPS 204) for digital signatures, supporting hybrid and strict PQC modes across all TLS connection types.

Key Changes

Configuration:

  • New namedGroups TLS field for key exchange algorithm control
  • Convenience pqc: hybrid / pqc: strict field preventing misconfiguration
  • Kubernetes CRD changes for VirtualKafkaCluster, KafkaService, admission webhook

API:

  • ClientTlsContext extensions: negotiatedNamedGroup() and isPqcConnection()
  • Filter visibility into negotiated TLS parameters for routing and enforcement

Metrics:

  • Augments existing *_connections_total metrics with pqc label (strict/hybrid/none)
  • New *_connections_failed_total metrics for migration planning

Dependencies:

  • BouncyCastle scope change: test → runtime (provides ML-KEM/ML-DSA until Java 28 LTS)

Motivation

Regulatory compliance: acroos government, banking, healthcare etc.

Harvest-now-decrypt-later threat: Quantum adversaries capturing traffic today for future decryption

Phased migration: Proxy decouples client-side and broker-side PQC adoption timelines

Related Issues

Addresses implementation planning for:

  • #4476 - PemUtils post-quantum key type parsing
  • #4477 - namedGroups configuration
  • #4478 - TlsHttpClientConfigurator named groups support
  • #4479 - Netty runtime named groups support
  • #4480 - TlsUtil post-quantum key validation
  • #4481 - Netty PEM parser post-quantum support
  • #4482 - Record validation filter schema registry TLS
  • #4483 - PQC TLS configuration documentation
  • #4484 - Kubernetes CRD named groups support
  • #4485 - Admission webhook TLS configuration
  • #4486 - System tests for PQC configuration
  • #4487 - AWS credential provider HTTP client TLS

Dependency on Proposal #94

This proposal does not block on #94 (TLS configuration refactoring). If this lands first, it's a minor addition to #94's work. If #94 lands first, the namedGroups and pqc fields migrate to whatever replaces Tls.


🤖 Generated with Claude Code

Adds proposal for post-quantum cryptography (PQC) support in
Kroxylicious. Proposal enables ML-KEM (Kyber) key exchange and ML-DSA
(Dilithium) digital signatures for TLS connections.

Key features:
- New `namedGroups` TLS configuration field for key exchange algorithm control
- Convenience `pqc` field (hybrid/strict modes) preventing misconfiguration
- ClientTlsContext API extensions exposing negotiated named group
- PQC support for client-to-proxy and proxy-to-broker connections
- Kubernetes CRD changes for operator deployments
- Directional connection metrics with PQC labeling

Motivation driven by regulatory compliance mandates and harvest-now-decrypt-later
quantum threats. Uses BouncyCastle 1.85 (test→runtime scope change) rather
than waiting for Java 28 LTS vendor convergence.

Proposal preserves phased migration: proxy can provide PQC to clients while
brokers remain classical, or vice versa. Topic-based routing enables granular
PQC enforcement.

Assisted-by: Claude Sonnet 4.5 <[email protected]>
Signed-off-by: Adam Pilkington <[email protected]>
@MarkNSweep
MarkNSweep requested a review from a team as a code owner September 27, 2026 18:01
Signed-off-by: Adam Pilkington <[email protected]>

@tombentley tombentley left a comment

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks @MarkNSweep to taking this on!


This multi-layered ecosystem has fragmented vendor support. JVM vendors have different roadmaps for PQC adoption. Algorithm backports for ML-KEM and ML-DSA to Java 21/17 are scheduled for 4Q26, hybrid TLS key exchange is targeted for 1H27. Java 28 is the next LTS where all vendors will support PQC.

The project already depends on BouncyCastle (test scope). BouncyCastle 1.85 provides production-ready ML-KEM and ML-DSA implementations for Java 21. This proposal changes BouncyCastle scope from test to runtime rather than waiting for JVM vendor convergence, enabling PQC support now rather than waiting for JDK adoption.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think we're happy to wait for the TLS support to be back-ported to Java 21. 1H27 is not really so very far away, and in the meantime we can do the work to make the named groups configurable on current releases of OpenJDK 21.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Can't you move to Java 25 which is LTS. It's getting the support for JEP-527 at the end of October. I already tested it with the early access release. You can find some references within the PQC-related proposal I opened for Strimzi here strimzi/proposals#249

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Aside: We already run CI on Java 21 and 25. We could start writing tests today and have them skip on Java 21 if that helps progress.

We haven't scheduled the move to Java 25 at the moment. We'd use the proposal process.
We'd be mindful of the fact that Enterprises often lag. We'd want to consider the impact to those users (filter authors and those running the proxy) before deciding to cut away from Java 21.

2027 isn't too far off.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

In my Strimzi proposal I am not proposing to move the compile source to Java 21 but the runtime for the images. The components will work on Java 25 runtime but with the addition of JEP-527. We don't need any additional feature at code level from Java 25.
Of course, Kroxy can be used on bare metal, like our HTTP bridge, where you could not have Java 25 but Java 21. In such case, I would expect just component to fall back using classical algorithms instead of the PQC ones.

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The reason for proposing BC was that I prefer control over being dependent on third party plans and roadmaps. However, in the scheme of things, the provider isn't that important. I'm assuming that we would want to work with JDK 21 ? @ppatierno suggests JDK 25 but I'm not sure what the process/consequences are to bump the JDK level.

@k-wall k-wall Oct 1, 2026 •

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I think we should assume JDK 21 for this work.

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

A clarification on this, when you say 'JDK 21' you mean OpenJDK 21 which as per the roadmap (https://www.java.com/en/jre-jdk-cryptoroadmap.html) will deliver JDK 21 hybrid key exchange support in 1H27 ?

- **Simplified configuration**: One field instead of coordinating three separate fields

**Mechanics:**
- The `pqc` field is mutually exclusive with manual `protocols`, `cipherSuites`, or `namedGroups` configuration

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I can see this is convenient for the person writing the configuration, and could prevent misconfiguration. But it's not clear to the person reading a configuration exactly what pqc: hybrid means in practice. They either have to refer to the documentation, the source code, or probe the endpoint with openssl s_client or similar to see what's allowed.

It also doesn't necessarily interact so well with other higher level things, such as FIPS. IIRC FIPS does not allow ChaCh20Poly1305, so a user who wanted FIPS + PQC couldn't use the pqc: hybrid shorthand, and would have to write it all out in full.

Some other reasons for not doing this:

  • simply the increased surface to have to test and document.
  • if/when NIST specify some additional quantum-safe algorithms, does the meaning of pqc change (thus users silently are opted into using algorithms they never wanted to use)? Or stay the same (thus users don't get to use new algorithms they did want to use)

Can you point to any other systems which are providing an equivalent to this pqc shortcut?

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for the clarification that ChaCha20 is not a FIPS compliant suite, however I'm not sure that I agree with your general conclusion.

The purpose of the pqc:strict/hybrid shortcut isn't primarily to save typing. It's to let the user declare the outcome they want and delegate endpoint management to the proxy. Most users know they want support 'PQC' in some form, but far fewer can correctly hand-assemble the cipher suites, protocols and named groups needed to deliver that. Security configuration is an area that users frequently get wrong and in this case getting it subtly wrong is worse than a shortcut as it fails silently (either in an invalid endpoint or no PQC guarantees).

With regards readability, I think that potential obscurity of the shortcut isn't an issue. If I write pqc:strict, as the author I know what I'm asking for. If I read pqc:strict and (assuming that I'm not a proxy author so it's unfamiliar to me) I'd look it up in the documentation. If I instead read an explicit list of cipher suites, protocols and named groups, I still have to go look up which combination gives which guarantee, only now across many values.

On "what if NIST standardises a new PQC algorithm", these took many years to arrive, so a potential new one will come with plenty of advance warning. If one does land, then it goes into the managed definition and that is what I would expect to happen as I'm delegating this to the proxy. Any change to what pqc: strict resolves to would ship with a proxy release and its documentation. A user who wants to pin an exact algorithm set can always use the explicit namedGroups/cipher-suite fields instead. The shorthand is for people who'd rather the proxy keep the endpoint current for them.

With regards to your point on prior art, yes there are equivalents one being Red Hat's system-wide cryptographic policies. Operators pick a named policy (DEFAULT, FUTURE, FIPS, LEGACY) and the system resolves it to concrete ciphers, protocols and groups. RHEL 9.7 added a composable PQ subpolicy (https://docs.redhat.com/en/documentation/red_hat_enterprise_linux/9/html/security_hardening/using-the-system-wide-cryptographic-policies_security-hardening#post_quantum_cryptography) that layers hybrid ML-KEM and ML-DSA on top, which is very close to what I'm proposing here. There are a number of other products that allow PQC to be controlled either through a single flag/shortcut or have the user provide the individual values e.g. AWS Transfer Family where you select a named PQC security policy (https://docs.aws.amazon.com/transfer/latest/userguide/post-quantum-security-policies.html)

I think that the prior art shows that it is common to offer shortcuts for people who do not want to manage the suites/protocols/named groups directly. To me this means that the proxy should add a fips:true/false flag to perform the same management of FIPS endpoints. The proxy can than take both the pqc and fips flags together to compute and apply the correct configuration. However, a fips flag is something that we can discuss separately.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The purpose of the pqc:strict/hybrid shortcut isn't primarily to save typing. It's to let the user declare the outcome they want and delegate endpoint management to the proxy. Most users know they want support 'PQC' in some form, but far fewer can correctly hand-assemble the cipher suites, protocols and named groups needed to deliver that. Security configuration is an area that users frequently get wrong and in this case getting it subtly wrong is worse than a shortcut as it fails silently (either in an invalid endpoint or no PQC guarantees).

I think this is an issue. Currently, we have the same tls profile information duplicated across the configuration file. It will be hard for the users to spot/prevent drift.

I don't have a fully thought through idea, but what if we made a tlsProfile a named object which are then referenced elsewhere in the config.

An admin could then declare the tlsProfiles they need, once

tlsProfile:
   name: mypqc
   tlsProtocols:
      ....
   namedGroups:
      ....

and other users could reference them:

kms: VaultKmsService
kmsConfig:
  vaultTransitEngineUrl: ....
  tls:
    tlsProfile: mytls

At a Kube Level, we could use the attachment idea from Kubernetes Gateway, so that profiles could be applied to the proxy, ingress, service, filter or router. Something like:

https://gateway.envoyproxy.io/docs/concepts/gateway_api_extensions/client-traffic-policy/

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@k-wall I agree that the duplication is going to be an annoyance. However, I don't think it's just TLS configs where this is a problem (though they're probably the worst case). At the proxy level there's a relatively simple way we could do something to improve the situation without needing a lot of machinery:

# Support a top level property for holding objects to be used elsewhere
buildingBlocks:
  # Each of them defines an anchor (`&default_tls`) 
  baseTls: &default_tls
    key: ...
    trust: ...
    tlsVersions: ...

# Anywhere where we want to reuse that object we just reference it (`*default_tls`)
kms: MyExampleKms
kmsConfig:
  tls: *default_tls

# YAML also supports merging and overriding properties (`<<: *default_tls`)
kms: MyExampleKms
kmsConfig:
  tls: 
    <<: *default_tls
    tlsVersions: ...

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@MarkNSweep I don't doubt that users want something simpler than having to use lists of weird strings they don't understand. The problem is how that logical scheme gets interpreted, and by whom.

  • Sure, it communicates intent between a human writer and a human reader of the config. But the proxy has to interpret what PQC means too. As an auditor, I need to know whether the proxy's interpretation of "PQC" is the same as my definition written in terms of the KEM/named groups.
  • My point about FIPs was intended to illustrate that more than one logical schema can be at play. FIPS ∩ PQC. That would be more involved to support.
  • The RHEL crypto policy example is a good one: Why are we proposing to build an application-level equivalent to something that's being provided by the platform? The correct way to address this is not at the application level, it's using Java's own security machinery, when running on RHEL, to prevent use use of Java algorithms which don't match with the OS crypto policy. Red Hat and AWS employ people specifically to be on top of on-going evolution of things like FIPS, new NIST standardized algorithms, other commonly used by not NIST-approved algorithms. By reinventing a separate mechanism in the application layer you're putting us on the decision path for keeping up to date with such on-going evolution. We don't have the bandwidth to do a good job of that.
  • The point about new algorithms being slow to standardize is true, but it's also a bit narrow. We should consider what would happen in a security conscious organization if an algorithm got broken. That's sort of thing can happen with much less warning (in theory even overnight). In some deployments the operating organization might want to act immediately (rather than wait for an updated FIPS or whatever). I think with PQC the chances of that is actually higher than we'd otherwise expect, because these new algorithms are cryptosystems are all still quite new and relatively understudied. So in a short amount of time an algorithm which was previously thought secure and PQC needs to be removed from the list -- a list which we've hard coded into our application.

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

For the auditing consideration, ideally the interpretation of 'pqc' would be written out in the audit log (but as #85 is paused) it can be written out in the normal stream. That would allow the auditor to compare what is running with what is required. Proxy audit logs should be ingested by a SIEM system which would flag non-compliance rather than relying on a manual inspection of either the yaml or the container logs (as that also assumes the auditor has the necessary access to do that).

I understood the point about FIPS showing the interplay of schemas, it just made me consider that my driver to provide a mechanism for users to simplify their endpoint management also applies to FIPS. In this case we could take an opinionated view and remove the non-FIPS allowed one, users still have the option to manually define all the entries if that doesn't suit their needs.

You are right that for FIPS / new NIST algorithms in general there would be an overhead which confirms that it belongs in its own proposal. Focussing solely on PQC, I think that PQC algorithms are different and have less of a maintenance overhead in keeping up with any new algorithms. A new algorithm also doesn't invalidate a currently valid PQC configuration and so there isn't a requirement to add that algorithm as soon as it's approved. The rate of approval for new PQC algorithms also reduces the impact on us to keep up without the need for dedicated resources.

I agree that there is a higher risk of the new PQC algorithms being broken which is why this proposal advocates the classical+pqc key exchange algorithm. When an algorithm is broken, there will be a security bulletin issued and someone is going to have to make a change to a deployed proxy in response. There are number of ways that this can be explicitly and implicitly supported.

  • The user can change from the pqc shortcut and explicitly define what is allowed, leaving out the broken algorithm.
  • Although this proposal advocates that use of shortcut is mutually exclusive with the individual protocol/cipher/groups trio, we can relax that to support only the disallowed field. A broken algorithm then gets added as a disallowed value. This provides a mechanism to explicitly remove the algorithm from the internal list, but is only something that needs to be used in this scenario.
  • For the overhead on us, I would expect the JDK vendor to remove the broken algorithm so that even if our internal list remains unchanged, it cannot be loaded and used. I would also expect us to be aware of a broken PQC algorithm via general industry bulletins.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Re r4159892074 I don't know that we want to get into the business of letting users compose arbitrary config YAML from building blocks, it feels like a step away from something we test with a schema in future.

On named presets, would it change anything for you @tombentley if they were externalized? So the proxy doesn't know what pqc/strict is, but the user can opt to load these TLS configurations by referencing a file we ship in the binary distribution, or applying an extra custom resource in k8s. So a supplementary thing that's still essentially on the user to review, auditors can inspect it, but it makes it easier for users to get the PQC combination.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It does make me wonder how we would ferry that information to Plugins in future (like Keith's example of outbound Vault TLS), I guess we would need an API so they can retrieve TLS presets? Doesn't need to be in this proposal, but would be good to have a feel for how Plugins would interact with presets.

Comment thread proposals/141-pqc-support.md Outdated

### ML-DSA certificate support

Existing configuration will work as is for ML-DSA (Dilithium) certificates as these certificates conform to the X.509 structure and require no special TLS handshake handling. However, certificate parsing will require updating to support Dilithium key types.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Suggested change
Existing configuration will work as is for ML-DSA (Dilithium) certificates as these certificates conform to the X.509 structure and require no special TLS handshake handling. However, certificate parsing will require updating to support Dilithium key types.
Existing configuration will work as-is for ML-DSA (Dilithium) certificates as these certificates conform to the X.509 structure and require no special TLS handshake handling. However, certificate parsing will require updating to support Dilithium key types.

Comment thread proposals/141-pqc-support.md Outdated

The directional metrics support migration planning. Proxy owners can observe PQC adoption separately for client-facing and broker-facing connections, contacting application developers or Kafka cluster owners independently as migration progresses. Failed connection metrics identify PQC compatibility issues.

## PQC topic filters

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I can see that the idea has some merit, but I could be inclined to keep this proposal smaller in scope. It seems like you could easily spin this out into a separate proposal, to be delivered on its own timeline.

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I take the point on keeping the scope tight. Following other review feedback I've moved the topic filter material into a 'Future work' section, so it's no longer part of the work being delivered here. The routing and redirect it depends on don't exist yet, so it couldn't be delivered on this proposal's timeline in any case.

I'd rather keep it in this proposal as future work than spin it out into a separate proposal now. It documents the direction that the in-scope PQC endpoint work enables and shows reviewers where this is heading, which is useful context while the foundational work is being agreed. When the capability is actually picked up it would get its own proposal with the detail it needs, at which point it can be delivered on its own timeline as you suggest.


Hard failure if peer doesn't support the named group - intentional by design.

Strict mode uses `X25519MLKEM768` (hybrid) rather than pure `ML-KEM-768`. While `X25519MLKEM768` includes a classical component, it can only be successfully negotiated if the client is PQC capable. It provides better protection against broken initial PQC algorithm implementations as both X25519 and ML-KEM must be broken. "Strict" refers to no classical fallback (the named group list has only one entry), not absence of classical components. The hybrid is strictly stronger than either pure classical or pure PQC alone.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

This definition of "strict" confuses me a bit. I would expect that the used names group would not be the hybrid one but containing only ML-KEM in order to force clients to use PQC algorithm and refuse the ones using classical. Isn't it the purpose of the "strict" definition here?

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

It does seem odd to have the classical component in there, but an exchange can only happen if the client is PQC capable and can fulfil the ML-KEM part as it uses both algorithms in the handshake. So strict does mean that a client has to be using PQC and non-PQC clients will not be able to connect.

Comment thread proposals/141-pqc-support.md Outdated
Comment thread proposals/141-pqc-support.md Outdated

### PQC topic enforcement

Not all topics contain data that need to be PQC protected and not all clients are PQC capable. The proxy is able to redirect clients to PQC endpoints based on the topic being accessed. This can also be combined with existing functionality such as replacing sensitive values and allowing that access to be over non-PQC endpoints.

@k-wall k-wall Sep 28, 2026 •

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'm not really following this paragraph. A router could redirect based on a topic being accessed, but we don't have that at the moment.

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I see what you mean, the paragraph reads as if this is something the proxy can do today, which it can't. The confusion is my fault as I put it under motivations, but it isn't really motivation for PQC itself (that is the compliance and harvest-now-decrypt-later case above). It's a future capability that PQC enables, where adoption can be made granular and enables sensitive topics to be moved to PQC endpoints while insensitive data (for example an event stream of fridge temperature readings) stays on legacy. In the case where sensitive data has to go over a legacy endpoint it could be combined with other filters to redact or encrypt parts of the message. That redirect capability, and the routing it depends on, does not exist yet, so I'll move it out of motivations into a future work section and reword it so it isn't stated as a current feature.

Comment thread proposals/141-pqc-support.md Outdated
Comment thread proposals/141-pqc-support.md Outdated
Comment thread proposals/141-pqc-support.md Outdated
Comment thread proposals/141-pqc-support.md
Comment thread proposals/141-pqc-support.md Outdated

### Record encryption filter

Already uses AES-256-GCM and does not need changing. However, the KMS exchange needs to be over PQC connections, so KMS provider implementations will need to be updated.

@k-wall k-wall Sep 28, 2026 •

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

We don't actually enforce anything about the KEK. We just use what the user provides to the filter. It is up the user to choose a key type that is PQC compliant.

There's nothing to stop a user creating a KEK (un-tested):

https://developer.hashicorp.com/vault/api-docs/secret/transit#aes128-gcm96

There's an issue suggesting we validate the KEK's characteristics, but it hasn't been prioritised. Maybe PQC is a good trigger to reevaluate kroxylicious/kroxylicious#783

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

You're correct, we don't enforce anything about the KEK today, we use whatever the user configures the filter with. I'd conflated two separate things in this section. The pqc setting controls the TLS connection to the KMS, so it governs how the DEK is protected in transit, however it has no relationship to the KEK and can't mandate its characteristics. I'll update the proposal to make this clear.

With regards to scope, I think the record encryption entry in this proposal should be limited to the connections that the filter makes, in other words ensuring the proxy to KMS exchange can be PQC. Full PQC protection also depends on the user choosing a quantum-resistant KEK, as the wrapped DEK (edek) is stored with the record.

I agree that PQC is a good trigger to revisit kroxylicious/kroxylicious#783 and enhance it with validating the KEK's characteristics, of which quantum resistance would be one case.

@k-wall k-wall Oct 1, 2026 •

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Let's raise an issue about this (or update #783) and flag it discuss-sync. It might be a good motivation for a separate proposal "PQC - data at rest"

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I agree and am happy to raise an issue to create a new proposal for adding PQC support to the record encryption filter which can then reference #783 as needing enhancement. One question, this issue is to raise a proposal, does that get raised in Kroxylicious or Kroxy-design ?

### Current Kroxylicious TLS support

Kroxylicious currently supports classical TLS only. Configuration permits selecting TLS versions and cipher suites but provides no control over key exchange algorithms (named groups in TLS 1.3). The proxy cannot negotiate ML-KEM key exchange or validate ML-DSA certificates.

@k-wall k-wall Sep 28, 2026 •

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The proxy also makes TLS connections in a few other places:

  • Sasl Termination - OAUTH server (bearer tokens go over the wire)
  • Record Validation - Apicurio Registry (schema ids/schema go over the wire)

Whilst I think a harvest now, decrypt later attack may be less of a concern, any audit may flag up these connections as not PQC compliant. I think we need to think about these interfaces too. I suspect it will just be a case of wiring up our config model to theirs

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

That's a good point and I agree we need to consider the SASL termination connection to the OAuth server, so I'll bring it into scope. More generally, the requirement for PQC support applies to all TLS connections the proxy makes, whether made by the runtime or filters, now or in the future, and will also need to be added to the developers guide (especially if proposal #94 is implemented and there is no longer a shared TLS config)

Comment thread proposals/141-pqc-support.md Outdated
Comment thread proposals/141-pqc-support.md Outdated
- Test clients that support PQC and ML-DSA certificates
- Certificate injection into test infrastructure

### Metrics and observability

@k-wall k-wall Sep 28, 2026 •

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I wonder if we really need metrics. We could wait for a user to make a feature request.

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Other reviewers have already commented on removing the new metrics and with the acceptable that 'what is PQC' at the connection level is hard I think I should remove the metrics section entirely. I still think that the observability is valuable, as anyone running a PQC rollout would want to understand the blast radius before moving to strict, however that isn't enough to justify adding it now. The problem is the same one we hit with isPqcConnection(): deciding whether a connection is PQC at runtime is hard, whereas at design and configuration time it's straightforward, and a pqc label would put the proxy back in that position. If a user requests this later then we can revisit it.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Sounds good to me.

Revise the PQC support proposal (kroxylicious#141) in response to review comments:

- Clarify scope up front: the proposal covers data in transit, and note
  the data-at-rest exception for keys protected by classical cryptography,
  deferring it to separate work (#783)
- Reframe the "Data encryption at rest" section as background rather than
  proposed work
- Bring all proxy TLS connections into scope, including the SASL
  termination connection to the OAuth server, and note the developers
  guide impact (especially under proposal kroxylicious#94)
- Move PQC topic enforcement out of Motivation into a "Future work"
  section and reword it as a future capability that does not exist yet
- State that the namedGroups field is optional and defaults to the JDK's
  standard groups when unset
- Document that the order of allowed named groups is significant
- Expose negotiatedNamedGroup(), negotiatedCipherSuite() and
  negotiatedProtocol() on ClientTlsContext; drop isPqcConnection() as the
  proxy should not classify connections as PQC at runtime
- Correct the record encryption section: the pqc setting controls the KMS
  connection, not the KEK; full protection also depends on the user
  choosing a quantum-resistant KEK; defer KEK validation to #783
- Remove the metrics and observability section
- Fix "as is" to "as-is"

Co-Authored-By: Claude Opus 4.8 <[email protected]>
Signed-off-by: Adam Pilkington <[email protected]>

PQC support allows the proxy to protect data in transit with encryption resistant to attacks by quantum computers.

This proposal addresses data in transit. Data at rest is largely out of scope: the symmetric encryption the proxy uses (AES-256) is weakened but not broken by Grover's algorithm. The exception is data at rest that is protected by classical asymmetric cryptography, such as a data encryption key wrapped by a classical KEK, which Shor's algorithm can break. Addressing that is out of scope here and is tracked separately (see the Record encryption filter section and issue #783), potentially as its own proposal.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

(Non blocking) I think having the analysis done about why record encryption is (mostly) not affected is really valuable. However, I'm wondering if we move that information out as a conclusion to an investigation issue. We can then just make this proposal exclusively about "PQC Support - data in transit". I think that'd make this proposal more digestable.

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

#141 (comment)

I was suggesting a more focussed issue on adding PQC support to the record encryption filter, but I can make the issue wider as you suggest to PQC support for data at res tin the proxy.


### Support phased adoption of PQC

Changing to use Dilithium certificates has a large blast radius. It's not just an algorithm name - certificates have to be created, injected into build/test systems, and capable of being read by various Java frameworks. Certificates are used to provide point-in-time authentication and do not suffer from the harvest now, decrypt later issue. They can be adopted at a later date. Changing the TLS key exchange algorithms is the priority.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

They can be adopted at a later date. Changing the TLS key exchange algorithms is the priority.

If I understand things right, kroxylicious/kroxylicious#4475 will mean that users will be blocked from using Dilithium certificates supplied in PEM format. The user would have to use a PKCS format keystore.

This will be something we need to document until Netty catch up.

# This secures the client-to-proxy connection.
virtualClusters:
demo:
targetCluster:

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

nit: Something's off with the configuration. Tls config for the downstream side is at the gateway level.
targetCluster is actually a deprecated field too (but you don't need the upstream in the example anyway).

    gateways:
    - name: mygateway
      # ...
      tls:
          # ....

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks, I'll tidy this up in the next revision.


This multi-layered ecosystem has fragmented vendor support. JVM vendors have different roadmaps for PQC adoption. Algorithm backports for ML-KEM and ML-DSA to Java 21/17 are scheduled for 4Q26, hybrid TLS key exchange is targeted for 1H27. Java 28 is the next LTS where all vendors will support PQC.

The project already depends on BouncyCastle (test scope). BouncyCastle 1.85 provides production-ready ML-KEM and ML-DSA implementations for Java 21. This proposal changes BouncyCastle scope from test to runtime rather than waiting for JVM vendor convergence, enabling PQC support now rather than waiting for JDK adoption.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The proposal selects BouncyCastle as the PQC algorithm provider — that's a reasonable near-term choice, but it reads as a hard requirement rather than one option in a provider landscape. The set of available algorithms (named groups, cipher suites) depends on which TLS provider is active at runtime, and there are several options in the Java/Netty ecosystem (JDK built-in, BouncyCastle, native OpenSSL via netty-tcnative, etc.), each with different algorithm support, FIPS certification status, and compliance profiles.

Today Kroxylicious doesn't expose any control over provider selection. Should it? And if so, should this proposal be the one to introduce it, given it's adding configuration that's provider-dependent?

This also affects the configuration surface. Algorithm names may not be consistent across providers — JSSE and OpenSSL may use different names for the same group. If the provider is variable, the set of valid values for namedGroups is also variable. Should the proxy validate at startup that the configured algorithms are recognised by the active provider? Without that, pqc: strict with no ML-KEM support means every handshake fails silently, and pqc: hybrid without ML-KEM silently degrades to classical-only.

A related question: should the ClientTlsContext expose which provider was used for the connection, alongside the negotiated parameters? That would let filters make decisions based on the provider, not just the negotiated algorithms.

These may be out of scope for this proposal — but the provider choice affects what the proposed configuration actually means at runtime, so it's worth at least acknowledging the dependency.

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

You're right in that I have positioned BC as a requirement to drive the conversation on whether it is better to use a known proven provider that already has full PQC support or wait for back porting by JDK vendors. You raise a good point that this has also triggered consideration of something I hadn't anticipated and that is the more general feature to be able to control providers. I'd be interested to know if people think this is something they'd want or is it only an issue where you have a case like this and an application uses a non-JDK supplied provider such as BC, OpenSSL etc. ? The desire is to stop / remove those providers rather than make a positive assertion to use a particular one.

For the names, I can't tell from the JDK roadmap (https://www.java.com/en/jre-jdk-cryptoroadmap.html) but my current understanding is that JDK 27 is the first place that all the necessary PQC names are added to the standard Java names e.g. the PQC named groups for hybrid key exchange are in https://docs.oracle.com/en/java/javase/27/docs/specs/security/standard-names.html but not https://docs.oracle.com/en/java/javase/26/docs/specs/security/standard-names.html. I'm not sure that there will be differences at the moment between the JSSE and provider name, but you're right that they may not always be the case.

After the discussions on this PR, I'm thinking that the proposal should be amended to go with the JDK option rather than BC, acknowledging the trade-off between control vs the consequences of moving to production use of BC.

@ppatierno ppatierno Oct 2, 2026 •

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@MarkNSweep I ran several tests within Strimzi and on bare metal by using a beta of Java 25 LTS downloaded from here https://wiki.openjdk.org/spaces/JDKUpdates/pages/170131468/JDK+25u. The OpenJDK 25.0.5 is coming on Octber 20, 2026 GA as you can see and it has JEP-527 implemented. This is also what I am proposing for Strimzi in this proposal strimzi/proposals#249 (where the usage of BC is a rejected alternative as well).

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

@ppatierno thanks, that's a very useful datapoint. It looks like using the JDK implementation is the emerging position and would be consistent with Strimzi. @k-wall earlier commented on using JDK 21, I'm not sure how the process works for bumping the JDK in the proxy, so it seems to be coming down to wait for JDK 21 or move to JDK 25 as part of this (which is presumably something that is going to happen at some point anyway).

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Within the proposal I also provided two links to some logging about my tests with Java 21 and Java 25 to highlight the differences:

Java 21: https://gist.github.com/ppatierno/88fde56295e7620cc84afbf943812a70
Java 25: https://gist.github.com/ppatierno/51dc41a79317a754546df0781eea3690


### TLS configuration changes

TLS configurations gain a new field `namedGroups` (allowed or disallowed) which controls the algorithms used during TLS key exchange.

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The current hot-reload model (Proposal #83) is drain-and-reconnect for all VC configuration changes. Since this proposal adds new VC-scoped configuration (namedGroups, pqc), it's worth stating explicitly that changing these parameters on a running proxy triggers a drain — operators planning a phased migration should expect that.

It might also be worth noting in the rejected alternatives (or future work) that TLS parameters are a candidate for a lighter-touch reload model. They only affect new handshakes, so in principle existing connections that are already compliant with the new policy don't need draining — e.g. a hybrid → strict transition could keep connections that already negotiated a PQC group and only drain classical ones. That would need a new kind of change detector ("are existing connections still compliant?") rather than just "did the config change?" — a different shape of problem to what Proposal #83's detectors handle today. Capturing this here makes it easier to find when we revisit the reload model.

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I'll update the future work section to note the impact on #83.


The order of the `allowed` named groups is significant. The first entry is the most preferred, with the remaining entries used as fallbacks in preference order. For example, `[X25519MLKEM768, X25519, secp256r1]` prefers the hybrid PQC group and falls back to classical groups only if the peer does not support it.

#### The `pqc` convenience field

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

The pqc field hardcodes the current best-practice algorithm selections — hybrid expands to X25519MLKEM768, X25519, secp256r1 etc. Have you given any thought to how these definitions evolve over time? If a stronger PQC algorithm emerges, does hybrid silently change its expansion, or does it stay frozen and a new mode gets added?

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I have touched on some of this in my comment to @tombentley regarding new or broken algorithms : #141 (comment)

The values of the pqc field stay as either strict or hybrid. My position is that the use of the pqc field is a signal from the user that the proxy takes an opinionated view of the PQC configuration. So if a new algorithm is released or a preferred hybrid emerges, then we update and doc it. If the user disagrees with our configuration choices made when the pqc flag is set, they can revert to manually specifying the values.

@ppatierno ppatierno Oct 2, 2026 •

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stupid question ... does "hybrid" maps directly to X25519MLKEM768? Because while, for example, this is the hybrid KEM enabled by default via JEP-527 there are also SecP256r1MLKEM768 and SecP384r1MLKEM1024 which are still hybrid but not enabled by default so you need to set them via named groups. At that point "hybrid" isn't unique anymore.

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

For me X25519MLKEM768 is not hybrid (and neither are the others) as they require a quantum resistant element of the key exchange (MLKEM768) as well as a classical one (X25519). Only a client supporting PQC will be able to make a connection and so it is strict. I think the confusion is in the name (and possibly means we need different values for hybrid and strict. 'Hybrid' in the context I am proposing is in the type of client able to connect rather than a hybrid algorithm that uses both classical and PQC parts. i.e.

  • Hybrid = allow both PQC and non-PQC clients to connect.
  • Strict = only PQC capable clients can connect

Considering @SamBarker 's comment below we could change the field to be something like

onlyAllowQuantumResistantConnections : true/false or quantumOnly : true/false

The intention would be to try and frame it as a true/false setting rather than a descriptive one.

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

For me X25519MLKEM768 is not hybrid (and neither are the others) as they require a quantum resistant element of the key exchange (MLKEM768) as well as a classical one (X25519).

X25519MLKEM768 is defined as "hybrid" by standard terminology because it means combining a classical algorithm (X25519) with a post-quantum one (ML-KEM-768).

Only a client supporting PQC will be able to make a connection and so it is strict

If the server supports only X25519MLKEM768 as named group then this is true but usually you also have X25519 in the list so that a server can fallback to classical if the client is non PQC capable. Of course, if you are interested to do so.

So yes, I agree that we should use different naming because "hybrid" would be easily interpreted as referring to the X25519MLKEM768 name group.

The order of the `allowed` named groups is significant. The first entry is the most preferred, with the remaining entries used as fallbacks in preference order. For example, `[X25519MLKEM768, X25519, secp256r1]` prefers the hybrid PQC group and falls back to classical groups only if the peer does not support it.

#### The `pqc` convenience field

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Minor naming question: pqc reads to me like a config section rather than a mode selector, and while I instinctively expand it to Post Quantum Crypto, I'm not sure how universal that is. This config is likely to be read cold by operations or platform teams outside the project. Is the abbreviation clear enough for that audience, or would something more explicit like postQuantumCryptoMode serve them better?

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I hadn't looked at it like that but I can see what you mean. A couple of alternatives I came up with are

  • pqcMode
  • pqcPolicy
  • pqcStrength

All of which keep the pqc acronym which is fine if, as you say, you already know what that means. I think expanding pqc might not be necessary and we could go with just using quantum

  • quantumReadiness
  • quantumMode
  • quantumPolicy

- **Simplified configuration**: One field instead of coordinating three separate fields

**Mechanics:**
- The `pqc` field is mutually exclusive with manual `protocols`, `cipherSuites`, or `namedGroups` configuration

Copy link
Copy Markdown
Member

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Is there a case for a classical mode alongside hybrid and strict? Today the absence of pqc means classical, but that's indistinguishable from "hasn't been configured yet." An explicit pqc: classical would make the intent auditable — useful for compliance reporting where you need to show which VCs have been deliberately assessed rather than just defaulted.

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

I did consider this but then decided against it, however I do agree that it presents a problem when trying to distinguish between omission and intent. My reasoning was that hybrid also covered allowing classical connections. I did think about pqc : none but then that would imply actively removing the PQC algorithms, but the use case is about moving forwards onto PQC rather than wanting to actively block it.

@SamBarker

Copy link
Copy Markdown
Member

Thanks for writing this up @MarkNSweep — thorough proposal and a lot of ground covered. I'm still digesting and mulling over the implications, so more comments may follow. The inline comments so far are the questions that jumped out on first read.

One thought I'm kicking around: defining TLS configuration inline on every virtual cluster risks divergence over time — one VC drifts from the intended policy because someone forgot to update it. I wonder if there's a case for a named TLS policy model (similar to named filter definitions) where you define the policy once and reference it by name. That might also allow a richer configuration model since it doesn't need to be repeated in a dozen places. Just spitballing.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

6 participants