Skip to content

Support TCP for protocol messages #3242

Description

@softins

What is the current behaviour and why should it be changed?

All Jamulus protocol (non-audio) messages are currently delivered over the same UDP channel as the audio. For most protocol messages, this is fine, but those that send a list of servers from a directory, or a list of clients from a server, can generate a UDP datagram that is too large to fit into a single physical packet. Physical packets are constrained by the MTU of the Ethernet interface (normally 1500 bytes or less), and further by any limitations in links between hops on the internet. Neither the client nor the server has any control over these limitation. It's also possible a large welcome message could require fragmentation.

The UDP protocol itself allows datagrams up to be up to nearly 65535 bytes in size, minus any protocol overhead. IPv4 will allow nearly all of this size to be used, in theory. If the IPv4 datagram being sent by a node (host or router) is too large to fit into a single packet on the outgoing interface, the IP protocol will fragment the packet into pieces that do fit, with IP headers that contain the information needed to order and reassemble the fragments into a single datagram at the receiving end. Normally intermediate hops do not perform any reassembly, but will further fragment an IP packet if it will not fit the MTU of the outgoing interface.

The receiving end needs to store all the received fragments as they arrive and can only reassemble them into the original datagram once all fragments have been received. The loss of even one fragment renders the whole datagram lost, and the remaining received fragments consume resources until they time out and are discarded. There are also possibilities for a denial of service attack if an attacker deliberately sends lots of fragments with one or more missing.

If a directory has more than around 35 servers registered (depending on the length of the name, city, etc.), the list of servers sent to a client when requested is certain to be fragmented. Similarly, if a powerful server has a lot of clients connected, e.g. a big band or large choir, the list of clients sent to each connected client can get fragmented. In either of these cases, a client that is unable to receive fragmented IP packets will show an empty list or an empty mixer panel.

There are several reasons that fragmented IP datagrams can fail to make it from server to client:

  • The configuration of a user's router, either accidentally or deliberately. Sometimes a user can be helped by a knowledgeable friend to check and fix this, but often not.
  • The configuration of an intermediate router along the path from server to client. This is fairly rare, but could be a carrier's deliberate choice to avoid the kind of DoS attack mentioned above. For whatever reason, it is outside the control of the user or server operator.
  • The IPv6 protocol deliberately has no provision for fragmentation of datagrams at the IP layer. So this is a complete show-stopper for the use of IPv6 in directories, as there is therefore no support at all for large UDP messages.

The IPv6 limitation means that resolving this issue is a prerequisite to implementing IPv6 support in directories as per the ongoing discussion in https://github.com/orgs/jamulussoftware/discussions/1950.

Describe possible approaches

There is a longstanding discussion at https://github.com/orgs/jamulussoftware/discussions/1058 about the problems this issue is intended to solve, and mentioning various approaches that have been tried or proposed.

  • Limiting the size of directories. This doesn't go far enough, and as mentioned above, a directory needs to be really small (less than 30 or so servers) to be sure of avoiding fragmentation.
  • Implementing "split" messages at the Jamulus protocol level using REQ_SPLIT_MESS_SUPPORT, SPLIT_MESS_SUPPORTED and SPECIAL_SPLIT_MESSAGE. I'm not sure whether such split messages are ever used in practice, and it appears that they only apply to connected messages, not the connectionless messages which are most at risk from fragmentation. In addition, the size of split parts is fixed, and not intelligently determined from any kind of path MTU discovery.
  • Having the directory also send a "reduced" server list with only the bare information of name, IP and port (CLM_RED_SERVER_LIST). This also fails to avoid the problem, as a directory list that may take around 7 fragments in its full form still takes around 3 fragments in its reduced form.
  • I experimented with zlib compression of servers lists (https://github.com/orgs/jamulussoftware/discussions/1058#discussioncomment-8354688), but it only provides around 40% compression, not enough to avoid fragmentation.

The only possible solution is to send some protocol messages using TCP instead of UDP, when talking to a compatible client. UDP would still be available for backward compatibility when talking to older clients or older servers.

There are two kinds of protocol message that each need to be handled differently:

  • Connectionless messages CLM_*. These are unrelated to a channel (with one exception). They are mainly used by a client to fetch information for the Connect dialog:
    • List of servers from a directory (CLM_REQ_SERVER_LIST).
    • List of connected clients from a server (CLM_REQ_CONN_CLIENTS_LIST).
    • Small messages such as requests, ping, version and OS, register and unregister server. These are small enough never to need fragmentation.
  • Channel-specific messages. These need to be related to a connected channel on the server. Currently, they are identified by the IP:port of the client end.

Connectionless Messages

For connectionless messages, the client can send a TCP connection request to the server, with a timeout. If the server supports TCP, this connection will be accepted and the client can then send the CLM_REQ_* message over the TCP connection. The server needs to interpret the message and send the response back over the same TCP connection. The client can then close the connection or leave it open for sending another message (tbd). If the TCP connection from the client is refused or times out (probably due to a firewall dropping the connect request), the client can fall back to the existing UDP usage to send the request. For this reason, the TCP connection timeout will need to be short, something like 2 seconds. This will be plenty of time for a compatible server to answer.

I have a branch that implements the server side of connectionless messages over TCP, currently just for CLM_REQ_SERVER_LIST and CLM_REQ_CONN_CLIENTS_LIST, but others could be added as needed. It can be seen at https://github.com/softins/jamulus/tree/tcp-protocol. It is necessary to pass the TCP socket pointer via the function calls, signals and slots, to the point at which the response message can be sent. If this socket pointer is nullptr, the response will be send over UDP as presently, otherwise it will be sent to the referenced TCP socket.

Note that due to the variable size of Jamulus protocol messages, and the stream-oriented nature of TCP sockets, it is necessary for the receiver at each end first to read the fixed-size header (9 bytes), determine from that header the payload size, and then read the payload, plus two more bytes for the CRC.

I have tested it using a Python client, based on @passing's jamulus-python project, but enhanced to support TCP. See https://github.com/softins/jamulus-python/tree/tcp-protocol.

The next step is to add to Jamulus the client side of using TCP for connectionless messages to fetch server and client lists.

Connected Channel Messages

For connected channel messages, the situation is a little more complicated. The following factors must be considered:

  • The list of connected clients is sent to a participating client using a connected channel message (CONN_CLIENTS_LIST).
  • Each time someone else connects to or disconnects from the server, the server sends an unsolicited CONN_CLIENTS_LIST to each client that is still connected. For a large busy server with many clients, this could be a long message and subject, presently, to UDP fragmentation. It should therefore be sent if possible over TCP.
  • A server cannot initiate a TCP connection to a client. Therefore the client needs to open a TCP connection to the server at the beginning of the session, and keep the connection open continuously until leaving the session. This connection should be used by the server to send the updated client lists to the client.
  • If a client has both a TCP and a UDP connection to the server, there is no way for the server to relate the two connections just by IP and port number, as the source ports will not be related to each other. Even if the client were to bind both TCP and UDP sockets to the same local port number, they could get mapped independently to different ports by a NAT router in the path.

My proposal to solve the last point above is as follows:

  • The client starts the session in the same was as at present, by sending an audio stream.
  • On receiving the new audio stream, the server searches by IP:port (CHostAddress) for a matching channel (in CChannel vecChannels[]), and on not finding a match, allocates a free channel in that array for the new client. It stores the CHostAddress value in the allocated channel, and returns the ID (index) of the new channel.
  • The server immediately sends this ID to the client as a CLIENT_ID connected channel message. This is all existing behaviour so far.
  • A TCP-enabled client, when it receives this CLIENT_ID message, will initiate a TCP connection to the server. If the connection attempt fails, the client will assume the server is not TCP-enabled and will not retry. Operation will continue over UDP only as at present.
  • If the TCP connection succeeds, the client will immediately send a CLIENT_ID message to the server, specifying the client ID that it had just received from the server. This will enable the server to associate that particular TCP connection with the correct CChannel, and the server will store the pTcpSocket pointer in the CChannel.
  • The server can then easily send the CONN_CLIENTS_LIST updates to the client over TCP if the socket pointer in the channel is not null, or otherwise over UDP. It could also send the welcome message over the same TCP socket, improving support for longer welcome messages.
  • Other connected channel messages that are not size-critical could be sent over either UDP as at present, or the TCP connection. This is open for discussion.
  • When the client wants to disconnect from the channel, it will send a CLM_DISCONNECTION the same as at present (over either UDP or TCP), but will also close any open TCP connection to the server.
  • Over UDP, connected channel messages need to be acked, and will be retried if the ack is not received. This is necessary due to the lack of guaranteed delivery in UDP. Over the TCP socket, which provides guaranteed delivery, it would be possible to send messages without queuing or needing acks, and this might simplify implementation. Comments?

I have not yet implemented any of this connected channel functionality, beyond adding the pTcpSocket pointer to the CChannel class.

This is currently a work in progress, as described above. The purpose of this issue is to allow input from other contributors on the technical details mentioned above, and to keep the topic visible until the code is ready for a PR.

The expectation is that all the public directories will support TCP connections. This will also need suitable firewall rules at the servers. However, clients implementing all the above will still be backward-compatible with older directories and servers run by third parties. Similarly, older clients connecting to newer directories and servers will continue to operate as a present over UDP, with no use of TCP required.

Has this feature been discussed and generally agreed?

See the referenced discussion at https://github.com/orgs/jamulussoftware/discussions/1058 for history. I would value comments within this Issue regarding the solution I am proposing above. @pljones @ann0see @hoffie @dtinth and any others interested.

Activity

  1. self-assigned this
    on Feb 27, 2024
  2. dtinth commented on Feb 27, 2024

    @dtinth
    Contributor

    I like the proposal and support this approach.

    I kinda wonder if it’s possible for another person to hijack the TCP connection. e.g. a client connects, receives a channel ID 5. Maybe the presence of a new channel is broadcasted, and a rogue client who sees this message quickly makes a TCP connection to hijack that ID. Maybe we don’t need to worry about this case, but it just kinda pops up.

  3. ann0see commented on Feb 27, 2024

    @ann0see
    Member

    I believe this is very much possible. There aren't any guarantees currently either way as we don't have encryption.

    What's the disadvantage of trying to connect via TCP first?

  4. softins commented on Feb 27, 2024

    @softins
    MemberAuthor

    I like the proposal and support this approach.

    Thanks! I used your RPC server as a template for the connection handler.

    I kinda wonder if it’s possible for another person to hijack the TCP connection. e.g. a client connects, receives a channel ID 5. Maybe the presence of a new channel is broadcasted, and a rogue client who sees this message quickly makes a TCP connection to hijack that ID. Maybe we don’t need to worry about this case, but it just kinda pops up.

    That's certainly a good observation, and worth considering, even if unlikely. I think it could largely be mitigated by only acting upon a CLIENT_ID message if it has come from the same IP address as the one already recorded in the channel it refers to. We can't compare the port of course, as it may be different as already mentioned, but I think it's extremely unlikely that the related UDP and TCP connections would come from different IPs. That would limit the scope for hijack to another client behind the same IP (e.g. same host or same NAT).

  5. softins commented on Feb 27, 2024

    @softins
    MemberAuthor

    What's the disadvantage of trying to connect via TCP first?

    I couldn't think of a compatible way to associate a subsequent UDP audio stream with a TCP connection that was made first. Especially if they had travelled through NAT. When I thought of starting with UDP as normal, and then sending back the CLIENT_ID message over a successful TCP connection, it was like a Eureka moment.

  6. pljones commented on Feb 28, 2024

    @pljones
    Collaborator

    It sounds pretty good. I'd be inclined only to move explicitly those messages we know cause problems. However, the "infrastructure" should be there to support other messages.

    I guess only the audio packets themselves need remain on UDP.

    TCP/IP keep alive will be running on the TCP/IP connection, right? At the moment, a UDP drop out isn't seen as a "connection failure" by the server for quite a large window (relatively). Would the TCP/IP connection remain "connected" or drop out if keep alive failed? How would a "continuous" UDP audio connection work in this situation?

  7. ann0see commented on Feb 28, 2024

    @ann0see
    Member

    I couldn't think of a compatible way to associate a subsequent UDP audio stream with a TCP connection that was made first.

    Seems like that's a fundamental problem.

    So to not break backwards compatibility the server must still respond on UDP audio messages for session creation.

    Could we send an empty audio message to the server to query the version/capabilities, then do something comparable to what syn cookies do (https://en.m.wikipedia.org/wiki/SYN_cookies) for some kind of authentication and then set up a TCP connection.
    The main idea cound also be to have some kind of "secret" stored on the server to verify the client.

  8. softins commented on Feb 28, 2024

    @softins
    MemberAuthor

    I couldn't think of a compatible way to associate a subsequent UDP audio stream with a TCP connection that was made first.

    Seems like that's a fundamental problem.

    A limitation, certainly, but I wouldn't call it a problem.

    So to not break backwards compatibility the server must still respond on UDP audio messages for session creation.

    Yes indeed.

    Could we send an empty audio message to the server to query the version/capabilities, then do something comparable to what syn cookies do (https://en.m.wikipedia.org/wiki/SYN_cookies) for some kind of authentication and then set up a TCP connection. The main idea cound also be to have some kind of "secret" stored on the server to verify the client.

    I don't really think that gains us anything except quite a lot of unneeded complexity. I think the source IP is an adequate enough "secret" to validate the TCP connection that follows the start of the UDP session.

  9. softins commented on Feb 28, 2024

    @softins
    MemberAuthor

    It sounds pretty good. I'd be inclined only to move explicitly those messages we know cause problems. However, the "infrastructure" should be there to support other messages.

    Yes, I agree.

    I guess only the audio packets themselves need remain on UDP.

    That's definitely true, also.

    TCP/IP keep alive will be running on the TCP/IP connection, right? At the moment, a UDP drop out isn't seen as a "connection failure" by the server for quite a large window (relatively). Would the TCP/IP connection remain "connected" or drop out if keep alive failed? How would a "continuous" UDP audio connection work in this situation?

    If I remember correctly (from a long time ago), TCP keepalive has a very long timeout, and it just intended to keep a session alive when there is no data to exchange. In the case of Jamulus, the server regularly sends the channel levels using CLM_CHANNEL_LEVEL_LIST, so if we make sure they go over TCP, that will be enough to keep the connection alive. If the actual connection fails, that will be picked up by TCP layer retries eventually giving up.

  10. pljones commented on Feb 29, 2024

    @pljones
    Collaborator

    OK, assuming that the TCP keepalive is longer than the Jamulus UDP audio "keep alive" time for a channel, we'd just need to ensure that part of channel clean up is to clear down the TCP socket, too?

  11. softins commented on Feb 29, 2024

    @softins
    MemberAuthor

    OK, assuming that the TCP keepalive is longer than the Jamulus UDP audio "keep alive" time for a channel, we'd just need to ensure that part of channel clean up is to clear down the TCP socket, too?

    Yes, absolutely.

  12. pljones commented on Feb 29, 2024

    @pljones
    Collaborator

    Let's say a client connects and establishes a TCP connection.

    At some point the server tries to send a message over the TCP socket and gets an error back indicating the socket is no longer valid.

    If the UDP audio stream continues, how does recovery work in this situation? Would the client get a CONN_FAIL, too, and know to re-initiate the TCP connection? Will the server know what its state was at the time the connection failed for recovery? I guess association with the same CChannel would continue.

  13. softins commented on Feb 29, 2024

    @softins
    MemberAuthor

    Well at the very least, the server will close its end of the TCP socket and set CChannel::pTcpSocket to nullptr. This would fall back to the UDP-only situation as per older versions.

    At some point the client would also notice the socket had failed, and would either also revert to UDP-only, or could try re-connecting TCP.

    I think it would be unlikely in practice that the TCP socket would fail while the UDP is still working. A typical network outage would affect both at once.

  14. pljones commented on Mar 1, 2024

    @pljones
    Collaborator

    OK, so as I see it, even if the client has established TCP for certain messages, it should handle them arriving over UDP as currently -- but use an "unexpected UDP fallback" to indicate the server can't send the message over TCP and re-initiation of the socket is probably needed.

  15. 10 remaining items

  16. softins commented on Sep 4, 2024

    @softins
    MemberAuthor

    I have picked this up again and revised the proposed sequence of operations for TCP support.

    The summary is that TCP need only be used as a fallback when it is determined that a UDP message failed due to fragmentation, and that the directory or server explicitly supports TCP.

    Current operation when client opens Connect dialog

    1. Client sends CLM_REQ_SERVER_LIST to the selected directory server to ask for a list of registered servers. It then starts a 2.5 sec re-request timer.

    2. Directory server fetches its internal list of registered servers, and sends a CLM_SEND_EMPTY_MESSAGE to each listed server, with the IP and UDP port of the requesting client as parameters.

    3. Directory server sends CLM_RED_SERVER_LIST (reduced server list) to client.

      a. If/when client receives CLM_RED_SERVER_LIST, it populates its list of servers with the reduced info, and sets an internal flag to say it has done so. It checks this flag to avoid processing a repeat reduced server list.

      b. If the list is large and fragmented, and the path does not correctly pass fragments, the client will not receive the list.

    4. Directory server sends CLM_SERVER_LIST to client. It does this immediately after sending the reduced list above.

      a. If/when the client receives CLM_SERVER_LIST, it populates its list of servers with the full info, replacing any existing list that contained reduced info. It then stops the 2.5 sec re-request timer mentioned above, so that the server list is not requested again.

      b. If the list is large and fragmented, and the path does not correctly pass fragments, the client will not receive the list, and the request time will be left running to retry.

    5. While client is displaying the server list, it periodically pings each server with CLM_PING_MS_WITHNUMCLIENTS including a timestamp in the message.

    6. Each pinged server, when it receives the ping, will create a CLM_PING_MS_WITHNUMCLIENTS in reply, containing a copy of the received timestamp, and the number of clients currently connected to that server.

    7. When the client receives the reply, it can calculate the round-trip time from the received timestamp and the current time.

    8. If the number of connected clients returned is different from the previously received number for that server, the client sends a CLM_REQ_CONN_CLIENTS_LIST to the server.

      a. The server responds with a list of clients in a CLM_CONN_CLIENTS_LIST. Most servers only have a small number of clients connected, and this message is not large enough to need IP fragmentation.

      b. If the server is a large one with many clients connected (e.g. for a choir, big band or WorldJam green room), the client list may be large enough to be fragmented by the IP layer. In that case, it might not be received by the requesting client.

    9. If/when the client receives the client list from the server, it can display the list of connected clients under the relevant server, if this is enabled in the GUI.

    10. The five steps above continue until the user clicks Connect or closes the dialog.

    Enhancement for TCP support

    1. A directory server will have a command-line option (and maybe a setting in the GUI and ini-file) to enable or disable TCP operation. Default TBD.

    2. If the directory server has TCP enabled, then after it has sent the CLM_RED_SERVER_LIST and CLM_SERVER_LIST by UDP, it will send a new message CLM_TCP_SUPPORTED to the client.

      a. An older version of client that does not support TCP will ignore this message and continue operating in the normal way just on UDP.

      b. A newer client that supports TCP should receive and process the CLM_TCP_SUPPORTED message after it has received and processed the UDP server list, unless fragmentation prevented it from doing the latter.

      c. If such a client has already processed a full server list from CLM_SERVER_LIST, it will have no need to open a TCP connection to the directory, so this will be skipped.

      d. If the client receives CLM_TCP_SUPPORTED having not received and processed a CLM_SERVER_LIST, it will open a TCP connection to the directory server, and request the server list again over the TCP connection.

    3. If the directory server accepts a TCP connection and receives a CLM_REQ_SERVER_LIST over it, it will process the request in the same way as for a UDP request, with the following differences:

      a. There is no need for the directory to send CLM_SEND_EMPTY_MESSAGE to the servers in the list, since that was already done in response to the original UDP request.

      b. There is no need for the directory to send CLM_RED_SERVER_LIST to the client, since the TCP connection is reliable, so the directory server just sends the CLM_SERVER_LIST over the TCP connection.

    4. When the client has received the CLM_SERVER_LIST over TCP, it closes the TCP connection, populates its list of servers in the connect dialog in the normal way and stops the 2.5 sec re-request timer.

    5. The client starts pinging each listed server as normal, using UDP, and the server responds with a ping including the timestamp and number of clients, as described above.

    6. As above, if the number of connected clients has changed, the client sends a CLM_REQ_CONN_CLIENTS_LIST over UDP in the normal way.

    7. If a server in the list supports TCP, when it has sent a reply to the client list request with CLM_CONN_CLIENTS_LIST, it will follow it immediately with a CLM_TCP_SUPPORTED.

      a. An older version of client that does not support TCP will ignore the CLM_TCP_SUPPORTED message and continue operating in the normal way just on UDP.

      b. A newer client that supports TCP should received the CLM_TCP_SUPPORTED message after it has received and processed the UDP client list, unless fragmentation prevented it from doing the latter.

      c. If such a client has already processed a client list from CLM_CONN_CLIENTS_LIST, it will have no need to open a TCP connection to the server, so this will be skipped.

      d. If the client receives CLM_TCP_SUPPORTED having not received and processed a CLM_CONN_CLIENTS_LIST, it will open a TCP connection to the server, and request the client list again over the TCP connection.

    8. If the server accepts a TCP connection and receives a CLM_REQ_CONN_CLIENTS_LIST over it, it will process the request in the same way as for a UDP request, but will send the reply over the TCP connection.

    9. When the client has received the CLM_CONN_CLIENTS_LIST over TCP, it closes the TCP connection and updates the list of clients for that server in the GUI.

      a. Consideration was given to keeping the TCP connection open for sending repeat requests, but this would require keeping track of potentially many open connections, one for each server, so closing the connection as soon as the response has been received is easier.

    Conclusion

    By sending the CLM_TCP_SUPPORTED message immediately after sending a potentially large list of servers or connected clients, it allows a client easily to determine whether or not it needs to fall back to TCP without the necessity of timeouts or other delays. It will only need to use TCP if it has not already succeeded in receiving the message over UDP.

    It is assumed that the server will only be configured to offer TCP fallback if the server operator has also configured any firewall to allow the inbound TCP connections.

    There is also no need to worry about TCP port numbers, as TCP is only used as a fallback when necessary, and all the port opening and pinging has already been set up with UDP as at present.

    If there are no problems seen with the above approach, I plan to get started with implementation over the next few weeks. As mentioned at the start of this Issue, I have had the TCP server side working in test mode for some months.

  17. ann0see commented on Sep 5, 2024

    @ann0see
    Member

    It is assumed that the server will only be configured to offer TCP fallback if the server operator has also configured any firewall to allow the inbound TCP connections.

    I didn't follow this through fully, but is it an issue if the firewall blocks (e.g all firewall traffic over TCP to be DROPped) TCP on the server side? What would happen? I assume the client will request a server list over TCP and wait until it times out.

  18. softins commented on Sep 5, 2024

    @softins
    MemberAuthor

    It is assumed that the server will only be configured to offer TCP fallback if the server operator has also configured any firewall to allow the inbound TCP connections.

    I didn't follow this through fully, but is it an issue if the firewall blocks (e.g all firewall traffic over TCP to be DROPped) TCP on the server side? What would happen? I assume the client will request a server list over TCP and wait until it times out.

    If a server offered TCP to the client, but the server's firewall didn't allow the connection, the client would indeed wait until its request times out.

    I think this has to be the responsibility of the server/directory operator, and why I think TCP operation should be controlled by a command-line option, rather than always enabled. The operator should only enable TCP in the Jamulus server if they know their environment has been configured to support it.

    This is not an issue that your average small server operator will need to be concerned with at all. The only server operators who will need to enable TCP support are those running large directories (e.g. Volker, Peter) or those running a large server designed to support many simultaneous client connections.

  19. softins commented on Sep 5, 2024

    @softins
    MemberAuthor

    I have picked this up again and revised the proposed sequence of operations for TCP support.

    The summary is that TCP need only be used as a fallback when it is determined that a UDP message failed due to fragmentation, and that the directory or server explicitly supports TCP.

    In addition to the above, which only covers operation of the Connect Dialog, we also still need to consider and define the use of TCP when a client is connecting to a server session.

    I have some more ideas on this, which I will write up and add here soon.

  20. ann0see commented on Sep 5, 2024

    @ann0see
    Member

    The only server operators who will need to enable TCP support are those running large directories (e.g. Volker, Peter) or those running a large server designed to support many simultaneous client connections.

    Fair point.

  21. ann0see commented on Oct 10, 2024

    @ann0see
    Member

    How is this processing? It would be good to have a draft PR for the branch you mentioned to have some more code discussion.

  22. softins commented on Mar 7, 2026

    @softins
    MemberAuthor

    In addition to the above, which only covers operation of the Connect Dialog, we also still need to consider and define the use of TCP when a client is connecting to a server session.

    I have some more ideas on this, which I will write up and add here soon.

    It's not exactly soon, but I have just found the notes I wrote on this 18 months ago, so am adding them here for the record. Now that we have almost got 3.12.0 finished, I want to pick this up again.

    Notes regarding TCP use in a connected session

    5 Sep 2024

    Current operation when client clicks on Connect

    All these steps use UDP.

    1. Client starts sending an audio stream to the server. This audio stream continues in parallel with the protocol exchange below.

      a. Server does not yet start sending an audio stream to the client.

    2. Server sees the audio stream and looks up the source IP:port in its channel table, finding no channel that matches.

    3. Server allocates a new channel in the channels array and stores the source IP:port in it. The channel index becomes the Channel ID of the connected client.

    4. Server sends a CLIENT_ID message to the client, containing the Channel ID mentioned above.

      a. Client replies with ACKN (CLIENT_ID). Note that all messages that do not begin CLM_ need to be acked by the receiving side. For clarity, these ACKN messages will not be mentioned below.

    5. Server sends CONN_CLIENTS_LIST to the client, containing a list of all current clients in the session.

    6. Server sends REQ_SPLIT_MESS_SUPPORT to ask the client if it supports split messages.

    7. Client sends back SPLIT_MESS_SUPPORTED immediately.

    8. Server sends REQ_NETW_TRANSPORT_PROPS to ask for the clients network transport parameters.

    9. Client sends NETW_TRANSPORT_PROPS containing the codec, packet size, number of channels, bitrate, etc.

    10. Server sends REQ_JITT_BUF_SIZE to ask for the client's required jitter buffer sizes.

    11. Client sends JIT_BUF_SIZE, containing the positions of the "server" jitter buffer slider in the Settings dialog. This is telling the server what size jitter buffer to use for receiving audio data from the client. (The position of the "client" jitter buffer slider is not needed by the server, as it is only used locally in the client).

    12. Server sends REQ_CHANNEL_INFOS to ask for the identity information for the channel.

    13. Client sends CHANNEL_INFOS containing the identity information from the user's profile settings in the client (country, instrument, skill level, name, city).

    14. Now that the server has received the CHANNEL_INFOS from the client, it starts to send the mixed audio stream to the client.

    15. Server sends CHAT_TEXT containing the server welcome message, if any. If there is none, this message is skipped.

    16. Server sends VERSION_AND_OS to tell the client the version of Jamulus on the server and the server platform.

    After this, messages are sent by either side when there is something to notify:

    • From server to client:

      • CLM_CHANNEL_LEVEL_LIST - list of audio levels for each channel. Sent every 250ms by a timer.
      • RECORDER_STATE - current state of the server-based recording. Sent when the state changes?
      • JITT_BUF_SIZE - the size of the receiving jitter buffer for this connection on the server. Sent in Auto mode when the value changes.
      • CLM_PING_MS - sent in response to a CLM_PING_MS received from the client. For client-side ping time calculation.
      • CONN_CLIENTS_LIST - list of connected clients. Sent when the list changes due to a client connecting or leaving. This message could be large on a server with many clients.
    • From client to server:

      • CLM_PING_MS - contains a timestamp and requests the server to send back the same timestamp, so that the round-trip time can be measured. Sent every 500ms by a timer.
      • NETW_TRANSPORT_PROPS - specifies codec, packet size, bitrate, etc. Sent when the user changes Audio Channels, Audio Quality, Buffer Delay or Small Network Buffers.
      • CHANNEL_GAIN - specifies the user's requested gain for a specific channel. Sent when the user moves a fader, but rate limited to avoid many changes in succession being sent.
      • JITT_BUF_SIZE - specifies the requested size of the server's jitter buffer, or Auto.
      • CLM_DISCONNECTION - sent when the user clicks on Disconnect, or connects to another server.

    Enhancement for TCP support

    1. As soon as the server sees audio from a new client and creates a channel for it, it will send CLM_TCP_SUPPORTED to the client.

    2. The server will then send the CLIENT_ID message as above, containing the channel ID that has been allocated.

    3. An older version of client that does not support TCP will ignore the CLM_TCP_SUPPORTED message and continue operating in the normal way just on UDP.

    4. A newer client that supports TCP will receive the CLM_TCP_SUPPORTED message, and will note that the server supports TCP.

    5. When a newer client receives the CLIENT_ID message from a server it knows supports TCP, the client will open a TCP connection to the server (on the same port number as UDP), and as its first message over the connection it will send a CLIENT_ID message containing the channel ID that it received over UDP.

    6. The server will accept the TCP connection, and will wait for the first message to arrive via that connection. This should be the CLIENT_ID message.

    7. The server will lookup the channel specified by the CLIENT_ID message, and will check that the IP address of the channel matches the remote address of the TCP connection. If it does not, it will close the connection. This prevents hijacking of a session by sending another client's ID.

    8. If the TCP connection matches the client channel, the socket descriptor will be stored in the channel.

    9. Any messages from the client that arrive over TCP will be handled in the same way as messages received over UDP. Responses will be send back over TCP too.

    10. Messages generated by the server will also be sent over TCP, if there is an active socket descriptor stored for the channel. If not, they will be sent over UDP in the normal way.

    11. Audio packets will continue to always use the UDP socket.

    NOTE THAT PING MESSAGES MUST STILL GO OVER UDP, SO THAT PING CALCULATIONS ARE NOT AFFECTED BU POTENTIAL TCP RETRIES. MAYBE ALSO THE LEVEL MESSAGES.

    Also say something about disconnecting.

    If the server receivs a disconnection of the TCP socket, it will revert to UDP. it could then send another CLM_TCP_SUPPORTED to invite the client to re-establish a TCP connection.

  23. softins commented on Mar 11, 2026

    @softins
    MemberAuthor

    How is this processing? It would be good to have a draft PR for the branch you mentioned to have some more code discussion.

    See #3636 which I have just posted as a starting point.

  24. linked a pull request that will close this issueSupport TCP for protocol messages #3636on Mar 16, 2026
  25. softins commented on Mar 29, 2026

    @softins
    MemberAuthor

    Revised enhancement for TCP support in connected mode

    With the experience gained in implementing the Connect Dialog support for TCP, I have revised the design for connected-mode TCP support, in place of that mentioned above in the previous comment.

    The connected-mode protocol messages sent over UDP are all sequence numbered and acknowledged, in order to be robust against potential packet loss. Over TCP, such packet loss will not occur, as sequencing and acknowledgement all happen at the TCP network layer.

    Consequently, TCP will not be used for connected-mode protocol messages.

    The reason for using a TCP connection in an active session is just to provide a reliable path for delivering a list of connected clients that could be large and subject to fragmentation (if it is sent over UDP). So in this design, the established TCP connection will only be used to deliver client lists, and not other protocol messages.

    Therefore, if the server has an active TCP connection for the client, it will use the connectionless CLM_CONN_CLIENTS_LIST message to deliver updates for the connected client list. If there is no active TCP connection, updates will be delivered using the connected-mode CONN_CLIENTS_LIST over UDP as at present.

    So the new sequence will be as follows:

    1. As soon as a TCP-enabled server sees audio from a new client and creates a channel for it, it will send CLM_TCP_SUPPORTED to the client.

    2. The server will then send the CLIENT_ID message as above, containing the channel ID that has been allocated.

    3. An older version of client that does not support TCP will ignore the CLM_TCP_SUPPORTED message and continue operating in the normal way just on UDP.

    4. A newer client that supports TCP will receive the CLM_TCP_SUPPORTED message, and will note that the server supports TCP. The client will open a TCP connection to the server (on the same port number as UDP).

    5. The server will accept the TCP connection, and will wait for the first message to arrive via that connection.

    6. When a newer client receives the CLIENT_ID message from a server it knows supports TCP, the client will send, as its first message over the connection, a CLM_CLIENT_ID message containing the channel ID that it received over UDP. (CLM_CLIENT_ID is a newly-defined connectionless message).

    7. The server will lookup the channel specified by the CLM_CLIENT_ID message, and will check that the IP address of the channel matches the remote address of the TCP connection. If it does not, it will close the connection. This prevents hijacking of a session by sending another client's ID.

    8. If the TCP connection matches the client channel, the socket descriptor will be stored in the channel, and the channel pointer will be stored in the TCP Connection instance.

    9. Any messages from the client that arrive over TCP will be handled in the same way as messages received over UDP. Responses will be send back over TCP too. At present, there are no such messages defined. Existing protocol messages will continue to use UDP.

    10. Updates to the Connected Clients List generated by the server will be sent over TCP as CLM_CONN_CLIENTS_LIST, if there is an active socket descriptor stored for the channel. If not, they will be sent over UDP as CONN_CLIENTS_LIST in the normal way.

    11. In order to keep the long-term TCP connection alive via firewalls, NAT routers, etc., the server and client will both start a periodic timer (e.g. 15 sec) to send a CLM_EMPTY_MESSAGE over the TCP connection.

    12. Audio packets will continue to always use the UDP socket.

    13. A client with an active TCP socket to the server could send a disconnection message over the TCP connection or via UDP (to be defined).

    If the server receives a disconnection of the TCP socket, it will revert to UDP for connected client updates. It could send another CLM_TCP_SUPPORTED to invite the client to re-establish a TCP connection.

  26. ann0see commented on Mar 30, 2026

    @ann0see
    Member

    CLM_EMPTY_MESSAGE

    Since it's a Keep alive message - why not call it CLM_KEEPALIVE_MESSAGE

  27. softins commented on Mar 30, 2026

    @softins
    MemberAuthor

    CLM_EMPTY_MESSAGE

    Since it's a Keep alive message - why not call it CLM_KEEPALIVE_MESSAGE

    Because CLM_EMPTY_MESSAGE is already defined (it's the one used for the automatic firewall/NAT hole punching), and there's no point duplicating it. It was designed for sending when there's no need to process it at the receiving end.

  28. moved this from Triage to In Progress in Trackingon May 30, 2026
  29. mcfnord commented on Jul 23, 2026

    @mcfnord
    Contributor

    Note

    📡 STAND BY FOR AN LLM-AUTHORED MESSAGE.

    @softins One design-level question on the connected-session part of the plan, after reading the two revisions here against the tcp-protocol branch.

    The Connect-dialog fallback and the connected-session fallback establish TCP on opposite triggers, and I want to check whether that asymmetry is intentional.

    For the Connect dialog, TCP is lazy: the client opens a TCP connection only when a UDP reply was demonstrably missed. In OnCLTcpSupportedReceived, the CLM_SERVER_LIST and CLM_CONN_CLIENTS_LIST cases retry over TCP only while the UDP request is still pending unanswered — that is, probably lost to fragmentation. This is the "only pay for TCP when UDP actually failed" property from your September 2024 write-up, and the code matches it.

    For a connected session, TCP is eager: a TCP-enabled server sends CLM_TCP_SUPPORTED(CLM_CLIENT_ID) to every new client in OnNewConnection, and the client opens the long-lived PROTO_TCP_LONG connection immediately, then holds it open for the whole session with 15-second keepalives — regardless of whether that session's client list is ever large enough to fragment. So every session on a TCP-enabled server carries a permanent TCP connection, keepalive traffic in both directions every 15 seconds, a server-side idle timer, and one extra file descriptor per client — whether or not that server's client list ever grows past one datagram. Since list size tracks server occupancy, the standing cost is paid even by small servers and by large servers outside their busy hours, where the fallback can never fire.

    I think there is a real reason for the difference: an in-session client-list update is an unsolicited server push, so the client has no "pending request that went unanswered" signal to drive lazy fallback the way the Connect dialog does. But the server does know the size — it is building CONN_CLIENTS_LIST and can compare it against a fragmentation threshold before sending. So the lazy trigger could move to the server's advertisement rather than the client's establishment: keep sending CONN_CLIENTS_LIST over UDP as today, and send CLM_TCP_SUPPORTED to invite a session TCP connection only the first time a list for that session actually crosses the threshold. Small servers would then never open a session TCP connection at all.

    The obvious cost of that is a one-time transient at the crossing point: the update that first exceeds the threshold would still go out over UDP and could fragment while the TCP connection is being set up. The server could bound that by re-sending the current list as CLM_CONN_CLIENTS_LIST the moment the connection is confirmed, so the staleness window is just connection-setup time. Whether that transient is worth trading against a permanent connection on every session is really your call — it may be that eager establishment is simply easier to keep correct, and you may already have weighed exactly this. I mainly want to confirm the asymmetry is deliberate before it settles into the protocol.

  30. softins commented on Jul 23, 2026

    @softins
    MemberAuthor

    Yes, the asymmetry is deliberate for the reason inferred. The problem with a fragmentation threshold is that we don't know what it should be. Links can fragment at different thresholds, and different routers may have different numbers of fragments they are prepared to reassemble. So it's hard for a client to detect a lost message when it doesn't know one is coming.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

Labels

Type

No type

Projects

  • Status
    In Progress

Relationships

None yet

Development

No branches or pull requests

Issue actions