10 min read

Deciding Steam session requests in Rust, before JavaScript can answer

ISteamNetworkingMessages expects a session to be accepted or rejected before its callback returns, and JavaScript can't run in time. So my steamworks.js fork decides from a Rust-side allow list that JavaScript fills in ahead of time.

One of the modules I added to my steamworks.js fork is networking_messages, which is about 370 lines of Rust, and nearly every design choice in it comes from a single constraint. When a peer opens a session with you, Steam asks whether you want to accept it, and it expects the answer before its callback returns. At that moment your JavaScript code has no chance to run, so whatever makes the decision has to already be sitting on the Rust side.

I am writing this up because the problem is not specific to networking. If you are wrapping any Steam API that expects you to make a decision inside a callback, you will run into the same thing, and the approach below should carry over.

How ISteamNetworkingMessages handles incoming sessions

ISteamNetworkingMessages is Valve’s connectionless messaging interface. Instead of opening a socket and managing a connection yourself, you send a message to a Steam ID, and Steam opens a session on demand the first time you do. The upstream steamworks.js only exposed the legacy ISteamNetworking, which Valve has deprecated, so wrapping the newer interface was one of the first things the fork added for Cozy Coast.

On the receiving side, the first message from a peer you have not talked to yet triggers a session request, and your code has to decide whether to let that peer in. In the steamworks Rust crate (version 0.13.1, which the fork builds on), you handle this by registering a closure:

pub fn session_request_callback(
    &self,
    mut callback: impl FnMut(SessionRequest) + Send + 'static,
)

The closure receives the SessionRequest by value, which means it owns the request and is the only place that can answer it. The request has accept() and reject() methods, and if the closure returns without calling either of them, the Drop impl rejects the session for you:

impl Drop for SessionRequest {
    fn drop(&mut self) {
        if !self.accepted {
            self.reject_inner();
        }
    }
}

reject_inner is a straight call to CloseSessionWithUser. In practice, this means you have to make the decision before the closure returns, because as soon as it returns the request is dropped and, unless you accepted it, the peer’s session is closed.

Why the JavaScript handler runs too late

To see why JavaScript can’t make that decision, it helps to look at how callbacks reach the closure in the first place. Steam does not push callbacks on its own, so something has to call SteamAPI_RunCallbacks on a regular basis. In the fork, init in index.js starts a timer that does exactly that:

runCallbacksInterval = setInterval(runCallbacks, 1000 / 30)

The timer fires 30 times a second, roughly every 33 ms, and because it is a setInterval, it runs on the JavaScript thread. The native function it calls is only two lines:

#[napi]
pub fn run_callbacks() {
    client::get_client().run_callbacks();
}

Now think about what happens during a single tick. The JS thread calls runCallbacks() and enters native code. Inside that call, Steam dispatches the pending session request, and the Rust closure runs, still on the JS thread and still inside the same native call. The only way for the closure to tell JavaScript about the request is through a napi ThreadsafeFunction, and calling one of those does not run the JavaScript function straight away. What it does is queue a call onto the event loop, to be run whenever the loop gets to it.

The problem is that the event loop is the thing currently executing runCallbacks(), so it cannot pick up anything from its queue until that call has finished. The queued handler only runs after the timer callback returns, which happens after the Steam callback has returned, which in turn happens after the SessionRequest has been dropped. By then the peer has already been rejected, so any “accept” that JavaScript computes would be for a request that no longer exists.

Where the session decision happens in one callback tick Inside a single runCallbacks call on the JavaScript thread, Steam dispatches a session request, a Rust closure checks the allow list, and the request is accepted or rejected. The JavaScript handler is only queued, and runs after runCallbacks returns. inside one runCallbacks() call, on the JS thread Steam dispatch session request Rust closure checks the allow list accept() or reject() decided here after the tick returns event loop queue NonBlocking call onSessionRequest (steamId64, accepted)
The accept or reject happens inside the tick. JavaScript hears about it only after runCallbacks has returned.

Keeping the allow list on the Rust side

Since JavaScript can’t answer the question at the moment Steam asks it, the answer has to be ready before the question comes up. What I did was keep the policy in Rust, in two statics that the closure can read synchronously while it is still holding the request:

static ALLOW_ALL: AtomicBool = AtomicBool::new(false);

lazy_static::lazy_static! {
    static ref ALLOWED: Mutex<HashSet<u64>> = Mutex::new(HashSet::new());
}

ALLOWED is a set of 64-bit Steam IDs behind a mutex, and ALLOW_ALL is a single flag that accepts everyone regardless of what is in the set. JavaScript manages the set through three functions: allowPeer inserts an ID, disallowPeer removes one, and clearAllowedPeers empties it. The closure that initSessionCallbacks registers then checks the same set whenever a request comes in:

let permitted =
    ALLOW_ALL.load(Ordering::Relaxed) || ALLOWED.lock().unwrap().contains(&remote);
if permitted {
    request.accept();
} else {
    request.reject();
}
requested.call(
    (BigInt::from(remote), permitted),
    ThreadsafeFunctionCallMode::NonBlocking,
);

The decision is made in Rust, against data that JavaScript wrote earlier, and JavaScript is told about it afterwards with the verdict passed along as the second argument. That means the onSessionRequest handler only works as a notification. You can use it to log the request or update the UI, but by the time it runs the session has already been accepted or rejected, so nothing the handler does can change the outcome.

Filling the allow list from lobby events

The next question is what should put peers into the set, and for Cozy Coast the answer is the lobby. Everyone who should be able to message you is, by definition, in your Steam lobby, and Steam already announces people joining and leaving through LobbyChatUpdate. So the simplest approach is to allow a peer when they enter the lobby and disallow them when they go, which the README wires up like this:

const nm = client.networking_messages

nm.initSessionCallbacks(
    (steamId64, accepted) => { /* peer requested a session; accepted per policy */ },
    (steamId64) => {
        // Session broke. Acknowledge it before sending to that peer again.
        nm.closeSessionWithUser(steamId64)
    },
)

client.callback.register(steamworks.SteamCallback.LobbyChatUpdate, ({ user_changed, member_state_change }) => {
    if (member_state_change === 'Entered') nm.allowPeer(user_changed)
    else nm.disallowPeer(user_changed)
})

That === 'Entered' comparison was an easy place to end up with a bug. Upstream’s callbacks.d.ts typed member_state_change as a numeric enum, but callback payloads go through serde, which serializes the variant as its name. So if you wrote your code against the types, you were comparing a string to a number, and the check never matched. To fix that, the fork now types the field as 'Entered' | 'Left' | 'Disconnected' | 'Kicked' | 'Banned', and the handler above treats anything other than 'Entered' as the peer leaving.

There is also setAllowAllSessions(true), which sets ALLOW_ALL so that every session request gets accepted without looking at the set. The doc comment on it says that a shipped host should allow only the peers it knows joined its lobby, because accepting everyone skips the lobby check entirely. I use it for private playtests and nowhere else.

The race on a peer’s first message, and how the sender recovers

The allow list approach has one gap, and I documented it in the source so that nobody has to discover it on their own. The decision happens when a peer’s first message arrives, so if that message gets there before your allowPeer call has run, the session is rejected once. This can happen when a peer joins the lobby and immediately sends a hello, before your LobbyChatUpdate handler has had the chance to add them to the set.

Recovering from it is the sender’s job: it sees the broken session through onSessionFailed, calls closeSessionWithUser for that peer to acknowledge it, and its next send opens a fresh session request. By then the receiver has had time to allow the peer, so the second request gets accepted. That is why the failed handler in the README does nothing except close the session. The two-machine test in test/networking_messages.js goes a step further and also closes the session when a member leaves the lobby.

Send errors that tell you what went wrong

sendMessageToUser used to return a bare bool, which told you something had gone wrong but nothing about what, so your code had no way to tell a broken session apart from a message that was simply too large. It now throws instead, and the error message is the EResult name:

crate::client::get_client()
    .networking_messages()
    .send_message_to_user(identity, flags, &data, channel)
    .map_err(|e| Error::from_reason(format!("{e:?}")))

{e:?} is the Debug form of the crate’s SteamError enum, so a broken session shows up in JavaScript as an error whose message is 'NoConnection', and you can branch on that string. The doc comment lists the ones worth handling: NoConnection means you should close the session and retry, LimitExceeded means the message is too big or too much is already queued, and InvalidParam means you passed something wrong.

try {
    nm.sendMessageToUser(peerId64, nm.MessageSendType.Reliable, Buffer.from(text), 0)
} catch (e) {
    if (e.message === 'NoConnection') nm.closeSessionWithUser(peerId64)
}

When you need more detail than an error name, getSessionConnectionInfo(steamId64) returns the session state, ping, local and remote delivery quality, bytes and packets per second, pending reliable and unreliable bytes, and whether the route is relayed. For this function I went below the crate’s wrapper and call SteamAPI_ISteamNetworkingMessages_GetSessionConnectionInfo directly through the raw sys bindings, then read the two structs it fills in field by field. Relay detection originally tested m_idPOPRelay != 0, but that only detects Steam Datagram Relay, so it now reads the Relayed connection flag instead, which covers TURN as well. One more detail is worth knowing about endReason: it is 0 both when the session is healthy and when Steam recorded no reason, and the doc comment notes that the second case happens on loopback pipe closes, so a 0 on its own does not prove the session is fine.

Receiving messages by polling

Incoming messages never go through the callback path at all. Instead, you pull them yourself on whatever schedule suits your game:

setInterval(() => {
    for (const { steamId, data, channel } of nm.receiveMessagesOnChannel(1, 32)) {
        // ...
    }
}, 50)

receiveMessagesOnChannel(channel, batchSize) copies each message’s bytes into a Node Buffer and returns the whole batch in one array. The benefit is that nothing gets queued across the napi boundary for each individual message, and your code reads the inbox when it is ready to handle what is in it. The README polls every 50 ms and the two-machine test every 66 ms, which are both slower than the 30 Hz callback pump. That works fine, because only session requests go through the pump, while the message payloads are read by these separate polls.

For testing, the single-machine smoke test (node test/smoke.js) tries to send a loopback message to your own Steam ID. If nothing comes back within 5 seconds, it marks that step as skipped instead of failed, with a note that Steam may not loop back to self, since a missing loopback does not necessarily mean the module is broken. The real check needs two machines in the same lobby, each running node test/networking_messages.js, and you can find the tests along with the rest of the code on GitHub.

Let's connect.

Always happy to talk shop, compare notes, or just say hi. Email or LinkedIn is the fastest way to reach me.

Get in touch