Skip to main content

Open navigation

Back to blog

MQTT Keep Alive: Standards, Broker Internals, and Timeout Tests

Follow MQTT Keep Alive through EMQX, HiveMQ CE, and Mosquitto source, then reproduce broker and client timeouts with real Mqttable packet evidence.

Author:
Mqttable Team
Published:
Hand-drawn conceptual cover of a device, server, and interrupted heartbeat; not experiment evidence

Setting Keep Alive to 10 seconds does not mean that a connection must receive a PINGRESP every 10 seconds. Nor does it guarantee that a visible DISCONNECT will arrive exactly 15 seconds later. The first assumption confuses a client's transmission duty with response waiting; the second overlooks activity accounting, timer granularity, and how the broker closes the connection.

This article starts with MQTT 3.1.1 and MQTT 5.0, follows configuration and timeout paths through three broker implementations, then reproduces both timeout directions using real Mqttable Clients, Proxy, and Trace. The source and lab versions are EMQX 6.0.0, HiveMQ Community Edition 2026.5, and Mosquitto 2.1.2. Commercial HiveMQ configuration is covered separately; CE source is not evidence of commercial internals.

The standard measures a client transmission interval

CONNECT carries Keep Alive as a two-byte unsigned integer measured in seconds. It limits the time from finishing transmission of one MQTT Control Packet to starting transmission of the next. It is not the permitted interval between application messages or a universal response deadline. MQTT 3.1.1 section 3.1.2.10 and MQTT 5.0 section 3.1.2.10 assign separate responsibilities to the sender and receiver.

Let the effective Keep Alive be K:

  • The client must keep its transmission interval within K. With a nonzero value and no other Control Packet to send, it must send PINGREQ.
  • If the broker receives no client Control Packet within 1.5 × K, it must close the Network Connection as though the network failed.
  • After sending PINGREQ, a client should close the connection if PINGRESP does not arrive within a reasonable time. The standard does not define that reasonable time as a universal 1.5 × K.

Sending and receiving are different observation points. A client SEND row cannot establish broker receipt: a proxy may discard the packet. When calculating a broker timeout, start at the last client packet the broker accepted, not the latest client SEND or the moment a Toxic was enabled.

Zero, the field maximum, and application traffic

K=0 disables MQTT Keep Alive. It does not guarantee an immortal connection: shutdown, administration, and other errors can still close it. The field maximum is 65535 seconds, or 18 hours, 12 minutes, and 15 seconds. That is an encoding limit, not a recommendation.

PUBLISH, SUBSCRIBE, and PUBACK are Control Packets too. Regular upstream application traffic can satisfy the transmission obligation without an additional ping, although an SDK may still send PINGREQ. In contrast, broker-to-client QoS 0 traffic does not directly satisfy the client's upstream duty. Downstream QoS 1/2 is different: the client's PUBACK or PUBREC becomes upstream activity. "The client is receiving lots of messages" and "the broker is receiving client packets" are not interchangeable statements.

MQTT 5: determine effective K from CONNACK

A MQTT 5 CONNACK can include Server Keep Alive. The client must use this instead of its requested CONNECT value. Only when the property is absent does the requested value remain effective. If CONNECT requests 30 but CONNACK returns 10, analyze subsequent timing against 10 seconds. A saved form that still displays 30 does not disprove the override. Server Keep Alive definition

MQTT 3.1.1 has no equivalent property. A broker that limits long intervals must apply its own policy; it cannot send a negotiation field that a MQTT 3 client does not understand. This is an important source of differences among the three implementations.

Schematic comparing a broker's upstream-idle deadline with client-side PINGRESP waiting for K=10 seconds

Figure 1. A schematic, not packet-capture data. The upper clock starts at the last packet accepted by the broker. The lower example shows this client implementation's checking cadence, not a standard-mandated 15-second response timeout.

EMQX 6.0.0: a multiplier and received-packet checks

Configuration separates override from timeout tolerance

The isolated baseline instance uses:

hocon
mqtt {
  server_keepalive = disabled
  keepalive_multiplier = 1.5
  keepalive_check_interval = 30s
}

server_keepalive defaults to disabled, so no replacement is specified in MQTT 5 CONNACK. When enabled it is a positive integer. keepalive_multiplier defaults to 1.5 and determines the tolerated idle duration. keepalive_check_interval defaults to 30 seconds, but that does not mean every connection is checked only once every 30 seconds. Pinned configuration schema

The override instance sets server_keepalive=10. In this version it is an override, not merely a maximum applied to oversized requests: requests of 0, 10, and 30 all receive Server Keep Alive 10. MQTT 3.1.1 retains its original value. CONNACK property generation

Received-packet counters, accumulated idle, and timers

The receive path increments recv_pkt for MQTT packets. emqx_keepalive:init retains a counter snapshot and check interval. On a check, a changed count resets accumulated idle to zero. An unchanged count adds one interval; reaching the threshold returns timeout. Receive counter, Keep Alive checker

The following is explanatory pseudocode, not runnable Erlang:

text
interval_ms = clamp(configured_check_ms, 1000, max(K_ms / 2, 1000))
max_idle_ms = ceil(K_ms * multiplier)

on_check(current_recv_pkt):
    if current_recv_pkt != previous_recv_pkt:
        idle_ms = 0
    else:
        idle_ms += interval_ms
        if idle_ms >= max_idle_ms:
            timeout()

For K=10, the configured 30-second interval is clamped to 5 seconds, while the threshold is 15 seconds. For K=1, the minimum check interval remains one second. This is discrete checking rather than a millisecond-accurate absolute deadline reset by every ordinary packet. The PINGREQ handler also checks the counter immediately and resets the Keep Alive timer. Reading only the periodic checker would miss this path.

Timeout exit and the legacy backoff field

A timeout result sends the Channel into its RC_KEEP_ALIVE_TIMEOUT disconnect path. MQTT 5 can expose a real DISCONNECT 0x8D. MQTT 3.1.1 has no server DISCONNECT packet and cannot expose the same wire reason code. Timeout exit

Older material often describes keepalive_backoff=0.75 and K × backoff × 2. The 6.0.0 schema retains a hidden compatibility field and converts it into a multiplier. Its default-multiplier branch can consult backoff, while an explicit nondefault multiplier is preserved. Use the current field in new configuration instead of setting both and guessing precedence from their names. Compatibility converter

Increasing the multiplier relaxes broker-side tolerance beyond the standard's 1.5 baseline. It does not change the client's protocol duty, restore a blocked link, or fix a stalled event loop.

HiveMQ: separate commercial documentation from CE internals

What the commercial configuration documents

The commercial HiveMQ documentation specifies a default max-keep-alive of 65535 and allow-unlimited=true. Oversized MQTT 5 requests receive a Server Keep Alive in CONNACK. MQTT 3 clients retain their own interval, and the allow-unlimited setting does not apply to them. Commercial configuration documentation

Commercial HiveMQ was not run in this lab, and its private internals were not inspected. The following call chain and measurements belong to HiveMQ CE 2026.5, even where their externally visible behavior matches the documentation.

xml
<hivemq>
  <mqtt>
    <keep-alive>
      <max-keep-alive>10</max-keep-alive>
      <allow-unlimited>true</allow-unlimited>
    </keep-alive>
  </mqtt>
</hivemq>

This is the MQTT fragment; a running configuration also needs a TCP listener. The evidence bundle contains the complete lab configurations.

CE does not simply reject a zero interval

For MQTT 5, CE substitutes the server maximum when the requested value exceeds that maximum, or when the request is zero and unlimited intervals are disabled. CONNACK advertises the substituted value. Otherwise it keeps the client's interval. MQTT 3.1.1 does not enter this override branch. CONNECT/CONNACK processing, effective timer value

Thus "allow-unlimited=false rejects clients requesting zero" is not an accurate description of this source. In the lab with a maximum of 10 and allow-unlimited=false, MQTT 5 with K=0 connects successfully and receives an override of 10. MQTT 3.1.1 with zero also connects successfully and retains zero.

A Netty read-idle handler and disconnect queue

ConnectHandler multiplies effective K by 1.5, converts that duration to integer seconds, and installs KeepAliveDisconnectHandler. This handler sits at the front of the Netty pipeline. It observes inbound reads and read completion, updates lastReadTime, and uses System.nanoTime() to calculate elapsed read-idle time. Read-idle handler

text
 timeout_seconds = integer(effective_K * 1.5)
 on_read_complete:
     last_read = monotonic_now()
 on_timeout_task:
     remaining = timeout - (monotonic_now() - last_read)
     if remaining > 0: reschedule(remaining)
     else: enqueue_disconnect(channel)

The tracked activity is Netty inbound reading. Do not translate that into "only completely decoded MQTT Control Packets reset the timer." Nor does a nanosecond monotonic clock make the resulting disconnect nanosecond-accurate: the caller already truncated the interval to integer seconds, and event-loop scheduling contributes additional delay.

Expired channels enter KeepAliveDisconnectService. The default batch size is 100; the first batch is scheduled after 100ms, subsequent batches also use 100ms intervals, and the actual disconnect is passed back to the channel's event loop with KEEP_ALIVE_TIMEOUT. Disconnect queue

These steps help explain an observed close later than a nominal 1.5K. They are not an extra grace period universally granted by the protocol. CE 2026.5 also lists the MQTT 5 timeout fix #675 in its release notes. This article uses that fixed version and does not turn an older defect into a permanent claim about HiveMQ.

Mosquitto 2.1.2: a maximum policy and a time wheel

A zero default maximum does not mean zero is accepted under every policy

The isolated baseline instance uses:

conf
listener 24883 127.0.0.1
allow_anonymous true
persistence false
max_keepalive 0

The listener binds only to loopback. Anonymous access is for this local lab, not a production recommendation. In 2.1.2, max_keepalive defaults to zero, imposing no additional maximum on the protocol field and allowing a client to disable Keep Alive. Do not carry an older version's default into this comparison. Configuration manual, 2.1.2 default

The restricted instance uses max_keepalive 10. Its CONNECT branch tests both a value greater than the maximum and a value of zero. MQTT 5 receives Server Keep Alive 10. MQTT 3.1.1 receives CONNACK 0x02, Identifier rejected, and is disconnected. CONNECT policy

In this controlled test, 0x02 is caused by Keep Alive policy; it does not establish a duplicated Client ID. A reason-code label is not a substitute for identifying its triggering condition.

How the time wheel schedules expiry

The default compile branch uses a ring of expiry buckets. A connection occupies the bucket given by:

text
 deadline = last_msg_in + floor(K * 3 / 2)
 slot = deadline % wheel_size

New activity removes the connection from its old bucket and reinserts it using the updated last_msg_in. As time advances, the main loop processes expired buckets. When max_keepalive is zero, the wheel is allocated using the protocol maximum; its length is not zero. Connections with K=0 do not participate in this expiry mechanism. Time-wheel implementation

The source's O(1) versus O(n) comparison is useful for maintaining an individual connection's expiry position without scanning all connected clients. It does not mean that processing a bucket containing many simultaneous expirations has constant total cost.

Receive updates, rounding, and the old branch

Normally, after a full packet is read and handled, the reader calls keepalive__update. There is also a partial-large-packet branch: on EAGAIN, when more than 1000 bytes remain to be received, it updates Keep Alive too. The source explains this as accommodating slow reception of large messages. Packet reader

The wheel uses integer seconds and integer K*3/2 arithmetic. Rounding is more visible for tiny or odd K values, and main-loop scheduling affects the final observation. The main lab therefore uses 10 seconds rather than generalizing from a one-second experiment.

The WITH_OLD_KEEPALIVE compile branch scans connections every five seconds and compares current time with last_msg_in. This article analyzes that branch but did not run a separately compiled old-mode binary. It is not a runtime multiplier option in mosquitto.conf. Old checker

The default timeout path calls do_disconnect with MOSQ_ERR_KEEPALIVE, logs exceeded timeout, and closes the transport. This TCP experiment received no MQTT DISCONNECT 0x8D. Correlate the broker log with TCP closure rather than inventing a reason-code packet to make the screenshots symmetrical. Disconnect path

Compare policies, not just option names

Scroll the table sideways to see all columns.

ItemEMQX 6.0.0HiveMQ CE 2026.5Mosquitto 2.1.2
Relevant defaultsServer Keep Alive disabled; multiplier 1.5; check interval 30sMaximum 65535; allow-unlimited truemax_keepalive 0
Configuration scopemqtt configuration and zones; lab uses the default zoneGlobal mqtt configuration; per-connection handlerGlobal max_keepalive; reloadable
MQTT 5 request over a 10s policyserver_keepalive=10 overrides even requests not over 10max=10 overrides oversized requestsmax=10 overrides oversized requests
MQTT 3.1.1 request 30No server_keepalive override; accepts 30No maximum override; accepts 30max=10 rejects with CONNACK 0x02
Default K=0MQTT Keep Alive checking disabledDisabled when unlimited is allowedDisabled with max=0
Restricted MQTT 5 K=0server_keepalive=10 overrides to 10true keeps zero; false overrides to 10max=10 overrides to 10
Restricted MQTT 3.1.1 K=0Accepts zeroAccepts zero, including allow-unlimited=falsemax=10 rejects with CONNACK 0x02
Activity accountingrecv_pkt and accumulated idle; PINGREQ can reset timerNetty read idle and scheduled tasklast_msg_in and time wheel; additional large-packet receive branch
Timing detailsMillisecond calculation, discrete checksInteger-second interval, monotonic clock, 100ms batchingInteger seconds, K*3/2 rounding, main-loop scheduling

The CONNECT policy results come from 24 local probes using unique Client IDs. Accepted connections were closed normally. These policy probes used a small standalone protocol probe, not the Mqttable built-in Client screenshot workflow. Commercial HiveMQ did not participate in this measured table.

Reproduce the complete timeout flow in Mqttable

Prepare one fault direction at a time

The path is Mqttable built-in Client → Mqttable Proxy → local broker. Baseline EMQX listens on 127.0.0.1:22883; Proxy listens on 127.0.0.1:25883 and forwards to it. The Client's broker address must use the Proxy port. HiveMQ CE and Mosquitto use 23883 → 25884 and 24883 → 25885 respectively.

Schematic of Mqttable's built-in Client, Proxy, local brokers, and the two TCP capture points

Figure 2. A schematic. A is Client-to-Proxy and B is Proxy-to-Broker. The broker-timeout case drops upstream PINGREQ while keeping downstream open. The client-timeout case drops only downstream PINGRESP.

Create a Client in Connections. Select MQTT 5, Keep Alive 10 seconds, Clean Start enabled, and Session Expiry 0. Leave Will Topic empty, add no Scheduled Messages, and enable no unrelated Toxic. Verify saved settings against the actual CONNECT; an unsaved form cannot establish the configuration of an active connection.

Mqttable Edit Client showing saved MQTT 5, Keep Alive 10 seconds, and Session Expiry 0

Figure 3. Callout 1 identifies the lab's 10-second Keep Alive; callout 2 identifies Session Expiry 0. The existing saved Client is inspected without changing or saving fields. Actual CONNECT and successful heartbeat evidence separately verifies the configuration in use.

Establish normal heartbeats and inspect an override

After connecting, enable Heartbeat in Trace and observe at least two PINGREQ/PINGRESP pairs. This client version sets force_ping=false. With no application traffic, the observed ping cadence is about 10 seconds; this is not a promise about every MQTT SDK.

Mqttable Trace showing actual PINGREQ and PINGRESP pairs at about ten-second intervals

Figure 4. SEND PINGREQ and RECV PINGRESP are distinct directional observations. Connected status proves the handshake succeeded, not that these heartbeat cycles succeeded.

Configure the independent EMQX override instance with server_keepalive=10 and create another Client requesting 30 seconds. Its real CONNECT requests 30, while CONNACK advertises 10; observed PINGREQ intervals are about 10 seconds. Changing the form to 10 before taking the image would hide the difference being demonstrated.

Mqttable showing the saved 30-second request beside an actual CONNACK Server Keep Alive of 10 seconds

Figure 5. Callout 1 marks the saved request of 30; callout 2 marks the broker's 10. This compares configuration with negotiation evidence. Saved table status is not a substitute for live network evidence.

Drop upstream PINGREQ: let the broker detect idle

Add Keepalive failure to the relevant Proxy. Select Direction Up, Action drop_pingreq, Toxicity 100%, and Schedule Always On. Keep downstream open and do not use a Disconnect Toxic to inject a packet. Delay=30000ms is another saved form parameter; the drop_pingreq branch does not use it. It is not a 30-second timeout setting.

The lab used Pro entitlement obtained through normal sign-in. Keepalive failure is a Pro feature in 1.5.0; do not modify permissions or fabricate entitlement to reproduce it. Free Packet loss supports directional and packet-type matching, but the evidence here comes from the Pro Keepalive failure implementation, not an untested substitute.

Mqttable Keepalive failure editor with upstream direction, drop_pingreq, 100 percent toxicity, and Always On

Figure 6. The existing Toxic is opened for inspection without changing it. This image was taken after the first timeout, so its current connection count is zero. It is not the connection state before fault activation.

In an additional reproduction for native 2× screenshot capture, the last successful broker-side PINGREQ was at 16:48:26.009034. The client sent another at 16:48:36.010573, but no corresponding packet appeared on the broker side. EMQX sent its real DISCONNECT 0x8D at 16:48:41.022589, about 15.014 seconds later. This calculation starts at the last successful broker-side PINGREQ, not the client's last received PINGRESP. The capture reproduction is retained separately from the three timing samples below.

Mqttable retaining the dropped client PINGREQ, last successful PINGRESP, and EMQX's actual DISCONNECT 0x8D

Figure 7. Callout 1 marks the unmodified broker response. Callout 2 is a SEND PINGREQ visible on leg A but absent on B. Callout 3 is the last successful PINGRESP before the fault. RECV and the reason code are actual packet fields, not generated demonstration data.

Mqttable showing Disconnected after a broker Keep Alive timeout with the real timeout Trace

Figure 8. The first observed state was Disconnected. The annotated 15.014 seconds is calculated from this separate capture, not a product timeout-meter field or a protocol accuracy guarantee.

Drop downstream PINGRESP: let the client time out

Remove the upstream Toxic, establish a new connection with two successful heartbeat cycles, then use only Down + drop_pingresp. PINGREQ now reaches the broker. Leg B records the broker's real PINGRESP, but the response does not reach the client on A.

The client closes TCP about 10.002 seconds after sending the unanswered PINGREQ. Mqttable records a SYSTEM EVENT with termination reason :pingresp_timeout. The first runtime observation is Reconnecting. No broker DISCONNECT packet or broker 0x8D was received.

Mqttable client lifecycle EVENT details showing pingresp_timeout rather than a broker DISCONNECT reason code

Figure 9. Callout 1 identifies EVENT; callout 2 identifies the actual client termination reason. This image is a separate native 2× retake under the same controlled conditions. After preserving the first result, the runner stops only this Client's reconnection. The image's final state is not presented as the first Reconnecting observation, which is recorded separately.

In this client dependency version, sending PINGREQ sets awaiting_pingresp. The next Keep Alive timer check ends the connection if the response is still missing. This explains the roughly one-K response wait in this example. The standard says a reasonable time; other SDKs need not use this checking policy. Pinned client dependency, response-state check.

Restore the path, send application traffic, and compare brokers

Delete the Toxic and reconnect the same Client through the original Proxy address, keeping effective K at 10 seconds. Verify that the scoped Toxic list is empty and that the new connection exchanges at least two successful heartbeat pairs. Connected status alone is not recovery proof.

Mqttable Proxy showing a new fault-free connection and restored heartbeat traffic

Figure 10. Callout 1 identifies the new connection and Keep Alive after fault removal; callout 2 identifies restored heartbeat traffic. Figure 4 is also a post-cleanup normal-baseline view. These are independently captured stages, not a claim that every screenshot shows one continuous connection. Original captures retain the initial success, failure, and recovery sequences separately.

The application-traffic control enables upstream drop_pingreq again and publishes QoS 0 with the same built-in Client roughly every three seconds. During 33.010 seconds, the broker-side capture contains 12 actual PUBLISH packets, the Client remains Connected, and it sends no PINGREQ in that window. This demonstrates sufficient upstream application activity in this implementation, not immediate ping suppression by every SDK or equivalent protection from downstream traffic.

HiveMQ CE's actual DISCONNECT 0x8D alongside original Trace evidence

Figure 11. HiveMQ CE 2026.5 returns an actual 0x8D. The selected packet comes from repeat two; all three timings are reported separately below. This is not commercial HiveMQ runtime evidence.

Mqttable SYSTEM EVENT after Mosquitto closes TCP, without a received MQTT DISCONNECT packet

Figure 12. Mosquitto 2.1.2 repeat three closes TCP and produces a :tcp_closed lifecycle reason. SYSTEM EVENT is not a broker MQTT packet; no 0x8D was inserted. After preserving comparison results, the runner explicitly stops each Client so reconnection traffic cannot contaminate later samples.

The table measures from the last successful broker-side PINGREQ to its first disconnect packet or TCP close. It includes the local proxy/Docker path, checking granularity, and scheduling; it is not a pure timer-error measurement.

Scroll the table sideways to see all columns.

BrokerRepeat 1 / sRepeat 2 / sRepeat 3 / sMedian / sRange / sClose evidence
EMQX15.03615.01415.02115.02115.014 to 15.036DISCONNECT 0x8D
HiveMQ CE15.16015.15815.12715.15815.127 to 15.160DISCONNECT 0x8D
Mosquitto16.07215.95915.17115.95915.171 to 16.072TCP close; no 0x8D

Each broker case was repeated three times. This is reproducibility evidence, not a throughput ranking or protocol certification. Explain timer, rounding, and queue differences with source evidence rather than treating one observed timestamp as a universally exact disconnect deadline.

Ask these five questions before changing the interval

  1. What is effective K? Inspect both CONNECT and CONNACK. Use Server Keep Alive when present instead of reading only the saved form.
  2. Which side decided to close first? A broker's 0x8D, a client's pingresp_timeout, and TCP closure are different observations.
  3. Where did the packet disappear? Client SEND, Proxy receive, forwarding, and broker receipt are not the same event.
  4. Was other upstream traffic still present? PUBLISH and acknowledgment packets can keep activity alive. Dropping only PINGREQ does not block every Control Packet.
  5. Was the observation window sufficient? Two successful heartbeat cycles, a controlled fault, the first close, and new successful heartbeats after recovery establish the full sequence.

Do not start diagnosis by simply increasing Keep Alive. A stalled client event loop, one-way proxy loss, and a broker that cannot answer have different causes. A larger value may merely postpone failure detection.

Frequently asked questions

How is MQTT Keep Alive different from TCP keepalive?

MQTT Keep Alive concerns MQTT Control Packets and client/broker activity. TCP keepalive is a transport-layer probe managed by the operating system. TCP acknowledgments and probes are not PINGREQ and do not establish receipt of an MQTT Control Packet by the broker.

Can a Keep Alive timeout happen without a visible 0x8D?

Yes. The broker must close an expired connection, but not every version and implementation first exposes a visible MQTT DISCONNECT reason code. MQTT 3.1.1 has no server DISCONNECT packet. The Mosquitto MQTT 5 TCP lab also shows transport closure and a timeout log rather than 0x8D. Never fill the missing reason code with a synthetic packet.

Does a timeout prove the session was lost?

No. It establishes timeout closure of the Network Connection. MQTT 5 Session Expiry governs retention after disconnection, while Clean Start governs clearing earlier state at connection time. The main experiment uses Session Expiry 0 and makes no claim about persistent-session recovery.

When will the Last Will be published?

A Keep Alive timeout can trigger a Will, but Will Delay, Session Expiry, reconnection, and whether a Will was actually configured all affect the result. The main lab deliberately has no Will, separating failure detection from Will timing. MQTT 5 session and Will definitions

What Keep Alive value should I use?

There is no universal value for every network, device, and broker. Choose based on acceptable failure-detection time, jitter, power consumption, server policy, and SDK behavior, then test it. Ten seconds is a quick, readable lab setting. The form's 300-second default is not a universal production recommendation either.

Use Brokers and connections to configure the Client, Proxy to control the fault direction, and PCAP Analysis to inspect what actually crossed the wire. Identify the timer that expired before deciding which setting to change.