Found during a targeted code review, not from a user report:
- _drain() could get permanently stuck: valid JSON that isn't an object
(e.g. a bare number or list, which json.loads() happily accepts) crashed
_dispatch()'s dict-oriented logic, and the exception escaped before
self._buf was advanced past the bad packet - so the same malformed bytes
sat at the front of the buffer and re-crashed every subsequent _drain()
call, forcing a reconnect each time. Now the buffer always advances (via
try/finally) and _dispatch() rejects non-dict payloads with a warning log
instead of crashing.
- _pending_msgid/_pending_report were mutated without synchronization
while the reader thread reads/resolves them in _dispatch() - a
check-then-set race on the shared report_key slot meant two concurrent
publish() calls for the same msg_type could have one silently miss its
reply. Added a dedicated lock around all registration/cleanup.
- The report-topic resolution path never verified a reply's msgid matched
the waiter it was about to resolve, so a late reply for an
already-timed-out request could be delivered to an unrelated, newer
caller waiting on the same report_key. Entries now carry their own msgid
for this comparison; replies without a msgid still resolve normally
(most printer push-reports don't carry one).
- upload_gcode() silently sent a request with an empty session token when
the upload URL was missing "?s=", instead of raising a clear error at
the actual point of failure. It also leaked the upload socket's file
descriptor on any send/recv failure other than a timeout, since there
was no try/finally around its lifetime.
New tests in tests/test_client_robustness.py cover all of the above.