[Client/Server/Protocol] Surface live metrics in the Developer tab (#7212)

* [Server] Instrument command processing, game starts, and event loops

Add a lock-free MetricsRegistry that accumulates per-command processing
times in preallocated histogram slots (one per protobuf command type,
bucketed at 1/5/10/25/50/100/250/500/1000/2500/5000 ms +Inf). The
hot-path observeCommand() uses only relaxed atomic adds — no locks,
no allocations, no cache-line ping-pong beyond the unavoidable counter
updates.

Wire the registry into AbstractServerSocketInterface::processCommandContainer()
so every processed command is attributed with its container's wall-clock
time. When a container exceeds metrics/slow_command_ms (default 500),
a warning is logged including the connected username.

Add an EventLoopWatchdog heartbeat that runs on every socket pool thread.
If a heartbeat overshoots metrics/stall_warn_ms (default 2000 ms), the
overshoot is recorded in atomic counters and a warning is logged. Both
thresholds are configurable in servatrice.ini; setting stall_warn_ms to 0
disables the watchdogs entirely.

Track game-start durations via a separate histogram in MetricsRegistry.
Server_Game::startGameNow() measures the time from zone creation through
player materialization and reports it via Server::observeGameStartDurationMs().

Add a live card-count gauge: Server_Game exposes getCardsInGame() and
Servatrice::getCardsInGamesTotal() sums across all running games under
the appropriate read locks.

Include a standalone metrics_registry_test (Google Test) that validates
empty registries, single/multi-sample histograms, kind encoding,
overflow-slot collapse, negative-duration clamping, gauge rendering,
and the game-start histogram separation.

Took 10 minutes

* [Client/Server/Protocol] Surface live metrics in the Developer tab

Extend Response_GetServerStats with live counters from the in-process
MetricsRegistry: cards in games, event loop stall totals/worst,
total commands processed, average command time, active command types,
and game-start count/duration. Add a repeated CommandStats message
carrying per-command breakdowns (kind, extension number, resolved
protobuf name, count, total ms) for every type that has seen at
least one sample.

Server-side cmdGetServerStats() populates all new fields after the
existing DB uptime snapshot query, resolving protobuf extension names
via the descriptor pool for human-readable labels like
session/Command_Ping.

Expand TabDeveloper with two tables: an overview section (existing
DB stats plus the new live metrics) and a per-command breakdown table
(Command / Count / Total ms / Avg ms) sorted by total_ms descending
so the hottest commands surface first.

Took 55 minutes

Took 47 seconds

* [Server] Drop dead Prometheus histogram, add developer command metrics, fix watchdog init order

- metrics_registry: remove toPrometheusText/appendCumulativeBuckets and the time-bucket histogram that nothing in production ever emitted (the future /metrics exporter can bring it back); keep counts/totals read by the Developer tab
- Fix +Inf bucket routing that never incremented, and its test that locked the bug in
- Instrument developer_command container (kind 6) in processCommandContainer and stats label resolution
- Read metrics/{slow_command_ms,stall_warn_ms} at the top of initServer() so stall_warn_ms=0 disables the watchdogs before pool threads start
- Shrink KindStride to 1280 (largest extension in use is 1206) with a static_assert; document scrape cost of getCardsInGamesTotal; note slow_command logging has no rate limit in servatrice.ini.example

* [Tests] Give metrics_registry_test an explicit main

* [Server] Record only the dispatched command family; drop unused totals

processCommandContainer recorded every family in a container even though
the base if/else-if dispatch processes at most one. An unauthenticated
client could batch a session command (login) with fabricated developer,
moderator, and admin entries and forge genuine-looking samples that were
never executed or authorized. Mirror the base's selection, skip when the
handler was already deleted, and skip entries whose extension number is
-1 (which would otherwise wrap into the previous kind's id range).

[Server] Drop dead process-lifetime byte/uptime counters

txBytesTotal/rxBytesTotal added an atomic RMW to every socket write and
read for counters nothing consumes (cmdGetServerStats fills tx_bytes,
rx_bytes, and uptime_secs from the DB snapshot). Remove the two atomics
and the getTxBytesTotal/getRxBytesTotal/getUptimeSeconds getters; the
incTxBytes/incRxBytes slots and mutexes remain for the ISL legacy
counters.

[Protocol] Document kind 5 as developer in CommandStats

NumKinds is 6 and the server emits kind_index = 5 for developer
commands; the comment stopped at 4.

* [Client] Togglable auto-refresh for Developer stats tab

* [Oracle] Fix clang-format alignment of card type priority list

---------

Co-authored-by: Lukas Brübach <Bruebach.Lukas@bdosecurity.de>
This commit is contained in:
BruebachL 2026-09-11 17:18:56 +02:00 committed by GitHub
parent d5d99e4dfb
commit 202a5ac958
No known key found for this signature in database
GPG key ID: B5690EEEBB952194
22 changed files with 818 additions and 7 deletions

View file

@ -66,6 +66,15 @@ Server_AbstractPlayer::Server_AbstractPlayer(Server_Game *_game,
Server_AbstractPlayer::~Server_AbstractPlayer() = default;
int Server_AbstractPlayer::getCardCount() const
{
int result = 0;
for (auto *zone : zones) {
result += zone->getCards().size();
}
return result;
}
void Server_AbstractPlayer::prepareDestroy()
{
delete deck;

View file

@ -43,6 +43,8 @@ public:
Server_AbstractUserInterface *_handler);
~Server_AbstractPlayer() override;
void prepareDestroy() override;
/// Total cards across all of this player's zones. The caller must hold the game's mutex.
int getCardCount() const;
const DeckList *getDeckList() const
{
return deck;

View file

@ -32,6 +32,7 @@
#include "server_spectator.h"
#include <QDebug>
#include <QElapsedTimer>
#include <QRegularExpression>
#include <QTimer>
#include <google/protobuf/descriptor.h>
@ -238,6 +239,17 @@ int Server_Game::getPlayerCount() const
return participants.size() - getSpectatorCount();
}
int Server_Game::getCardsInGame() const
{
QMutexLocker locker(&gameMutex);
int result = 0;
for (auto *player : getPlayers()) {
result += player->getCardCount();
}
return result;
}
int Server_Game::getSpectatorCount() const
{
QMutexLocker locker(&gameMutex);
@ -330,6 +342,9 @@ void Server_Game::doStartGameIfReady(bool forceStartGame)
}
}
// Only actual starts are timed. The early returns above are no-ops.
QElapsedTimer startupTimer;
startupTimer.start();
players = getPlayers(); // players could have been kicked, get new list of players
if (lifecycleStrategy->onGameStarting(this) == Server_GameLifecycleStrategy::StartAction::Handled) {
locker.unlock();
@ -373,6 +388,7 @@ void Server_Game::doStartGameIfReady(bool forceStartGame)
activePlayer = -1;
nextTurn();
room->getServer()->observeGameStartDurationMs(startupTimer.nsecsElapsed() / 1000000);
locker.unlock();

View file

@ -123,6 +123,8 @@ public:
return gameStarted;
}
int getPlayerCount() const;
/// Total cards across all players' zones. Takes gameMutex itself.
int getCardsInGame() const;
int getSpectatorCount() const;
QMap<int, Server_AbstractPlayer *> getPlayers() const;
Server_AbstractPlayer *getPlayer(int id) const;

View file

@ -180,6 +180,11 @@ public:
{
return false;
}
/// Called once per actual game start with how long bringing every player's
/// zones online took, so servers can spot deck sizes that wedge threads.
virtual void observeGameStartDurationMs(qint64 /* elapsedMs */)
{
}
Server_DatabaseInterface *getDatabaseInterface() const;
int getNextLocalGameId()

View file

@ -136,7 +136,7 @@ public:
return timeRunning - lastDataReceived;
}
bool addSaidMessageSize(int size);
void processCommandContainer(const CommandContainer &cont);
virtual void processCommandContainer(const CommandContainer &cont);
void sendProtocolItem(const Response &item);
void sendProtocolItem(const SessionEvent &item);