Cockatrice/servatrice/src/metrics_registry.h
BruebachL 202a5ac958
[Client/Server/Protocol] Surface live metrics in the Developer tab (#7212)
* [Server] Instrument command processing, game starts, and event loops

Add a lock-free MetricsRegistry that accumulates per-command processing
times in preallocated histogram slots (one per protobuf command type,
bucketed at 1/5/10/25/50/100/250/500/1000/2500/5000 ms +Inf). The
hot-path observeCommand() uses only relaxed atomic adds — no locks,
no allocations, no cache-line ping-pong beyond the unavoidable counter
updates.

Wire the registry into AbstractServerSocketInterface::processCommandContainer()
so every processed command is attributed with its container's wall-clock
time. When a container exceeds metrics/slow_command_ms (default 500),
a warning is logged including the connected username.

Add an EventLoopWatchdog heartbeat that runs on every socket pool thread.
If a heartbeat overshoots metrics/stall_warn_ms (default 2000 ms), the
overshoot is recorded in atomic counters and a warning is logged. Both
thresholds are configurable in servatrice.ini; setting stall_warn_ms to 0
disables the watchdogs entirely.

Track game-start durations via a separate histogram in MetricsRegistry.
Server_Game::startGameNow() measures the time from zone creation through
player materialization and reports it via Server::observeGameStartDurationMs().

Add a live card-count gauge: Server_Game exposes getCardsInGame() and
Servatrice::getCardsInGamesTotal() sums across all running games under
the appropriate read locks.

Include a standalone metrics_registry_test (Google Test) that validates
empty registries, single/multi-sample histograms, kind encoding,
overflow-slot collapse, negative-duration clamping, gauge rendering,
and the game-start histogram separation.

Took 10 minutes

* [Client/Server/Protocol] Surface live metrics in the Developer tab

Extend Response_GetServerStats with live counters from the in-process
MetricsRegistry: cards in games, event loop stall totals/worst,
total commands processed, average command time, active command types,
and game-start count/duration. Add a repeated CommandStats message
carrying per-command breakdowns (kind, extension number, resolved
protobuf name, count, total ms) for every type that has seen at
least one sample.

Server-side cmdGetServerStats() populates all new fields after the
existing DB uptime snapshot query, resolving protobuf extension names
via the descriptor pool for human-readable labels like
session/Command_Ping.

Expand TabDeveloper with two tables: an overview section (existing
DB stats plus the new live metrics) and a per-command breakdown table
(Command / Count / Total ms / Avg ms) sorted by total_ms descending
so the hottest commands surface first.

Took 55 minutes

Took 47 seconds

* [Server] Drop dead Prometheus histogram, add developer command metrics, fix watchdog init order

- metrics_registry: remove toPrometheusText/appendCumulativeBuckets and the time-bucket histogram that nothing in production ever emitted (the future /metrics exporter can bring it back); keep counts/totals read by the Developer tab
- Fix +Inf bucket routing that never incremented, and its test that locked the bug in
- Instrument developer_command container (kind 6) in processCommandContainer and stats label resolution
- Read metrics/{slow_command_ms,stall_warn_ms} at the top of initServer() so stall_warn_ms=0 disables the watchdogs before pool threads start
- Shrink KindStride to 1280 (largest extension in use is 1206) with a static_assert; document scrape cost of getCardsInGamesTotal; note slow_command logging has no rate limit in servatrice.ini.example

* [Tests] Give metrics_registry_test an explicit main

* [Server] Record only the dispatched command family; drop unused totals

processCommandContainer recorded every family in a container even though
the base if/else-if dispatch processes at most one. An unauthenticated
client could batch a session command (login) with fabricated developer,
moderator, and admin entries and forge genuine-looking samples that were
never executed or authorized. Mirror the base's selection, skip when the
handler was already deleted, and skip entries whose extension number is
-1 (which would otherwise wrap into the previous kind's id range).

[Server] Drop dead process-lifetime byte/uptime counters

txBytesTotal/rxBytesTotal added an atomic RMW to every socket write and
read for counters nothing consumes (cmdGetServerStats fills tx_bytes,
rx_bytes, and uptime_secs from the DB snapshot). Remove the two atomics
and the getTxBytesTotal/getRxBytesTotal/getUptimeSeconds getters; the
incTxBytes/incRxBytes slots and mutexes remain for the ISL legacy
counters.

[Protocol] Document kind 5 as developer in CommandStats

NumKinds is 6 and the server emits kind_index = 5 for developer
commands; the comment stopped at 4.

* [Client] Togglable auto-refresh for Developer stats tab

* [Oracle] Fix clang-format alignment of card type priority list

---------

Co-authored-by: Lukas Brübach <Bruebach.Lukas@bdosecurity.de>
2026-09-11 17:18:56 +02:00

120 lines
No EOL
3.7 KiB
C++

/**
* @file metrics_registry.h
* @ingroup Servatrice
*/
#ifndef METRICS_REGISTRY_H
#define METRICS_REGISTRY_H
#include <QList>
#include <QString>
#include <array>
#include <atomic>
/**
* @brief Lock-free accumulation of command processing statistics.
*
* observeCommand() is called once per processed command from whichever socket
* thread handled it. It uses relaxed atomic adds on preallocated storage only,
* so it introduces no locks, allocations, or shared cache-line ping-pong
* beyond the unavoidable counter updates.
*
* Reading happens rarely (metrics scraping), accepts momentary tears between
* related counters, and therefore also needs no synchronization.
*
* Only counts and totals are retained. An earlier Prometheus-style cumulative
* histogram (per-type, time-bucketed) was cut because nothing in the server
* ever wrote it out; it belongs to the future /metrics exporter that needs it.
*/
class MetricsRegistry
{
public:
/**
* Extension numbers are only unique per command kind, so recorded ids
* combine the kind index with the protobuf extension number.
*
* The stride is only as wide as it needs to be: 1280 is the first round
* number above the largest extension actually in use (ModeratorCommand =
* 1206) and keeps the preallocated TypeStats array small. Bump it if a new
* command exceeds it.
*/
static constexpr int KindStride = 1280;
static constexpr int NumKinds = 6;
static constexpr const char *KindNames[NumKinds] = {"session", "room", "game", "moderator", "admin", "developer"};
/// Upper bound on distinct command type ids (see typeIdFor).
static constexpr int MaxTypes = NumKinds * KindStride;
/// Guard against typeIdFor() overflowing into the neighbouring kind's slots.
static_assert(KindStride > 1206, "KindStride must exceed the highest command extension number in use");
static int typeIdFor(int kindIndex, int extensionNumber)
{
return kindIndex * KindStride + extensionNumber;
}
void observeCommand(int typeId, qint64 elapsedMs);
/**
* Records how long one game start took to bring every player's zones
* online. Kept separate from command timings because it is triggered by
* the server itself and can dwarf any single command when decks are huge.
*/
void observeGameStartDurationMs(qint64 elapsedMs);
/// Total number of observed commands across all types.
qint64 totalCommands() const
{
return totalCommandsCounter.load(std::memory_order_relaxed);
}
/// Cumulative processing milliseconds across all types.
qint64 totalTimeMs() const
{
return totalTimeCounter.load(std::memory_order_relaxed);
}
/// Number of distinct type slots that have seen at least one sample.
int activeTypeCount() const;
struct ActiveTypeStats
{
int typeId;
qint64 count;
qint64 totalMs;
};
/**
* Returns stats for every type slot that has seen at least one sample.
* Callers resolve the numeric type id to a human-readable label via
* typeIdFor()/KindNames as needed.
*/
QList<ActiveTypeStats> collectActiveStats() const;
struct GameStartSnapshot
{
qint64 count;
qint64 totalMs;
};
GameStartSnapshot getGameStartSnapshot() const;
private:
struct TypeStats
{
std::atomic<qint64> count{0};
std::atomic<qint64> totalMs{0};
};
TypeStats &slotFor(int typeId);
const TypeStats &slotFor(int typeId) const;
std::array<TypeStats, MaxTypes> typeSlots{};
TypeStats gameStartStats{};
std::atomic<qint64> totalCommandsCounter{0};
std::atomic<qint64> totalTimeCounter{0};
};
#endif