mirror of
https://github.com/Cockatrice/Cockatrice.git
synced 2026-09-21 00:55:09 -07:00
* [Server] Instrument command processing, game starts, and event loops
Add a lock-free MetricsRegistry that accumulates per-command processing
times in preallocated histogram slots (one per protobuf command type,
bucketed at 1/5/10/25/50/100/250/500/1000/2500/5000 ms +Inf). The
hot-path observeCommand() uses only relaxed atomic adds — no locks,
no allocations, no cache-line ping-pong beyond the unavoidable counter
updates.
Wire the registry into AbstractServerSocketInterface::processCommandContainer()
so every processed command is attributed with its container's wall-clock
time. When a container exceeds metrics/slow_command_ms (default 500),
a warning is logged including the connected username.
Add an EventLoopWatchdog heartbeat that runs on every socket pool thread.
If a heartbeat overshoots metrics/stall_warn_ms (default 2000 ms), the
overshoot is recorded in atomic counters and a warning is logged. Both
thresholds are configurable in servatrice.ini; setting stall_warn_ms to 0
disables the watchdogs entirely.
Track game-start durations via a separate histogram in MetricsRegistry.
Server_Game::startGameNow() measures the time from zone creation through
player materialization and reports it via Server::observeGameStartDurationMs().
Add a live card-count gauge: Server_Game exposes getCardsInGame() and
Servatrice::getCardsInGamesTotal() sums across all running games under
the appropriate read locks.
Include a standalone metrics_registry_test (Google Test) that validates
empty registries, single/multi-sample histograms, kind encoding,
overflow-slot collapse, negative-duration clamping, gauge rendering,
and the game-start histogram separation.
Took 10 minutes
* [Client/Server/Protocol] Surface live metrics in the Developer tab
Extend Response_GetServerStats with live counters from the in-process
MetricsRegistry: cards in games, event loop stall totals/worst,
total commands processed, average command time, active command types,
and game-start count/duration. Add a repeated CommandStats message
carrying per-command breakdowns (kind, extension number, resolved
protobuf name, count, total ms) for every type that has seen at
least one sample.
Server-side cmdGetServerStats() populates all new fields after the
existing DB uptime snapshot query, resolving protobuf extension names
via the descriptor pool for human-readable labels like
session/Command_Ping.
Expand TabDeveloper with two tables: an overview section (existing
DB stats plus the new live metrics) and a per-command breakdown table
(Command / Count / Total ms / Avg ms) sorted by total_ms descending
so the hottest commands surface first.
Took 55 minutes
Took 47 seconds
* [Server] Drop dead Prometheus histogram, add developer command metrics, fix watchdog init order
- metrics_registry: remove toPrometheusText/appendCumulativeBuckets and the time-bucket histogram that nothing in production ever emitted (the future /metrics exporter can bring it back); keep counts/totals read by the Developer tab
- Fix +Inf bucket routing that never incremented, and its test that locked the bug in
- Instrument developer_command container (kind 6) in processCommandContainer and stats label resolution
- Read metrics/{slow_command_ms,stall_warn_ms} at the top of initServer() so stall_warn_ms=0 disables the watchdogs before pool threads start
- Shrink KindStride to 1280 (largest extension in use is 1206) with a static_assert; document scrape cost of getCardsInGamesTotal; note slow_command logging has no rate limit in servatrice.ini.example
* [Tests] Give metrics_registry_test an explicit main
* [Server] Record only the dispatched command family; drop unused totals
processCommandContainer recorded every family in a container even though
the base if/else-if dispatch processes at most one. An unauthenticated
client could batch a session command (login) with fabricated developer,
moderator, and admin entries and forge genuine-looking samples that were
never executed or authorized. Mirror the base's selection, skip when the
handler was already deleted, and skip entries whose extension number is
-1 (which would otherwise wrap into the previous kind's id range).
[Server] Drop dead process-lifetime byte/uptime counters
txBytesTotal/rxBytesTotal added an atomic RMW to every socket write and
read for counters nothing consumes (cmdGetServerStats fills tx_bytes,
rx_bytes, and uptime_secs from the DB snapshot). Remove the two atomics
and the getTxBytesTotal/getRxBytesTotal/getUptimeSeconds getters; the
incTxBytes/incRxBytes slots and mutexes remain for the ISL legacy
counters.
[Protocol] Document kind 5 as developer in CommandStats
NumKinds is 6 and the server emits kind_index = 5 for developer
commands; the comment stopped at 4.
* [Client] Togglable auto-refresh for Developer stats tab
* [Oracle] Fix clang-format alignment of card type priority list
---------
Co-authored-by: Lukas Brübach <Bruebach.Lukas@bdosecurity.de>
120 lines
No EOL
3.7 KiB
C++
120 lines
No EOL
3.7 KiB
C++
/**
|
|
* @file metrics_registry.h
|
|
* @ingroup Servatrice
|
|
*/
|
|
|
|
#ifndef METRICS_REGISTRY_H
|
|
#define METRICS_REGISTRY_H
|
|
|
|
#include <QList>
|
|
#include <QString>
|
|
#include <array>
|
|
#include <atomic>
|
|
|
|
/**
|
|
* @brief Lock-free accumulation of command processing statistics.
|
|
*
|
|
* observeCommand() is called once per processed command from whichever socket
|
|
* thread handled it. It uses relaxed atomic adds on preallocated storage only,
|
|
* so it introduces no locks, allocations, or shared cache-line ping-pong
|
|
* beyond the unavoidable counter updates.
|
|
*
|
|
* Reading happens rarely (metrics scraping), accepts momentary tears between
|
|
* related counters, and therefore also needs no synchronization.
|
|
*
|
|
* Only counts and totals are retained. An earlier Prometheus-style cumulative
|
|
* histogram (per-type, time-bucketed) was cut because nothing in the server
|
|
* ever wrote it out; it belongs to the future /metrics exporter that needs it.
|
|
*/
|
|
class MetricsRegistry
|
|
{
|
|
public:
|
|
/**
|
|
* Extension numbers are only unique per command kind, so recorded ids
|
|
* combine the kind index with the protobuf extension number.
|
|
*
|
|
* The stride is only as wide as it needs to be: 1280 is the first round
|
|
* number above the largest extension actually in use (ModeratorCommand =
|
|
* 1206) and keeps the preallocated TypeStats array small. Bump it if a new
|
|
* command exceeds it.
|
|
*/
|
|
static constexpr int KindStride = 1280;
|
|
|
|
static constexpr int NumKinds = 6;
|
|
|
|
static constexpr const char *KindNames[NumKinds] = {"session", "room", "game", "moderator", "admin", "developer"};
|
|
|
|
/// Upper bound on distinct command type ids (see typeIdFor).
|
|
static constexpr int MaxTypes = NumKinds * KindStride;
|
|
|
|
/// Guard against typeIdFor() overflowing into the neighbouring kind's slots.
|
|
static_assert(KindStride > 1206, "KindStride must exceed the highest command extension number in use");
|
|
|
|
static int typeIdFor(int kindIndex, int extensionNumber)
|
|
{
|
|
return kindIndex * KindStride + extensionNumber;
|
|
}
|
|
|
|
void observeCommand(int typeId, qint64 elapsedMs);
|
|
|
|
/**
|
|
* Records how long one game start took to bring every player's zones
|
|
* online. Kept separate from command timings because it is triggered by
|
|
* the server itself and can dwarf any single command when decks are huge.
|
|
*/
|
|
void observeGameStartDurationMs(qint64 elapsedMs);
|
|
|
|
/// Total number of observed commands across all types.
|
|
qint64 totalCommands() const
|
|
{
|
|
return totalCommandsCounter.load(std::memory_order_relaxed);
|
|
}
|
|
|
|
/// Cumulative processing milliseconds across all types.
|
|
qint64 totalTimeMs() const
|
|
{
|
|
return totalTimeCounter.load(std::memory_order_relaxed);
|
|
}
|
|
|
|
/// Number of distinct type slots that have seen at least one sample.
|
|
int activeTypeCount() const;
|
|
|
|
struct ActiveTypeStats
|
|
{
|
|
int typeId;
|
|
qint64 count;
|
|
qint64 totalMs;
|
|
};
|
|
|
|
/**
|
|
* Returns stats for every type slot that has seen at least one sample.
|
|
* Callers resolve the numeric type id to a human-readable label via
|
|
* typeIdFor()/KindNames as needed.
|
|
*/
|
|
QList<ActiveTypeStats> collectActiveStats() const;
|
|
|
|
struct GameStartSnapshot
|
|
{
|
|
qint64 count;
|
|
qint64 totalMs;
|
|
};
|
|
|
|
GameStartSnapshot getGameStartSnapshot() const;
|
|
|
|
private:
|
|
struct TypeStats
|
|
{
|
|
std::atomic<qint64> count{0};
|
|
std::atomic<qint64> totalMs{0};
|
|
};
|
|
|
|
TypeStats &slotFor(int typeId);
|
|
const TypeStats &slotFor(int typeId) const;
|
|
|
|
std::array<TypeStats, MaxTypes> typeSlots{};
|
|
TypeStats gameStartStats{};
|
|
std::atomic<qint64> totalCommandsCounter{0};
|
|
std::atomic<qint64> totalTimeCounter{0};
|
|
};
|
|
|
|
#endif |