Conversation
The only per-task data for the workers was dramatiq's built-in duration histogram. Releases and errata releases take 2-6 minutes on average, but there was no way to tell whether that time goes to waiting in the queue, the database, Pulp, or our own code. TaskMetricsMiddleware gives each message a timing context held in a context variable. dramatiq calls before_process_message and the actor in the same worker thread, and run_until_complete and asyncio.gather copy the context into every task, so all of them add to the same object. - SQLAlchemy cursor events on the Engine class time every statement and count queries, split into the ALBS DB and the Pulp DB. - PulpClient.request records HTTP time and the wait for the client semaphore separately; wait_for_task polling is accounted as pulp_task_wait. get_repo_modules_yaml, which bypasses request(), is timed too. - When the message finishes, the middleware publishes queue wait, total duration (buckets 0.5s-1h, so 1-10 minute tasks are no longer interpolated across a single 60s-600s bucket), per-component time, business time (the remainder) and the query count. - Named stages: class_measure_work_time_async now also reports its ~16 release planner steps, and the release and errata flows mark their main steps (signature checks, Pulp package lookup, modify and publish, updateinfo preparation, OVAL generation, release log). - albs_pulp_request_seconds records Pulp latency per method, templated endpoint and status. The middleware is a separate class from dramatiq's own Prometheus middleware, which the broker already installs: it has no forks, uses albs_* names, and its samples are served by the existing :9191 exposition server. A test guards that Prometheus stays registered once. prometheus_client is imported lazily. It picks in-memory or multiprocess storage on first import, and dramatiq only sets PROMETHEUS_MULTIPROC_DIR in after_process_boot, after alws.dramatiq is imported. An eager import made every worker metric, dramatiq_* included, disappear from :9191; a test now checks that importing alws.dramatiq does not import prometheus_client. Metrics failures are swallowed everywhere, so they cannot fail a query, a Pulp request or a release. Concurrent calls are summed per component, so components can exceed wall time; business is clamped at zero. Related to AlmaLinux/build-system#558
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Part of AlmaLinux/build-system#553. Fixes AlmaLinux/build-system#558.
The only per-task data for the dramatiq workers was the built-in
dramatiq_message_duration_millisecondshistogram. Over the last weekexecute_release_planaveraged 164s,release_new_errata173s andbulk_errata_release358s, but there was no way to tell whether that time goes to the queue, the database, Pulp, or our own code.Changes
New
alws/utils/task_metrics.py.TaskMetricsMiddlewaregives each message a timing context held in a context variable. Hooks add the time they spend to it, and the middleware publishes it when the message finishes:albs_task_queue_wait_secondsactoralbs_task_duration_secondsactoralbs_task_component_secondsactor,componentdb,pulp_db,pulp_http,pulp_semaphore,pulp_task_wait,business(the remainder)albs_task_db_queriesactoralbs_task_stage_secondsactor,stagealbs_pulp_request_secondsmethod,endpoint,status{id}/{n}alws/dramatiq/__init__.py: registers the middleware and SQLAlchemy cursor hooks on theEngineclass (covers the sync engine inside each async engine).alws/utils/pulp_client.py:request()records HTTP time and semaphore wait separately;wait_for_task()polling counts aspulp_task_wait;get_repo_modules_yaml(), which bypassesrequest(), is timed too.alws/utils/measurements.py:class_measure_work_time_asyncalso reports its ~16 release planner steps as stages.alws/release_planner.py,alws/crud/errata.py: stages for signature checks, CAS, module lookup, Pulp modify/publish, Pulp package lookup, updateinfo preparation, package matching, OVAL generation, release log and GitHub issue creation.Things worth reviewing
Prometheusmiddleware. Registering dramatiq'sPrometheus()again duplicated collectors and conflicted on:9191(c35f672). This is a separate class with noforksand onlyalbs_*names; the existing exposition server already serves everything in the multiprocess directory. A test assertsPrometheusstays registered exactly once.prometheus_clientis imported lazily. It picks in-memory or multiprocess storage on first import, and dramatiq only setsPROMETHEUS_MULTIPROC_DIRinafter_process_boot, afteralws.dramatiqis imported. An eager import made every worker metric,dramatiq_*included, disappear from:9191. The histograms are now created on first use, so no compose/env change is needed, and a test checks that importingalws.dramatiqdoes not importprometheus_client.asyncio.gatherover Pulp calls), so components can exceed wall time;businessis clamped at zero. Stage timings are wall-clock and exact.Testing
tests/test_unit/test_task_metrics.py(21 tests): middleware accounting and skip path, context shared acrossgather, stages inside and outside tasks,class_measure_work_time_async, query counting and recovery after a failed statement, ALBS vs Pulp DB classification, realPulpClientrequests against a local aiohttp server (semaphore split, polling accounted as task wait, error status label), endpoint templating, the single-Prometheusguard, the import-order guard, and metrics failures not reaching callers.pytest --ignore tests/test_ovalagainst Postgres 13).dramatiq --processes 2, both with and withoutPROMETHEUS_MULTIPROC_DIRin the environment: allalbs_task_*series are served on:9191,dramatiq_messages_totalis 3 for 3 messages (no double counting), and there are no port conflicts.Follow-ups
process_errata_release_for_repos,update_errata_references_in_pulp), so a publication finishing in 31s is noticed after 60s.pulp_task_waitshould make this visible.albs_task_component_secondsper actor once this is deployed.