feature: local-usage-stats (4/4) - #1134
Conversation
…cit-any Add new test file to eslint-suppressions.json with count of 26 no-explicit-any suppressions. These are standard test patterns (mock objects, private property access via 'as any') consistent with other test files in the suppressions list. Fixes CI lint failure in PR #25 compile (lint) job.
…cing - Remove UTF-8 BOM (U+FEFF) from costRecalculation.ts and costRecalculation.spec.ts - Fix qwenCodeModels pricing: qwen3-coder-plus inputPrice 0->1.0, outputPrice 0->5.0 - Fix qwenCodeModels pricing: qwen3-coder-flash inputPrice 0->0.3, outputPrice 0->1.5 Fixes invisible-chars CI check and 3 failing costRecalculation tests
…exactly-once recorder - UsageRecorder: per-task exactly-once usage event recording with endpoint domain extraction - costRecalculation: compute effective cost from token deltas and model pricing - Provider usage deltas: moonshot, openai, openai-codex, vscode-lm yield cumulative usage; Task diffs and records - Task finalization: flush pending usage events on abort/complete - ClineProvider: initialize UsageStatsService, expose getUsageStatsService, forward usageStatsChanged to webview - types: add usage-stats schemas and usageStatsChanged ExtensionMessage type
…proper types, fix run->start renames, add UsageEventStore import
The B15 usage-capture cherry-pick was authored against an older base and reverted newer upstream/base behavior in several files, causing e2e-mock subtask timeouts (7 tests) and unit-test failures. Restore clobbered base behavior while keeping B15's genuine usage/cost capture additions: - Task.ts: restore run() + _runPromise/_isHistoryTask, safeEnsureModelFetched (def + 3 call sites), abort-aware ask wait, resume_completed_task via initialStatus, and t() i18n in sayAndCreateMissingParamError. - ClineProvider.ts: scheduler gates on task.run() (completion promise) instead of fire-and-forget task.start(). This is the root cause of the subtask/resume e2e timeouts. - openai-codex.ts: restore service-tier feature alongside cost capture. - moonshot.ts, vscode-lm.ts, vscode-lm-format.ts, eslint-suppressions.json: revert to base (pure clobber, no genuine B15 content). - task-run-dispatch.spec.ts: bind run() (not start()). - openai-usage-tracking.spec.ts: assert totalCost from cost capture.
|
Important Review skippedToo many files! This PR contains 198 files, which is 48 over the limit of 150. To get a review, reduce the PR to 150 files or fewer by splitting it into smaller PRs or changing its base branch. Upgrade to a paid plan to raise the limit. This review couldn't start because sufficient usage credits or metered capacity aren't available. Add credits or update usage-based reviews in the billing tab, then retry. ⚙️ Run configurationConfiguration used: defaults Review profile: CHILL Plan: Pro Plus Run ID: 📒 Files selected for processing (198)
You can disable this status message by setting the Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
Codecov Report❌ Patch coverage is 📢 Thoughts on this report? Let us know! |
f398395 to
2fe7eb2
Compare
The 100K/1M-event shape tests timed out under CI coverage instrumentation (run 30751922341): bulkAppend performs an INSERT OR IGNORE plus a per-row seq SELECT and 4 rollup updates per event, so 100K/1M rows exceeded the 120s/600s per-test timeouts on the instrumented runner. The whole suite took 33 minutes. These are result-shape assertions, not wall-clock benchmarks. Reduce to 1K (100 sessions) and 5K (1000 sessions) events so they assert the same shape/counts but complete deterministically on any runner. The full spec now runs in ~19s locally.
- vscode mock: make EventEmitter a constructable class (was an arrow factory, breaking 'new vscode.EventEmitter' in TaskHistoryStore and DashboardTaskCatalog under the global vitest alias) - specs constructing ClineProvider: add EventEmitter with a disposable- returning event accessor to vscode mocks, and onDidChange to TaskHistoryStore mocks, so DashboardTaskCatalog wiring and dispose work - mimo.spec: align stale parallel-tool-calls expectation with stream-level suppression (only the first call is emitted) - DashboardView.spec: remove unused getByRole destructure - eslint-suppressions: prune stale mimo.ts entry (no longer occurs)
…igration migrateFromGlobalState only wrote per-task files and the index; delegated parents introduced by the migration kept their stale state until the next restart. Run the idempotent reconcileDelegationState() pass after a migration that changed entries, outside the write lock to avoid deadlock.
clearStats() only cleared the NDJSON store, so the dashboard kept showing cleared data from the SQLite projection. After the store clear, reset the stream generation via the coordinator (which pushes a reset snapshot to subscribers) or clear the database generation directly when no coordinator exists. Projection failures are logged, never thrown to the caller. Also replace two pre-existing 'as any' catalog doubles in the service spec with typed doubles (lint gate).
Both terminal finalize paths in Task (completed and failed/cancelled) built the UsageRecordingContext without rootTaskId even though the field exists on Task and is documented on the context, so sub-task usage was never grouped into the parent session. Pass this.rootTaskId at both sites.
The updateRollup fallback recorded uncached input as all-or-nothing (inputTokens when no cache reads, 0 otherwise), so events with cache reads left no base for the dashboard cacheRatio simulation. Compute the uncached input per event from its inclusion semantics — input minus cacheRead/cacheWrite when those are included in input (OpenAI-style), full input when excluded or unknown (Anthropic-style) — and thread it through every rollup path (live append, bulk append, v2/v3/rebuild). Also make upsertSession keep last_activity_ms monotonic (MAX of existing and new) so a backfilled older event no longer moves a session's last activity backward.
doInitialize() runs database.initialize() (whose v4 migration flips timezone_offset_minutes for rows already in SQLite) before the NDJSON migration copies legacy events verbatim with the old inverted sign, so pre-fix NDJSON events stayed wrong forever. Apply the same sign correction to migrated events. Post-fix events are unaffected: they were dual-written to SQLite by UsageEventStore and are skipped by INSERT OR IGNORE. Documented in code why no per-event discriminator exists.
…eam sink on webview cleanup
… labels from b14 merge
…ures, and detail The Dashboard preset/custom range flowed into the stats subscription but the Tasks section ignored it: pages came from the full History catalog, per-task totals were all-time (task_usage_metadata), and task details returned every event. - Add statsQueryRange module: single source for StatsQuery -> half-open [fromMs, toMs) bounds (presets via startOfDayInTimezone, custom from/to ISO, "all" unbounded); UsageStatsService.filterEventsByQuery now uses it too so export and task bounds cannot drift - DashboardTaskCatalog.getPage: optional range filters membership on HistoryItem.ts; totalEstimate becomes the filtered count; (ts DESC, id DESC) revision-tagged cursor semantics unchanged - UsageStatsDatabase: queryTaskUsageByTaskIds/queryEventsByTaskIds take an optional range; bounded aggregation reads usage_events with ms bounds and mirrors upsertTaskUsage semantics (cancelled included, getEffectiveCost, model/provider from the latest in-range event); unbounded keeps the metadata fast path - DashboardTaskProjection: computeTaskPage/computeTaskSummaries/ computeTaskDetail thread the range (membership by creation ts, figures and detail events by occurredAt) - UsageStatsStreamCoordinator: resolves the range per subscription for snapshot pages and drain upserts; new getSubscription(sink) lets the message handler align one-off task page/detail reads with the active stream subscription (unbounded fallback) - DashboardView: drop the range-bound task detail cache on preset/custom range change so expansions refetch against the new range
…task range resolution
…d height With only maxHeight set, the Virtuoso scroller's height:100% resolves against an auto-height parent, collapses to 0px, and deadlocks (zero viewport -> zero rendered items -> zero content height), so the Tasks header showed a count but no rows ever rendered. Drive an explicit height from totalListHeightChanged (capped at 400px) and bootstrap measurement with initialItemCount clamped to the task count (a larger fixed value crashes itemContent with undefined items). Adds Playwright CT regression tests (jsdom mocks Virtuoso and cannot catch this) and switches the dashboard i18n imports to the @src spelling so the CT harness can stub the TranslationContext.
Clicking the active range preset re-armed the resyncing banner without triggering a resubscription, so no snapshot ever arrived to clear it and the indicator spun forever (e.g. on double-click). Gate the banner on an actual preset change and clear it on the custom-range early return.
…ubtasks The Tasks list paged every History task, so subtasks appeared as sibling rows even though each parent row already aggregates its whole subtree (double-counted visually, detached from the summary cards). - Catalog pages root tasks only; bounded-range membership is subtree-based (a root is listed when the root or any descendant was created in range), orphans promote to roots. - DashboardTaskSummary gains childTaskIds; DashboardTaskPage gains childTasks carrying direct children of the page's roots. - Reducer keeps childTasks/subtask upserts out of the visible root order while storing them in the normalized map. - TaskList renders roots; expanding a root with subtasks shows an indented subtask list, and each subtask toggles its own API-call detail. Childless roots expand directly into their detail. - Adds Playwright CT coverage for the expand interaction (jsdom mocks react-virtuoso and cannot exercise it).
The root-only task page reads childrenByParentId for childTasks; the handler/routing specs' catalog stubs predated that index.
b138930 to
37b8788
Compare
…re coverage for codecov/patch
Stack Position
feature/local-usage-statsDescription
https://www.youtube.com/shorts/UHnnOCM1_f0
Full Feature Description
feature/local-usage-statsusage-stats.ts,src/services/stats, the provider/task capture pathsTask.ts, the stats IPCusageStatsMessageHandler.ts, and the UIDashboardView.tsxanduseDashboardStatsStream.ts.Why Split Into 17 PRs
Instead of submitting this feature as a single unified PR, it was split into individual PRs because as code size grows, safely reviewing a PR becomes very difficult. The feature was broken into mutually exclusive individual PRs so that each can be reviewed independently.
What This PR Specifically Changes
Adds SQLite projection, transactional/idempotent migration, local-day rollup, rebuild/query/stream IPC, epoch guard, dashboard summary/session/heatmap/loading/retry UI, and visual coverage. Removes other feature files.
Included Files
src/services/stats/UsageStatsDatabase.tssrc/services/stats/UsageStatsMigration.tssrc/services/stats/UsageStatsProjection.tssrc/services/stats/UsageStatsStreamCoordinator.tssrc/core/webview/usageStatsMessageHandler.tswebview-ui/src/components/dashboard/DashboardView.tsxwebview-ui/src/components/dashboard/useDashboardStatsStream.tswebview-ui/src/components/dashboard/DashboardView.visual.tsxExclusion Scope