live1,247 agents deployed
← All skillsSign up to install

skill-analytics

General↓ 0 installsUpdated 105d ago
Curatedaaronjmars

Fleet-level skill-run analytics — ranks skills by 7d run count, surfaces success rates, exit-taxonomy distribution, and anomaly flags (significance-gated)

SKILL.md preview

---
name: skill-analytics
description: Fleet-level skill-run analytics — ranks skills by 7d run count, surfaces success rates, exit-taxonomy distribution, and anomaly flags (significance-gated)
var: ""
tags: [meta]
---

> **${var}** — Window in hours (default: 168 = 7 days). Pass an integer like "72" for a shorter window.

Today is ${today}. Generate a fleet-level performance view of every Aeon skill that has run in the window. **The point of this skill is to answer four questions in one report:** which skills run most, which fail most, which are silently skipping (new exit taxonomy from the autoresearch-evolution rewrites), and which scheduled skills haven't fired at all. heartbeat gives binary ok/not-ok per run; skill-health audits one skill at a time. This is the only place the operator can see the entire fleet ranked side-by-side.

## Why this exists

`heartbeat` runs three times daily and emits a per-skill ✓/✗. `skill-health` files issues for skills that breach degradation thresholds. Neither produces a ranked, fleet-wide view. The 80 autoresearch-evolution rewrites (aeon PRs #46–#136) introduced new exit taxonomies — `SKIP_UNCHANGED`, `NEW_INFO`, `SKIP_QUIET` — that classify quiet-but-correct runs separately from failures. Existing health checks treat any non-`*_OK` exit as worth attention; the analytics widget makes the actual distribution visible so a skill running mostly `SKIP_UNCHANGED` reads as healthy-quiet, not silently broken.

## Steps

### 1. Determine the window

- Default: 168 hours (7 days). If `${var}` parses as a positive integer, use that many hours instead. Cap at 720 (30 days) — anything longer slows the `gh api` paginate.
- Compute `WINDOW_HOURS=N` and `WINDOW_LABEL` (e.g. `"last 7d"` or `"last 72h"`).

### 2. Pull the run snapshot

```bash
./scripts/skill-runs --json --hours $WINDOW_HOURS > .outputs/skill-analytics-runs.json 2>/dev/null
```

If the script fails (auth, rate limit, sandbox block) or the JSON is empty:
- Log `SKILL_ANALYTICS_NO_DATA — skill-runs returned empty (gh api / sandbox block?)` to `memory/logs/${today}.md` and stop with **no notification**. A silent fleet view is correct on data-fetch failure — fall back rather than guess.

The script's JSON shape (see `scripts/skill-runs`):
```json
{
  "period": {"since": "...", "until": "...", "hours": 168},
  "summary": {"total": N, "succeeded": N, "failed": N, "cancelled": N, "in_progress": N},
  "skills": [{"skill": "name", "total": N, "success": N, "failure": N, "cancelled": N, "in_progress": N, "last_run": "...", "last_conclusion": "..."}],
  "anomalies": {"duplicates": [...], "failing": [...]}
}
```

### 3. Cross-reference with cron schedule

Read `aeon.yml` and build `SCHEDULED_SKILLS`: dict `{skill_name -> {enabled: bool, schedule: str}}` for every entry under `skills:`. Treat `schedule: "workflow_dispatch"` and `schedule: "reactive"` as exempt from the "no runs in window" anomaly — those are dispatched on demand, not by cron.

For every skill in `SCHEDULED_SKILLS` where `enabled: true` AND schedule is a valid cron expression AND the skill is **not** present in the snapshot's `skills` array, mark `silent_scheduled: true` (zero runs in window despite an active schedule).

### 4. Cross-reference with cron-state.json

Load `memory/cron-state.json` if present (missing → empty dict, not failure). For each skill in the snapshot, attach:
- `consecutive_failures` (0 if missing)
- `last_status` (`"unknown"` if missing)

Used to compute the consecutive-failure anomaly without a second `gh api` round-trip.

### 5. Mine exit taxonomy from logs

For each daily log file `memory/logs/YYYY-MM-DD.md` whose date falls in the window, scan for these markers (one match per skill section):
- `_OK` → success (excluding `_OK_SILENT`)
- `_OK_SILENT` / `_QUIET` / `SKIP_QUIET` → quiet-success
- `SKIP_UNCHANGED` → skip-unchanged (autoresearch-evolution exit)
- `NEW_INFO` → new-info (autoresearch-evolution exit)
- `_SKIP*` (other) → skip-other
- `_ERROR` / `_FAILE

…