Run run_full_audit first, then classify every ad against these EXACT rules (the same ones AutoAdy's engine applies — be transparent about the numbers, never hand-wave).
GATES (applied before any verdict) - Skip ads under €5 spend — too little data, any verdict is noise. - Untracked-conversions guardrail: if an ad shows more conversions through an event we don't count yet (untrackedConversions > leads — almost always a custom conversion) it is NEVER kill or scale. Flag WATCH "confirm your conversion" so the user maps their real result event. This is the rule that stops a working ad from being false-killed.
KILL (highest score wins) - Zero leads & spend ≥ €20 -> burning budget, no conversions (score ~90). - CPL > 3x account avg & ≥1 lead -> extremely inefficient (score ~85). - Frequency > 3.5 & CPL > 1.5x avg -> audience fatigued, costs rising (score ~80). - CTR < 0.5% & spend > €20 -> creative not resonating (score ~70). Pause at the AD level only when the ad itself is the problem; never pause an ad inside an adset that's hitting KPI (you break a sequence Meta won't show you).
SCALE (act via the scale skill — verify eligibility there first) - CPL < 0.7x avg & CTR > 1.5% & freq < 3 & ≥3 leads -> strong performer (score ~90). - ≥5 leads & CPL < avg & freq < 2.5 -> room to scale (score ~85). Default move is +20% on the adset budget — staged, not a jump.
WATCH (not critical yet — monitor) - Warming up (spend < €20). Moderate frequency 2.5-3.5. CPL 1-2x avg. CPM > €30.
TRACKING & PIXEL HYGIENE (audit_tracking auto-checks pixel freshness + server/CAPI events + UTMs) - Pixel firing recently (stale signal invalidates every decision below it). - Conversions API / server events present (detected from pixel stats by source; browser-only loses signal to ad-blockers + iOS). - Event Match Quality — surface what the Graph API exposes; do not invent a score (verify manually in Events Manager). - UTMs present and consistent on ad URLs (broken UTMs = blind attribution). - For ecommerce: Shopify/catalog sync fresh. A tracking break looks exactly like a performance drop — rule it out FIRST (see the diagnose skill's "tracking" theory).
OUTPUT: lead with the one or two things actually costing money, quantify the recoverable daily spend, then the watch list. Don't bury the decision under a metrics dump.