--- name: daily-finds description: Capture-only daily sweep for unique.rzen.dev. Scans every source in sources.json, writes one find file per qualifying item to src/content/find/, marks the day's shortlist, adds Amazon affiliate links where the product is confidently matched, performs mechanical source maintenance, and nominates new sources. Does NOT write bundles or reviews — that's /build-bundle and /build-review (or /daily-pipeline for the whole chain). Use when the user asks for "today's finds", "morning sweep", "capture", or runs /daily-finds. disable-model-invocation: true --- # daily-finds Pure capture plus triage: every item that passes the bar becomes a `find` file, and the run ends by ranking the day's captures into a **shortlist** (up to 12) that drives downstream automation. Whether a find appears in a `bundle` or graduates to a review is still a separate step — but the shortlist is the pipeline's selection signal. This skill is designed to run unattended. Never pause for confirmation mid-run; make the calls, log them, and surface anything odd in the end-of-run digest. Invoke from inside the `unique.rzen.dev` repo. ## Inputs - `sources.json` at the project root — curated source list with per-source fetch strategy (`feed` / `webfetch` / `curl` / `skip`), optional `feedUrl` (RSS/Atom) and `fetchUrl`. - `daily-finds.log.md` — running log of every find slug ever written (used to skip duplicates). - `src/content/find////index.mdx` — already-captured finds (every existing slug is a "do not re-suggest" item; slugs are globally unique across years/months even though folders are nested by date). - `src/content/reviews//index.mdx` — already-reviewed items (do not re-suggest). - `candidate-sources.md` — newly nominated sources awaiting user vetting. ## The bar A find qualifies if and only if it is **(a)** delightful, unusual, or distinctive in a way the audience would care about *and* **(b)** usable daily (relax to "memorable for a visitor" for travel). Drop generic mass-market items, ad-copy listicles, speculative / not-yet-shipping products, self-promo, and anything not traceable to a real currently-buyable / currently-visitable thing. **Hard exclusion — no tobacco or alcohol.** The site does not promote either in any form. Drop any item whose point is consuming, preparing, serving, storing, or celebrating alcohol or tobacco: spirits, wine, beer, cocktail kits, barware, decanters, wine openers, home-bar and whisky gear; cigars, pipes, vapes, ashtrays, smoking accessories. This is categorical — no exception for beautiful design, craftsmanship, or gift appeal, and it overrides every other signal. Borderline cases (dual-use glassware, non-alcoholic "spirits", brewery-adjacent merch) get skipped too: when in doubt, skip. (Restates `EDITORIAL.md` "Content rules" — the charter is the source of truth.) Crucially: do **not** filter by "is it the best of the day." Capture broadly. Selection happens in the shortlist step, and editorial treatment later. ## Procedure ### 1. Scan every source For each source in `sources.json`: - **`fetch: "feed"` (preferred)**: `curl -sL --max-time 25 -A "Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/605.1.15 (KHTML, like Gecko) Version/17.0 Safari/605.1.15" ""` and parse the RSS/Atom entries (title, link, summary, date). Consider only entries newer than the source's last productive run (when in doubt, the last ~30 days). Feeds bypass Cloudflare/HTML breakage — when a `webfetch`/`curl` source keeps failing, look for a feed and switch it over (see step 6). - **`fetch: "webfetch"` or no field**: WebFetch against `fetchUrl` if present, else `url`. Extract recently-featured items (name, link, one-line description). - **`fetch: "curl"`**: shell out to curl with the Safari UA above against `fetchUrl|url` and parse the HTML (or JSON — e.g. Shopify `products.json` endpoints). Treat HTTP 403 with a non-empty HTML body as a content response — only treat empty bodies or Cloudflare challenge markers as blocked. - **`fetch: "skip"`**: log the source as skipped (with `skipReason`) and move on. Do not retry (re-probing skipped sources is a step-6 quarterly job, not a per-run one). Respect `cadence`: sources marked `weekly` / `monthly` / `as-needed` only need scanning when their cadence window has elapsed since the last productive scan — check the log; skip (and log as `cadence-skipped`) otherwise. This cuts run cost and rate-limit exposure. Parallelize fetches in batches of ~8 concurrent requests; finish all sources before proceeding. For each source record: items extracted, or `empty` / `blocked` / `skipped` / `cadence-skipped`. ### 2. Filter For each extracted item, drop: - Anything whose slug already exists anywhere under `src/content/find/` (find slugs are globally unique — check the leaf folder name across all `//` buckets). - Anything whose name appears anywhere in `daily-finds.log.md`. - Anything reviewed in `src/content/reviews/` (compare by `name` field). - Items that fail the bar above. ### 3. Write find files For every survivor, write one file at `src/content/find////index.mdx` where `` and `` come from today's date. Hero image (when present) sits next to it as `src/content/find////hero.`. The leaf folder name is the slug; URLs stay flat at `/find//` because the content-collection loader uses a `leafSlug` id derivation. - Slug: kebab-case form of the name (lowercase, ASCII, no leading numbers). - Frontmatter: `name`, `subtitle` (optional, one-liner), `date` (today), `link` (external URL — required), `linkText` (optional), `source` (the source name from sources.json — required, exact match), `topics` (1-3 values chosen from the controlled vocabulary in `src/lib/topics.ts` — the `TOPIC_META` keys; judge from the item itself, the source's `topics` in sources.json are hints only and may contain retired values; the build fails on anything outside the vocabulary; `[]` if nothing genuinely fits), `tags` (free editorial tags inferred from the item — keep to 2-4; add `giftable` when the item would make a genuinely good gift), `description` (optional meta override). - Body: 2-3 sentence note (40–90 words), written to `EDITORIAL.md` — concise, and delightful through specificity: name a *specific* virtue. Bad: "great design". Good: "the silicone bumper means it survives drops onto concrete; the lanyard hole is positioned where your thumb naturally rests". - No daily cap — write every survivor. If `link` for a Reddit-discovered item points at a v.redd.it / i.redd.it URL, fetch the post's selftext or comments to find the actual product URL before writing the file. #### Supporting media For every find, try to retrieve a single representative image from the linked page so the find has visual support if it later graduates to a review (or gets surfaced on a tile). Heroes also gate bundling — a bundle needs at least 2 hero-bearing items (aiming for ~4) before it can ship — so a find without a hero can hold a whole theme back: try a second source (developer site, App Store, Wikimedia) before giving up. The quality bar below still wins, though — a bad image is worse than none. - **Where to look (in priority order):** `og:image` meta tag → JSON-LD `image` field → main product photo / hero image markup → first article image. - **Where to save:** `src/content/find////hero.` — colocated next to `index.mdx`. Preserve the original extension (`.jpg`, `.png`, `.webp`, `.gif`, `.avif`). One file per find — `hero` is the canonical name. - **What to skip:** images with marketing/sales overlay text (Amazon-style "BEST PRESS / NO GRIT" tags), tiny thumbnails (<400px on the long edge), generic site logos, paywalled CDN URLs that 403 to curl. If the only available shots are marketing collages, skip rather than save a noisy image — a missing hero is fine, a bad one isn't. - **For Mac/iOS apps:** prefer the App Store hero or developer-site screenshot. For physical products: prefer a clean studio shot or lifestyle photo without overlay copy. For travel destinations: prefer a Wikimedia Commons / Wikipedia image (free-to-use only). For text-only finds (quotes, dad jokes, essays): no media expected — skip. - **Fetch strategy:** `curl -sL --max-time 25 -A "Mozilla/5.0 (...Safari...)" "" -o `. WebFetch is for HTML, not binaries. Log image-fetch failures (404, 403, marketing-only) as a one-line note in the capture-run section of `daily-finds.log.md` so we can revisit broken media-source patterns later. ### 4. Score and shortlist After all finds are written, rank the day's captures and mark the shortlist — the up-to-a-dozen finds the rest of the pipeline should act on. - Score each of today's finds on: **delight** (would the audience light up), **specificity of virtue** (does the body name something concrete), **media** (clean hero present), and **variety** (the shortlist should not be twelve variations of one topic — spread across topics when quality allows). - Set `shortlist: true` in the frontmatter of the top finds. Aim for 6–12 on a normal day; fewer on a thin day. **Never pad** — a 4-item shortlist beats a 12-item one with filler. - Record the shortlist in the log, ranked, with a one-line rationale each. Note the top 1–3 as **review candidates** (strong media + strong story) — `/build-review` (or the pipeline) picks from these. #### Amazon affiliate lookup For every **shortlisted physical product** (tools, gifts, clothing, kitchen, EDC — not apps, not travel, not quotes/jokes/web pages): - Check whether the exact product is sold on Amazon: WebSearch `site:amazon.com " "`, or curl the Amazon search page. - Only when the hit is unambiguously the **same product** (same maker, same model — not a lookalike, not a marketplace knockoff), add `amazonLink: https://www.amazon.com/dp/` to the find's frontmatter. Use the canonical `/dp/` form with no query parameters — the site's affiliate layer (`src/lib/affiliate.ts`) appends the tracking ID, badge, and disclosure at render time. - When in doubt, omit. A wrong affiliate link is worse than none. The site's focus is interesting items, not monetizable ones — this is a bonus, not a filter. - Non-shortlisted physical finds may get the same treatment when the match is trivially confident; don't burn time on them otherwise. ### 5. Nominate new sources Watch for new sources organically and via search. No cap on nominees per run. - **From inline references in source posts.** When a post on a known source credits or links out to another publication as where they originally discovered the thing, capture that publication as a candidate source. Trigger phrases: *"via X"*, *"first seen on Y"*, *"spotted by Z"*, *"h/t X"*, *"originally posted on Y"*, or any byline / sidebar link pointing to a different domain that itself looks like a curated feed. Cool Tools, Recomendo, Kottke, and Daring Fireball especially tend to credit their finds. - **From general observation.** When a publication is mentioned in a way that suggests it'd fit the brief (curated, taste-driven, focused on delightful or unusual items), nominate it. - **From targeted search.** If the run produced few or no organic candidates, run one `WebSearch` for `"best newsletter"` or `" curator blog"`. Append every nominee to `candidate-sources.md` under today's date heading. For each: name, URL, a one-line case for adding, and how it was surfaced. Adding new sources to `sources.json` remains a **user decision** — never auto-add. ### 6. Source maintenance (mechanical edits — apply, don't recommend) This skill has authority over mechanical `sources.json` hygiene. Past runs accumulated "overdue" recommendations nobody executed; that pattern is over. Apply these directly during the run: - **Hard-blocked source** (Cloudflare challenge, empty 403s, dead DNS) for the 3rd consecutive run: before skipping, probe for a feed (`` on the homepage, or common paths `/feed`, `/rss.xml`, `/atom.xml`, `/index.xml`). A working feed rescues the source — set `fetch: "feed"` + `feedUrl`. Otherwise set `fetch: "skip"` with a dated `skipReason`. - **Dormant source** (live but no new content across 3+ runs): demote `cadence` to `as-needed`; if the newest content is over a year old, `fetch: "skip"` with reason "dormant since ". - **Better fetch path discovered** (feed URL, Shopify `/collections/new/products.json`, article-level URL pattern that bypasses an index-level block): record it in `fetch`/`feedUrl`/`fetchUrl` immediately. - **Quarterly re-probe**: if any `fetch: "skip"` entry has a `skipReason` dated 90+ days ago, re-test it this run; blocks get lifted. Remove the skip when the source is back. - Update the top-level `updated` field when you touch the file. Every mechanical change gets one line in the log's maintenance section. Only *new* sources need the user (step 5). Note on Reddit: anonymous JSON/RSS endpoints have been blocked since ~June 2026. If `REDDIT_CLIENT_ID` / `REDDIT_CLIENT_SECRET` are available in the environment (or `.env`), obtain an app-only OAuth token (`https://www.reddit.com/api/v1/access_token`, grant_type=client_credentials) and fetch `https://oauth.reddit.com/r//top?t=week` with it. If no credentials exist, log the Reddit sources as blocked-pending-oauth and move on — don't retry anonymous endpoints every run. ### 7. Update the log Append to `daily-finds.log.md`: ``` ## YYYY-MM-DD — capture run Sources scanned: total — hits, empty, blocked, skipped, cadence-skipped. Productive sources: . Captured finds: - — source: - ... Shortlist (, ranked): 1. [review candidate] 2. ... Amazon links added: (, , ...) — or "none". Source maintenance: - — or "none". New source nominees: (see candidate-sources.md). ``` ### 8. Report to the user Short digest in chat (or, when running as part of `/daily-pipeline`, feed this into the pipeline digest): ``` Captured finds across productive sources; shortlisted . Review candidates: , . Amazon links: . Source maintenance: changes. New nominees: . Blocked: . ``` ## Notes - "Why it fits" / find descriptions should name a *specific* virtue. Bad: "great design". Good: "magnetic connector means it never strains the cable when a foot catches it". - Travel destinations relax the daily-use criterion — a place qualifies if a visitor would remember it years later. - If a source's homepage is too noisy, prefer adding a `feedUrl`/`fetchUrl` to `sources.json` rather than working around it inside the skill. - Never invent slugs that collide with existing finds or reviews. If a name would produce a colliding slug, append a year or distinguishing word. - Source attribution must be exact — the `source` field in each find file must match a `name` in `sources.json` so the `/sources` page cross-reference works. - For large sweeps, fan out: scanner subagents per source cluster → central dedup/bar/shortlist pass → writer subagents for files + heroes. Keep the shortlist ranking in the main context — it's the editorial judgment call of the run.