feat(kb): daily feeds brief — sweep, classify, personalized rank, dual-write (K-H)#286
Conversation
Design of record for three KB capabilities on top of v1.0 (#235-#242): - K1 paste any link (web / 公众号 / B站 / YouTube) -> Markdown source - K2 a watchlist -> daily fetched, classified, summarized brief - K3 [[slug]] parsed into a real link graph; index.md becomes a ranked map-of-content; memory stores [[kb:slug]] links instead of content Review of the shipped v1.0 code found five gaps this plan closes, one of them a confirmed bug: both TF-IDF tokenizers (kb/search.ts:45 and memory/vector.ts:212) split on whitespace, so a whole CJK run collapses into a single token and Chinese content only matches an exact whole-string query. Also: normalizeSlug strips non-ASCII so every Chinese title falls back to entry-<timestamp>; web_fetch is lossy plain text with no metadata and no store; [[slug]] is documented in SCHEMA.md but nothing parses it; and there is no KB scheduler at all. Includes the pro/con debate for each decision (self-built extractor vs deps vs remote reader, immutable sources vs re-ingest, autonomous ingest as a prompt-injection path, consent, ASCII vs CJK slugs, brief storage, and RSSHub for feeds without one) and a nine-PR ship order. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
Three foundation fixes the rest of the KB v2 work stands on
(docs/PLAN_KNOWLEDGE_BASE_v2.0.md K-B). All three are things that would
have corrupted or lost data once link ingest starts pouring Chinese
content in, so they land before the firehose.
1. CJK search was broken. Both TF-IDF tokenizers (kb/search.ts,
memory/vector.ts) carried the same three lines: lowercase, strip
non-word chars, split on whitespace. Chinese has no whitespace, so a
whole clause survived as ONE token:
doc "这篇公众号文章讲的是知识库的设计" -> one token
query "知识库" -> one token
intersection = empty -> zero hits
i.e. CJK content was only findable by an exact whole-clause query.
Replaced by a shared src/tokenize.ts that emits character bigrams for
CJK runs (plus the run itself, so an exact phrase still ranks
highest), segments at latin/CJK boundaries so "gpt模型" indexes as
"gpt" + CJK bigrams, and now keeps kana and CJK Ext-A, which the old
character class dropped entirely. Latin tokenization is unchanged.
2. Chinese titles produced throwaway slugs. normalizeSlug strips every
non-[a-z0-9] char, so a CJK title normalized to "" and addSource fell
back to entry-<Date.now()> — opaque, unstable, un-dedupable. New
kb/slug.ts mints <date>-<sha256/8> for titles with no usable latin,
keeping slugs ASCII on purpose (macOS NFD vs git NFC filenames, URLs,
git paths) with the real title in frontmatter. canonicalUrl() drops
utm_*/known tracking params so the same article shared twice
canonicalizes — and hashes — identically.
3. Entries can now carry arbitrary frontmatter via KbEntry.extra,
round-tripped on read/write. This is where ingest provenance goes
(url/site/author/published/hash/via) without the store knowing about
any of it, and it means keys a future version adds survive a
round-trip through an older one. Values are whitespace-collapsed on
write so a newline can't forge frontmatter lines on the next read.
Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
…d MOC v1.0 told Lisa to cross-link wiki pages with [[slug]] (SCHEMA.md) and recorded each page's sources: in frontmatter — but nothing ever read either one. The KB was a pile of files with decorative links. src/kb/links.ts builds the graph from three edge sources ([[slug]], [[slug|alias]], [[kb:slug]] in bodies; a wiki page's sources: frontmatter) and derives what you can't get from search: backlinks, hubs, orphans, and links pointing at pages that don't exist yet. index.md is now a ranked map of content rather than a flat listing. It is injected into EVERY system prompt and hard-capped there (prompt.ts:165), so with a flat list the truncation lands at an arbitrary point once the KB grows and whatever fell below the line stops existing as far as Lisa is concerned. Ordering wiki pages by (1 + backlinks) x recency means the cut always eats the least-connected tail. Orphans and broken links are listed as the tending to-do list, and every truncation says how much it dropped instead of trailing off. Sources are now listed by TITLE ONLY. Layer 1 is raw captured material — after link ingest lands that means whole web pages — and index.md is always-on prompt text, so an excerpt there is a direct "any page on the internet -> Lisa's system prompt" path. This is a stronger form of the containment the plan assigned to K-I (it covers every source, not just origin: web), and it costs nothing: a title is all a map needs, and the body is one kb_read away. Wiki gists stay — Lisa wrote those herself. Also: kb/index.json (the machine-readable graph, for the web UI and tools), a kb_links tool (links to / linked from / shares tags with), and kb_read now takes a bare slug or a [[wikilink]] with the layer optional, and appends the pages that link to what you just read. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
K-A/K-B/K-C are open as #278/#279/#280. This is the pickup document for whoever continues: the branch stack and how not to break it, the APIs the first three PRs landed, a file-by-file spec for each remaining PR, the repo conventions that are easy to trip over (LISA_HOME must be set before import, the tool-subset triple, the giant template literals in lisa-client.ts), and the eight decisions that are already settled so they don't get re-litigated. Two things it corrects against the plan: the index-side injection containment assigned to K-I already shipped in K-C, and in a stronger form (every source is title-only in index.md, not just origin: web), so K-I is down to the watchlist restriction and the kb_read fence. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>
…rvative autolink (K-D) Closes the "memory stores links, not content" loop from the v2.0 plan (user need ③). Four pieces: - prompt.ts inlines page titles for [[kb:slug]] pointers found in MEMORY.md / USER.md at assembly time ([[kb:oauth]] → [[kb:oauth]](OAuth 与 PKCE)), so a pointer is meaningful without a kb_read. Titles come from the already- generated kb/index.json behind a stat-fingerprint cache — the prompt builds every turn, so no full-KB read on this path. Unresolvable links are kept verbatim: a dangling pointer is a "write this page" signal, not noise. - memory tool description now steers knowledge OUT of memory (4KB cap) and into the KB with a [[kb:slug]] pointer. - The idle tending prompt closes the loop from the other side: after distilling a wiki page, back-link it from any related memory entry. - kb/autolink.ts: conservative title → [[link]] pass applied on kb_write. Whole-word title matches only, first occurrence, longest-first, never in code / existing links, short titles skipped — a wrong edge pollutes the (backlinks × recency) hub ranking the always-on index sorts by, so misses are preferred over false links. Emits [[slug|Title]] when the slug differs so CJK prose keeps reading as prose. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…K-E) kb_ingest(url) turns a link into a searchable, citeable source entry (v2.0 user need ②). New src/kb/ingest/, all zero-dependency (D1: what the user reads never leaves the machine): - html-to-md.ts: lenient HTML tree parser + markdown serializer (headings, nested lists, fenced/inline code, blockquotes, tables, links, lazy-loaded images, emphasis). Prose is markdown-escaped BEFORE syntax is added, so page text can't forge emphasis/links/wikilinks into the KB; code bodies can't terminate their own fence; javascript: hrefs are dropped. - readability.ts: small Readability — score candidates by text mass, <p> count, link density, class/id hints; fall back to <body> when unsure (noise is recoverable by distillation, clipped content is not). - provenance.ts: og:* / JSON-LD Article / <title> / rel=canonical / lang → frontmatter (url · site · author · published · lang · hash · via). - dedupe.ts: kb/.ingested.json (hash → slug), rebuilt from the sources' own hash: frontmatter when missing/corrupt. Same URL returns the existing entry (D2, immutable sources); force re-ingests + records supersedes:. - index.ts: ingestUrl orchestration with an adapter seam for K-F. ALL fetching goes through web_fetch's fetchFollowingSafeRedirects (per-hop private-host validation) plus an up-front assertAllowedUrl — no second fetch path, no second SSRF surface. kb_ingest is REMOTE_BLOCKED (a channel message must not make Lisa fetch an attacker URL into the KB) but NOT autonomous-blocked — K-I scopes autonomous use to the feeds.json domain watchlist (D3) instead of a blanket ban. Tests are offline-fixture only: converter syntax coverage, chrome-stripping, provenance forms, dedupe/force/ledger-rebuild, SSRF rejection before fetch, content-type gates. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Three IngestAdapter implementations behind the K-E seam, for the sites where generic readability extraction can't reach the content: - wechat: article body from #js_content, account from #js_name, epoch publish time from inline `var ct`; lazy-loaded images already handled via data-src. The human-verification interstitial throws a loud, actionable error (open the share link on the phone, paste the text, kb_add) — never a silently captured blank page. - bilibili: metadata via the public view API; b23.tv short links resolve through the guarded fetch. Subtitles are cookie-gated (SESSDATA, which the user may volunteer in kb/feeds.json) — absent that, degrade. - youtube: oEmbed as the reliability floor + one InnerTube player POST for description/captions; captionTracks → &fmt=json3, manual tracks over ASR. Handles the documented failure modes (PO token, 200-with-empty-body, datacenter throttling) by degrading. Fixed subtitle layering (handoff: do not reorder): built-in API → yt-dlp (--dump-single-json only when installed; discovered subtitle URLs are still fetched through the guarded path) → metadata + description. A MISSING TRANSCRIPT IS NOT A FAILURE: the entry records `transcript: unavailable (<reason>)` and the tool reply tells the user they can paste one via kb_add. Supporting changes: fetchFollowingSafeRedirects gains an optional SafeFetchInit (method/headers/body) so adapter API calls stay on the single SSRF-guarded fetch path; ingest types moved to types.ts to break the adapter↔index cycle; the min-body-length gate now applies only to the generic path so degraded video captures can be legitimately short. All tests offline-fixture; degradation paths asserted explicitly. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
… kb CLI (K-G)
Four entry points onto the K-E/K-F engine, one per place a link shows up:
- POST /api/kb/ingest ({url, title?, tags?, force?} — shaped so a future
share-sheet client can call it directly). Ingest failures return 422 with
the engine's actionable message (verification page, paywall, bad type)
instead of a generic 500.
- Knowledge view: a paste-a-link bar above search — progress state, saves,
then opens the new entry; dedupe and no-transcript degradation are surfaced
in the status line.
- Chat: a user message containing a bare URL gets a one-tap 存入知识库 chip
under the bubble (same interaction shape as the KB capture bar; kbToast is
now shared via window.lisaKbToast).
- `lisa kb add|list|search|brief` (src/cli/kb.ts), registered like mail:
passthrough subargs so --title/--tags/--force never collide with global
flags. `brief` prints the newest sources/brief-<date>.md — the file K-H
starts writing — and points at feeds.json until then.
The client edits live inside the MAIN_CLIENT_JS template literal (regex
backslashes double-escaped — the html-syntax test caught the first attempt);
lisa-html-snapshot constants recomputed accordingly.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
…l-write (K-H)
User need ①: a daily, personalized brief over the user's own feed watchlist.
New src/kb/feeds/, deliberately shaped after the mail module (the one place
this problem is already solved end-to-end here):
- store.ts: ~/.lisa/kb/feeds.json ({feeds:[{id,url,tags,max,weight}],
briefHour, budgetTokens}) — user-authored; its existence IS the consent
(D4, no new consent signal; empty/missing = the capability is fully inert).
chmod 0600 on every read + kb/.gitignore coverage, since the file may carry
a bilibili SESSDATA. Machine state (seen ids, last brief date) lives apart
in kb/feeds/.state.json and is rebuild-safe.
- rss.ts: zero-dep RSS2/Atom parsing (CDATA, entities, content:encoded,
atom link@href) — lenient, an unparseable field degrades to blank.
- classify.ts: mail/classify.ts's batch shape over {category, importance 0-3,
oneLine} with the same injection stance (fenced untrusted data, closed
taxonomy validation, heuristic fallback). The daily budgetTokens gate
(default 120k) stops model batches at the ceiling WITH a log line —
remaining items get neutral grades, never dropped silently.
- brief.ts: isBriefDue (isDigestDue's twin), personalized scoring —
(1+importance) × watchlist weight × (1 + interest/wiki term-overlap, both
saturating) — buildBrief + formatBriefText, all pure.
- service.ts: incremental sweep (per-feed seen-id state; a feed failing
degrades, ALL feeds failing doesn't burn the day) → classify → rank →
top-3 full-text ingest via the K-E engine → the D7 dual write:
kb/feeds/<date>.json for the UI and sources/brief-<date>.md as a real
Layer-1 entry (searchable, linkable, distillable; `lisa kb brief` from K-G
picks it up as-is).
Delivery mirrors mail: a 30-min server timer + 20s restart catch-up,
broadcast (kb_brief_update + idle_message source:kb) and a new
pushBridge.onKbBrief behind a new `brief` push pref (default on). Plus
GET /api/kb/brief serving the latest JSON.
Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
|
Reviewed as the penultimate PR of the KB v2 stack. The shape (mail-module mirror), the injection stance in Blocking
Bugs
Security
Nits: Personalization-by-overlap-count (no memory text in the output), the 0600 + gitignore handling of a SESSDATA-bearing |
# Conflicts: # src/kb/ingest/html-to-md.test.ts # src/kb/ingest/html-to-md.ts # src/kb/ingest/index.ts # src/kb/links.test.ts # src/kb/links.ts # src/kb/store.test.ts # src/kb/store.ts # src/tools/web_fetch.ts # src/web/server.ts
|
Reconciled onto main + landed the review blockers: (1) brief-loss fixed — |
Part of the knowledge-base v2.0 stack (plan: #278, handoff: #281). Stacked on #285 (K-G) — merge order: #278 → … → #285 → this.
差距
需求 ①(信息日报)完全缺位:没有订阅、没有分类、没有"跟用户相关"的排序,也没有让日报进入知识系统的通道。
做了什么
新目录
src/kb/feeds/,整体照抄 mail 模块的形状(交接文档指定):isBriefDueisDigestDuesetInterval+ 20 秒重启补跑,unref()classify.tsmail/classify.ts{category, importance 0-3, oneLine},同样的注入防护(数据围栏 + 封闭类目校验 + 启发式兜底)buildBrief/formatBriefTextbuildDigest/formatDigestTextonMailDigestpushBridge.onKbBrief(新briefpush 偏好,默认开)+broadcast({type:'idle_message', source:'kb'})三条不能省的,逐条落实:
feeds.json不存在/为空时,任何入口直接 return null:不碰网络、不调模型;文件的存在即同意,不新增 consent signal。文件每次读取 chmod 0600 并确保进kb/.gitignore(可能含 SESSDATA)。kb/feeds/<date>.json给 UI(配GET /api/kb/brief),sources/brief-<date>.md走addSource进 Layer 1(可搜、可链、可蒸馏;K-G 的lisa kb brief直接读到)。budgetTokens默认 120k/天,分类批次到顶即停并打日志,剩余条目降级为中性评级而不是被丢;Top-N 全文摄取 N=3(走 K-E 引擎,feed 即用户自己的 watchlist,D3 语义自洽)。个性化排序(相对通用 RSS 阅读器的差异点):
(1+importance) × watchlist 权重 × (1 + 兴趣/wiki 词重叠)——兴趣项来自 MEMORY.md/USER.md 分词,wiki 项来自已有页面标题+标签,重叠计数饱和封顶防关键词堆砌刷分。增量抓取:按 feed 记 seen-id 环(封顶 300);单 feed 失败降级为该 feed 无新条目;全部 feed 失败不烧掉当天(下个 tick 重试);成功但无新条目才算完成当天。
测了什么
13 例全离线:RSS2(CDATA/guid/content:encoded)与 Atom 解析、isBriefDue 四态、评分三信号各自抬升、buildBrief 排序 + 文本渲染(含 ingested 反链)、分类校验(未知类目拒收/importance 钳位/空 oneLine 兜底)、模型失败兜底、预算闸门批次截停、端到端(sweep→classify→rank→topN ingest→两份产物+gitignore+0600)、当天去重、全失败不烧天、pickNewItems 去重封顶。
npm run typecheck干净;npm test1268 pass / 0 fail。🤖 Generated with Claude Code