基于 Postgres 全文检索的站内搜索:按语言词干、前缀匹配、安全摘要和错字容错
代码8 个文件上下文约 511 个 token扫描通过
安装
$
genpm add @core/search你将获得
- 源代码位于 src/lib/search/,共 8 个文件。 (13.2 kB)
- AI 规则位于 src/lib/search/AGENTS.md,另附 IDE 规则文件。
- 自动为你解析 @core/contracts, @core/db, @core/jobs。
README
此包没有 README。
约 511 个 token→ src/lib/search/AGENTS.md→ .cursor/rules/genpm-core-search.mdc
这正是你的 AI 在 src/lib/search 中工作时读取的内容。不会向其上下文添加其他任何内容。
@core/search — rules for AI agents
Purpose
Site search on Postgres full-text search: one search_documents table fed by any SearchSource (@core/contracts),
weighted title/body vectors with the stemmer of each document's language, prefix matching for search-as-you-type,
snippets returned as plain parts (never HTML) and, if the pg_trgm extension exists, a typo-tolerant fallback on
titles. No external engine, no semantic search.
Map
index.ts— public API:search,registerSearchSource,reindex,indexDocument,removeDocument,reindexJob.search.ts— indexing and querying.schema.ts— table and GIN index.adapters/hono.ts—searchRoutes().adapters/next.ts—searchRoute(GET).
Integration
- Generate and apply migrations (see
src/lib/db/AGENTS.md). Optional typo tolerance: runcreate extension if not exists pg_trgm;once (available on Neon, Supabase, RDS). - Register sources in the project file
src/genpm/search.ts, imported at startup:registerSearchSource(contentSearchSource('pages', { path: ({ slug }) =>/${slug}})). - Build the index once and then daily:
await reindexJob.enqueue({})andschedule('search.reindex', '0 3 * * *')(@core/jobs). - Query:
const { hits } = await search(q, { locale }), or mountsearchRoutes()/searchRouteat/api/search. - Render snippets by mapping parts:
hit.snippet.map(p => p.match ? <mark>{p.text}</mark> : p.text). - Verify: index a document and find it by the first letters of a word in its title.
Conventions
- Sources yield plain text bodies (strip Markdown/HTML first, e.g. with
toPlainTextof @core/rich-text). - Pass
localewhen you know it: the query uses that language's stemmer and the GIN index. - Only index published, public content; private data never goes into
search_documents.
Don't
- Don't build SQL from the query string;
search()already sanitizes it to words. - Don't render snippets with
dangerouslySetInnerHTML. - Don't call
reindex()inside a request: use the job.
# @core/search — rules for AI agents
## Purpose
Site search on Postgres full-text search: one `search_documents` table fed by any `SearchSource` (@core/contracts),
weighted title/body vectors with the stemmer of each document's language, prefix matching for search-as-you-type,
snippets returned as plain parts (never HTML) and, if the `pg_trgm` extension exists, a typo-tolerant fallback on
titles. No external engine, no semantic search.
## Map
- `index.ts` — public API: `search`, `registerSearchSource`, `reindex`, `indexDocument`, `removeDocument`, `reindexJob`.
- `search.ts` — indexing and querying. `schema.ts` — table and GIN index.
- `adapters/hono.ts` — `searchRoutes()`. `adapters/next.ts` — `searchRoute` (GET).
## Integration
1. Generate and apply migrations (see `src/lib/db/AGENTS.md`). Optional typo tolerance: run
`create extension if not exists pg_trgm;` once (available on Neon, Supabase, RDS).
2. Register sources in the project file `src/genpm/search.ts`, imported at startup:
`registerSearchSource(contentSearchSource('pages', { path: ({ slug }) => `/${slug}` }))`.
3. Build the index once and then daily: `await reindexJob.enqueue({})` and `schedule('search.reindex', '0 3 * * *')` (@core/jobs).
4. Query: `const { hits } = await search(q, { locale })`, or mount `searchRoutes()` / `searchRoute` at `/api/search`.
5. Render snippets by mapping parts: `hit.snippet.map(p => p.match ? <mark>{p.text}</mark> : p.text)`.
6. Verify: index a document and find it by the first letters of a word in its title.
## Conventions
- Sources yield plain text bodies (strip Markdown/HTML first, e.g. with `toPlainText` of @core/rich-text).
- Pass `locale` when you know it: the query uses that language's stemmer and the GIN index.
- Only index published, public content; private data never goes into `search_documents`.
## Don't
- Don't build SQL from the query string; `search()` already sanitizes it to words.
- Don't render snippets with `dangerouslySetInnerHTML`.
- Don't call `reindex()` inside a request: use the job.
应用 .genpmignore 后将被注入的确切目录树。固定于
// Tablas de @core/search. Las recoge drizzle-kit vía src/lib/db/drizzle.config.ts.
import { customType, index, pgTable, primaryKey, text, timestamp } from 'drizzle-orm/pg-core';
const tsvector = customType<{ data: string }>({ dataType: () => 'tsvector' });
/** Índice de búsqueda: una fila por documento (`type` + `id`) con su vector ponderado (título A, cuerpo B). */
export const searchDocuments = pgTable(
'search_documents',
{
type: text('type').notNull(),
id: text('id').notNull(),
locale: text('locale').notNull(),
url: text('url').notNull(),
title: text('title').notNull(),
body: text('body').notNull(),
/** Configuración de texto de Postgres usada (`spanish`, `english`, `simple`…). */
config: text('config').notNull(),
tsv: tsvector('tsv').notNull(),
indexedAt: timestamp('indexed_at', { withTimezone: true, mode: 'date' }).notNull().defaultNow(),
},
(t) => [
primaryKey({ columns: [t.type, t.id] }),
index('search_documents_tsv_idx').using('gin', t.tsv),
index('search_documents_locale_idx').on(t.locale, t.type),
],
);
export type SearchRow = typeof searchDocuments.$inferSelect;
此包未声明 MCP 服务器。
| 版本 | 提交 | 发布时间 | 扫描 |
|---|---|---|---|
| 1.0.0 | ae66128 | 3小时前 | 扫描通过 |
- npm
- zod ^4.0.0
- 建议
- GenPM 会给出 npm 命令建议,只有你同意时才会运行。
- 被以下包使用(1)
- @core/kit-cms ^1.0.0
- 扫描
- 扫描通过 · 0 个问题
- 提交
- v1.0.0 → ae6612826aed11b87040adc5f891b5bf6e858c58 · 获取后已校验
- 脚本
- 无。GenPM 从不运行包中的代码。
- 许可证
- MIT
- 质量
- 100/100
- 可识别的许可证已满足
- AGENTS.md 说明了用途已满足
- AGENTS.md 包含集成步骤已满足
- AGENTS.md 列出约定或禁止事项已满足
- 包含测试已满足
- 通过安全扫描已满足
- 最近 6 个月内发布已满足
- 已验证的发布者已满足
- 摘要和关键词已满足
- 举报
- 发现问题了吗?