A RAG index has a shelf life. The moment a source changes, the index starts drifting: a pricing page is updated, a documentation page is rewritten, a URL is retired. A pipeline that crawls and embeds once has no way to notice any of it, and the model keeps answering from a version of the web that no longer exists.
The obvious fix is to crawl everything again and rebuild every embedding. It works, but it treats an unchanged page and a rewritten page as the same problem. On a real corpus most recrawled pages are identical to what is already indexed, and you still pay to chunk and embed them. This post builds the alternative, an incremental RAG index: the Crawlbase Enterprise Crawler recrawls on a schedule and delivers each page to a webhook, a content hash decides whether anything changed, only changed pages are re-embedded into pgvector, and every citation carries the time its source was last verified.
- Recrawling is how you learn what changed. Re-embedding is the part you can skip when nothing did.
- Claim each delivery's
ridbefore doing any work, so a retried webhook never indexes the same page twice. - Hash normalized Markdown, not raw HTML, and compare it with the stored hash. Same hash: touch a timestamp. New hash: re-chunk and re-embed.
- Treat 404 and 410 as deletions. A page that is gone should leave the index, not linger in it.
- Keep the webhook healthy. A failed delivery is crawled and billed again, and a failing endpoint pauses the crawler.
The runnable code lives in ScraperHub/incremental-rag-index-with-crawler-webhooks-and-pgvector, with staged checkpoints under steps/ and the complete application under final/. Every snippet below is quoted from final/.
Why full re-indexing does not scale
A knowledge base is a changing dataset, not a snapshot. Rebuilding the whole index on every crawl solves freshness at the wrong cost: the crawl has to revisit the corpus anyway, but there is no reason to regenerate embeddings for a page whose content is identical to last time, and every rebuild rewrites the vector index for nothing. The useful abstraction is an update loop:
- Recrawl each URL on a schedule.
- Decide whether its indexed content actually changed.
- Re-embed the pages that changed.
- Leave unchanged pages alone, but record that they were verified.
- Remove pages that no longer exist.
The crawl establishes the current state of the source. Everything after it decides whether that state requires an index change. That split is the whole design: crawl cost follows the size of the corpus, while embedding cost follows the rate of change.
Architecture
rid and acknowledges at once; ingestion runs in the background and decides whether pgvector changes.Initial seed URLs are pushed with python push.py; a scheduled job later selects live URLs from the pages table and pushes them again. Both use the same Enterprise Crawler with crawler=NAME and callback=true. A push returns an rid immediately; the crawler owns the queue, concurrency, retries and delivery.
Each delivery arrives at a FastAPI webhook as a POST whose body is the page and whose headers carry the metadata: rid, url, original_status (what the site answered) and cb_status (Crawlbase's result). The webhook claims the rid in PostgreSQL, returns 200, and hands the page to a background task. Ingestion tombstones 404 and 410 pages, skips any non-200 cb_status, hashes the rest, and re-embeds only on a hash change. The /query endpoint searches pgvector and returns citations with last_verified_at.
The data model
PostgreSQL 16 with the pgvector extension holds three tables. deliveries records every rid so a retry can be acknowledged without being processed twice. pages keeps one row per URL with the state the hash gate needs: content_hash, last_verified_at, deleted_at and original_status. chunks holds the text slices and their embeddings.
-- final/sql/schema.sql (excerpt) CREATE TABLE IF NOT EXISTS chunks ( id BIGSERIAL PRIMARY KEY, url TEXT NOT NULL REFERENCES pages (url) ON DELETE CASCADE, chunk_index INTEGER NOT NULL, content TEXT NOT NULL, embedding vector(1536) NOT NULL, created_at TIMESTAMPTZ NOT NULL DEFAULT now(), UNIQUE (url, chunk_index) ); CREATE TABLE IF NOT EXISTS deliveries ( rid TEXT PRIMARY KEY, url TEXT, original_status INTEGER, cb_status INTEGER, received_at TIMESTAMPTZ NOT NULL DEFAULT now() );
The embedding column is vector(1536) because the example uses OpenAI's text-embedding-3-small, and similarity search runs on an HNSW index with vector_cosine_ops. The schema also carries a recrawl_every column that the current code does not use yet; recrawling runs on one global interval, and per-URL cadence comes up at the end.
One property of the model matters later: chunks.content holds overlapping retrieval slices, not a copy of the page. You cannot rebuild the original Markdown from them, so changing the chunking strategy or the embedding model means getting the source again.
Pushing URLs to the Enterprise Crawler
Create a named crawler in the Crawlers console with webhook delivery and a public HTTPS callback URL. For local development, a tunnel such as ngrok or Cloudflare Tunnel exposes the FastAPI app. Pick the token type for the target: the JavaScript token when pages need rendering or page_wait and ajax_wait.
# final/app/crawler.py def push_options() -> dict: options = { "crawler": settings.crawler_name, "callback": "true", "format": "md", "md_readability": "true", } if settings.page_wait is not None: options["page_wait"] = str(settings.page_wait) if settings.ajax_wait is not None: options["ajax_wait"] = str(settings.ajax_wait) return options def push_url(url: str) -> str: """Enqueue one URL. Returns the Crawlbase rid.""" api = _client() res = api.get(url, push_options()) rid = _rid_from_response(res if isinstance(res, dict) else {}) if not rid: raise RuntimeError(f"Crawler push failed for {url!r}: {res!r}") return rid
The published crawlbase Python package exposes this through CrawlingAPI.get: a crawler push is a Crawling API request with crawler and callback=true. format=md is applied on crawler pushes, so the webhook receives GitHub Flavored Markdown. The readability pass is not part of the crawler's Markdown conversion today, so expect the full page, navigation and footer included, and plan the hash step around that.
The crawler owns retries and pacing, so the app should not add its own per-URL sleep or backoff. Push within the documented rate limits and let the queue drain. The Enterprise Crawler documentation covers configuration, delivery and the management endpoints.
An idempotent webhook
The webhook is the boundary between the asynchronous crawler and ingestion, and it has three jobs: decode the body, recognise health probes, and claim the delivery before any real work starts. Deliveries are gzip-compressed with Content-Encoding: gzip, Markdown included, and a Markdown delivery also carries Content-Type: text/markdown; charset=utf-8.
# final/app/webhook.py (inside the /webhook handler) raw = await request.body() markdown = decode_body(raw, request.headers.get("content-encoding")) if is_monitor_probe(request.headers.get("user-agent", ""), markdown): return Response(status_code=200) if settings.webhook_token and token != settings.webhook_token: raise HTTPException(status_code=401, detail="invalid webhook token") # ... read rid, url, original_status and cb_status from the headers ... if not claim_delivery(rid, url, original_status, cb_status): return Response(status_code=200) background_tasks.add_task( ingest_delivery, rid, url or "", original_status, cb_status, markdown ) return Response(status_code=200)
claim_delivery runs INSERT INTO deliveries ... ON CONFLICT (rid) DO NOTHING and checks the row count: one row inserted means this request owns the delivery, zero means the rid was already processed. Treat delivery as at-least-once. A retry after a timeout carries the same rid, and claiming it before scheduling ingestion means two copies of one delivery can never both write chunks.
Health probes need their own path. Crawlbase checks a callback roughly every five minutes with a request from User-Agent: Crawlbase Monitoring Bot 1.0; it is not a page delivery, so the handler answers 200 and stops. Only 200, 201 or 204 count as healthy. If the probe keeps failing, the crawler stops picking up work and resumes on its own once the endpoint recovers, which keeps a broken deploy from draining the queue into a dead endpoint.
Two operational points follow. First, authenticate deliveries with a token in the callback URL (?token=..., set as WEBHOOK_TOKEN) rather than an IP allowlist. Second, keep the acknowledgment fast and move embedding out of the request: a delivery that fails is re-queued, crawled again and billed again, so a slow handler turns an application problem into crawl spend. FastAPI BackgroundTasks is enough for the example; a persistent worker queue is the production upgrade.
Hash-gated re-embedding
This is where the savings come from. The page is normalized, hashed with SHA-256 and compared with the stored hash before any embedding call is made.
# final/app/ingest.py (inside ingest_delivery) with get_conn() as conn: if original_status in (404, 410): tombstone_page(conn, url, original_status) conn.commit() return "deleted" if cb_status is not None and cb_status != 200: log.warning("skip rid=%s: cb_status=%s (not embedding)", rid, cb_status) conn.commit() return "skipped" normalized = normalize_markdown(markdown) content_hash = content_sha256(normalized) page = fetch_page(conn, url) if page and page["content_hash"] == content_hash and page["deleted_at"] is None: touch_page(conn, url, original_status) conn.commit() return "unchanged" texts = chunk_markdown(normalized) vectors = embed_texts(texts) upsert_page(conn, url, content_hash, original_status) replace_chunks(conn, url, texts, vectors) conn.commit() return "embedded"
cb_status never reaches the embedder, and an unchanged hash never calls it.Normalization is deliberately small: Unicode NFC, Windows line endings converted, trailing whitespace stripped from every line, and runs of blank lines collapsed to at most two. That removes whitespace noise without guessing at meaning. Changed pages are split with tiktoken into chunks of about 512 tokens with a 64-token overlap, embedded, and swapped in for the page's previous chunks.
Because crawler pushes deliver the full page, anything that changes on every render, such as a footer date, a session banner or a rotating promo block, changes the hash and triggers a re-embed. If your sources do that, strip the known boilerplate before hashing. It is a few lines in normalize_markdown, and it is what turns "the page was fetched" into "the content changed".
The payoff is straightforward. Recrawling still uses Crawlbase requests, but embedding and index writes follow the change rate: recrawl 1,000 documents where 30 changed and you make embedding calls for 30.
Deletions, status codes and redirects
Deletion is decided before anything else. A page whose site answers 404 or 410 is tombstoned: deleted_at is set, content_hash is cleared, and its chunks are deleted, so it can no longer surface in search. Keep the two statuses apart: original_status is what the site said, while cb_status is whether Crawlbase got a usable answer. A response can carry original_status 200 alongside a non-200 cb_status, and that body must not be embedded.
Redirects need one decision. When the crawler follows an HTTP redirect, the url header carries the final URL and original_status carries the 3xx code. A push for http://example.com/doc can therefore come back as https://example.com/doc/ and create a second pages row. The example keys pages on the delivered URL and recrawls that. If you need stable identity across redirects, store a separate pushed_url and manage the relationship explicitly rather than letting one page silently replace another.
Scheduled recrawls and crawler state
Recrawling reuses the same push path as the seed. The job reads live URLs from pages, skips tombstoned rows, and never re-reads the seed file:
# final/app/recrawl.py def run_recrawl() -> dict[str, str]: results: dict[str, str] = {} for url in list_live_urls(): try: results[url] = push_url(url) log.info("recrawl queued url=%s rid=%s", url, results[url]) except Exception: log.exception("recrawl push failed url=%s", url) results[url] = "error" return results
APScheduler runs it every RECRAWL_INTERVAL_HOURS inside the app (0 disables it); final/recrawl.py runs the same job from cron, and POST /recrawl triggers it on demand. For queue health, GET /crawler-stats proxies the crawler stats endpoint, which lists every crawler on the token with its waiting count, concurrency, latency and a paused flag. That flag is the first thing to check after a deploy:
curl "https://api.crawlbase.com/crawler/YOUR_TOKEN/stats"
Freshness-aware answers
The query path embeds the question, runs cosine similarity against chunks, and joins pages so each hit carries its verification time:
# final/app/query.py cur.execute( """ SELECT c.content, c.url, p.last_verified_at FROM chunks c JOIN pages p ON p.url = c.url WHERE p.deleted_at IS NULL ORDER BY c.embedding <=> %s::vector LIMIT %s """, (qvec, k), )
The excerpts go to the chat model as context, and the API returns the answer with citations holding the URL, the excerpt and last_verified_at. Attaching freshness at retrieval time is the point: a source verified an hour ago and one verified six weeks ago carry different weight, even when their content hashes match.
Push URLs, get each page delivered to your webhook as Markdown, and let the crawler own the queue, retries and pacing. Failed crawls are not billed. Start with up to 5,000 free requests, no card.
Running it in production
Keep a copy of the source if you may re-chunk. Because chunks cannot rebuild the page, a new chunking strategy or embedding model means fetching every page again, unless you kept them. Adding store=true to a webhook push also keeps the raw HTML of each successfully delivered page in Cloud Storage, at half a credit per stored page per month, so a later re-chunk can convert from storage instead of recrawling. A crawler can instead be created in Storage mode, which delivers to storage rather than to a webhook; a crawler uses one delivery mode or the other.
Give each URL its own cadence. The scheduler uses one global interval, but the schema already has pages.recrawl_every. A per-URL scheduler can recrawl when last_verified_at + recrawl_every has passed, so a pricing page is checked hourly and an archived changelog monthly.
Size the savings honestly. With 800 pages at about 2,000 embedding tokens each, a full nightly re-index embeds roughly 1.6 million tokens. If 4% of pages change, hash gating cuts that to about 64,000 tokens, while the crawl volume stays the same. Set the recrawl interval by how stale a source is allowed to get: crawling pays for checking freshness, and hashing keeps embedding proportional to change.
Conclusion
An index goes stale the day it is built, and a full rebuild buys freshness at a price that grows with the corpus. Splitting the problem in two fixes the cost: the Enterprise Crawler recrawls and delivers on a schedule, and the ingestion layer decides from a content hash whether the index needs to change at all. Unchanged pages cost a timestamp, changed pages cost their embeddings, and deleted pages leave.
The full implementation is in ScraperHub/incremental-rag-index-with-crawler-webhooks-and-pgvector. To run it, create a free Crawlbase account and set up a webhook crawler.
Frequently Asked Questions (FAQs)
Does incremental indexing remove the need to recrawl?
No. Pages still have to be recrawled to learn whether they changed or disappeared. Incremental indexing removes the embedding and index writes for pages that did not change.
Why hash normalized Markdown instead of raw HTML?
Markdown drops markup that changes without the content changing, and normalization removes whitespace noise on top of that, so the hash tracks what is actually indexed. Crawler pushes deliver the full page, so strip boilerplate that changes on every render, such as footer dates, before hashing.
What happens when a page has not changed?
The new hash matches pages.content_hash, the existing chunks stay, no embedding call is made, and last_verified_at is updated to record the check.
How are deleted pages removed from the index?
A delivery with original_status 404 or 410 tombstones the page: it is marked deleted, its hash is cleared, and its chunks are removed, so it can no longer appear in search results.
What does a failed webhook delivery cost?
A failed delivery is re-queued and the page is crawled again, and each attempt is billed. If the monitoring probe keeps failing, the crawler pauses until the endpoint is healthy. Acknowledge quickly and do the heavy work in the background.
Crawl any site at scale, without fighting infrastructure.
Crawlbase handles proxies, fingerprints, and CAPTCHAs so your team ships data pipelines instead of maintaining crawl plumbing. Up to 5,000 requests free, no card required.
