Publishing content into the webGCP catalog

Audience: engineering teams who want their skills, tools, or extracted vocabulary discoverable through the webGCP catalog. This is the practical companion to the webGCP v0.1 specification — the spec is normative; this guide is not. For the other direction — querying the catalog on your own bill — see /onboarding/.


What publishing gets you

Your entities land in the shared catalog and become discoverable: semantically searchable through the catalog's public query surfaces (the local_search tool, the webGCP /v0/query envelope, and the registry browse endpoints), and queryable in place via the public BigQuery projection. Every entity carries typed provenance and an explicit truth class, so consumers always know what kind of claim they are reading.

The publishing model, in five rules

1. One door, and registration first. Publishing goes through the catalog's ingest service, and a new source is registered before it writes — an ownership claim reviewed by the catalog maintainers, never a self-service toggle. One source system has exactly one writer; nothing else may write rows under your name. Contact the maintainers (see /contact/) to begin registration; expect to complete a short corpus scorecard describing your content's shape before an adapter is built.

2. Your content's origin decides its truth class — not you. Human-authored content that a machine extracts from (transcripts, articles, documentation) lands as inferred. Generated taxonomies and vocabularies land as synthetic. The classification is enforced at write time and cannot be overridden by the publisher: attempting to smuggle a truth class into a payload is rejected. The canonical class is reserved for the catalog's own evaluation pipelines — no publisher can produce it.

3. Provenance is mandatory and four-part. Every batch carries source_uri, source_kb, retriever_version, and retrieved_at. A missing member is rejected by name. Bump retriever_version when your extraction shape changes, not on every deploy.

4. Ids are yours, and they are literal. You publish under your own stable ids; the catalog derives its entity ids deterministically from them and keeps an alias record for every mapping. Renaming an item mints a new entity rather than moving one — choose natural keys that survive refactors. Duplicate ids within one batch are refused, not silently merged.

5. Never supply your own vectors. Embeddings are produced catalog-side in one pinned embedding space. A vector from any other space would not error — it would rank as confident nonsense, which is worse. Publishers send text; the catalog embeds it.

The call pattern: plan, inspect, commit

Every ingest route supports a dry-run mode that validates your batch and reports what would be written — counts, estimated embedding volume, and the entity ids that would result — without writing anything. The pattern that works:

  1. Plan — send the batch with store: false. Inspect the response.
  2. Bail on any validation error. Validation fails whole, never partially — a rejected batch wrote nothing.
  3. Commit — resend the identical batch with store: true.

Two operational notes from teams that have done this: verify your row count after a second publish (idempotency comes from the catalog's merge, not from your ids alone), and schedule recurring publishes off-peak.

What is coming, honestly

A document-ingestion lane — where you submit documents and the catalog's extraction chain mines the skills for you, rather than you publishing pre-extracted vocabulary — has been ruled and built, and is in staged rollout. It is not yet generally available; publishers today submit extracted entities. Ask the maintainers about current availability rather than coding against it.

Known Limitations and Deferred Work