Audience: engineering teams who want their skills, tools, or extracted vocabulary discoverable through the webGCP catalog. This is the practical companion to the webGCP v0.1 specification — the spec is normative; this guide is not. For the other direction — querying the catalog on your own bill — see /onboarding/.
Your entities land in the shared catalog and become discoverable: semantically searchable
through the catalog's public query surfaces (the local_search tool, the webGCP /v0/query
envelope, and the registry browse endpoints), and queryable in place via the public BigQuery
projection. Every entity carries typed provenance and an explicit truth class, so consumers always
know what kind of claim they are reading.
1. One door, and registration first. Publishing goes through the catalog's ingest service, and a new source is registered before it writes — an ownership claim reviewed by the catalog maintainers, never a self-service toggle. One source system has exactly one writer; nothing else may write rows under your name. Contact the maintainers (see /contact/) to begin registration; expect to complete a short corpus scorecard describing your content's shape before an adapter is built.
2. Your content's origin decides its truth class — not you. Human-authored content that a
machine extracts from (transcripts, articles, documentation) lands as inferred. Generated
taxonomies and vocabularies land as synthetic. The classification is enforced at write time
and cannot be overridden by the publisher: attempting to smuggle a truth class into a payload is
rejected. The canonical class is reserved for the catalog's own evaluation pipelines — no
publisher can produce it.
3. Provenance is mandatory and four-part. Every batch carries source_uri, source_kb,
retriever_version, and retrieved_at. A missing member is rejected by name. Bump
retriever_version when your extraction shape changes, not on every deploy.
4. Ids are yours, and they are literal. You publish under your own stable ids; the catalog derives its entity ids deterministically from them and keeps an alias record for every mapping. Renaming an item mints a new entity rather than moving one — choose natural keys that survive refactors. Duplicate ids within one batch are refused, not silently merged.
5. Never supply your own vectors. Embeddings are produced catalog-side in one pinned embedding space. A vector from any other space would not error — it would rank as confident nonsense, which is worse. Publishers send text; the catalog embeds it.
Every ingest route supports a dry-run mode that validates your batch and reports what would be written — counts, estimated embedding volume, and the entity ids that would result — without writing anything. The pattern that works:
store: false. Inspect the response.store: true.Two operational notes from teams that have done this: verify your row count after a second publish (idempotency comes from the catalog's merge, not from your ids alone), and schedule recurring publishes off-peak.
A document-ingestion lane — where you submit documents and the catalog's extraction chain mines the skills for you, rather than you publishing pre-extracted vocabulary — has been ruled and built, and is in staged rollout. It is not yet generally available; publishers today submit extracted entities. Ask the maintainers about current availability rather than coding against it.