Consuming the webGCP catalog — a quick-start

Audience: teams who want to query the catalog — semantic search, registry browse, or query-in-place over the public projection. Companion to the webGCP v0.1 specification; the spec is normative.


The three doors

1. Semantic search over HTTP — the catalog service at https://catalog.myparallel.dev (always use this origin; underlying hostnames are implementation details that change). It serves the standard discovery documents (/.well-known/webgcp, /.well-known/mcp.json) and two MCP tools — local_search (semantic search over the catalog, filterable by entity type) and read_surface (discovery-first reading of an external page). Plain HTTP works too:

curl -s https://catalog.myparallel.dev/api/v1/cognition-engine/local_search \
  -H 'Content-Type: application/json' \
  -d '{"query": "solar forecasting", "entity_type": "skill", "limit": 5}'

2. The webGCP query envelope — POST /webgcp/v0/query with a §6.4 envelope returns a §6.5 knowledge bundle. Failures are typed statuses on a 200 (manifest_mismatch, out_of_scope), never bare errors — an out-of-scope query is an answer, not a fault. Results carry authority: "derived" always: embedding-ranked retrieval is a query-dependent view, never a canonical assertion.

curl -s https://catalog.myparallel.dev/webgcp/v0/query -H 'Content-Type: application/json' -d '{
  "bundle": {"contract_uri": "urn:webgcp:catalog.myparallel.dev:entity-retrieval/v0.1"},
  "filter": {"query": "vector search skill", "max_results": 3},
  "scope":  {"allow_kbs": ["urn:webgcp:catalog.myparallel.dev:*"]}
}'

3. Query in place via BigQuery — the public projection is published as an Analytics Hub listing (webGCP Public Catalog) that any authenticated Google Cloud user may subscribe to. A subscription materializes a linked dataset in your project; you never hold a grant on the publisher's project, and it is a view, not a snapshot — corrections appear to subscribers immediately.

What is in the catalog — and what is deliberately not

The projection carries the marketplace-safe entity types (skills and agencies today; more as they are ruled non-PII). Identity-bearing types, internal work items, and all operational data are withheld by design, not deferred — the projection is an allowlist, so a new sensitive type can never leak in silently. Consumers never see the canonical truth class; public rows are synthetic or inferred, and every row says which.

The one rule that fails silently

Since 2026-08-24 the projection no longer ships raw vectors — semantic search over the corpus routes through the hosted query door (capped, receipted), or you embed the shipped embed_source_text yourself, in your own project, on the pin. Either way the space rule below still governs, because the pin is what makes your vectors and the network's comparable.

The catalog's space is one pinned embedding space, published in the projection's embedding_space table. A query embedded in any other space — a different model, a preview variant, even the same dimensionality from another family — does not error. It returns confidently ranked nonsense. Read the space id from the table, embed in exactly that space, and note the similarity floor in this space is high (~0.44–0.49): a "0.5 threshold" matches almost everything. Verify your setup by embedding a string already in the corpus and checking self-similarity ≈ 1.0 against your own re-embedding of the same text.

Fair use

The public endpoints are unauthenticated with per-IP rate limits (currently 20/min on search and query routes — back off on 429, honor Retry-After). Service liveness is probed at /readyz; there is no /health route, and a 404 there is not an outage.

Known Limitations and Deferred Work