Onboarding into TheNetwork — the self-serve path

Changed 2026-08-24. This page previously offered BigQuery Data Viewer on the raw datasets. That offer was never exercised and is superseded by the current one below — a curated public projection, subscribed per-party through BigQuery Analytics Hub. Everything else about the deal (you query from your own project; your queries bill you, not us) is unchanged. Per §8, changes land here first.

1. Getting access

  1. Have a GCP project with billing enabled. That project is where your queries run and where their cost lands (BigQuery on-demand: per TB scanned, see §6).
  2. Send your Google identity — a service account email (...@your-project.iam.gserviceaccount.com), a Google group, or a Workspace domain — to ops@webgcp.org with the domain you operate and what you intend to build.
  3. You are added as a named subscriber on the public Analytics Hub listing (webGCP Public Catalog, exchange webgcp_catalog, project webgcporg, us-central1). Subscribe from your project: BigQuery console → Analytics Hub → Search listings, or:
bq add-data-exchange-subscription \
  --location=us-central1 \
  --data-exchange=projects/webgcporg/locations/us-central1/dataExchanges/webgcp_catalog \
  --listing=catalog_public \
  --destination-dataset=webgcp_catalog_public

Subscribing creates a linked dataset in YOUR project (name it as you like; the examples below assume webgcp_catalog_public). You query it in place — no cross-project grant exists in either direction, nothing is copied, and bigquery.jobs.create + the scan bytes come from your project.

-- prove access, from your own project (bills ~MBs to you):
SELECT entity_type, COUNT(*) AS n
FROM `webgcp_catalog_public.entities`
GROUP BY entity_type ORDER BY n DESC;

2. The surface — what the projection contains

One dataset, webgcporg.webgcp_catalog_public, delivered through the listing. It is a designed projection, not raw table access: filters and column selection are part of the contract.

The corpus

view what
entities ~100k marketplace entities (skills, agencies, and more types as they land): entity_id, entity_type, name, payload (JSON body), truth_class (labeled, never hidden — §3), tags, embed_source_text, embedder_version, embedding_dim, timestamps. Soft-deleted rows are already filtered. Raw embedding vectors do not ship — see §5.
embedding_space the pin: gemini-embedding-2 · 3072 dims · COSINE. Embed your own text on this exact space for any semantic work.

The capability graph

view what
skill_records the operational skill registry (~700 rows): skill_id, type, status, truth_class, side_effect_class, requires_human_approval, description, last_updated_at
role_records the role registry (~70 rows): role_id, display_name, descriptions, applicable_skills, tags, active, template_version, timestamps
agency_records, sector_records agency/sector registries — schema-first: 0 rows today; global rows appear as they land
role_requires_skill role → skill requirement edges — schema-first: 0 rows today
agent_provides_skill, agent_holds_role agent capability edges (agent_id is an opaque key) — schema-first: 0 rows today

The graph views publish the global scope only. Per-tenant graphs exist and never publish — that is a hard filter in the view definitions, not a policy statement.

The evidence surface

Proof lives in append-only evidence lanes (evaluation runs + real production outcomes). Raw evidence rows are not part of the self-serve surface — what you get is derived:

object what
evidence_coverage (view) per skill_id: does canonical proof exist, in which lane, a banded count, and the latest evidence month. The free proven-vs-claimed discriminator.
evidence_distribution (view) network-level banded score distribution (groups under 5 suppressed)
coverage_summary (view) one row: how much of the public catalog carries proof, as of when
evidence_for(skill_id, task_class_id) (table function) per-claim point lookup: up to 50 rows, score/price/time bands (never raw values), 30-day lag, month-resolution dates. Pass NULL as the second argument to skip the task-class filter.
evidence_summary_for(skill_id) (table function) the same, grouped per task class

Fresh, exact, signed verdicts are the hosted attestation door's job (§7) — the self-serve evidence surface is deliberately coarse and 30 days stale. If the table functions are not visible in your linked dataset, use the hosted doors for point lookups and tell us.

Contract vs plumbing: the objects named above are the contract. Anything else you may ever see is internal plumbing and can change or vanish without notice. Build only against the tables above.

3. Truth-class semantics — the one rule you must not skip

Every row tells you what kind of claim it is. The vocabulary, in descending strength: canonical (ground truth from the evaluation or production-outcome lanes) · derived (computed over canonical) · inferred (extracted/self-declared) · synthetic (seeded) · unlabeled (the source table carries no marker; the view says so rather than inventing one).

4. Query patterns (plain SQL, no services)

-- skills matching a keyword, with provenance
SELECT entity_id,
       JSON_VALUE(payload, '$.name')        AS name,
       JSON_VALUE(payload, '$.skill_kind')  AS kind,
       truth_class
FROM `webgcp_catalog_public.entities`
WHERE entity_type = 'skill'
  AND LOWER(JSON_VALUE(payload, '$.name')) LIKE '%column%masking%';

-- a role's required skills (the capability graph; returns rows as the
-- global graph populates -- schema-first today, see section 2)
SELECT r.role_id, r.display_name, rrs.skill_id
FROM `webgcp_catalog_public.role_records` r
JOIN `webgcp_catalog_public.role_requires_skill` rrs USING (role_id);

-- proven vs claimed skills (the evidence discriminator, verbatim)
SELECT s.skill_id,
       c.skill_id IS NOT NULL   AS proven,
       c.has_production_evidence,
       c.evidence_band
FROM `webgcp_catalog_public.skill_records` s
LEFT JOIN `webgcp_catalog_public.evidence_coverage` c USING (skill_id)
WHERE s.status = 'active';

-- evidence detail for one skill (banded, lagged, capped)
SELECT * FROM `webgcp_catalog_public.evidence_for`('SKILL_ID_HERE', NULL);

The corpus is embedded on gemini-embedding-2, 3072 dimensions, COSINE (the embedding_space view is authoritative; embedder_version rides every row). The raw vectors are our embedding spend and do not ship — two supported paths:

  1. The hosted query door — POST /webgcp/v0/query (§7): embed-and-rank on our side, capped at 50 results, receipted. This is the drop-in replacement for local vector search over the corpus.
  2. Bring your own vectors — embed_source_text ships precisely so you can embed any slice of the corpus yourself, in your own project, on the pin the embedding_space view declares. Your spend, your vectors, full control — and comparable with ours because the space matches.

If you fetch text and embed it elsewhere, match the pin exactly: cross-space cosine returns confidently ranked noise, never an error.

6. What it costs you

7. The optional hosted surfaces

Everything above needs none of these. When convenience beats self-sufficiency:

8. Stability contract

Contract objects (§2) evolve additively; renames/removals are announced via this page and the descriptor before they land — exactly as this page's 2026-08-24 change was. Schema-first views gaining rows is not a change event; a column or object disappearing is. The truth-class vocabulary (§3) is governed by TheNetwork's decision log and does not change silently. Questions, access requests, and breakage reports: ops@webgcp.org.