API

The JSON API for building websets, plus export and run-event endpoints.

Dataembed's API is the same one the web app uses. There are no API keys: anonymous websets need no credentials, and owned websets use the better-auth session cookie.

Access follows one rule everywhere: a webset owned by a user is accessible only with that user's session cookie; a webset created anonymously has no owner and is open to anyone holding its (unguessable) UUID — link-sharing semantics.

JSON API

The machine-readable spec is at /openapi.json (OpenAPI 3). Every operation is a POST to https://dataembed.com/api/v1/<operation> with a JSON body, and every error is JSON:

{ "code": "NOT_FOUND", "message": "Search session not found", "defined": false }

Build a webset without an account:

curl -sX POST https://dataembed.com/api/v1/search/createSession \
  -H 'content-type: application/json' \
  -d '{"originalQuery": "Seed-stage climate startups in Europe", "category": "companies"}'
# → { "session": { "id": "<uuid>", ... } }

curl -sX POST https://dataembed.com/api/v1/search/startRun \
  -H 'content-type: application/json' -d '{"sessionId": "<uuid>"}'

curl -sX POST https://dataembed.com/api/v1/search/getSession \
  -H 'content-type: application/json' -d '{"sessionId": "<uuid>"}'
# poll until session.runAlive is false, then export (below)

Anonymous creation is rate limited per IP (429 TOO_MANY_REQUESTS); sign in for higher limits. Browser requests from another origin are refused with 403.

Webset endpoints

Two further endpoints per webset stream or download it outside the JSON API.

Export a webset

GET /api/sessions/{sessionId}/export?format=csv|json

Downloads the full webset as an attachment (webset-<id-prefix>.csv / .json). format defaults to csv.

StatusMeaning
200Export body with content-disposition: attachment
400format was something other than csv or json
403Webset is owned by another user
404No webset with that id
500Export failed

Error responses are JSON: { "error": "..." }.

CSV shape

One row per result. Columns, in order: URL, Title, one column per enrichment (extracted value or empty), one column per criterion (verdict: match, partial, no_match, needs_review, or empty), and Overall (the row's combined verdict).

JSON shape

{
  "id": "…",
  "query": "the original request",
  "category": "…",
  "status": "completed",
  "createdAt": "…",
  "criteria": [{ "label": "…", "description": "…" }],
  "enrichments": [{ "label": "…", "description": "…", "format": "text" }],
  "results": [
    {
      "url": "…",
      "title": "…",
      "status": "completed",
      "matchVerdict": "match",
      "enrichments": { "Name": "…", "Founded": null },
      "criteria": {
        "HQ in Europe": { "verdict": "match", "evidence": "verbatim quote from the page" }
      }
    }
  ]
}

results[].enrichments maps enrichment label → extracted value (or null). results[].criteria maps criterion label → { verdict, evidence } (or null when not yet evaluated).

Stream run events

GET /api/sessions/{sessionId}/stream?startIndex={n}

An authenticated NDJSON proxy of the webset's eve session event stream (content-type: application/x-ndjson) — one JSON event per line, covering the agent's lifecycle, tool calls, and narration text as the run progresses.

Treat events as doorbells, not data: the stream carries no webset rows, and the event shape is eve-internal rather than a stable contract. The app's own client refetches the session over oRPC whenever an event arrives.

  • startIndex is a reconnect cursor. Events are indexed; on reconnect, pass the index after the last event you consumed to resume without replaying.
  • Responses stream for up to 5 minutes per connection; reconnect with startIndex to continue following a long run.
  • Client disconnects propagate upstream, so closing the connection stops the proxying immediately (the run itself continues server-side).
StatusMeaning
200NDJSON stream
404Webset not found, not accessible with your cookie, or no run has been started yet
502The upstream eve stream could not be reached

On this page