Schedules & Runs API
A read-only API for your organization’s schedules (your cron-scheduled spiders) and runs (every job, with its metrics) — the same data the dashboard shows, available to your own systems and to your spiders.
Use it to build dashboards, alerting, or a metrics-export spider that ships run stats to your own warehouse.
The credentials on this page — sp_ API keys and the token handed to each
job — are scoped to your organization and grant read access only.
You see your own schedules and runs, never anyone else’s, and nothing here
can start, stop, or change a spider. If you want an AI agent to read this
data (and, with a separate write-scoped token, act on it), see
Connect your AI agent (MCP).
Authentication
Two ways to authenticate, depending on where you’re calling from.
From your own systems — use a deploy API key. Create one in Deploy →
API Keys (it’s shown once, starts with sp_). Send it as a bearer token:
curl -H "Authorization: Bearer sp_your_key" \
"https://<your-arachne-host>/api/v1/client/schedules"X-API-Key: sp_your_key works too. The key returns all of your
organization’s schedules and runs; narrow it with ?project= when you want
one project.
From inside a spider — no setup. Every job is handed a short-lived, org-scoped, read-only token and the base URL as environment variables:
import os, urllib.request, json
base = os.environ["INSIGHT_METRICS_API_URL"].rstrip("/")
token = os.environ["INSIGHT_METRICS_API_TOKEN"]
req = urllib.request.Request(
f"{base}/api/v1/client/jobs?limit=20",
headers={"Authorization": f"Bearer {token}"},
)
runs = json.load(urllib.request.urlopen(req, timeout=30))["jobs"]See Inside a Job for the full list of injected variables.
Endpoints
All endpoints are GET under /api/v1/client and return JSON with a
"status": "ok" field.
| Endpoint | Returns |
|---|---|
GET /api/v1/client/schedules | Your schedules (schedules[], total). |
GET /api/v1/client/schedules/{id} | One schedule (schedule). |
GET /api/v1/client/schedules/{id}/executions | A schedule’s recent fires (executions[]). |
GET /api/v1/client/jobs | Runs, newest first (jobs[], total). |
GET /api/v1/client/jobs/{id} | One run (job). |
GET /api/v1/client/jobs/{id}/stats | A run’s detailed stats (stats, or null until it finishes). |
GET /api/v1/client/jobs/stats | Run counts by status. |
GET /api/v1/client/metrics/dashboard | Counts + active schedules + online nodes. |
Query parameters
schedules and jobs are paginated and sortable:
| Parameter | Applies to | Notes |
|---|---|---|
limit, offset | both | Page size (default 50, no hard cap) and offset. Responses include total. |
project | both | Exact project name. |
spider | jobs | Substring match on the spider name. |
status | jobs | Exact run status: pending / running / finished / error / cancelled. |
search | both | Substring over the spider name (and, on either endpoint, the schedule name). |
sort, dir | both | Sort column (see below) + asc / desc. |
Sortable columns — schedules: next_run_at, last_run_at, run_count,
name, project, spider, status, cron_expression, timezone,
created_at. jobs: created_at (default, newest first), start_time,
end_time, items_scraped, pages_scraped, error_count, status,
priority, project, spider, schedule_name.
The schedules endpoint has no status filter (only jobs does). To
see just active or just paused schedules, filter client-side on the
status field — see the note under The next scheduled spiders.
The next scheduled spiders
Sort schedules by next_run_at ascending to get what’s coming up:
GET /api/v1/client/schedules?sort=next_run_at&dir=asc&limit=30{
"status": "ok",
"total": 42,
"schedules": [
{
"id": 7,
"name": "Nightly catalogue",
"project": "shop",
"spider": "catalogue",
"cron_expression": "0 3 * * *",
"timezone": "America/Mexico_City",
"status": "active",
"next_run_at": "2026-07-08T09:00:00Z",
"last_run_at": "2026-07-07T09:00:00Z",
"run_count": 128
}
]
}Each schedule also carries args, settings, tags, priority, version,
jitter_seconds, misfire_grace_seconds, coalesce, max_instances,
retry_on_failure, retry_count, start_date, end_date, and pause_reason.
Paused schedules can sort first. A paused schedule keeps its last
next_run_at — a time in the past that doesn’t advance while paused. So
sorting by next_run_at ascending can surface paused schedules ahead of the
ones that will actually fire. Since the endpoint has no status filter, keep
only status == "active" client-side to get “what will run next”:
page = get("/api/v1/client/schedules", sort="next_run_at", dir="asc", limit=200)
upcoming = [s for s in page["schedules"] if s["status"] == "active"][:30]Recent runs & their metrics
Runs come back newest-first by default:
GET /api/v1/client/jobs?limit=20{
"status": "ok",
"total": 5310,
"jobs": [
{
"id": "b3f1a2c4-...",
"project": "shop",
"spider": "catalogue",
"status": "finished",
"items_scraped": 12840,
"pages_scraped": 1310,
"error_count": 2,
"warning_count": 5,
"start_time": "2026-07-07T09:00:03Z",
"end_time": "2026-07-07T09:14:51Z",
"finish_reason": "finished",
"schedule_id": 7,
"schedule_name": "Nightly catalogue"
}
]
}status is one of pending, running, finished, error, cancelled. The
per-run counters (items_scraped, pages_scraped, error_count,
warning_count) update live while a job runs and are final once it’s finished.
schedule_name is present when the run came from a schedule.
Each run returns only these client-relevant fields (id, project, spider,
version, status, priority, args, settings, tags, the counters,
start_time, end_time, finish_reason, error_message, schedule_id,
schedule_name, created_at, updated_at) — internal execution details are
not exposed.
For the deeper breakdown of a single run — response codes, retries, dropped
items, runtime — call its stats:
GET /api/v1/client/jobs/{id}/stats{
"status": "ok",
"stats": {
"runtime_seconds": 888.4,
"items_scraped": 12840,
"items_dropped": 3,
"pages_crawled": 1310,
"response_200": 1290,
"response_404": 12,
"response_429": 4,
"response_500": 0,
"retry_count": 18,
"error_count": 2,
"finish_reason": "finished"
}
}stats is null until a run finishes (the detailed row is written at job
end). While a job is running, read the live counters from the run object
itself (/jobs/{id}).
Build a metrics-export spider
Put the two together to export everything on a schedule of your own: a spider
that reads your schedules + runs and writes them wherever you want. A working
end-to-end example ships in the reference project as metrics_export_demo —
it logs the next 30 scheduled spiders and the last 20 runs with their metrics.
The core of it:
import os, urllib.request, urllib.parse, json
import scrapy
class MetricsExportSpider(scrapy.Spider):
name = "metrics_export"
def __init__(self, *a, **k):
super().__init__(*a, **k)
self.base = os.environ["INSIGHT_METRICS_API_URL"].rstrip("/")
self.token = os.environ["INSIGHT_METRICS_API_TOKEN"]
def get(self, path, **params):
url = self.base + path
if params:
url += "?" + urllib.parse.urlencode(params)
req = urllib.request.Request(url, headers={"Authorization": f"Bearer {self.token}"})
return json.load(urllib.request.urlopen(req, timeout=30))
async def start(self):
upcoming = self.get("/api/v1/client/schedules", sort="next_run_at", dir="asc", limit=30)
for s in upcoming["schedules"]:
self.logger.info("next: %s %s @ %s", s["spider"], s["cron_expression"], s["next_run_at"])
runs = self.get("/api/v1/client/jobs", limit=20)
for r in runs["jobs"]:
self.logger.info("run: %s %s items=%s errors=%s",
r["spider"], r["status"], r["items_scraped"], r["error_count"])
# …ship `r` to your own warehouse / dashboard here.
return
yield # marks start() as an async generatorBecause the token is injected automatically, the same spider works unchanged on
any schedule — point it at your destination in settings.py (see Getting your
data out).