Key/Value Cache
Every organization with the Key/Value Cache addon gets a private,
Redis-compatible cache. Connect with any Redis client —
redis-py , ioredis, go-redis or redis-cli — to
memoise lookups, deduplicate work and share state between runs and parallel
jobs. Your spiders get the connection URL with zero configuration.
It’s a cache, not a database: every key expires within 24 hours, and a
key can be evicted earlier. Write your spider so that a miss simply means “do
the work”. For data you need to keep, use Storage or the
postgres add-on.
Why use it
A long-running spider walks categories and stores, and for every product it needs one extra request just to get the EAN. The same product shows up in many categories and stores, so that request repeats — and Scrapy’s duplicate filter drops the repeat, taking the item that was waiting for its EAN with it. The usual workaround is a giant dict in memory: it grows for the whole run, dies with the job, and isn’t shared with the jobs crawling in parallel.
With the cache, you look the EAN up first, fetch it only on a miss, and store it for the next product, the next run, and every parallel job. Typical uses:
- Memoise enrichment lookups — EANs, geocoding, one field from a detail page.
- Deduplicate work across runs and parallel jobs — “has anyone handled this URL today?”
- Serve repeated downloads from a shared HTTP cache.
- Counters and small shared state —
INCR, hashes, sets.
Enabling it
- Enable Key/Value Cache under Billing → Addons in the dashboard.
- In Crawlers → Projects, turn on the
cacheadd-on for each project that should use it. - Add a Redis client to your image — for Python,
redis>=5:
RUN pip install --no-cache-dir "scrapy>=2.12,<3" "redis>=5"Every new job of that project then finds the connection in its environment.
Connecting
| Variable | What it is |
|---|---|
INSIGHT_CACHE_URL | redis://username:password@host:6379/0 — your organization’s cache |
REDIS_URL | The same URL, under the name many libraries look for |
If your project has a secret named REDIS_URL, your secret wins;
INSIGHT_CACHE_URL always points at the cache.
import os
import redis
r = redis.Redis.from_url(os.environ["INSIGHT_CACHE_URL"], decode_responses=True)
r.set("hello", "world", ex=3600)
print(r.get("hello")) # world- Any client, default settings. RESP2 and RESP3 are negotiated
automatically, so redis-py, ioredis, go-redis, node-redis and
redis-cliall work out of the box. - Database
0only.SELECT 0works; any other number is rejected. - No TLS. The cache is reachable only from inside the platform (your jobs)
and over the VPN for local development, so the URL is
redis://, notrediss://. - One client per process. A client pools its connections; don’t create one per request. Your organization can hold 64 connections at once, across all jobs and machines.
Example: memoise a lookup
The EAN case from above. Look in the cache first; on a miss, fetch the EAN and store it for everyone else:
import os
import redis
import scrapy
class ProductsSpider(scrapy.Spider):
name = "products"
def __init__(self, *args, **kwargs):
super().__init__(*args, **kwargs)
# One client per spider, shared by every callback.
self.cache = redis.Redis.from_url(
os.environ["INSIGHT_CACHE_URL"], decode_responses=True
)
def parse_product(self, response):
item = {
"sku": response.css("[data-sku]::attr(data-sku)").get(),
"name": response.css("h1::text").get(),
}
ean = self.cached_ean(item["sku"])
if ean:
item["ean"] = ean
yield item
else:
yield scrapy.Request(
f"https://api.example.com/products/{item['sku']}/ean",
callback=self.parse_ean,
cb_kwargs={"item": item},
dont_filter=True, # never let the duplicate filter drop this item
)
def parse_ean(self, response, item):
item["ean"] = response.json()["ean"]
try:
self.cache.set(f"ean:{item['sku']}", item["ean"], ex=86400)
except redis.RedisError:
pass # budget exhausted or cache unreachable: the item is still complete
yield item
def cached_ean(self, sku):
try:
return self.cache.get(f"ean:{sku}")
except redis.RedisError:
return None # treat any cache problem as a miss
def closed(self, reason):
self.cache.close()Inside the platform each cache call is a sub-millisecond network round
trip, so the regular (blocking) client is fine in ordinary callbacks.
dont_filter=True matters: the same SKU can be looked up twice before the first
answer lands in the cache, and without it the duplicate filter would drop the
second request — and its item.
With async def callbacks
If your project runs the asyncio reactor
(TWISTED_REACTOR = "twisted.internet.asyncioreactor.AsyncioSelectorReactor",
the default since Scrapy 2.13), use the asyncio client and await it:
import os
import redis.asyncio as aioredis
import scrapy
class ProductsSpider(scrapy.Spider):
name = "products"
def __init__(self, *args, **kwargs):
super().__init__(*args, **kwargs)
self.cache = aioredis.Redis.from_url(
os.environ["INSIGHT_CACHE_URL"], decode_responses=True
)
async def parse_product(self, response):
item = {"sku": response.css("[data-sku]::attr(data-sku)").get()}
ean = await self.cache.get(f"ean:{item['sku']}") # wrap in try/except as above
if ean:
item["ean"] = ean
yield item
else:
yield scrapy.Request(
f"https://api.example.com/products/{item['sku']}/ean",
callback=self.parse_ean,
cb_kwargs={"item": item},
dont_filter=True,
)
async def parse_ean(self, response, item):
item["ean"] = response.json()["ean"]
await self.cache.set(f"ean:{item['sku']}", item["ean"], ex=86400)
yield itemExample: deduplicate across runs and jobs
SET key 1 NX EX 86400 writes the key only if it doesn’t exist yet — and tells
you whether you were first. It’s a single atomic command, so two jobs racing for
the same URL can’t both win:
def parse_listing(self, response):
for href in response.css("a.product::attr(href)").getall():
url = response.urljoin(href)
# True only for the first job (or run) to claim this URL in the last 24h.
if self.cache.set(f"claimed:{url}", 1, nx=True, ex=86400):
yield scrapy.Request(url, callback=self.parse_product)A claim means “someone took it”, not “it’s done”: if a job dies right after claiming, that URL is skipped until the key expires. To record completion instead, set the key after the item has been saved.
Locks
redis-py’s Lock helper (r.lock(...)) doesn’t work here. It acquires with
SET … NX PX, but it releases with a Lua script (EVALSHA), which the cache
refuses with ERR command 'evalsha' is not available on the InsightScrap cache —
and the lock key then stays put until its timeout. Use the same claim pattern
instead: SET … NX EX to take the lock, DEL to release it.
import uuid
token = uuid.uuid4().hex # who holds the lock, handy when debugging
if r.set("lock:daily-export", token, nx=True, ex=600): # held for 10 minutes at most
try:
run_export()
finally:
r.delete("lock:daily-export")Give the lock an expiry longer than the work takes. If it expires first, another
job can take the lock, and your delete would release theirs.
Example: a shared HTTP cache for Scrapy
Scrapy’s HTTP cache can keep responses anywhere. Drop this storage backend into your project and repeated requests are answered from the cache instead of the network — within the run, in later runs, and in parallel jobs.
# myproject/cache_storage.py
import json
import logging
import os
import redis
from scrapy.exceptions import NotConfigured
from scrapy.http import Headers
from scrapy.responsetypes import responsetypes
logger = logging.getLogger(__name__)
MAX_TTL = 86400 # the cache never keeps a key longer than 24 hours
MAX_BODY = 900 * 1024 # values are capped at 1 MiB; leave room for the rest
class InsightCacheStorage:
"""Scrapy HTTP cache storage backed by the InsightScrap Key/Value Cache."""
def __init__(self, settings):
self.url = os.environ.get("INSIGHT_CACHE_URL") or os.environ.get("REDIS_URL")
if not self.url:
raise NotConfigured("INSIGHT_CACHE_URL is not set")
expiration = settings.getint("HTTPCACHE_EXPIRATION_SECS")
self.ttl = min(expiration, MAX_TTL) if expiration > 0 else MAX_TTL
def open_spider(self, spider):
# decode_responses=False: bodies are stored and returned as raw bytes.
self.r = redis.Redis.from_url(
self.url, decode_responses=False, socket_timeout=2, socket_connect_timeout=2
)
self.fingerprinter = spider.crawler.request_fingerprinter
def close_spider(self, spider):
self.r.close()
def _key(self, spider, request):
return f"httpcache:{spider.name}:{self.fingerprinter.fingerprint(request).hex()}"
def retrieve_response(self, spider, request):
try:
data = self.r.hgetall(self._key(spider, request))
if not data:
return None
url = data[b"url"].decode()
status = int(data[b"status"])
headers = Headers({
k.encode("latin-1"): [v.encode("latin-1") for v in values]
for k, values in json.loads(data[b"headers"]).items()
})
body = data[b"body"]
except (redis.RedisError, KeyError, ValueError) as e:
logger.debug("httpcache read failed, treating as a miss: %s", e)
return None
respcls = responsetypes.from_args(headers=headers, url=url, body=body)
return respcls(url=url, headers=headers, status=status, body=body)
def store_response(self, spider, request, response):
if len(response.body) > MAX_BODY:
return
# latin-1 round-trips any header byte; repeated headers (Set-Cookie) are kept.
headers = {
k.decode("latin-1"): [v.decode("latin-1") for v in values]
for k, values in response.headers.items()
}
key = self._key(spider, request)
try:
pipe = self.r.pipeline() # MULTI ... EXEC
pipe.delete(key) # a re-store starts a fresh lifetime
pipe.hset(key, mapping={
"status": response.status,
"url": response.url,
"headers": json.dumps(headers),
"body": response.body,
})
pipe.expire(key, self.ttl)
pipe.execute()
except redis.RedisError as e:
logger.debug("httpcache write failed, skipping: %s", e)Turn it on in settings.py:
# settings.py
HTTPCACHE_ENABLED = True
HTTPCACHE_STORAGE = "myproject.cache_storage.InsightCacheStorage"
HTTPCACHE_EXPIRATION_SECS = 86400 # anything longer is capped at 24h anywayThen yield repeated requests with dont_filter=True. The duplicate filter runs
before the HTTP cache, so a filtered request never gets the chance to be
answered from it:
yield scrapy.Request(ean_url, callback=self.parse_ean, cb_kwargs={"item": item}, dont_filter=True)Responses served from the cache carry "cached" in response.flags.
- Every downloaded response is cached and counts toward your
write budget. Keep big or one-off pages out with
meta={"dont_cache": True}, and don’t cache blocks:HTTPCACHE_IGNORE_HTTP_CODES = [403, 429, 500, 502, 503, 504]. - Bodies over ~900 KB are skipped (a single value can be at most 1 MiB).
- If the cache is unreachable or your budget is used up, every lookup is a miss and the spider downloads as usual — the crawl never fails because of the cache.
- Without
INSIGHT_CACHE_URL(a local run with no.env), the backend switches itself off and Scrapy runs without an HTTP cache.
From your laptop
Reveal your credentials in the dashboard under Proxy Users → Cache
(Redis-compatible). The card gives you the URL, host, port, username and
password, a ready-to-paste .env block (INSIGHT_CACHE_URL=… and
REDIS_URL=…), and a redis-cli command. Your machine must be connected to
the VPN (WARP), as for the other local-development credentials.
export INSIGHT_CACHE_URL='redis://username:password@host:6379/0' # from the dashboard
redis-cli -u "$INSIGHT_CACHE_URL" PING # PONG
redis-cli -u "$INSIGHT_CACHE_URL" SET greeting hello EX 300
redis-cli -u "$INSIGHT_CACHE_URL" GET greeting # "hello"
redis-cli -u "$INSIGHT_CACHE_URL" TTL greeting # (integer) 300
redis-cli -u "$INSIGHT_CACHE_URL" --scan --pattern 'ean:*' | headUse redis-cli 6 or newer (older versions don’t send the username). Your laptop
sees the same keyspace as your production jobs, so prefix the keys you write
while experimenting (for example dev:).
Keys expire within 24 hours
The rule is simple: data lives at most 24 hours after it was written.
| You do | What happens |
|---|---|
| Write a key without a TTL | It gets 24 hours. |
Write with a TTL over 24 hours (SET … EX/PX/EXAT/PXAT, SETEX, PSETEX) | Silently clamped to 24 hours. |
Create a key with HSET, SADD, RPUSH, INCR, APPEND, MSET, SETNX, GETSET or SET … KEEPTTL | It gets 24 hours if it has no TTL; an existing TTL is left alone. |
EXPIRE, PEXPIRE, EXPIREAT, PEXPIREAT | Can only shorten the TTL (like the LT flag); anything over 24 hours is clamped first. A request to lengthen returns 0 and changes nothing. |
EXPIRE … GT or PERSIST | Rejected. |
> SET a 1
OK
> TTL a
(integer) 86400
> SET b 1 EX 604800 # asked for 7 days
OK
> TTL b
(integer) 86400 # clamped to 24h
> EXPIRE b 3600
(integer) 1 # shortened
> EXPIRE b 7200
(integer) 0 # can't lengthenWhat that means in practice:
- Overwriting a key with
SETstarts a fresh lifetime (up to 24 hours). - Adding to a hash, set or list does not extend it. The key still expires 24
hours after it was created. For rolling data, overwrite it or use per-day keys
(
seen:2026-10-05).
Write budget
Each organization can write up to 1 GiB per rolling 24 hours — counted as the bytes of keys, fields and values you send in write commands, plus a small fixed overhead for what storing them really costs: about 100 bytes per key and 16 bytes per field, member or value. Reads and deletes don’t count against it.
In practice that overhead only matters for very small entries: a million
SET product:123 <13-digit EAN> writes cost about 140 MB of budget, not the 24 MB of raw key and value bytes.
- Over budget, write commands fail with an error starting with
OOM(OOM cache write budget exceeded for this organization …). In redis-py that’sredis.exceptions.OutOfMemoryError, a kind ofResponseError. - Reads and deletes keep working, and the budget frees up as older hours roll out of the 24-hour window.
- Deleting doesn’t give budget back. The budget counts bytes written in
the window, not bytes stored, so
DEL,FLUSHDBor overwriting a key frees none of it — only time does. - Because nothing lives longer than 24 hours, the budget is also the most data your organization can hold at once.
- Need more? Contact us and we’ll raise it for your organization.
Treat an OOM like any other cache failure: skip the write and carry on. The
addon is billed monthly, with usage metered by data written — see Billing →
Addons for current pricing.
Limits
| Limit | Value |
|---|---|
| Key lifetime | 24 hours at most |
| Write budget | 1 GiB per rolling 24 hours per organization (can be raised) |
| Largest key, field or value | 1 MiB |
| Largest command | 8 MiB |
| Largest reply | 32 MiB (read bigger collections with HSCAN, SSCAN, LRANGE or GETRANGE) |
| Simultaneous connections | 64 per organization |
| Databases | 0 only |
| Protocol | RESP2 and RESP3 (negotiated automatically) |
| Encryption in transit | None — reachable only inside the platform and over the VPN |
Errors
| Error | What it means | What to do |
|---|---|---|
OOM cache write budget exceeded … | You’ve written 1 GiB in the last 24 hours. | Skip the write. Reads and deletes still work (deleting doesn’t free budget); writes resume as the window rolls. |
ERR value too large … | A key, field or value over 1 MiB, or a command over 8 MiB. | Store less (split or compress) or skip it. The connection stays usable (only a command of tens of MiB closes it). |
ERR max number of clients reached | Your organization already has 64 open connections. | Share one client per process; cap the pool (max_connections). |
ERR command '…' is not available on the InsightScrap cache | The command isn’t supported. With evalsha, it’s usually redis-py’s r.lock() releasing. | See supported commands; for locks, see Locks. |
ERR DB index is out of range | SELECT with a number other than 0. | Use database 0. |
ERR GT is not supported … | EXPIRE … GT — TTLs can only shrink. | Drop GT; set a new lifetime by overwriting the key. |
WRONGPASS … / NOAUTH … | Missing, wrong or rotated credentials. | Connect with the full INSIGHT_CACHE_URL; reveal it again for local use. |
ERR cache access was revoked for this credential | The credential has been revoked. | Check the addon is still enabled, then contact us. |
ERR reply too large … | The answer would be over 32 MiB — usually HGETALL, SMEMBERS or LRANGE 0 -1 on a very large collection. The connection stays usable. | Read it in pieces with HSCAN, SSCAN or LRANGE ranges. |
ERR cache busy, retry / ERR Protocol error: server busy, retry | Too many very large requests or replies at the same moment across the platform. The second form closes the connection. | Retry with a short backoff, or treat it as a miss. |
ERR cache temporarily unavailable | A brief interruption on our side; the connection is closed. | Retry, or treat it as a miss. |
ERR … inside a transaction / … inside MULTI | A connection command (AUTH, HELLO, SELECT, CLIENT, COMMAND, INFO) or FLUSHDB/FLUSHALL inside MULTI (PING and ECHO are fine). | Run it outside the transaction. |
Supported commands
| Group | Commands |
|---|---|
| Connection & server | AUTH, HELLO, PING, ECHO, QUIT, SELECT 0, CLIENT SETNAME / GETNAME / SETINFO / ID, COMMAND (returns an empty list), INFO (minimal) |
| Strings & counters | GET, SET (EX / PX / EXAT / PXAT / KEEPTTL / NX / XX / GET), SETNX, SETEX, PSETEX, MGET, MSET, MSETNX, GETDEL, GETSET, GETRANGE, STRLEN, APPEND, INCR, INCRBY, INCRBYFLOAT, DECR, DECRBY |
| Keys | DEL, UNLINK, EXISTS, TYPE, TTL, PTTL, EXPIRETIME, PEXPIRETIME, EXPIRE, PEXPIRE, EXPIREAT, PEXPIREAT, TOUCH, SCAN (MATCH / COUNT / TYPE), FLUSHDB / FLUSHALL (delete your organization’s keys only) |
| Hashes | HGET, HSET, HSETNX, HMSET, HMGET, HDEL, HEXISTS, HLEN, HSTRLEN, HKEYS, HVALS, HGETALL, HINCRBY, HINCRBYFLOAT, HSCAN |
| Sets | SADD, SREM, SISMEMBER, SMISMEMBER, SCARD, SMEMBERS, SSCAN |
| Lists | LPUSH, RPUSH, LPOP, RPOP, LLEN, LRANGE, LINDEX, LTRIM |
| Transactions & pipelines | MULTI, EXEC, DISCARD, WATCH, UNWATCH — redis-py’s pipeline() works with its default transaction=True |
Not available: Lua scripting and functions (so redis-py’s r.lock() can’t
release — see Locks), pub/sub, blocking commands
(BLPOP and friends), KEYS (use SCAN), DBSIZE, sorted sets, streams,
admin commands (CONFIG, DEBUG, MONITOR, …), PERSIST, RENAME,
OBJECT, DUMP / RESTORE. They fail with
ERR command '…' is not available on the InsightScrap cache.
Listing keys: use SCAN
KEYS isn’t available — use SCAN. The cache walks a shared keyspace and
filters it down to your keys, so a page can come back empty while the cursor
is still non-zero. Always loop until the cursor returns 0 (redis-py’s
scan_iter does that for you) and ask for big pages with COUNT 1000:
for key in r.scan_iter(match="ean:*", count=1000):
print(key)To wipe everything your organization has stored, use FLUSHDB — it only
touches your keys.
FAQ
Can a key disappear before its TTL? Yes. It’s a cache: under memory pressure on the platform, keys may be evicted early. Always treat a miss as normal — fetch the data and write it back.
Is it consistent? A write is visible to every connection in your
organization as soon as the command returns. Single commands are atomic (INCR,
SET … NX, HSET, …), and MULTI/EXEC with WATCH gives you optimistic
transactions.
Does data survive platform maintenance? Treat it as best-effort. The cache isn’t persisted, so a platform restart can empty it. Never keep the only copy of anything here.
Can I keep a key longer than 24 hours? No — PERSIST and longer TTLs aren’t
available. Use Storage or the postgres add-on for durable data.
Can other organizations see my keys? No. Keys are private to your
organization: you only ever see, SCAN and flush your own.
Can I use it as a job queue, a Celery/RQ broker or for pub/sub? No. Blocking commands, pub/sub, Lua scripts and sorted sets aren’t available, so queue and broker libraries (and distributed schedulers built on them) won’t work. Lists are fine for simple non-blocking push/pop.
Do my local runs share the cache with production jobs? Yes — same organization, same keys. Prefix experimental keys so they don’t mix with real data.