Legal · 4 September 2026

Data source policy

Effective 4 September 2026. How Scrapely is allowed to get records, and what this deployment actually does.

This deployment

Scrapely searches 64 public platforms you check — social, video, forums, news, reviews, and developer sites. Open sources (Reddit RSS, Hacker News, Wikipedia, GitHub, GitLab, Stack Overflow, Mastodon tags, and others) are fetched natively. YouTube Data API and X (lobstr.io) use keys each signed-in user pastes in Settings — never another account’s keys. Bluesky (public actor/feed AppView), Telegram public channels, Kick channel VODs, and Twitch Helix (operator env credentials) have native adapters where open. Locked networks (LinkedIn, Meta, TikTok, Rumble HTML, Truth Social) stay on the public web/news index — still not a paid firehose. Every result keeps a source URL. Empty means no public hit, not a complete firehose.

Each stored record keeps platform, source id, source URL, provider, and retrieval time so it cannot honestly be passed off as something we authored.

Permitted sources (product rules)

Official APIs and public feeds with a valid developer agreement or public access, licensed firehoses, user-supplied exports the user has a right to upload, and public news indexes.

Forbidden sources

CAPTCHA bypass, private-account access, session hijacking, malware, buying stolen cookies, or presenting scraped private content as public.

The architecture is not designed to circumvent platform rules. If a provider forbids a use, we do not ship a workaround as a feature.

One API

There is one public HTTP API for every platform: GET /api/v1/search. Filter with PLATFORM: or &platform=. There is no per-network API.

Platform filters change which records match. They are not separate products, keys, or billing meters.

These pages are the live contract for this instance. Have qualified counsel review them before you take real payments or process real personal data at scale. No document can waive fraud or rights a statute says you cannot sign away.