Bright Data
Search, Crawl and Scrape any site, at scale, without getting blocked
0.5.1Bright Data is a web data platform that provides proxy infrastructure and scraping services at scale. This toolkit enables Arcade agents to search the web, scrape any URL, and extract structured data from major platforms without getting blocked.
Capabilities
- Web scraping: Fetch any public webpage and receive clean Markdown-formatted content, suitable for downstream LLM processing.
- Multi-engine search: Query Google, Bing, or Yandex with configurable parameters including result count, country code, and search type (web, images, etc.).
- Structured data extraction: Pull normalized, structured JSON from 20+ supported source types across Amazon, LinkedIn, Instagram, Facebook, X, Zillow, Booking.com, YouTube, and ZoomInfo — no parsing required.
Secrets
This toolkit requires two secrets to authenticate with Bright Data's API.
-
BRIGHTDATA_API_KEY— Your Bright Data API key, used to authenticate all requests. Obtain it from the Bright Data dashboard under Account Settings → API Token. A paid Bright Data account is required; free trials may have limited access. -
BRIGHTDATA_ZONE— The name of a Bright Data zone (proxy zone or dataset zone) that the toolkit routes requests through. Create or find zone names in the Bright Data dashboard under Proxies & Scraping Infrastructure → Zones. The zone type must be appropriate for the operations you intend to use (e.g., a Web Unlocker or Scraping Browser zone for scraping, a dataset zone forWebDataFeed). The zone name is the string identifier shown in the zone list, not a URL or ID number.
For help configuring secrets in Arcade, see the Arcade secrets guide or manage secrets directly at https://api.arcade.dev/dashboard/auth/secrets.
Available tools(3)
| Tool name | Description | Secrets | |
|---|---|---|---|
Scrape a webpage and return content in Markdown format using Bright Data.
Examples:
scrape_as_markdown("https://example.com") -> "# Example Page
Content..."
scrape_as_markdown("https://news.ycombinator.com") -> "# Hacker News
..."
| 2 | ||
Search using Google, Bing, or Yandex with advanced parameters using Bright Data.
Examples:
search_engine("climate change") -> "# Search Results
## Climate Change - Wikipedia
..."
search_engine("Python tutorials", engine="bing", num_results=5) -> "# Bing Results
..."
search_engine("cats", search_type="images", country_code="us") -> "# Image Results
..."
| 2 | ||
Extract structured data from various websites like LinkedIn, Amazon, Instagram, etc.
NEVER MADE UP LINKS - IF LINKS ARE NEEDED, EXECUTE search_engine FIRST.
Supported source types:
- amazon_product, amazon_product_reviews
- linkedin_person_profile, linkedin_company_profile
- zoominfo_company_profile
- instagram_profiles, instagram_posts, instagram_reels, instagram_comments
- facebook_posts, facebook_marketplace_listings, facebook_company_reviews
- x_posts
- zillow_properties_listing
- booking_hotel_listings
- youtube_videos
Examples:
web_data_feed("amazon_product", "https://amazon.com/dp/B08N5WRWNW")
-> "{"title": "Product Name", ...}"
web_data_feed("linkedin_person_profile", "https://linkedin.com/in/johndoe")
-> "{"name": "John Doe", ...}"
web_data_feed(
"facebook_company_reviews", "https://facebook.com/company", num_of_reviews=50
) -> "[{"review": "...", ...}]" | 2 |