Discourse Forum Scraper
Extract topics and full post text from any Discourse forum. Clean plain-text output for LLM training, RAG corpora and community research.
14%
users
Apify builder
Neil Sangwaiya I build scrapers and data pipelines that return clean, structured output.
Leaderboard position
out of 6,961
New actors
#342
Total actors
#1,284
↓ -80
Monthly users
#1,732
↓ -96
Total users
#3,005
↓ -109
Monthly runs
#2,495
↑ +2,551
Total runs
#4,922
↑ +712
Portfolio stats
New actors
7
Total actors
7
+0
Monthly users
7
+0
Total users
14
+0
Monthly runs
89
↑ +64
Total runs
106
↑ +69
New actors
Each bar shows one day in the last 30 days.
Portfolio concentration
6 of 7 actors cover 80%+ of monthly users.
Extract topics and full post text from any Discourse forum. Clean plain-text output for LLM training, RAG corpora and community research.
14%
users
Scrape any documentation site to clean markdown. Works on Docusaurus, Mintlify, GitBook, MkDocs, ReadTheDocs and more. Preserves code blocks for RAG and LLM training.
14%
users
Turn any public GitHub repository into a clean dataset of code and documentation files. One download, no API token, no rate limits. Built for AI coding assistants and RAG.
14%
users
Extract every table from any web page into clean rows, JSON and markdown. Correctly handles colspan, rowspan and stacked headers that break other extractors.
14%
users
Convert PDFs to clean markdown with real reading order. Handles two-column layouts, detects headings and paragraphs. Built for RAG pipelines and LLM ingestion.
14%
users
Scrape posts, pages and content from any WordPress site via the REST API. Clean plain text with authors, categories and tags, ready for LLM training and RAG.
14%
users
ActorHunt MCP
ActorHunt gives paid members MCP access to keyword data, actor evidence, and Store movement for deciding what to build next.