io.github.ArkNill/markgrab is an MCP server that universal web content extraction — any URL to LLM-ready markdown. HTML, YouTube, PDF, DOCX. Its tool list has not been published yet over sse, requires no API key, and scores 51/100 on MCPpedia's security, maintenance and efficiency rubric.
Config is the same across clients — only the file and path differ.
{
"mcpServers": {
"io-github-arknill-markgrab-quartzunit": {
"args": [
"markgrab"
],
"command": "uvx"
}
}
}Are you the author?
Add this badge to your README to show your security score and help users find safe servers.
Universal web content extraction — any URL to LLM-ready markdown.
Run this in your terminal to verify the server starts. Then let us know if it worked — your result helps other developers.
uvx 'markgrab' 2>&1 | head -1 && echo "✓ Server started successfully"
After testing, let us know if it worked:
Five weighted categories — click any category to see the underlying evidence.
No known CVEs.
Checked markgrab against OSV.dev.
Be the first to review
Have you used this server?
Share your experience — it helps other developers decide.
Sign in to write a review.
Others in writing / browser
Chrome DevTools for coding agents
Monitor browser logs directly from Cursor and other MCP compatible IDEs.
🔥 Official Firecrawl MCP Server - Adds powerful web scraping and search to Cursor, Claude and any other LLM clients.
MCP server paired with a browser extension that enables AI agents to control the user's browser.
MCP Security Weekly
Get CVE alerts and security updates for io.github.ArkNill/markgrab and similar servers.
Start a conversation
Ask a question, share a tip, or report an issue.
Sign in to join the discussion.
Universal web content extraction — any URL to LLM-ready markdown.
from markgrab import extract
result = await extract("https://example.com/article")
print(result.markdown) # clean markdown
print(result.title) # "Article Title"
print(result.word_count) # 1234
print(result.language) # "en"
pip install markgrab
Optional extras for specific content types:
pip install "markgrab[browser]" # Playwright for JS-rendered pages
pip install "markgrab[youtube]" # YouTube transcript extraction
pip install "markgrab[pdf]" # PDF text extraction
pip install "markgrab[docx]" # DOCX text extraction
pip install "markgrab[all]" # everything
import asyncio
from markgrab import extract
async def main():
# HTML (auto-detects content type)
result = await extract("https://example.com/article")
# YouTube transcript
result = await extract("https://youtube.com/watch?v=dQw4w9WgXcQ")
# PDF
result = await extract("https://arxiv.org/pdf/1706.03762")
# Options
result = await extract(
"https://example.com",
max_chars=30_000, # limit output length (default: 50K)
use_browser=True, # force Playwright rendering
stealth=True, # anti-bot stealth scripts (opt-in)
timeout=60.0, # request timeout in seconds
proxy="http://proxy:8080",
)
asyncio.run(main())
markgrab https://example.com # markdown output
markgrab https://example.com -f text # plain text
markgrab https://example.com -f json # structured JSON
markgrab https://example.com --browser # force browser rendering
markgrab https://example.com --max-chars 10000 # limit output
result.title # page title
result.text # plain text
result.markdown # LLM-ready markdown
result.word_count # word count
result.language # detected language ("en", "ko", ...)
result.content_type # "article", "video", "pdf", "docx"
result.source_url # final URL (after redirects)
result.metadata # extra metadata (video_id, page_count, etc.)
flowchart TD
A["🔗 URL Input"] --> B{"Content\nType?"}
B -->|"HTML"| C["HTTP fetch\n(httpx)"]
C --> D{"JS\nrequired?"}
D -->|"no"| E["HTML Parser\n→ clean markdown"]
D -->|"yes"| F["Playwright\nfallback"]
F --> E
B -->|"YouTube"| G["Transcript API\n→ timestamped markdown"]
B -->|"PDF"| H["PDF Parser\n→ structured markdown"]
B -->|"DOCX"| I["DOCX Parser\n→ markdown"]
E --> J["✅ LLM-ready\nMarkdown"]
G --> J
H --> J
I --> J
For HTML pages, if the initial httpx fetch yields fewer than 50 words, MarkGrab automatically retries with Playwright to handle JavaScript-rendered content.
This software is provided for legitimate purposes only. By using MarkGrab, you agree to the following:
robots.txt. Users are solely responsible for check