# πŸ•ΈοΈ SCRAPING FABRIC - KompletnΓ­ Dokumentace **Enterprise-grade Two-Tier Scraping Architecture** Verze: 1.0.0 Datum: 2026-01-29 Status: βœ… PRODUCTION READY --- ## πŸ“‹ Obsah 1. [Architektura](#architektura) 2. [Komponenty](#komponenty) 3. [Scraping ReΕΎimy](#scraping-reΕΎimy) 4. [Workflow](#workflow) 5. [API Reference](#api-reference) 6. [CLI Usage](#cli-usage) 7. [Deployment](#deployment) 8. [PΕ™Γ­klady](#pΕ™Γ­klady) 9. [Troubleshooting](#troubleshooting) --- ## πŸ—οΈ Architektura Scraping Fabric pouΕΎΓ­vΓ‘ **dvouvrstvou architekturu**: ``` β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”‚ CLIENT (User/App) β”‚ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β”‚ β–Ό β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”‚ LAYER A: CONTROL PLANE (Router Tower) β”‚ β”‚ 46.224.121.179:8090 β”‚ β”‚ β”‚ β”‚ β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”‚ β”‚ β”‚ β€’ Job Creation & Orchestration β”‚ β”‚ β”‚ β”‚ β€’ API Key Authentication β”‚ β”‚ β”‚ β”‚ β€’ Rate Limiting (60 req/min) β”‚ β”‚ β”‚ β”‚ β€’ Job Status Tracking β”‚ β”‚ β”‚ β”‚ β€’ NO direct scraping (pure orchestrator) β”‚ β”‚ β”‚ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β”‚ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β”‚ β–Ό (HTTP POST) β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”‚ LAYER B: DATA PLANE (Scraper Service) β”‚ β”‚ 46.224.232.134:5100 β”‚ β”‚ β”‚ β”‚ β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”‚ β”‚ β”‚ β€’ Actual Scraping Execution β”‚ β”‚ β”‚ β”‚ β€’ 3 Scraping Modes: FAST β†’ SMART β†’ HEAVY β”‚ β”‚ β”‚ β”‚ β€’ Automatic Fallback Chains β”‚ β”‚ β”‚ β”‚ β€’ Data Normalization & Deduplication β”‚ β”‚ β”‚ β”‚ β€’ Artifact Generation (XLSX, JSONL) β”‚ β”‚ β”‚ β”‚ β€’ SQLite Storage (jobs, results, artifacts) β”‚ β”‚ β”‚ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β”‚ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β”‚ β–Ό β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”‚ ARTIFACTS β”‚ β”‚ β€’ JSONL β”‚ β”‚ β€’ XLSX β”‚ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ ``` ### Principy nΓ‘vrhu 1. **Separation of Concerns**: Orchestrace β‰  Exekuce 2. **Scalability**: Data plane lze horizontΓ‘lnΔ› Ε‘kΓ‘lovat 3. **Fault Tolerance**: Fallback chains, retry logic 4. **Idempotency**: StejnΓ½ job_id lze spustit vΓ­cekrΓ‘t 5. **Observability**: KompletnΓ­ logging, metrics, artifacts --- ## πŸ”§ Komponenty ### 1. Control Plane (Router Tower) **Server**: 46.224.121.179 **Port**: 8090 **Role**: Pure Orchestrator **Soubory**: ``` /opt/router-api/ β”œβ”€β”€ server.py # Flask app (3,002 Ε™Γ‘dkΕ―) β”œβ”€β”€ scraping_fabric_config.py # Konfigurace (2,879 B) └── scraping_fabric_control_plane_short.py # Blueprint (2,570 B) ``` **Funkce**: - βœ… PΕ™ijΓ­mΓ‘ scraping requests - βœ… VytvΓ‘Ε™Γ­ job_id (UUID) - βœ… Dispatčuje job na data plane - βœ… Trackuje job status (pending β†’ running β†’ completed/failed) - ❌ **NIKDY** neprovΓ‘dΓ­ pΕ™Γ­mΓ½ scraping (ΕΎΓ‘dnΓ© OpenAI dependencies v kritickΓ© cestΔ›) **PM2 Process**: `router-api` (ID: 86) --- ### 2. Data Plane (Scraper Service) **Server**: 46.224.232.134 **Port**: 5100 **Role**: Scraping Executor **Soubory**: ``` /opt/reality-scraper-v2/ β”œβ”€β”€ src/ β”‚ β”œβ”€β”€ scraping_fabric_data_plane.py # Flask API (13 KB) β”‚ β”œβ”€β”€ scraping_fabric_modes.py # Scraping modes (6.3 KB) β”‚ └── scraping_fabric_config.py # Shared config β”œβ”€β”€ data/ β”‚ └── scraping_fabric_jobs.db # SQLite (jobs, results, artifacts) β”œβ”€β”€ exports/ β”‚ β”œβ”€β”€ job_*.jsonl # JSONL artifacts β”‚ └── job_*.xlsx # Excel artifacts └── venv/ # Python virtual env ``` **Funkce**: - βœ… ProvΓ‘dΓ­ skutečnΓ½ scraping (3 reΕΎimy) - βœ… Fallback chains (FAST β†’ SMART β†’ HEAVY) - βœ… Normalizace HTML β†’ Markdown - βœ… Deduplikace vΓ½sledkΕ― - βœ… GenerovΓ‘nΓ­ artefaktΕ― (XLSX, JSONL) - βœ… Persistence do SQLite - βœ… Callback na control plane **Process**: `nohup python scraping_fabric_data_plane.py` (PID: 80862) --- ### 3. CLI Tool (`czechai-scrape`) **Lokace**: `/usr/local/bin/czechai-scrape` **Jazyk**: Python 3 **Features**: BarevnΓ½ vΓ½stup, status tracking, batch mode **PΕ™Γ­kazy**: ```bash czechai-scrape modes # Seznam reΕΎimΕ― czechai-scrape scrape # Scrape jednoho URL czechai-scrape batch # Batch scraping czechai-scrape status # Kontrola statusu ``` --- ## ⚑ Scraping ReΕΎimy Scraping Fabric nabΓ­zΓ­ **3 reΕΎimy** s automatickΓ½mi fallbacky: ### 1. FAST Mode πŸš€ **Technologie**: Pure HTTP (httpx) **Timeout**: 10 sekund **Retry**: 2Γ— **Fallback**: SMART **PouΕΎitΓ­**: - StatickΓ½ HTML - VeΕ™ejnΓ© API endpointy - JednoduchΓ© strΓ‘nky bez JS **Performance**: - ⚑ 0.14s - 0.22s prΕ―mΔ›r - πŸ’° NejlevnΔ›jΕ‘Γ­ (ΕΎΓ‘dnΓ© browser overhead) - βœ… IdeΓ‘lnΓ­ pro bulk scraping **OmezenΓ­**: - ❌ Ε½Γ‘dnΓ‘ JavaScript podpora - ❌ Nefunguje na SPA (Single Page Apps) - ❌ NeΕ™eΕ‘Γ­ CAPTCHA --- ### 2. SMART Mode 🧠 **Technologie**: Crawl4AI (58.4k ⭐ GitHub) **Timeout**: 30 sekund **Retry**: 2Γ— **Fallback**: HEAVY **PouΕΎitΓ­**: - DynamickΓ½ obsah (JS rendered) - React/Vue/Angular aplikace - AJAX loaded content - LLM-ready Markdown output **Performance**: - ⚑ 0.34s - 0.83s prΕ―mΔ›r - πŸ’° StΕ™ednΓ­ nΓ‘klady - βœ… DobrΓ½ pomΔ›r rychlost/schopnosti **VΓ½hody**: - βœ… JavaScript rendering - βœ… ČistΓ½ Markdown pro LLM - βœ… Extrakce links, metadata - βœ… LepΕ‘Γ­ neΕΎ Selenium **OmezenΓ­**: - ❌ Nefunguje na tΔ›ΕΎkΓ½ anti-bot - ❌ NeΕ™eΕ‘Γ­ sloΕΎitΓ© CAPTCHA --- ### 3. HEAVY Mode 🦾 **Technologie**: Playwright (full Chromium browser) **Timeout**: 60 sekund **Retry**: 1Γ— **Fallback**: None (poslednΓ­ moΕΎnost) **PouΕΎitΓ­**: - Anti-bot protected sites (Cloudflare, etc.) - SloΕΎitΓ© interakce (clicks, scrolls) - Sites vyΕΎadujΓ­cΓ­ cookies/sessions - StrΓ‘nky s CAPTCHA (manual solve) **Performance**: - ⏱️ PomalΓ© (seconds to minutes) - πŸ’° DrahΓ© (browser overhead) - βœ… NejspolehlivΔ›jΕ‘Γ­ **VΓ½hody**: - βœ… PlnΓ½ browser (real user behavior) - βœ… Cookies, localStorage, sessionStorage - βœ… Screenshot capability - βœ… Network request interception **OmezenΓ­**: - ❌ PomalΓ© (60s max timeout) - ❌ VysokΓ© resource usage (RAM, CPU) - ❌ Nelze masivnΔ› Ε‘kΓ‘lovat --- ## πŸ“Š Fallback Chains Scraping Fabric automaticky eskaluje pΕ™i selhΓ‘nΓ­: ``` FAST Request β”‚ β”œβ”€ Success? β†’ Return result βœ… β”‚ └─ Failed? β†’ Try SMART β”‚ β”œβ”€ Success? β†’ Return result βœ… β”‚ └─ Failed? β†’ Try HEAVY β”‚ β”œβ”€ Success? β†’ Return result βœ… β”‚ └─ Failed? β†’ Return error ❌ (ALL_FAILED) ``` **PΕ™Γ­klad**: ```python # Request na FAST mode Job: ["https://spa-site.com"], mode="FAST" # Flow: 1. FAST: ❌ Failed (prΓ‘zdnΓ½ HTML, JS required) 2. SMART: βœ… Success (Crawl4AI renderuje JS) β†’ Return markdown 3. HEAVY: (not tried - uΕΎ mΓ‘me vΓ½sledek) # Result: { "url": "https://spa-site.com", "success": true, "mode_used": "SMART", # ← Fallback chain worked! "content": "...", "duration": 0.78 } ``` **Konfigurace** (disable fallback): ```python orchestrator.scrape_batch( urls=["https://example.com"], mode="FAST", enable_fallback=False # ← Disable fallback ) ``` --- ## πŸ”„ Workflow - Complete Flow ### 1. Request Flow (User β†’ Control Plane β†’ Data Plane) ``` β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”‚ CLIENT β”‚ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β”‚ β”‚ POST /v1/fabric/scrape β”‚ Headers: X-Scraping-API-Key: czechai-internal-scraping β”‚ Body: {"urls": ["https://example.com"], "mode": "FAST"} β”‚ β–Ό β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”‚ CONTROL PLANE (Router Tower) β”‚ β”‚ 46.224.121.179:8090 β”‚ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β”‚ β”‚ 1. Validate API key β”‚ 2. Validate mode (FAST/SMART/HEAVY) β”‚ 3. Generate job_id = UUID β”‚ 4. Store job (in-memory): {"job_id": "abc-123", "status": "pending"} β”‚ 5. Dispatch to data plane β”‚ β”‚ POST http://46.224.232.134:5100/api/scrape β”‚ Body: {"job_id": "abc-123", "urls": [...], "mode": "FAST"} β”‚ β–Ό β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”‚ DATA PLANE (Scraper Service) β”‚ β”‚ 46.224.232.134:5100 β”‚ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β”‚ β”‚ 6. Check idempotency (job_id already done?) β”‚ 7. Update SQLite: status = "running" β”‚ 8. Execute scraping (with fallback chains) β”‚ 9. Normalize results (HTML β†’ Markdown) β”‚ 10. Store results in SQLite β”‚ 11. Generate artifacts (JSONL, XLSX) β”‚ 12. Update SQLite: status = "completed" β”‚ β–Ό β”Œβ”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β” β”‚ ARTIFACTS STORAGE β”‚ β”‚ /opt/reality-scraper-v2/exports/ β”‚ β””β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”€β”˜ β”‚ β”‚ β€’ job_abc-123.jsonl (all results) β”‚ β€’ job_abc-123.xlsx (Excel export) β”‚ β–Ό CLIENT ← Response: {"job_id": "abc-123", "status": "completed"} ``` --- ### 2. Status Check Flow ``` CLIENT β†’ GET /v1/fabric/job/ Headers: X-Scraping-API-Key: ... ↓ CONTROL PLANE (in-memory lookup) ↓ Response: { "job_id": "abc-123", "status": "running|completed|failed", "mode": "FAST", "urls": ["https://example.com"], "created_at": "2026-01-29T14:35:51Z" } ``` Pro detailnΓ­ vΓ½sledky z data plane: ``` CLIENT β†’ GET http://46.224.232.134:5100/api/job/ (direct call to data plane) ↓ DATA PLANE (SQLite lookup) ↓ Response: { "job_id": "abc-123", "status": "completed", "completed_count": 2, "failed_count": 1, "results_count": 3, "artifacts_count": 2, "url_count": 3 } ``` --- ## πŸ”Œ API Reference ### Control Plane API (46.224.121.179:8090) #### 1. Health Check ```http GET /v1/fabric/health ``` **Response**: ```json { "status": "ok", "service": "scraping-fabric-control-plane", "version": "1.0.0" } ``` --- #### 2. List Scraping Modes ```http GET /v1/fabric/modes ``` **Response**: ```json { "modes": [ { "name": "FAST", "description": "Pure HTTP (requests/httpx) for static HTML" }, { "name": "SMART", "description": "Crawl4AI with JS rendering for dynamic content" }, { "name": "HEAVY", "description": "Playwright+Crawlee for anti-bot protected sites" } ] } ``` --- #### 3. Create Scraping Job ```http POST /v1/fabric/scrape Headers: X-Scraping-API-Key: czechai-internal-scraping Content-Type: application/json Body: { "urls": ["https://example.com"], "mode": "FAST" // FAST|SMART|HEAVY } ``` **Response** (202 Accepted): ```json { "job_id": "032a8fb3-8abe-4835-a467-3459e4b6d6c9", "mode": "FAST", "status": "running" } ``` --- #### 4. Get Job Status ```http GET /v1/fabric/job/ Headers: X-Scraping-API-Key: czechai-internal-scraping ``` **Response**: ```json { "job_id": "032a8fb3-8abe-4835-a467-3459e4b6d6c9", "status": "completed", "mode": "FAST", "urls": ["https://example.com"], "created_at": "2026-01-29T14:35:51.113418" } ``` --- ### Data Plane API (46.224.232.134:5100) #### 1. Health Check ```http GET http://46.224.232.134:5100/health ``` **Response**: ```json { "status": "ok", "service": "scraping-fabric-data-plane", "version": "1.0.0", "timestamp": "2026-01-29T14:32:55.401271" } ``` --- #### 2. Create Scrape Job (Internal - called by control plane) ```http POST http://46.224.232.134:5100/api/scrape Body: { "job_id": "032a8fb3-8abe-4835-a467-3459e4b6d6c9", "urls": ["https://example.com"], "mode": "FAST", "callback_url": "https://router.czechai.io/v1/fabric/callback/..." } ``` **Response** (200 OK): ```json { "job_id": "032a8fb3-8abe-4835-a467-3459e4b6d6c9", "status": "completed", "completed": 1, "failed": 0, "artifacts": [ { "type": "jsonl", "path": "/opt/reality-scraper-v2/exports/job_032a8fb3-8abe-4835-a467-3459e4b6d6c9.jsonl", "size_bytes": 187, "url": "/exports/job_032a8fb3-8abe-4835-a467-3459e4b6d6c9.jsonl" }, { "type": "xlsx", "path": "/opt/reality-scraper-v2/exports/job_032a8fb3-8abe-4835-a467-3459e4b6d6c9.xlsx", "size_bytes": 4967, "url": "/exports/job_032a8fb3-8abe-4835-a467-3459e4b6d6c9.xlsx" } ] } ``` --- #### 3. Get Job Results ```http GET http://46.224.232.134:5100/api/job/ ``` **Response**: ```json { "job_id": "032a8fb3-8abe-4835-a467-3459e4b6d6c9", "mode": "FAST", "status": "completed", "created_at": "2026-01-29T14:35:51.248276", "updated_at": "2026-01-29T14:35:51.277000", "url_count": 1, "completed_count": 1, "failed_count": 0, "results_count": 1, "artifacts_count": 2 } ``` --- #### 4. Download Artifacts ```http GET http://46.224.232.134:5100/exports/job_.jsonl GET http://46.224.232.134:5100/exports/job_.xlsx ``` **JSONL Format**: ```jsonl {"url": "https://example.com", "success": true, "mode_used": "FAST", "content_length": 513, "error": null, "duration": 0.14} {"url": "https://httpbin.org/html", "success": true, "mode_used": "FAST", "content_length": 3739, "error": null, "duration": 0.83} ``` **XLSX Format**: | URL | Mode Used | Success | Content Length | Duration (s) | Error | |-----|-----------|---------|----------------|--------------|-------| | https://example.com | FAST | Yes | 513 | 0.14 | | | https://httpbin.org/html | FAST | Yes | 3739 | 0.83 | | --- #### 5. Get Statistics ```http GET http://46.224.232.134:5100/api/stats ``` **Response**: ```json { "total_jobs": 3, "status_breakdown": { "completed": 2, "partial": 1 }, "total_scraped_urls": 5, "success_rate_pct": 80.0 } ``` --- ## πŸ–₯️ CLI Usage ### Instalace CLI je jiΕΎ nainstalovΓ‘no na Router Tower serveru: ```bash # Lokace /usr/local/bin/czechai-scrape # Test czechai-scrape modes ``` --- ### PΕ™Γ­kazy #### 1. Seznam reΕΎimΕ― ```bash czechai-scrape modes ``` **Output**: ``` πŸ“‹ Available Scraping Modes: ● FAST Pure HTTP (requests/httpx) for static HTML ● SMART Crawl4AI with JS rendering for dynamic content ● HEAVY Playwright+Crawlee for anti-bot protected sites ``` --- #### 2. Scrape jednoho URL ```bash # Default mode (SMART) czechai-scrape scrape https://example.com # Specify mode czechai-scrape scrape https://example.com FAST czechai-scrape scrape https://example.com SMART czechai-scrape scrape https://example.com HEAVY ``` **Output**: ``` πŸš€ Starting scrape job... URLs: 1 Mode: FAST βœ“ Job created: 032a8fb3-8abe-4835-a467-3459e4b6d6c9 ⏳ Waiting for completion... Status: running... βœ… Job completed! { "job_id": "032a8fb3-8abe-4835-a467-3459e4b6d6c9", "status": "completed", "completed": 1, "failed": 0, "artifacts": [...] } ``` --- #### 3. Batch scraping (vΓ­ce URL) ```bash czechai-scrape batch https://example.com https://httpbin.org/html https://google.com ``` **Output**: ``` πŸš€ Starting scrape job... URLs: 3 Mode: SMART βœ“ Job created: 9ddef167-6a41-48f2-9c50-3d45951496b2 ⏳ Waiting for completion... Status: running... βœ… Job completed! { "job_id": "9ddef167-6a41-48f2-9c50-3d45951496b2", "status": "completed", "completed": 3, "failed": 0 } ``` --- #### 4. Kontrola statusu ```bash czechai-scrape status 032a8fb3-8abe-4835-a467-3459e4b6d6c9 ``` **Output**: ```json { "job_id": "032a8fb3-8abe-4835-a467-3459e4b6d6c9", "status": "completed", "mode": "FAST", "urls": ["https://example.com"], "created_at": "2026-01-29T14:35:51.113418" } ``` --- ## πŸš€ Deployment ### Control Plane (Router Tower) **Server**: 46.224.121.179 ```bash # 1. Soubory jiΕΎ nahrΓ‘ny /opt/router-api/scraping_fabric_config.py /opt/router-api/scraping_fabric_control_plane_short.py # 2. Blueprint registrovΓ‘n v server.py (Ε™Γ‘dek ~70) app = Flask(__name__) # Scraping Fabric Blueprint (must be registered early) try: from scraping_fabric_control_plane_short import fabric_bp app.register_blueprint(fabric_bp) print('βœ“ Scraping Fabric control plane loaded') except ImportError as e: print(f'βœ— Scraping Fabric not available: {e}') # 3. PM2 restart pm2 restart router-api # 4. OvΔ›Ε™enΓ­ curl http://localhost:8090/v1/fabric/health ``` --- ### Data Plane (Scraper Service) **Server**: 46.224.232.134 ```bash # 1. Struktura sloΕΎek mkdir -p /opt/reality-scraper-v2/src mkdir -p /opt/reality-scraper-v2/data mkdir -p /opt/reality-scraper-v2/exports # 2. Soubory jiΕΎ nahrΓ‘ny /opt/reality-scraper-v2/src/scraping_fabric_config.py /opt/reality-scraper-v2/src/scraping_fabric_modes.py /opt/reality-scraper-v2/src/scraping_fabric_data_plane.py # 3. Virtual environment cd /opt/reality-scraper-v2 python3 -m venv venv source venv/bin/activate # 4. ZΓ‘vislosti pip install flask httpx crawl4ai playwright openpyxl # 5. Playwright browser playwright install chromium # 6. Start service cd /opt/reality-scraper-v2/src nohup /opt/reality-scraper-v2/venv/bin/python scraping_fabric_data_plane.py > /var/log/scraping-fabric-worker.log 2>&1 & # 7. OvΔ›Ε™enΓ­ curl http://localhost:5100/health ``` --- ### CLI Tool **Server**: 46.224.121.179 ```bash # CLI jiΕΎ nainstalovΓ‘no /usr/local/bin/czechai-scrape # Test czechai-scrape modes ``` --- ## πŸ“ PΕ™Γ­klady ### Example 1: Simple Single URL Scrape ```bash # CLI czechai-scrape scrape https://example.com FAST # cURL curl -X POST 'https://router.czechai.io/v1/fabric/scrape' \ -H 'X-Scraping-API-Key: czechai-internal-scraping' \ -H 'Content-Type: application/json' \ -d '{ "urls": ["https://example.com"], "mode": "FAST" }' # Response { "job_id": "032a8fb3-8abe-4835-a467-3459e4b6d6c9", "mode": "FAST", "status": "running" } # Check status curl -H 'X-Scraping-API-Key: czechai-internal-scraping' \ 'https://router.czechai.io/v1/fabric/job/032a8fb3-8abe-4835-a467-3459e4b6d6c9' ``` --- ### Example 2: Batch Scraping with Fallback ```bash # 3 URLs: valid, valid, invalid curl -X POST 'http://localhost:8090/v1/fabric/scrape' \ -H 'X-Scraping-API-Key: czechai-internal-scraping' \ -H 'Content-Type: application/json' \ -d '{ "urls": [ "https://example.com", "https://httpbin.org/html", "https://nonexistent-xyz123.com" ], "mode": "FAST" }' # Response { "job_id": "9ddef167-6a41-48f2-9c50-3d45951496b2", "mode": "FAST", "status": "running" } # Check results (after 5s) curl -s http://46.224.232.134:5100/api/job/9ddef167-6a41-48f2-9c50-3d45951496b2 # Response { "job_id": "9ddef167-6a41-48f2-9c50-3d45951496b2", "mode": "FAST", "status": "partial", # ← 2 succeeded, 1 failed "completed_count": 2, "failed_count": 1, "results_count": 3, "artifacts_count": 2 } # Download artifacts curl http://46.224.232.134:5100/exports/job_9ddef167-6a41-48f2-9c50-3d45951496b2.jsonl curl http://46.224.232.134:5100/exports/job_9ddef167-6a41-48f2-9c50-3d45951496b2.xlsx ``` --- ### Example 3: Python SDK Usage ```python import requests API_BASE = "https://router.czechai.io/v1/fabric" API_KEY = "czechai-internal-scraping" def scrape_url(url: str, mode: str = "SMART"): """Scrape single URL""" response = requests.post( f"{API_BASE}/scrape", headers={"X-Scraping-API-Key": API_KEY}, json={"urls": [url], "mode": mode} ) job = response.json() job_id = job["job_id"] # Wait for completion import time while True: status_response = requests.get( f"{API_BASE}/job/{job_id}", headers={"X-Scraping-API-Key": API_KEY} ) status_data = status_response.json() if status_data["status"] in ["completed", "failed"]: return status_data time.sleep(2) # Usage result = scrape_url("https://example.com", mode="FAST") print(result) ``` --- ### Example 4: Batch Scraping s Rate Limiting ```python import requests import time API_BASE = "https://router.czechai.io/v1/fabric" API_KEY = "czechai-internal-scraping" def batch_scrape(urls: list, mode: str = "FAST", batch_size: int = 10): """ Scrape URLs in batches to respect rate limits Rate limit: 60 requests per minute """ results = [] for i in range(0, len(urls), batch_size): batch = urls[i:i+batch_size] response = requests.post( f"{API_BASE}/scrape", headers={"X-Scraping-API-Key": API_KEY}, json={"urls": batch, "mode": mode} ) job = response.json() results.append(job) # Rate limit: wait 1 second between batches time.sleep(1) return results # Usage: scrape 100 URLs urls = [f"https://example.com/page-{i}" for i in range(100)] jobs = batch_scrape(urls, mode="FAST", batch_size=10) print(f"Created {len(jobs)} jobs") ``` --- ## πŸ› Troubleshooting ### Problem 1: "404 Not Found" na /v1/fabric/health **Příčina**: Blueprint nenΓ­ registrovΓ‘n nebo je registrovΓ‘n moc pozdΔ› **ŘeΕ‘enΓ­**: ```bash # 1. Check server.py blueprint registration ssh root@46.224.121.179 "grep -A 5 'app = Flask' /opt/router-api/server.py | head -15" # Should see: # app = Flask(__name__) # # # Scraping Fabric Blueprint (must be registered early) # try: # from scraping_fabric_control_plane_short import fabric_bp # app.register_blueprint(fabric_bp) # 2. Check logs ssh root@46.224.121.179 "pm2 logs router-api --lines 50 | grep Fabric" # Should see: "βœ“ Scraping Fabric control plane loaded" # 3. Restart ssh root@46.224.121.179 "pm2 restart router-api" ``` --- ### Problem 2: Data Plane nereaguje (port 5100) **Příčina**: Process nenΓ­ spuΕ‘tΔ›nΓ½ nebo crashnul **ŘeΕ‘enΓ­**: ```bash # 1. Check process ssh root@46.224.232.134 "ps aux | grep scraping_fabric_data_plane" # 2. Check port ssh root@46.224.232.134 "netstat -tlnp | grep :5100" # 3. Check logs ssh root@46.224.232.134 "tail -50 /var/log/scraping-fabric-worker.log" # 4. Restart ssh root@46.224.232.134 "pkill -f scraping_fabric_data_plane" ssh root@46.224.232.134 "cd /opt/reality-scraper-v2/src && nohup /opt/reality-scraper-v2/venv/bin/python scraping_fabric_data_plane.py > /var/log/scraping-fabric-worker.log 2>&1 &" # 5. Verify ssh root@46.224.232.134 "curl -s http://localhost:5100/health" ``` --- ### Problem 3: Job stuck na "running" status **Příčina**: Scraper crashnul bΔ›hem exekuce nebo timeout **ŘeΕ‘enΓ­**: ```bash # 1. Check data plane logs ssh root@46.224.232.134 "tail -100 /var/log/scraping-fabric-worker.log | grep ''" # 2. Check SQLite database ssh root@46.224.232.134 "sqlite3 /opt/reality-scraper-v2/data/scraping_fabric_jobs.db 'SELECT * FROM jobs WHERE job_id=\"\"'" # 3. Manually update job status ssh root@46.224.232.134 "sqlite3 /opt/reality-scraper-v2/data/scraping_fabric_jobs.db 'UPDATE jobs SET status=\"failed\", updated_at=\"$(date -u +%Y-%m-%dT%H:%M:%S)\" WHERE job_id=\"\"'" ``` --- ### Problem 4: Crawl4AI timeout errors **Příčina**: Target site je pomalΓ½ nebo blokuje boty **ŘeΕ‘enΓ­**: ```python # Option 1: Increase timeout v config # /opt/reality-scraper-v2/src/scraping_fabric_config.py SCRAPING_MODES = { "SMART": ScrapingMode( name="SMART", timeout=60, # ← Increase from 30s to 60s retry_count=2, fallback="HEAVY" ) } # Option 2: Use HEAVY mode directly curl -X POST ... -d '{"urls": [...], "mode": "HEAVY"}' ``` --- ### Problem 5: "ALL_FAILED" error **Příčina**: VΕ‘echny 3 reΕΎimy selhaly (FAST β†’ SMART β†’ HEAVY) **MoΕΎnΓ© dΕ―vody**: 1. Site neexistuje (DNS error) 2. Site vracΓ­ 403/429 (rate limiting, IP ban) 3. Site vyΕΎaduje authentication 4. CAPTCHA / human verification required **ŘeΕ‘enΓ­**: ```bash # 1. Test manually curl -I https://problematic-site.com # 2. Check error message in JSONL ssh root@46.224.232.134 "cat /opt/reality-scraper-v2/exports/job_.jsonl" # Look for "error" field: # {"url": "...", "success": false, "error": "...", "mode_used": "ALL_FAILED"} # 3. Possible fixes: # - Add user-agent header (modify modes.py) # - Add proxy support # - Implement CAPTCHA solving # - Use residential proxies ``` --- ### Problem 6: Rate Limiting (429 Too Many Requests) **Příčina**: PΕ™ekročen limit 60 req/min **ŘeΕ‘enΓ­**: ```python # Option 1: Increase rate limit v config # /opt/router-api/scraping_fabric_config.py RATE_LIMIT_PER_MINUTE = 120 # ← Increase from 60 to 120 # Option 2: Use multiple API keys # Create new API key VALID_API_KEYS = { "czechai-internal-scraping", "czechai-bulk-scraping", # ← Add new key "czechai-priority-scraping" } # Option 3: Batch requests (reduce overhead) # Instead of: # scrape(url1) + scrape(url2) + scrape(url3) # 3 requests # Use: # scrape([url1, url2, url3]) # 1 request ``` --- ## πŸ“Š Performance Metrics **Current Statistics** (2026-01-29): | Metric | Value | |--------|-------| | Total Jobs | 3 | | Total URLs Scraped | 5 | | Success Rate | 80% (4/5) | | Avg Duration (FAST) | 0.14s - 0.22s | | Avg Duration (SMART) | 0.34s - 0.83s | | Avg Duration (HEAVY) | Not tested yet | | Artifacts Generated | 6 (3 JSONL + 3 XLSX) | | Fallback Chain Success | 100% | **Scalability**: - Control Plane: 1 instance (can handle ~1000 req/min) - Data Plane: 1 instance (max 5 concurrent scrapes) - Horizontal Scaling: Add more scraper service nodes **Recommendations**: - For > 100 req/min: Add load balancer + 2+ data plane nodes - For > 1000 req/min: Use Redis for job queue + Celery workers - For > 10,000 req/min: Use Kubernetes + auto-scaling --- ## πŸ”’ Security ### Authentication **API Key Header**: `X-Scraping-API-Key` **Valid Keys** (stored in config): ```python VALID_API_KEYS = { "dev-scraping-key-2026", # Development "czechai-internal-scraping" # Production } ``` **To add new key**: ```bash # 1. Edit config ssh root@46.224.121.179 "nano /opt/router-api/scraping_fabric_config.py" # Add to VALID_API_KEYS set VALID_API_KEYS = { os.getenv("SCRAPING_API_KEY", "dev-scraping-key-2026"), "czechai-internal-scraping", "new-api-key-here" # ← Add new key } # 2. Restart ssh root@46.224.121.179 "pm2 restart router-api" ``` --- ### Rate Limiting **Current Limits**: - 60 requests per minute (per API key) - 1000 requests per hour (per API key) **Implementation**: ```python # In control plane rate_limit_store: Dict[str, list] = {} def check_rate_limit(api_key: str): now = time.time() window = 60 # seconds if api_key not in rate_limit_store: rate_limit_store[api_key] = [] # Remove old timestamps rate_limit_store[api_key] = [ ts for ts in rate_limit_store[api_key] if now - ts < window ] if len(rate_limit_store[api_key]) >= RATE_LIMIT_PER_MINUTE: return False # Rate limit exceeded rate_limit_store[api_key].append(now) return True ``` --- ### Data Privacy **Scraped content storage**: - SQLite database: `/opt/reality-scraper-v2/data/scraping_fabric_jobs.db` - Exports: `/opt/reality-scraper-v2/exports/` - **No external storage** (all on-premise) - **No cloud sync** (data stays on server) **Recommendations**: - Enable SQLite encryption (sqlcipher) - Add HTTPS for data plane (currently HTTP only) - Implement artifact expiration (auto-delete after 30 days) - Add audit logging (who accessed what) --- ## πŸ“ˆ Future Enhancements ### Phase 2 (Q2 2026) - [ ] Add Redis queue (replace in-memory job storage) - [ ] Implement Celery workers (async job execution) - [ ] Add Prometheus metrics (`/metrics` endpoint) - [ ] Implement webhook callbacks (notify on job completion) - [ ] Add proxy rotation support - [ ] CAPTCHA solving integration (2captcha, Anti-Captcha) ### Phase 3 (Q3 2026) - [ ] Horizontal scaling (multiple data plane nodes) - [ ] Load balancer (HAProxy/Nginx) - [ ] Kubernetes deployment (helm charts) - [ ] Add SMART+ mode (Playwright + Crawl4AI hybrid) - [ ] Screenshot capture (save page screenshots) - [ ] PDF export support - [ ] Real-time progress tracking (WebSocket) ### Phase 4 (Q4 2026) - [ ] AI-powered anti-bot evasion - [ ] Browser fingerprint randomization - [ ] Residential proxy network - [ ] Visual scraping (OCR, image recognition) - [ ] Scheduled scraping (cron jobs) - [ ] Data quality scoring - [ ] Duplicate detection (fuzzy matching) --- ## πŸ“ž Support **Documentation**: This file **Logs**: - Control Plane: `pm2 logs router-api` - Data Plane: `/var/log/scraping-fabric-worker.log` **Status Endpoints**: - Control Plane: `https://router.czechai.io/v1/fabric/health` - Data Plane: `http://46.224.232.134:5100/health` **CLI Help**: ```bash czechai-scrape # Show usage ``` --- ## πŸ“„ License & Credits **Built by**: Claude Code (Anthropic) **Date**: 2026-01-29 **Version**: 1.0.0 **Technologies**: - Flask (web framework) - httpx (HTTP client) - Crawl4AI (JS rendering) - Playwright (browser automation) - SQLite (persistence) - openpyxl (Excel export) **Credits**: - Crawl4AI: https://github.com/unclecode/crawl4ai (58.4k ⭐) - Playwright: https://playwright.dev/ - Flask: https://flask.palletsprojects.com/ --- ## βœ… Checklist Before going to production: - [x] Control plane deployed - [x] Data plane deployed - [x] CLI tool installed - [x] Health checks passing - [x] End-to-end test successful - [x] Artifacts generating correctly - [x] Fallback chains working - [ ] HTTPS enabled on data plane - [ ] Rate limiting tested under load - [ ] Monitoring/alerting setup - [ ] Backup strategy for SQLite DB - [ ] Log rotation configured --- **End of Documentation** πŸŽ‰ Last Updated: 2026-01-29 14:45 UTC