Published on Aug 29, 2026 · We confirmed on Aug 30, 2026 that it's still live
₹ 12.500 – ₹ 37.500 per project
Project Title: Senior Data Acquisition & Document Intelligence Engineer We are looking for a highly experienced Senior Data Acquisition & Document Intelligence Engineer to build, operate, and continuously improve a large-scale production data acquisition and document processing system. Project Type: Long-term / Ongoing Engagement: Full-time Experience Required: 5+ years Location: On-site / Hybrid About the Project We collect large volumes of publicly available documents and structured records from hundreds of external web sources every day. The sources are highly inconsistent and frequently change their website structure, formats, URLs, APIs, and document layouts. A significant portion of the data also comes from scanned PDFs, images, and other difficult-to-process documents. We need someone who can take end-to-end ownership of the entire data acquisition pipeline, including web scraping, crawling, document downloading, OCR, data extraction, structuring, validation, quality assurance, monitoring, and ongoing maintenance. This is not a one-time scraping project. We are looking for someone with proven experience building and maintaining production-grade extraction pipelines at scale. Key Responsibilities 1. Web Scraping & Data Acquisition Build and maintain large-scale crawlers across hundreds of websites and external data sources. Handle dynamic websites, JavaScript rendering, pagination, sessions, cookies, tokens, forms, and file downloads. Work with technologies such as Python, Scrapy, Playwright, Selenium, Puppeteer, Requests/httpx, etc. Implement rate limiting, retries, backoff, and request management. Build incremental crawling and change-detection mechanisms. Detect website/source changes and quickly fix broken extractors. Ensure data is collected completely, accurately, and on schedule. 2. OCR & Document Processing Process large volumes of native PDFs, scanned PDFs, images, and other document formats. Build and optimize OCR pipelines for poor-quality documents. Handle skewed/rotated pages, faded text, stamps, seals, watermarks, handwriting, and complex layouts. Apply image preprocessing such as deskewing, denoising, binarization, cropping, and image enhancement. Extract structured information from complex tables, including merged cells, multi-page tables, borderless tables, and nested headers. Process multilingual documents, including Indian regional languages. Benchmark and select appropriate OCR engines based on document type and accuracy. 3. Data Extraction & Structuring Convert unstructured web/document content into structured datasets. Design and maintain data schemas. Use rules, regex, layout-aware extraction, NLP/NER, and LLM-assisted extraction where appropriate. Normalize dates, amounts, units, names, addresses, identifiers, and other inconsistent values. Implement deduplication and entity resolution. Maintain taxonomies, classification logic, and reference/master data. 4. Data Quality & Validation Build automated quality checks for completeness, accuracy, freshness, consistency, uniqueness, and validity. Implement validation and anomaly-detection mechanisms. Maintain gold-standard datasets and benchmark extraction/OCR accuracy. Create human-in-the-loop review processes for low-confidence results. Reconcile extracted data against the original source. Maintain data lineage and audit trails so records can be traced back to the original source document and extraction run. 5. Production Operations Build and operate reliable, scheduled production pipelines. Use Airflow, Prefect, Dagster, or similar orchestration tools. Implement monitoring, logging, alerting, retries, and failure recovery. Track uptime, throughput, latency, coverage, cost, and data quality. Use Docker and CI/CD. Troubleshoot production issues and prevent recurring failures. Document the system, extraction logic, and operational procedures. Must-Have Skills 5+ years of hands-on experience in production data extraction, web scraping, document processing, or data acquisition. Strong Python programming and production engineering practices. Strong experience with large-scale web scraping/crawling. Experience with Scrapy, Playwright, Selenium, Puppeteer, Requests/httpx, or equivalent. Deep practical OCR experience with Tesseract, PaddleOCR, EasyOCR, Google Document AI, AWS Textract, Azure Document Intelligence, or similar. Strong PDF/document processing experience with PyMuPDF, pdfplumber, pdfminer, Camelot, Tabula, or equivalent. Strong SQL and relational database knowledge. PostgreSQL or equivalent database experience. Experience with Airflow, Prefect, Dagster, or similar workflow orchestration. Experience with AWS and/or GCP. Experience with Docker, CI/CD, logging, monitoring, and alerting. Strong understanding of data quality and automated testing. Experience maintaining and improving production datasets over time. Understanding of responsible and lawful web data collection, including robots.txt, terms of service, privacy, and applicable data protection requirements. Good to Have OpenCV and image processing NLP / NER LLM-based document extraction Multilingual OCR Indian regional language OCR Entity resolution / record linkage Data lineage and provenance Great Expectations, Soda, Pandera, or similar dbt and dbt tests BigQuery AWS S3, Lambda, Batch GCP Cloud Storage Distributed/batch processing Experience with government, legal, financial, regulatory, or public-sector documents Ideal Candidate We are not looking for someone who has only built basic scraping scripts. The ideal candidate should have experience operating a production system involving: Hundreds of Sources → Crawling → Document Acquisition → OCR → Data Extraction → Structuring → Validation → QA → Production Dataset You should understand common failure points such as website changes, broken selectors, missing documents, OCR errors, duplicate records, incorrect table extraction, silent pipeline failures, source downtime, and data quality degradation—and k
Create a free account to see the full job and apply.