Publicado el 02 sept 2026 · Lo confirmamos en el momento en que este empleo fue agregado
US$ 30 – US$ 80 por proyecto
# Financial Statement Data Extraction Specialist – 800+ Scanned PDFs I am looking for an experienced **Data Extraction / OCR / Data Engineering specialist** to extract financial data from approximately **800 phone-scanned financial statements** for academic research. The documents cover **multiple companies and multiple years**, and the final output must be structured as a **panel dataset in Excel** using the template I provide. ### Pilot Test I have uploaded **5 financial statements as a pilot sample**. The freelancer should extract the pilot documents into my Excel template exactly as the full project would be processed. The pilot will be used to evaluate the extraction method and accuracy before awarding the full project. ### What Must Be Extracted For each company/year: * All required financial statement numbers * Correct financial statement line-item/account mapping * Current-year and comparative-year values * Company name * Reporting/financial year * Auditor / audit firm name * Audit report date * **Source PDF filename** The Excel template I provide will define the required structure and line items. ### Accuracy Requirement The target is **more than 98% field-level accuracy**. Each required numeric or metadata field will be checked individually against the original PDF. For numeric fields, the extracted value must have the correct: * Number * Sign * Year/column * Account mapping * Scale/unit * Company Errors include wrong values, missing values, wrong signs, wrong years, wrong account mapping, wrong auditor, wrong audit report date, or values assigned to the wrong company. I will manually verify the pilot to assess the accuracy. ### Quality Control For traceability, I only require the **source PDF filename** to be linked to each extracted company-year record. I do **not require** page numbers, bounding-box images, OCR confidence scores, screenshots, or raw OCR text in the final dataset. The freelancer may use any additional OCR validation or manual review procedures internally to achieve the required accuracy. Uncertain values should **not be guessed** and should be clearly flagged for review. ### When Applying Please briefly explain: * What OCR / Document AI / Python tools you will use * How you will achieve and measure **98%+ accuracy** * How you will validate financial numbers and comparative-year columns * How you will handle low-confidence or unreadable values * Your experience with similar financial-document extraction * Your estimated price for approximately **800 PDFs** **Accuracy, data integrity and consistency are more important than speed.**
Crea una cuenta gratis para ver el empleo completo y postularte.