ZentixSoft is looking for two Senior Data Ingestion Engineers!
Format: Direct Contract Engagement (1-Year Contract, 100% Full-Time Remote).
Why ZentixSoft Partner Network?
Transparent Compensation: Direct payroll arrangement with $40 USD/hour base compensation.
Work-Life Balance: Sustainable engineering workflows with predictable long-term deliverables.
Trust & Transparency: Zero micromanagement and full autonomy over your technical pipeline architecture.
Culture & Growth: Long-term project stability within large-scale financial and insurance data domains.
Responsibilities:
Pipeline Design & Engineering: Design and deploy scalable data ingestion pipelines for high-volume structured/unstructured documents (PDFs, scans, emails, Word, Excel, PowerPoint).
Document Extraction & OCR: Implement OCR and document processing workflows using AWS Textract (or equivalent) for text extraction, cleaning, normalization, and metadata tagging.
RAG & Vector Storage Prep: Execute semantic chunking, metadata extraction, vector storage schema design, and retrieval mechanism preparation for downstream AI models.
Integrations & Connectors: Build robust connectors with enterprise sources, including SharePoint, email servers, and public cloud repositories.
Validation & Observability: Implement automated error monitoring, OCR extraction validation, and CI/CD automated testing using AWS Step Functions and CloudWatch.
Our Perfect Match (Requirements):
Senior Data Engineering: 7+ years of commercial Data Engineering experience, with strong proficiency in Python and SQL.
AWS Expertise: 5+ years of hands-on experience in AWS environments, specifically AWS S3, Step Functions, CloudWatch, and public cloud data processing.
Unstructured Data & OCR: Proven background in building document extraction pipelines (handling PDFs, scanned images, emails, Office documents) and utilizing OCR technology (AWS Textract or similar).
Source Connectors & Pipeline Operations: Practical experience integrating sources like SharePoint and email, performing text normalization, metadata tagging, and maintaining CI/CD/Git testing standards.
EU Residency & Location: Candidates MUST reside in an EU member country (with active legal residency; citizenship can be non-EU/global).
Language & Communication: C1 Advanced English (verbal and written) for direct technical collaboration.
Good to Have:
Experience with Vector Databases, RAG architectures, semantic chunking, and vector retrieval mechanisms.
Background in Insurance or Financial Services industries handling enterprise security standards.
Exposure to Azure or Databricks platform components.
Project Details:
Duration: 1 Year (Full-time contract).
Target Start: End of October.
Location: 100% Remote (Must be physically located in an EU member state with valid residency).
Сродна праця в Telegram
Двічі на тиждень — віддалена вакансія, розібрана людською мовою: що робити, кому сродна, чесно про мінуси. А в коментарях — Сковорідка (ШІ), яка допоможе з резюме.
Підписатися