Data Engineer (On-Site) | Engenheiro de Dados (Remoto)
Back to Jobs
DadosferaGet Smart Job AI Coach in the appFree on iOS and Android

Data Engineer (On-Site) | Engenheiro de Dados (Remoto)
Location
Remote (Brazil)
Experience
Mid
Posted
Jul 18, 2026
Apply by
August 17, 2026
Applicants
0
Early applicantEasy applyContractWork from Home
Job Description
Company Overview
Dadosfera is transforming the data landscape with a platform that delivers advanced Data, AI, and Analytics capabilities—previously available only to tech giants like Meta, Amazon, Alphabet, and Microsoft—into the hands of small and medium-sized businesses. By democratizing access to these technologies, we empower our clients to accelerate business opportunities through AI-powered Data Apps.
Our platform minimizes the time spent on cloud management and streamlines the entire data lifecycle—from collection and exploration to processing and interaction—allowing clients to focus on their business goals. Built on a high-growth SaaS model, Dadosfera offers tailored solutions that meet real business needs.
What sets Dadosfera apart is our innovative focus on speed, scalability, and flexibility, combining big data storage with advanced analytics to drive strategic decision-making. Leveraging years of experience with data platform implementations on foreign public clouds, we now deliver a cost-effective, locally adapted solution tailored for both Brazilian and global markets. By blending deep expertise with localized infrastructure, Dadosfera provides a powerful, efficient tool that optimizes corporate data management and drives data-driven transformation.
About the Job
We are looking for a Data Engineer to join our engineering team and help design, develop, and evolve scalable data platforms and pipelines on AWS.
In this role, you will be responsible for building robust solutions for data ingestion, processing, integration, and delivery, ensuring reliable data for analytics, data products, and business decision-making.
You will also develop solutions that leverage Artificial Intelligence to improve data quality and enrichment, including automated attribute extraction, information classification, entity recognition, record deduplication, and match-and-merge processes.
You will collaborate with cross-functional teams, participate in architecture discussions, and contribute to building modern, cloud-native, and scalable data ecosystems while following engineering, governance, and security best practices.
Responsibilities
- Design, develop, and maintain scalable ETL/ELT pipelines using Python, SQL, and AWS services.
- Design and evolve layered Data Lake architectures following the Medallion Architecture pattern.
- Develop incremental data pipelines with checkpointing, idempotency, failure recovery, and reprocessing strategies.
- Build solutions for ingesting, processing, and integrating structured, semi-structured, and unstructured data.
- Develop automated mechanisms to extract attributes and relevant information from documents, text, and multiple data sources.
- Apply Artificial Intelligence, Machine Learning, or Large Language Models (LLMs) to classify, standardize, validate, and enrich data.
- Implement entity resolution, record linkage, deduplication, and match-and-merge processes to identify related records and build trusted, unified datasets.
- Define similarity criteria, business rules, confidence thresholds, and review workflows for record-matching processes.
- Optimize SQL queries and storage structures for analytical workloads and large-scale datasets.
- Work with AWS services including Amazon S3, AWS Lambda, AWS Glue, AWS Step Functions, Amazon Athena, and Amazon DynamoDB.
- Ensure the quality, reliability, performance, security, and observability of data pipelines and data products.
- Implement validation rules, monitoring, metrics, and alerting mechanisms to detect failures and data inconsistencies.
- Collaborate with software engineers, data analysts, business stakeholders, and clients to translate business requirements into scalable technical solutions.
- Participate in code reviews, architecture discussions, and continuous improvement initiatives.
- Document data flows, processing rules, technical solutions, and architectural decisions.
Minimum Qualifications
- Solid experience with AWS services, including:
- - Amazon S3
- AWS Lambda
- AWS Step Functions
- AWS Glue
- Amazon Athena
- Amazon DynamoDB
- Strong experience with Python for data engineering, including ETL/ELT and batch processing.
- Advanced SQL skills, including data modeling, incremental processing, and query optimization.
- Experience designing and implementing layered Data Lake architectures.
- Experience working with columnar storage formats, especially Parquet.
- Experience building incremental data pipelines with checkpoint management and reprocessing capabilities.
- Knowledge of data quality, data cleansing, and standardization techniques.
- Experience integrating multiple data sources and identifying duplicate or related records.
- Experience with Git, source code versioning, and software engineering best practices.
Preferred Qualifications
- Experience applying Artificial Intelligence or Machine Learning to data engineering and data quality challenges.
- Experience with Large Language Models (LLMs), Generative AI APIs, or Natural Language Processing (NLP) libraries.
- Knowledge of entity extraction, text classification, and structured information extraction techniques.
- Experience with entity resolution, record linkage, entity matching, deduplication, and match-and-merge processes.
- Knowledge of probabilistic matching, fuzzy matching, text similarity techniques, and master record creation.
- Experience with Apache Iceberg or transactional table formats such as Delta Lake or Apache Hudi.
- Experience with Terraform or other Infrastructure as Code (IaC) tools.
- Experience with Docker and Amazon ECS Fargate or equivalent container platforms.
- Experience with streaming technologies such as Amazon Kinesis Firehose.
- Knowledge of Change Data Capture (CDC) architectures and ingestion patterns.
- Experience with orchestration, observability, and data quality tools.
At Dadosfera, we celebrate diversity in all its forms and are committed to fostering an inclusive environment every day. We have zero tolerance for any form of discrimination. Do you share our values? Then don’t wait—apply today!
Key Responsibilities
- Design, develop, and maintain scalable ETL/ELT pipelines using Python, SQL, and AWS services.
- Design and evolve layered Data Lake architectures following the Medallion Architecture pattern.
- Develop incremental data pipelines with checkpointing, idempotency, and failure recovery strategies.
- Build solutions for ingesting, processing, and integrating structured, semi-structured, and unstructured data.
- Develop automated mechanisms to extract attributes and relevant information from documents and text.
- Apply AI, Machine Learning, or LLMs to classify, standardize, validate, and enrich data.
- Implement entity resolution, record linkage, deduplication, and match-and-merge processes.
- Optimize SQL queries and storage structures for analytical workloads.
- Ensure the quality, reliability, performance, security, and observability of data pipelines.
- Implement validation rules, monitoring, metrics, and alerting mechanisms.
- Collaborate with cross-functional teams to translate business requirements into technical solutions.
- Participate in code reviews, architecture discussions, and continuous improvement initiatives.
- Document data flows, processing rules, and architectural decisions.
Skills Required
AWSAmazon S3AWS LambdaAWS Step FunctionsAWS GlueAmazon AthenaAmazon DynamoDBPythonSQLETLELTData ModelingParquetGitData Lake ArchitectureMedallion ArchitectureCollaborationCommunicationProblem solvingArtificial IntelligenceMachine LearningLarge Language Models (LLMs)Generative AI APIsNatural Language Processing (NLP)Entity extractionText classificationStructured information extractionEntity resolutionRecord linkageEntity matchingDeduplicationMatch-and-mergeProbabilistic matchingFuzzy matchingText similarityApache IcebergDelta LakeApache HudiTerraformInfrastructure as Code (IaC)DockerAmazon ECS FargateAmazon Kinesis FirehoseChange Data Capture (CDC)Orchestration toolsObservability toolsData quality tools
App exclusive · Free
Smart Job AI Coach
Your personal interview coach on every job — readiness tips, profile improvements, and role-specific prep. Available only in the Pulse Job app.
Interview readiness
See how prepared you are and what to improve for each role.
Personalized tips
Actionable suggestions based on your profile and the job.
After you apply
Keep coaching momentum from job detail through application success.
Similar roles for you
Matched using this role's title and skills. Open the job search anytime to see every listing.
React JS Developer
Qtechus, Inc.
$60–$80 / Hour
ContractEasy applyWork from Office
Austin1 applicant
Data Engineer
L&t Technology Services
₹12,00,000–₹20,00,000 / Year
Full-timeEasy applyWork from Office
Bellary Rd
Gen AI / Data Science Engineer
L&t Technology Services
₹14,00,000–₹24,00,000 / Year
Full-timeEasy applyWork from Office
Bellary Rd
Backend Node.js Developer
IITIL
₹60,000–₹1,00,000 / Month
Full-timeEasy applyWork from Home
Remote
QA Engineer
Sriven Pros Inc
$85–$120 / Hour
ContractEasy applyHybrid
Chicago
Machine Learning Engineer
Trilyon, Inc.
$60–$85 / Hour
ContractEasy applyWork from Home
Remote