Skip to content
AI Trainer Jobs

Big Data Engineer

micro1 · 100% remote · Contract · Posted

Pay
$30–80/hr
Location
Worldwide
Languages
English
Hours
Flexible
Openings
50
Level
Expert
Apply for this position

Summary

Design and maintain large-scale data pipelines and distributed data systems in a remote contract role supporting AI training. Use Python and big data technologies to process data, improve reliability, and enforce data quality, security, and governance.

What you'll do

  • Design, build, and maintain scalable data pipelines and architectures.
  • Collaborate with teams to define and meet data requirements.
  • Implement data integration, transformation, and processing with Python and big data technologies.
  • Develop and optimize distributed databases and storage systems.
  • Monitor, troubleshoot, and improve data availability and performance.
  • Enforce data quality, security, and governance standards.
  • Document solutions and explain technical concepts.

Requirements

  • Proven big data engineering expertise and hands-on experience with large-scale data pipelines.
  • Advanced Python proficiency for data processing, automation, and integration.
  • Deep knowledge of relational and NoSQL database management and optimization.
  • Experience with distributed processing frameworks such as Hadoop, Spark, or Flink.
  • Knowledge of data modeling, ETL, and data warehousing.
  • Clear written and verbal communication with technical and non-technical stakeholders.
  • Attention to detail and ability to work independently remotely.
  • Preferred: Global team experience, cloud big data platforms, machine learning operations, and data science workflows.

Skills

  • Big Data
  • Python
  • Databases
  • Data Engineering

Full description

Job Title: Big Data Engineer


Job Type: Contractor


Location: Remote


Job Summary: In this role, you'll apply your expertise to help train next-generation AI systems. Your work will shape how models learn, reason, and perform through high-quality, real-world input.


Key Responsibilities:

  1. Design, build, and maintain scalable big data pipelines and architectures to support robust data solutions.
  2. Collaborate with cross-functional teams to understand and deliver on data requirements and business objectives.
  3. Implement data integration, transformation, and processing solutions using Python and relevant big data technologies.
  4. Develop, manage, and optimize distributed databases and storage systems for efficiency and reliability.
  5. Monitor, troubleshoot, and enhance data systems to ensure high availability and performance.
  6. Enforce data quality, security, and governance standards across all solutions.
  7. Document solutions and communicate complex technical concepts effectively, both in writing and verbally.


Required Skills and Qualifications:

  1. Proven expertise in big data engineering with hands-on experience building and maintaining large-scale data pipelines.
  2. Advanced proficiency in Python for data processing, automation, and integration.
  3. Deep understanding of relational and NoSQL databases, including optimization and management techniques.
  4. Experience with distributed data processing frameworks (e.g., Hadoop, Spark, Flink).
  5. Strong foundation in data modeling, ETL processes, and data warehousing principles.
  6. Excellent written and verbal communication skills, with the ability to convey technical ideas clearly to both technical and non-technical stakeholders.
  7. Detail-oriented, proactive, and self-motivated, thriving in remote and autonomous work environments.


Preferred Qualifications:

  1. Prior experience in fast-paced or startup-like environments supporting global teams.
  2. Expertise with cloud-based big data platforms (e.g., AWS, GCP, Azure).
  3. Familiarity with machine learning operations and data science workflows is a plus.

Similar jobs

View all jobs