About this opportunity
Role Overview
CNTXT AI is seeking a Data Engineer to build, optimize, and maintain scalable data infrastructure supporting analytics, AI, and machine learning products. The role focuses on data pipelines, data integration, quality, infrastructure, and accessibility.
Key Responsibilities
Design and maintain scalable ETL/ELT pipelines for structured and unstructured data.
Build and manage data warehouses and data lakes.
Develop batch and real-time data processing pipelines.
Integrate data from APIs, databases, third-party platforms, and cloud services.
Ensure data quality, security, integrity, and governance.
Optimize data models and database performance.
Monitor and troubleshoot data pipelines.
Collaborate with data scientists, ML engineers, product managers, and software engineers.
Implement CI/CD and infrastructure automation for data workflows.
Document data architecture, pipelines, and engineering practices.
Keep up with emerging cloud and big data technologies.
Requirements
Bachelor’s degree in Computer Science, Software Engineering, Information Systems, or a related field.
4+ years of Data Engineering or Backend/Data Platform development experience.
Strong Python and SQL skills.
Experience with ETL/ELT orchestration tools such as Airflow, Prefect, or Dagster.
Knowledge of relational and NoSQL databases, including PostgreSQL, MySQL, MongoDB, or Cassandra.
Experience with AWS, Azure, or Google Cloud Platform.
Experience with cloud data warehouses such as Snowflake, BigQuery, Redshift, or Synapse.
Experience with Apache Spark.
Familiarity with Kafka or RabbitMQ.
Experience with Docker, Kubernetes, and CI/CD pipelines.
Strong understanding of data modeling, partitioning, indexing, and performance optimization.
Experience with Git and collaborative software development workflows.
Preferred Qualifications
Experience building data platforms for AI/ML workloads.
Knowledge of Delta Lake, Apache Iceberg, or Apache Hudi.
Experience with dbt.
Familiarity with Terraform or Infrastructure as Code.
Knowledge of data governance, metadata management, and data cataloging.
Experience in high-growth technology or AI companies.
Nice to Have
Generative AI or LLM application experience.
Experience with vector databases such as Pinecone, Weaviate, or Milvus.
Familiarity with data observability platforms such as Monte Carlo or Great Expectations.
Experience with event-driven architectures and real-time analytics.
Exposure to MLOps platforms and feature stores.