2 days ago
Data Engineer
Coretek Services · Kondapur, Telangana, India
WorkableApply on company site
Full Time
Hybrid
Role overview
Coretek is looking for a Data Engineer to build and operate the pipelines and data models that the rest of the business runs on. You'll own ingestion from source systems through to curated, well-documented datasets that analysts, data scientists, and application teams depend on. This is a hands-on engineering role: you'll write production code, design schemas, and be accountable for the reliability and cost of what you ship.
Responsibilities
- Design, build, and maintain batch and streaming data pipelines that are idempotent, observable, and recoverable.
- Model data for analytics (dimensional models, semantic layers, and curated marts), balancing query performance against maintainability.
- Integrate data from operational databases, SaaS APIs, files, and event streams, including handling schema drift and late-arriving data.
- Build data quality checks (freshness, volume, uniqueness, referential integrity) into pipelines rather than bolting them on afterward, and define how failures alert and escalate.
- Own pipelines in production: monitoring, on-call rotation for data incidents, root-cause analysis, and backfills.
- Tune performance and cost (partitioning, clustering, file sizing, warehouse and cluster sizing) and make the tradeoffs explicit.
- Apply engineering discipline to data: version control, code review, CI/CD, automated testing, and infrastructure as code.
- Implement access controls, PII handling, retention, and lineage and audit requirements in partnership with security and compliance.
- Partner with analysts, data scientists, and product engineers to turn ambiguous requirements into durable data contracts.
- Maintain data dictionaries, lineage, and pipeline runbooks so consumers can find a dataset, understand what each field means and how current it is, and use it correctly without having to ask the team that built it.
Requirements
- 5+ years building production data pipelines.
- Strong hands-on Python development for data engineering, with real testing, packaging, and code review practice, not scripting alone.
- Working knowledge of PySpark: DataFrame and SQL APIs, joins and aggregations at scale, partitioning and shuffle behavior, and the ability to read a Spark UI to diagnose a slow or failing job.
- Strong SQL: window functions, query plans, and performance tuning, not just SELECTs.
- Hands-on experience with the Azure data platform: Data Factory, Databricks, Synapse/Fabric, and ADLS.
- Solid data modeling fundamentals: normalization, star schemas, slowly changing dimensions.
- Git-based workflow and experience shipping through CI/CD.
- Excellent communication skills, with the ability to debug a failing pipeline end to end and articulate the impact to diverse audiences, including non-technical stakeholders.
- Exceptional analytical and problem-solving skills, with the judgment to find the root cause of a data issue rather than patching the symptom.
- Strong knowledge and experience in working with customers in a consultative approach in a technical environment.
Additional Qualifications
- Streaming experience (Kafka, Event Hubs).
- Lakehouse formats: Delta Lake, Iceberg.
- Infrastructure as code (Terraform, Bicep) and containerization (Docker, Kubernetes).
- Experience in a regulated environment (HIPAA, SOC 2, PCI, GDPR): auditability, encryption, data residency.
- Experience building data platforms for ML or supporting feature pipelines.
- Proven ability to manage multiple client projects and deliver high-quality results on time.
- Experience in Azure DevOps or GitHub for source control and pipelines.
Responsibilities
1Design, build, and maintain batch and streaming data pipelines that are idempotent, observable, and recoverable.
2Model data for analytics (dimensional models, semantic layers, and curated marts), balancing query performance against maintainability.
3Integrate data from operational databases, SaaS APIs, files, and event streams, including handling schema drift and late-arriving data.
4Build data quality checks (freshness, volume, uniqueness, referential integrity) into pipelines rather than bolting them on afterward, and define how failures alert and escalate.
5Own pipelines in production: monitoring, on-call rotation for data incidents, root-cause analysis, and backfills.
6Tune performance and cost (partitioning, clustering, file sizing, warehouse and cluster sizing) and make the tradeoffs explicit.
7Apply engineering discipline to data: version control, code review, CI/CD, automated testing, and infrastructure as code.
8Implement access controls, PII handling, retention, and lineage and audit requirements in partnership with security and compliance.
Requirements
15+ years building production data pipelines.
2Strong hands-on Python development for data engineering, with real testing, packaging, and code review practice, not scripting alone.
3Working knowledge of PySpark: DataFrame and SQL APIs, joins and aggregations at scale, partitioning and shuffle behavior, and the ability to read a Spark UI to diagnose a slow or failing job.
4Strong SQL: window functions, query plans, and performance tuning, not just SELECTs.
5Hands-on experience with the Azure data platform: Data Factory, Databricks, Synapse/Fabric, and ADLS.
6Solid data modeling fundamentals: normalization, star schemas, slowly changing dimensions.
7Git-based workflow and experience shipping through CI/CD.
8Excellent communication skills, with the ability to debug a failing pipeline end to end and articulate the impact to diverse audiences, including non-technical stakeholders.
9Exceptional analytical and problem-solving skills, with the judgment to find the root cause of a data issue rather than patching the symptom.
10Strong knowledge and experience in working with customers in a consultative approach in a technical environment.
11Streaming experience (Kafka, Event Hubs).
12Lakehouse formats: Delta Lake, Iceberg.
13Infrastructure as code (Terraform, Bicep) and containerization (Docker, Kubernetes).
14Experience in a regulated environment (HIPAA, SOC 2, PCI, GDPR): auditability, encryption, data residency.
Skills and tags
INPythonSQLAzureDockerKubernetesDevOpsAuditData AnalysisLogisticsWarehouseCompliance