Posted 7 days ago

Data Engineer

1074 Amgen Technology Pvt Ltd. India, India - Hyderabad
Onsite Full Time

Job description

Career Category Engineering Job Description ABOUT AMGEN Amgen harnesses the best of biology and technology to fight the world’s toughest diseases, and make people’s lives easier, fuller and longer. We discover, develop, manufacture and deliver innovative medicines to help millions of patients. Amgen helped establish the biotechnology industry more than 40 years ago and remains on the cutting-edge of innovation, using technology and human genetic data to push beyond what’s known today. Role Description Let’s do this. Let’s change the world. We are seeking an experienced Data Engineer to design, develop, and support scalable data pipelines and integration solutions. This role will work with large and complex datasets to ensure data is reliable, accessible, secure, and available for analytics, reporting, and AI use cases. The ideal candidate has strong hands-on experience with Databricks, Apache Spark, Python, SQL, cloud technologies, data modeling, and ETL/ELT processes. The candidate will collaborate with data architects, business teams, data scientists, product teams, and DevOps teams to deliver high-quality data solutions.

Experience

in biotechnology, pharmaceutical, manufacturing, life sciences, or another regulated industry is preferred. Roles and

Responsibilities

Design, develop, test, and maintain scalable data pipelines and data-integration solutions. Build ETL/ELT processes for structured, semi-structured, and unstructured data. Integrate data from enterprise applications, databases, APIs, cloud platforms, and third-party systems. Contribute to the technical design and implementation of end-to-end data solutions. Develop reusable and maintainable Python, PySpark, Spark SQL, and SQL components. Implement data-quality checks, validation rules, reconciliation processes, logging, and exception handling. Optimize Spark workloads, SQL queries, partitioning, and data-processing performance. Develop and maintain data models, data dictionaries, mappings, and technical documentation. Implement data-security, privacy, governance, and role-based access requirements. Support workflow orchestration, scheduling, monitoring, alerting, and recovery processes. Contribute to CI/CD pipelines, automated testing, version control, and deployment processes. Troubleshoot data-pipeline failures, performance issues, and data-quality problems. Collaborate with data architects, business SMEs, analysts, data scientists, product teams, and DevOps teams. Participate in sprint planning, backlog refinement, technical estimation, and delivery activities. Take ownership of assigned data-engineering work from development through deployment and production support. Evaluate new technologies and recommend improvements to data-engineering processes and platform performance. Follow coding, testing, documentation, security, and reusable-development standards. Participate in operational support activities, including occasional off-hours support. Basic

Qualifications

Master’s or Bachelor’s degree in Computer Science, Information Technology, Engineering, Data Science, or a related field, with 5–8 years of relevant professional experience. Must-Have Skills Hands-on experience with Databricks, Apache Spark, PySpark, Spark SQL, Python, and SQL.

Experience

designing, developing, and supporting production ETL/ELT pipelines.

Experience

with workflow orchestration and scheduled data-processing workloads. Strong understanding of data modeling, data warehousing, lakehouse architecture, and data-integration concepts.

Experience

working with large and complex datasets.

Experience

with Spark and SQL performance tuning.

Experience

with AWS or another major cloud platform.

Experience

with Git, CI/CD, automated testing, and production deployment.

Experience

implementing data-quality, metadata, lineage, and governance controls. Understanding of role-based access control, data privacy, security, and compliance requirements. Strong analytical, troubleshooting, communication, and collaboration skills.

Experience

working in Agile delivery environments.

Preferred Qualifications

Experience

in biotechnology, pharmaceutical, life sciences, manufacturing, or another regulated industry.

Experience

with Delta Lake, Unity Catalog, Databricks Workflows, or comparable technologies.

Experience

with AWS data, storage, integration, security, and monitoring services.

Experience

developing APIs or data services for downstream consumers.

Experience

with relational, NoSQL, analytical, or vector databases.

Experience

with data visualization tools such as Tableau or Power BI.

Experience

supporting machine-learning or AI data pipelines.

Experience

designing or developing Generative AI solutions using LLMs, Retrieval-Augmented Generation, embeddings, vector search, prompt engineering, or governed enterprise data. Familiarity with AI-assisted development tools such as GitHub Copilot, OpenAI Codex, or equivalent platforms. Preferred Certifications Databricks Certified Data Engineer Associate. AWS Certified Data Engineer or another relevant cloud certification. Relevant data engineering, analytics, or Agile certification. Soft Skills Excellent critical-thinking and problem-solving skills. Strong verbal and written communication skills. Ability to work effectively with global and cross-functional teams. High degree of initiative, ownership, and attention to detail. Ability to manage multiple priorities and meet delivery commitments. Strong teamwork and collaboration skills. Ability to clearly present technical information to technical and non-technical audiences. . Amgen is committed to unlocking the potential of biology for patients suffering from serious illnesses by discovering, developing, manufacturing and delivering innovative human therapeutics. This approach begins by using tools like advanced human genetics to unravel the complexities of disease and understand the fundamentals of human biology. Amgen focuses on areas of high unmet medical need and leverages its biologics manufacturing expertise to strive for solutions that improve health outcomes and dramatically improve people's lives. A biotechnology pioneer since 1980, Amgen has grown to be one of the world's leading independent biotechnology companies, has reached millions of patients around the world and is developing a pipeline of medicines with breakaway potential. For more information, visit www.amgen.com and follow us on www.twitter.com/amgen

Skills and functions

  • Aws
  • Data Engineering
  • Data Science
  • Machine Learning
  • Python
  • Spark
  • Sql