Data Engineer - Pharma R&D
Roche Diagnostics India Pvt LtdJob Description
Data Engineer - Pharma R&D
At Roche you can show up as yourself, embraced for the unique qualities you bring. Our culture encourages personal expression, open dialogue, and genuine connections, where you are valued, accepted and respected for who you are, allowing you to thrive both personally and professionally. This is how we aim to prevent, stop and cure diseases and ensure everyone has access to healthcare today and for generations to come. Join Roche, where every voice matters.
The Position
Job Description:
The Data Engineer – Clinical Study Design sits at the intersection of data architecture, clinical science, and technology delivery, helping build the data foundation behind Study Designer, a digital product transforming how studies are designed. This role combines strong data engineering skills, an understanding of clinical study data and workflows, and technical curiosity to translate complex clinical data structures into reliable, scalable pipelines and models that power the platform's insights. The ideal candidate is curious, collaborative, AI-minded, and passionate about using data to enable smarter, faster, and more effective clinical study design.
Key Responsibilities
Build ingestion pipelines for clinical trial protocols, ICF documents, SmPCs, CSRs, and published articles (PubMed, CTIS, ClinicalTrials.gov) - handling PDF parsing, text extraction, and structured data normalization
Design and implement data models in Amazon Aurora (relational) and GraphDB (knowledge graph) to represent trial design entities: endpoints, eligibility criteria, study arms, interventions, therapeutic areas, and their relationships
Develop embedding and vectorization pipelines to prepare extracted clinical text for RAG-based retrieval in LangGraph agentic workflows - chunking strategies, metadata enrichment, and vector store population
Build and maintain ETL/ELT workflows that transform unstructured clinical content into queryable, linked data across both relational and graph stores
Implement data quality validation specific to clinical data - protocol section classification accuracy, entity extraction completeness, cross-reference integrity (NCT IDs, EudraCT numbers, MeSH terms)
Build data serving APIs (Python/FastAPI) that expose curated datasets to the Angular frontend and LangGraph agent layer
Set up data lineage tracking and audit trails to support regulatory traceability of AI-generated trial design recommendations
Preferred Qualifications:
Education: Bachelor's degree in Computer Science, Data Engineering, or a related discipline.
5-8 years of experience building production grade data platforms and pipelines.
Experience with biomedical knowledge graphs (e.g., linking drugs -> targets -> diseases -> trials)
Prior work with PubMed/MEDLINE data, ClinicalTrials.gov API, or EMA/CTIS data
Apache Spark or Databricks for batch processing of large document
Required Skills
Python — Primary language; experience with PDF/document parsing libraries (PyMuPDF, pdfplumber, unstructured.io, or similar)
SQL — Advanced PostgreSQL-compatible SQL (Aurora); schema design, migrations, query optimization, indexing strategies for clinical data volumes
Graph Databases — Hands-on with Neptune, Neo4j, or similar; SPARQL or Cypher query language; ontology/knowledge graph modeling for biomedical entities
AWS — Aurora (PostgreSQL), S3, Lambda, Step Functions, SQS/SNS, IAM; infrastructure for data pipeline orchestration
NLP / Document Processing — Text extraction from PDFs, section classification, named entity recognition for clinical/biomedical text; familiarity with embedding models and vector stores (OpenSearch, pgvector, or Pinecone)
FastAPI — Building data serving endpoints; async patterns; integration with the application backend
AI/ML Data Infrastructure — Preparing data for LangChain/LangGraph consumption; RAG pipeline design (chunking, retrieval, reranking); prompt-data integration patterns
Pipeline Orchestration — Experience with workflow orchestration tools (Airflow, Prefect, Step Functions, or Temporal); designing DAGs for multi-stage data pipelines with dependency management, retry logic, and monitoring
CI/CD & IaC — Terraform or CDK, Docker, Git; automated pipeline testing and deployment on AWS
#Hyderabad2026
Who we are
A healthier future drives us to innovate. Together, more than 100’000 employees across the globe are dedicated to advance science, ensuring everyone has access to healthcare today and for generations to come. Our efforts result in more than 26 million people treated with our medicines and over 30 billion tests conducted using our Diagnostics products. We empower each other to explore new possibilities, foster creativity, and keep our ambitions high, so we can deliver life-changing healthcare solutions that make a global impact.
Let’s build a healthier future, together.
Roche is an Equal Opportunity Employer.
Experience Level
Senior LevelJob role
Job requirements
About company
Similar jobs you can apply for
Manufacturing / ProductionQA / QC Manager
Vision Engineers
Web Developer
Ametecs India Private Limited
Quality Assurance Executive
Dr. Reddy's Foundation
Associate Software Engineer
Elevate Global Technologies Private Limited
FULL STACK ENGINEER – MERN STACK
The Hard Cash
Quality Control Engineer
CNC TechnicsYou can expect a minimum salary of 0 INR. The salary offered will depend on your skills, experience and performance in the interview.
The candidate should have completed the required education and people who have 5 to 8 years are eligible to apply for this job. You can apply for more jobs in Hyderabad to get hired quickly.
The candidate should have sound communication skills and sound communication skills for this job.
Both Male and Female candidates can apply for this job.
No, it's not a work from home job and can't be done online. You can explore and apply for other work from home jobs in Hyderabad at apna.
No work-related deposit needs to be made during your employment with the company.
Go to the apna app and apply for this job. Click on the apply button and call HR directly to schedule your interview.
The last date to apply for this job is . For more details, download apna app and find Full Time jobs in Hyderabad . Through apna, you can find jobs in 64 cities across India. Join NOW!