Round 1: Technical Screening
These questions will check your baseline fit — your skills in SQL, Python, ETL and cloud (especially Azure) as per the job spec.
SimplyHired
+3
careers.specsavers.com
+3
careers.specsavers.com
+3
Sample questions:
What SQL functions or window functions have you used recently, and what problem did they solve?
In Python, how would you remove duplicates from a list of dictionaries (or JSON records)?
Explain what an ETL pipeline is and describe one you’ve built.
What experience do you have with Azure (e.g., Azure Data Factory, Azure Databricks, Azure SQL)?
careers.specsavers.com
+1
How do you ensure data quality and trustworthiness in a pipeline?
Round 2: Deep Technical
This round delves deeper into architecture, performance, scalability, and your hands-on skills. Expect scenario or problem-solving questions.
Sample questions:
Walk through how you’d design a data pipeline on Azure to ingest, transform and load customer behavioural data into a data warehouse.
What schema design would you choose for a high-volume fact table and why (star vs snowflake)?
How do you optimise performance in Spark / Databricks (or Azure Databricks)?
Given a table of timestamped events, write a SQL query to compute the 7-day moving average of events per user.
You discover the pipeline output has 10% unexpected NULLs — how do you debug root cause and fix it?
Round 3: Technical + Behavioural
Here the interview will combine advanced technical questions and behavioural questions about collaboration, stakeholder management, culture fit (reflecting Specsavers’ values and teamwork emphasis).
Sample technical questions:
How would you migrate an on-premises legacy database to Azure, while minimising downtime and ensuring data integrity?
Compare different table storage formats (e.g., Parquet vs CSV vs ORC) and when you’d choose each.
Describe how you would build a monitoring/alerting system for data pipelines in production.
Sample behavioural questions:
Tell me about a time you worked with a non-technical stakeholder to understand their requirements and translated that into a data solution.
Describe a situation where you found a data issue late in a project — how did you handle it and what did you learn?
How do you stay up to date with new data technologies and ensure your team adopts good practices?