Job Description
Key responsibilities
- Design, build, and maintain ETL pipelines for large volumes of imagery and video data.
- Conduct exploratory data analysis to understand dataset quality, coverage, and gaps.
- Clean, structure, and organise datasets so they’re readily usable by the ML team.
- Build fast prototype visualisations to surface data issues and communicate findings to the team.
- Write and optimise SQL to query, transform, and manage data at scale.
- Collaborate closely with ML engineers to understand what “training-ready” data actually requires.
- Maintain documentation of datasets, pipelines, and data quality processes.
Your experience and skills
- Tertiary qualification in computer science, engineering, data science, or a related discipline.
- 3+ years of experience in data engineering, data analysis, or a related field, ideally working with imagery or video data.
- Demonstrated experience building ETL pipelines, ideally with image or video data.
- Strong SQL skills, with experience managing and querying large datasets.
- Comfortable prototyping quickly in Python, using libraries such as Pandas, Matplotlib, or similar for exploratory analysis and visualisation.
- Experience working with unstructured or semi-structured data (images, video, sensor data).
- Understanding of what makes a dataset ML-ready — labelling, balance, quality checks.
- Eligible for security clearance.
Desirable:
- Exposure to computer vision or ML data pipelines specifically.
- Experience with cloud data storage/processing (e.g. S3, BigQuery, or similar).
- Familiarity with OpenCV or similar image-processing libraries.
- Experience with data versioning or annotation tooling.
Are you interested in this position?
Apply by clicking on the “Apply Now” button below!
#GraphicDesignJobsOnline
#WebDesignRemoteJobs
#FreelanceGraphicDesigner
#WorkFromHomeDesignJobs
#OnlineWebDesignWork
#RemoteDesignOpportunities
#HireGraphicDesigners
#DigitalDesignCareers
# Dynamicbrand guru