Data Scientist – RL (Representation Learning)
Organisation Overview
Our Client is a geospatial intelligence business focused on turning high-resolution aerial imagery into meaningful, actionable insights. By combining advanced machine learning with large-scale geospatial data, they help organisations find signals, patterns, and opportunities hidden in complex environments.
This is an energetic, technically driven environment where continuous improvement matters. You’ll be working within a cross-functional team culture that values strong engineering practices, measurable experimentation, and building reusable capabilities that scale.
Role Summary
Our Client is expanding its core representation learning capability to support multiple downstream use cases, including search, similarity matching, clustering, retrieval, anomaly detection, and predictive modelling.
As a Senior Data Scientist in this area, you’ll play a strategic role in designing embedding models that capture semantic, spatial, and structural information from large imagery and geospatial datasets. Your work will help establish foundational infrastructure across products-bridging data science, engineering, and product teams to deliver high-impact capabilities that improve performance and accelerate innovation.
This is a hands-on role for someone who can translate deep representation learning ideas into robust, production-ready systems.
Responsibilities
- Lead the end-to-end lifecycle of embedding-based systems, from problem definition and dataset curation through model design, training, evaluation, and iteration
- Develop embedding models using self-supervised, contrastive, metric learning, or multimodal learning approaches
- Create spatially informed, semantically meaningful representations from aerial imagery and structured geospatial data
- Build and optimise representation learning architectures such as CNNs, Vision Transformers (ViTs), multimodal encoders, and graph-based models
- Improve embedding quality through loss function design (for example triplet loss, contrastive loss, InfoNCE) and effective sampling strategies
- Define and implement quantitative evaluation approaches (for example retrieval accuracy, clustering quality, and downstream task uplift)
- Run rigorous experimentation, including ablation studies and representation analysis, to understand what drives performance
- Optimise embeddings for dimensionality, storage efficiency, and inference latency
- Collaborate with engineering teams to productionise embedding models and expose them through reliable APIs
- Support scalable indexing and similarity search solutions (for example ANN-based systems and vector search approaches)
- Help shape standards for reproducibility, monitoring, and embedding drift detection
- Partner with product and domain stakeholders to identify high-value embedding-driven opportunities and convert them into measurable learning objectives
- Stay up to date with advances in representation learning, foundation model adaptation, multimodal embeddings, and large-scale retrieval systems
Essential Skills & Experience
- Degree level qualification (BS/MS or PhD) in Computer Science, Machine Learning, Applied Mathematics, or a closely related quantitative discipline
- 5+ years’ experience applying advanced machine learning, including at least 2+ years in a senior or lead capacity
- Proven experience building embedding models using contrastive, metric learning, or self-supervised approaches
- Strong hands-on expertise with deep learning frameworks such as PyTorch or TensorFlow
- Experience working with large-scale image and/or multimodal and geospatial datasets
- Strong foundation in linear algebra, optimisation, probability, and statistical modelling
- Experience working with engineering teams to deploy models into production environments
Desirable Skills & Experience
- Experience creating multimodal embeddings (for example combining imagery with text and/or structured spatial data)
- Familiarity with large-scale similarity search tooling and vector databases (for example FAISS and similar approaches)
- Experience adapting or fine-tuning foundation models for domain-specific embedding tasks
- Knowledge of GIS concepts and geospatial data structures
- Publications or meaningful involvement in representation learning or computer vision research communities
Call to Action
If you’re excited by the chance to build reusable, embedding-first infrastructure that powers multiple high-impact applications, we’d love to hear from you. Please submit your CV for consideration.