Data Scientist
Yoseph F. Lemma · Ethiopia
Job Description
About the role
We provide data annotation and data cleaning services to AI teams building machine learning models. Our annotators label images, text, and audio at scale, and our pipelines clean and validate client datasets before delivery.
We're hiring a Data Scientist to make that work faster and more accurate. This is not a model-building role in the usual sense — you'll be building the systems that measure annotation quality, automate the repetitive parts of labeling, and catch problems in client data before the client does.
What you'll do
- Design and monitor quality metrics for annotation projects — inter-annotator agreement, gold-standard test sets, sampling-based audits
- Build pre-labeling models that give annotators a first-pass suggestion instead of a blank canvas, cutting time per item
- Implement active learning to prioritize which unlabeled items get sent to annotators next
- Build automated data-cleaning pipelines: deduplication, outlier and anomaly detection, format normalization, PII detection and redaction
- Work with client teams to translate vague requirements into concrete labeling taxonomies, annotation guidelines, and acceptance criteria
- Analyze annotator throughput and error patterns to inform training, task routing, and pricing
- Run evaluation on delivered datasets and produce quality reports clients can act on
Requirements
- 3+ years in data science, ML engineering, or a closely related role
- Strong Python — pandas, NumPy, scikit-learn — and comfortable SQL
- Practical experience with at least one of PyTorch or TensorFlow
- Solid grasp of evaluation methodology: precision/recall tradeoffs, class imbalance, sampling, agreement statistics such as Cohen's or Fleiss' kappa
- Experience working with messy, real-world data — not just clean benchmark datasets
- Able to explain technical findings clearly to non-technical clients and operations staff
- BSc in Computer Science, Statistics, Mathematics, or equivalent practical experience
Preferred
- Prior work at a data annotation, BPO, or data-services company
- Experience with annotation tooling (Label Studio, CVAT, Prodigy, or in-house equivalents)
- Weak supervision or programmatic labeling experience
- NLP or computer vision depth in at least one area
- Amharic language data experience — a real advantage for our local-language projects
- Cloud experience (AWS/GCP) and containerization
What success looks like in the first 6 months
- A live quality dashboard covering every active project
- At least one pre-labeling model deployed and measurably reducing time per item
- A repeatable data-cleaning pipeline replacing ad-hoc scripts