About the Role
You will train and adapt models using Everfit’s product data and feedback, then take them into production. You will own the technical work from problem definition and data preparation through training, evaluation, release, and monitoring.
Some problems call for a small classifier or ranking model. Others may benefit from adapting a pretrained language or vision model. You will assess where training can improve quality, personalization, speed, or cost, recommend an approach, and test it against the current product.
You will work with Applied AI Engineers, Backend Engineers, Data Engineers, and product partners. Your focus will be training data, learned representations, model adaptation, evaluation, and the model lifecycle. Applied AI colleagues contribute the surrounding context, tools, and application workflows. Together, you will make sure model improvements reach users and solve the intended problem.
Responsibilities
Choose the problem and approach
- Work with product and domain partners to define the user problem, the decision a model should support, and how to measure a better result.
- Investigate whether errors come from the model, labels, missing context, reference data, or product integration.
- Compare the current workflow with simple baselines, existing models, and training-based approaches. Recommend the approach the evidence supports.
Build data that supports learning
- Work with data engineers and domain experts to prepare training and evaluation datasets. Define labels, review disagreements, and check missing records, duplicates, and uneven coverage.
- Turn product feedback into useful training examples. Check that a coach edit, click, purchase, or correction means what the training objective assumes.
- Keep training and evaluation separate. Use time, user, coach, or workspace boundaries where needed, and prevent future information from entering earlier examples.
- Test unfamiliar users and items, sparse histories, and important failure cases. Distinguish missing activity records from inactivity when working with engagement data.
Train, adapt, and evaluate models
- Build task-specific models and adapt pretrained models where appropriate. Methods may include supervised fine-tuning (SFT), Direct Preference Optimization (DPO), distillation, and parameter-efficient fine-tuning (PEFT), including low-rank adaptation (LoRA).
- Choose metrics that match the task. Examine ranking quality, per-category classification errors, language quality, or recognition and portion errors as appropriate.
- Check uncertainty and define when a model should abstain, request more information, or use a fallback. Review the consequences of different mistakes with product partners.
- Version datasets, labels, transformations, model artifacts, and evaluation settings so another engineer can reproduce the result.
Deliver and improve the product
- Work with backend, data, and DevOps colleagues to deploy and serve models in batch or online services. Keep training and serving transformations consistent and test the full path to the result users receive.
- Compare model quality with response time, memory use, and operating cost. Diagnose failures and make improvements without hiding regressions in an average score.
- Run controlled product experiments with the team. Measure useful outcomes alongside offline scores, and document what the evidence supports.
- Monitor model and data performance after release. Validate updates before deployment, maintain a rollback path, and turn confirmed failures into regression tests.
- Protect coach and client data in datasets, inference, and logs. Follow workspace access boundaries and approved retention and deletion rules.
Requirements
- Good English communication - Clear communication about technical choices and limitations.
- 5+ years of relevant experience in Machine Learning / Applied ML / ML Engineering, with strong hands-on experience building and deploying ML models to production.
- Strong Python and practical SQL skills.
- Strong foundations in probability, statistics, and machine learning. You can apply regularization, recognize overfitting and sampling bias, and explain how training objectives affect model behavior.
- Experience building and tuning conventional ML models, such as linear or tree-based models, and comparing them with neural or pretrained approaches. You can explain why the complexity of a model is justified.
- Experience preparing training data and developing features or embeddings, including work with missing values, imbalanced labels, noisy feedback, and limited examples.
- Experience fine-tuning Transformer models (e.g., via Hugging Face) for downstream tasks and domain-specific datasets.
- Experience choosing baselines, validation splits, and task-appropriate metrics; preventing data leakage; and analyzing errors across important user groups and cases. You can assess whether an improvement is meaningful for the product.
- Experience training or adapting models with tools such as PyTorch, TensorFlow, or scikit-learn, taking them into production, and measuring their performance after release. You can explain the decisions you owned, the failures you investigated, and the results.
- Strong practical depth in at least one relevant area: recommendation and ranking, classification, natural-language processing and language-model adaptation, or computer vision and multimodal learning.
- Experience with reproducible training pipelines, experiment tracking, dataset and model versioning, and containerized deployment using Docker or equivalent tools. You can work with automated testing and release workflows.
- Good software engineering habits, including version control, readable code, tests, and documentation.
- Turn ambiguous product problems into practical experiments, define data and evaluation requirements, and make evidence-based decisions about deployment.
- Own model quality and production performance, balancing quality, latency, cost, and maintainability while driving follow-up improvements.
- Review technical work, mentor earlier-career engineers, and lead the investigation and resolution of model or data failures in your area.
Preferred Skills
- Multi-label classification, image segmentation, pose estimation, or learning from time-series data.
- Efficient inference, model compression, or cloud-based training and deployment on platforms such as AWS, Google Cloud, or Azure.
- Designing human evaluation and annotation workflows for subjective or ambiguous labels.
- Health, fitness, nutrition, wearable, or other longitudinal user data.
- Using AI coding tools effectively and checking their output.
- Relevant research or open-source contributions.
Benefits