Overview
We are actively seeking a highly skilled and experienced Senior AI/ML Engineer with a focus on MLOps to join our innovative team. If you have 6 to 10 years of hands-on experience in the AI/ML space and a passion for driving technological advancements, this role is for you.
Key Technologies: Python, NumPy, Pandas, PyTorch, Docker, Kubernetes, Git, Jenkins, Azure DevOps, AWS SageMaker, Prometheus, Grafana
Specific Skills
- Python Expertise: proficiency in Object-Oriented Python
- Data Science (Jupyter Notebooks): demonstrated expertise in data science, including analysis and modeling using Jupyter Notebooks
- Deep Learning (PyTorch): proven experience in deep learning, particularly with PyTorch, and familiarity with other frameworks
- Good LLM knowledge: solid understanding of Natural Language Processing (NLP) and Language Models (LLM)
- Any successful implementation of GenAI (LLMs) on custom data is preferred
- Bachelor's/Master's in Data Science is preferred
Responsible For
- MLOps Implementation (Docker, Kubernetes, Azure DevOps, AWS SageMaker): lead the implementation of MLOps practices, ensuring seamless integration of machine learning models into production systems using containerization and orchestration, and MLOps tooling from both Azure and AWS
- Code Development (Python, NumPy, Pandas): develop and maintain scalable and efficient Python code for machine learning applications, using NumPy and Pandas for effective data manipulation and analysis
- Collaboration (Git): collaborate with cross-functional teams to understand business requirements and seamlessly integrate machine learning solutions into software applications, using Git for version control
- DevOps Integration (Jenkins, GitLab): work closely with DevOps teams to streamline deployment processes, implementing CI/CD practices with tools like Jenkins or GitLab
- Observability (Prometheus, Grafana, Azure Monitor, AWS CloudWatch): focus on fine-tuning models and identifying data anomalies, implementing observability tooling for monitoring and troubleshooting across Azure and AWS
- Model Evaluation (TensorBoard): implement model evaluation tools such as TensorBoard to ensure models meet performance criteria
- Documentation (Confluence, Markdown): create comprehensive documentation for code, models, and deployment processes
- Training and Knowledge Sharing: provide training and knowledge-sharing sessions to team members on best practices in MLOps and Python coding
Job NatureFull Time
Experience6 to 10 years
Job LocationRemote, USA
Job LevelSr. Position