Remote
(Canada) Principal ML System Engineer
About this role
Team Summary This team will serve as the product owner for the machine learning platform capabilities within PointClickCare, working closely with other engineering teams across the organization to identify, build and support traditional machine learning (ML) and hybrid ML/LLMsolutions. This centralized team with deep specialization will closely integrate with key horizontal partners to ensure delivery of safe, scalable, and high-impact AI products.
Job Summary The Principal AI Machine Learning Platform Engineer will set the technical vision and strategy for the machine learning platform that powers ML and generative AI development across PointClickCare, partnering with Product and Engineering leadership to align that direction with product and business goals. As the technical authority for the ML platform, the Principal Engineer will define the reference architectures, standards, and roadmap for the pipelines, tooling, and infrastructure used for model training, deployment, serving, and monitoring, and will provide technical leadership and mentorship to raise the engineering bar for ML systems company-wide.
Key Responsibilities Partner with product and engineering leadership to translate business and product objectives into a multi-quarter technical strategy and roadmap for the ML platform. Define the reference architectures and standards for scalable data and ML pipelines spanning model training, evaluation, deployment, and serving that engineering teams across the organization build upon. Set the direction and best practices for MLOps across the company — including CI/CD for models, model registry, feature stores, and experiment tracking — and drive build-vs-buy decisions for core platform components.
Establish the practices and architecture for reliability, observability, and performance of ML systems in production, including monitoring, alerting, and automated remediation. Establish the security architecture for the ML platform, including authentication, role-based access control, audit logging, and compliance monitoring, and ensure adoption across teams. Define secure, cost-efficient integration and infrastructure patterns for connecting the platform with existing systems, APIs, and data sources at scale.
Provide technical leadership and mentorship across engineering teams, guiding senior engineers and influencing the org-wide technical roadmap for ML infrastructure. Qualifications & Skills Expert level in Python and Java with strong software engineering fundamentals. Deep experience designing and building ML platforms and ML Ops workflows at scale, familiarity with tools such as MLFlow, Kubeflow, Ray, and model-serving frameworks or equivalents.
Extensive experience with cloud platforms (AWS, Azure, and/or GCP), containerization, and orchestration (Docker, Kubernetes). Demonstrated track record of setting technical direction and driving org-wide technical initiatives across multiple teams. Preferred Bachelor’s degree or higher in Computer Science, Machine Learning, or a related field. Sufficient familiarity with Azure Machine Learning components, Databricks processing and serverless environments, and ML Frameworks to support strategic decision making Demonstrable history of leading and sustaining build out of critical cross-team systems Experience implementing security at scale including role-based access control, multi-factor authentication, network security best practices, and compliance monitoring.
Experience optimizing large model training and inference (including LLM serving) for performance and cost. #LI-AJ1 #LI-remote
Source listing: lever_pointclickcare