Wealthsimple Technologies

Software Development Manager, Observability Platform

Location

Remote (Canada)

Workplace

Remote

Type

Full Time

Salary

CAD 160,000 – 220,000

Level

Manager

Role

Engineering Manager

Posted

Jul 23, 2026

Full TimeRemoteManager

The role

Summary

Lead the founding Software Development Manager role for Wealthsimple's Observability Platform team, responsible for building production telemetry infrastructure that empowers all engineers and AI agents with logging, metrics, and tracing capabilities. This position combines strategic platform vision-setting with hands-on technical leadership, requiring deep expertise in observability architecture, event-based systems, and OpenTelemetry standards. You'll manage and grow an engineering team while defining multi-year roadmaps, establishing golden paths for observability best practices, and architecting systems that support both human and AI-driven incident investigation.

What you'll do

Team Leadership and Development: Hire, mentor, and manage software engineers on the Observability Platform team. Own career development, performance management, coaching, and foster a culture of technical excellence, collaboration, and continuous improvement while building an inclusive, high-performing engineering organization.
Observability Vision and Strategy: Define observability standards and direction for Wealthsimple, articulate a multi-year technical roadmap, and align engineering teams and leadership behind a cohesive vision. Translate strategic objectives into sequenced platform capabilities and measurable outcomes.
Events-First Architecture Design: Lead the design and implementation of a wide-event data model built on open standards such as OpenTelemetry. Architect shared SDKs, libraries, and instrumentation defaults that make consistent, rich telemetry the path of least resistance for all engineering teams across the organization.
Reliability and Detection Excellence: Drive the company's detection and reliability initiatives by reducing mean time to detect on critical user flows. Implement SLI and SLO frameworks that catch production issues before customers are impacted, moving from machine-state monitoring to customer-experience-signal-driven alerts.
AI-Agent Integration and High-Cardinality Querying: Build fast, high-cardinality query foundations and access patterns including Model Context Protocol that enable AI agents to investigate incidents, verify deployment changes, and operate in tight feedback loops alongside human engineers for accelerated incident resolution.
Pilot Validation and Experimentation: Establish a test-and-learn methodology by running pilots with select teams to validate observability approaches before organization-wide scaling. Use data-informed decision-making to prove value, iterate on implementations, and measure impact before full deployment.
Tool and Vendor Strategy: Own build-vs-buy-vs-broker decisions across the observability stack. Maximize value from existing tools like Datadog while reducing vendor lock-in, managing costs, and ensuring the technology portfolio supports long-term strategic flexibility and scalability.
Hands-On Technical Contribution: Remain actively involved in critical architectural decisions and complex technical design reviews. Contribute directly to the codebase when necessary to unblock teams, prototype innovative approaches, and maintain technical credibility while demonstrating platform engineering best practices.
AI-Assisted Development Leadership: Fluently leverage modern AI coding tools such as Claude Code and large language models to prototype solutions, navigate complex systems, and accelerate development velocity. Lead by example in raising the team's capability and comfort with AI-assisted software development practices.
Cross-Functional Collaboration: Build trust and influence across engineering teams, peer managers, and senior leadership. Communicate observability vision and technical decisions clearly to both technical and non-technical stakeholders, ensuring alignment and buy-in across organizational boundaries.

What we look for

Technical

Platform Engineering and Observability ArchitectureDemonstrated expertise in designing and implementing observability or telemetry platforms within engineering organizations. You have built or implemented the equivalent of commercial products like Datadog or Honeycomb internally, understanding logging, metrics, and distributed tracing infrastructure at architectural scale.
Event-Based Systems and High-Cardinality DataDeep hands-on experience with wide structured events, high-cardinality data modeling, columnar or streaming backends, sampling strategies, context propagation, and tag-based cardinality management. You understand the architectural trade-offs that enable asking unanticipated questions of production data.
OpenTelemetry and Open StandardsPractical, hands-on fluency with OpenTelemetry standards, instrumentation patterns, and SDK development. Experience designing SLI and SLO frameworks that drive reliability engineering practices and inform product decisions.
SDK and Platform Library DevelopmentTrack record of shipping SDKs, shared libraries, paved-road frameworks, and instrumentation defaults that other engineers build upon. Experience treating internal engineers as customers and designing APIs that prioritize developer experience and adoption.
Datadog and Observability ToolsHands-on experience with Datadog or similar enterprise observability platforms, including configuration, optimization, cost management, and strategic decisions around tool selection and vendor relationships in observability stacks.
AI and Modern Development ToolingFluency with AI coding assistants (Claude Code, GitHub Copilot, etc.) and large language models for software development. Ability to leverage AI tooling for rapid prototyping, code navigation, and problem-solving in complex systems.

Education

Bachelor's Degree in Computer Science or Related FieldFormal education in computer science, software engineering, electrical engineering, mathematics, or equivalent practical experience demonstrating equivalent knowledge in software engineering fundamentals and system design.

Experience

Software Engineering Leadership (3+ years)Minimum three years of hands-on engineering management experience with a track record of developing software engineers, delivering complex technical initiatives, and leading teams through organizational growth and technical change.
Core Software Engineering (7+ years)Seven or more years of software engineering experience with a strong background building software platforms and tools, not operations or systems administration. You apply software engineering rigor including testing, design patterns, and architectural principles to platform development.
Observability Platform ImplementationDirect hands-on experience building or implementing observability and telemetry platform capabilities. You have executed this type of roadmap at least once and understand the problem space deeply rather than learning it on the job.
Platform Foundations EngineeringExperience in platform or foundations teams shipping shared infrastructure, SDKs, and internal developer tooling. You have built the libraries and frameworks that other engineers depend on and understand the customer-centric mindset required for internal platforms.

Skills

Required skills

Software Architecture and Systems DesignExpert-level ability to architect scalable, distributed systems. Strong grasp of design patterns, trade-offs between consistency and availability, and the capacity to make sound architectural decisions that balance performance, maintainability, and cost.
Observability and Telemetry ExpertiseDeep understanding of logging, metrics, and tracing paradigms. Proficiency in designing high-cardinality telemetry collection, understanding cardinality explosion problems, sampling strategies, and query optimization for production observability systems.
Engineering Management and LeadershipProven ability to hire, mentor, and develop software engineers. Experience with performance management, clear feedback delivery, career coaching, and creating inclusive, psychologically safe team environments that attract and retain talent.
Strategic Vision and ExecutionCapability to define technical direction, create multi-year roadmaps, sequence work intelligently, and bring diverse stakeholders along through clear communication and demonstrated impact. Ability to translate vision into actionable work and measurable outcomes.
OpenTelemetry and Instrumentation StandardsWorking knowledge of OpenTelemetry standards, semantic conventions, and instrumentation best practices. Experience designing SDKs and instrumentation libraries that encourage adoption and consistent telemetry collection across teams.
Hands-On Coding and Technical DepthAbility to write production code in modern programming languages, review complex code for correctness and design, and jump into codebases to prototype solutions or unblock teams. Maintains technical credibility and current skills despite management responsibilities.
Data-Driven Decision MakingComfort framing work as testable hypotheses, designing experiments and pilots to validate approaches, and making decisions based on evidence. Willingness to iterate, measure, and adjust rather than commit everything upfront.
Cross-Functional CommunicationAbility to communicate complex technical concepts clearly to non-technical audiences including executives, product managers, and business stakeholders. Strong interpersonal skills for building trust and influencing across organizational boundaries.

Nice to have

Regulated Industry and Fintech ExperiencePrior experience working in regulated environments such as fintech, financial services, healthcare, or compliance-heavy industries. Understanding of regulatory considerations, security requirements, and risk management in production systems.
Kubernetes and Progressive DeliveryHands-on experience with Kubernetes orchestration and progressive delivery frameworks such as Argo Rollouts. Knowledge of deployment patterns, canary releases, and automated rollback strategies for production reliability.
Columnar Data Stores and AnalyticsPractical experience with columnar databases such as ClickHouse, Apache Druid, or similar high-cardinality analytics platforms. Understanding of query optimization, data compression, and time-series analysis at scale.
LLM and AI Agent ObservabilityExperience designing observability solutions for large language models and AI agent workloads. Understanding of unique observability challenges in generative AI systems, token tracking, and cost monitoring for LLM applications.
Model Context Protocol (MCP)Familiarity with Model Context Protocol standards and patterns for enabling AI agents to interact with external systems and data. Experience building or integrating MCP-compatible tools for agent-based workflows.
Argo and GitOps PatternsExperience with Argo CD and GitOps deployment patterns, including declarative infrastructure management, automated synchronization, and continuous deployment practices.

Compensation & benefits

Salary

CAD 160,000 – 220,000 (annual)


Apply for this position

You'll be redirected to the company's application page