Wealthsimple Technologies

Staff Software Engineer, Observability Platform

Location

Remote (Canada)

Workplace

Remote

Type

Full Time

Salary

CAD 210,000 – 280,000

Level

Staff

Role

Staff Software Engineer

Posted

Jul 23, 2026

Full TimeRemoteStaff

The role

Summary

Staff Software Engineer, Observability Platform at Wealthsimple Technologies is a senior technical leadership role focused on architecting and building the foundational observability infrastructure serving over 4 million users. This position requires 8+ years of hands-on software engineering experience with deep expertise in observability platforms, high-cardinality telemetry systems, and open standards like OpenTelemetry, combined with demonstrated proficiency in modern development practices and AI-assisted coding tools. The ideal candidate will design event-first architectures, establish instrumentation standards, and mentor senior engineers while building the systems that make production environments legible to both human engineers and AI agents.

What you'll do

Design and build events-first foundation: Own the architecture of the wide-event data model and design high-throughput ingest, storage, and query pipelines. Build high-cardinality and columnar or streaming systems that enable engineers to ask arbitrary questions about production environments, ensuring both performance and flexibility at scale.
Architect and ship instrumentation SDKs: Design and deliver instrumentation libraries, shared SDKs, and sensible defaults that make rich, consistent telemetry the default path for every engineering team. Drive adoption through golden paths and best-practice patterns that reduce engineering friction and standardize observability practices across the organization.
Establish technical standards and conventions: Define instrumentation conventions, naming schemes, tagging strategies, sampling approaches, and context-propagation practices using open standards such as OpenTelemetry. Codify these standards so humans and AI agents share one common language for understanding system behavior and performance.
Enable AI agent-driven incident investigation: Build fast query foundations and access patterns including protocols such as MCP that allow AI agents to investigate incidents independently, verify their own changes, and operate in tight feedback loops alongside human engineers. Design for both human and machine consumption of observability data.
Drive adoption of AI-assisted development: Fluently use AI coding tools such as Claude and modern LLMs to prototype, build, and navigate large systems. Lead by example in demonstrating how to maintain quality while accelerating development velocity, and help the team raise its engineering standards for building with AI tooling.
Navigate and improve complex codebases: Jump into unfamiliar systems, quickly form accurate mental models, and make significant, well-reasoned changes with measurable impact. Apply sophisticated debugging and refactoring techniques to codebases you did not build, reducing technical risk while improving system reliability.
Validate approaches through experimentation: Pilot new observability approaches with one or two partner engineering teams, measure quantifiable results, and scale successful patterns across the organization. Use data-driven decision making to prioritize platform investments and prove business value before full-scale rollouts.
Mentor and elevate engineering capability: Mentor senior engineers, review and provide feedback on complex system designs, and lead cross-team technical initiatives through influence and technical credibility rather than formal authority. Raise the overall engineering bar through knowledge sharing and collaborative problem-solving.

What we look for

Technical

Observability and telemetry platform architectureDirect, hands-on experience designing and building complete observability platforms including instrumentation libraries, data pipelines, and query systems. Shipped production SDKs and platforms that other engineers depend on, with deep understanding of architectural tradeoffs.
High-cardinality event systemsExpert-level understanding of wide structured events, columnar or streaming backends, cardinality management, sampling strategies, tagging schemes, and context propagation. Experience ingesting and querying telemetry at scale with awareness of performance, cost, and usability tradeoffs.
OpenTelemetry and open standardsHands-on proficiency with OpenTelemetry specification or equivalent instrumentation standards. Experience designing SLIs and SLOs, implementing instrumentation conventions, and working with vendor-agnostic observability approaches that minimize lock-in.
Multi-language software developmentCurrent hands-on coding ability in multiple programming languages such as Kotlin and Ruby. Comfortable working across diverse technology stacks and regularly shipping production code, not just architecture design.
Large-scale data systemsProduction experience building and operating systems that handle high-volume event ingestion, storage optimization, and query performance. Understanding of streaming architectures, columnar data structures, and query optimization for analytical workloads.
AI coding tools and modern LLMsFluent daily use of AI coding assistants such as Claude Code or similar tools. Demonstrated ability to maintain code quality while leveraging AI for acceleration, with clear judgment about appropriate use cases and limitations.

Education

Software engineering fundamentalsStrong foundation in computer science, software engineering, or equivalent practical experience. Understanding of systems design, distributed systems principles, data structures, and software architecture patterns.

Experience

Senior software engineering experienceTypically 8+ years of professional software development experience with a strong background in building production systems, platforms, and tools. Experience should emphasize software engineering rigor rather than operations or systems administration.
Observability platform expertiseMultiple years working on observability, monitoring, or telemetry systems. Experience shipping logging, metrics, or distributed tracing solutions that served as internal platforms for other engineering teams.
Navigating ambiguous systemsProven ability to enter unfamiliar codebases, quickly understand complex system behavior, and make substantial, safe changes with high impact. Comfortable with codebases you did not author and systems with incomplete documentation.
Cross-functional influenceExperience working with multiple engineering teams, communicating complex technical tradeoffs clearly, and driving adoption of new approaches through influence and credibility rather than hierarchical authority.

Skills

Required skills

Observability platform designArchitecting end-to-end observability systems including data ingestion, storage, and query layers
Event-driven architectureDesigning systems around high-volume structured events and wide-cardinality data models
OpenTelemetryImplementing instrumentation standards and conventions using OpenTelemetry specification
Distributed systemsUnderstanding of distributed tracing, context propagation, and trace correlation across service boundaries
High-performance database systemsExperience with columnar stores, time-series databases, or streaming data systems at scale
SDK and library designBuilding reusable instrumentation SDKs, shared libraries, and developer-focused platform tooling
Python or Kotlin or RubyMulti-language programming proficiency with current production coding experience
AI-assisted developmentEffective use of Claude, GitHub Copilot, or similar AI coding tools in daily workflow
Systems design and architectureDesigning scalable, reliable systems with attention to performance, cost, and maintainability
Technical communicationExplaining complex system designs and tradeoffs clearly to both technical and non-technical audiences

Nice to have

ClickHouse or columnar database experienceHands-on experience with ClickHouse, Apache Druid, or similar columnar stores for analytics workloads
Kubernetes and progressive deliveryFamiliarity with Kubernetes operations, Argo Rollouts, Flux, or other progressive deployment patterns
Fintech or regulated environment experienceBackground working in financial services, cryptocurrency, or other regulated industries with compliance requirements
LLM and agent observabilityExperience building observability solutions specifically for large language models, AI agents, or GenAI workloads
SLO and SLI designPractical experience designing Service Level Objectives and Indicators aligned with business metrics
Model Context Protocol (MCP)Familiarity with Model Context Protocol or similar standards for AI agent-system integration
Query optimizationDeep expertise in query optimization, cost reduction, and cardinality management for high-volume systems

Compensation & benefits

Salary

CAD 210,000 – 280,000 (annual)

Stock options

Available

Benefits

Comprehensive health and life insurance

Top-tier health benefits coverage including dental, vision, and comprehensive life insurance protection for you and your family

Long-term employer-matched savings

Group savings program through Wealthsimple for Business with employer matching contributions to help you build long-term financial security

Flexible time off

20 vacation days annually, 4 dedicated wellness days, and unlimited sick and mental health days to prioritize your wellbeing

Work anywhere for 90 days

Flexibility to work outside of Canada for up to 90 days per year, enabling remote work arrangements and international mobility

Employee resource groups

Active community groups including Rainbow (2SLGBTQ+), Women of Wealthsimple, and Black at Wealthsimple for connection and support

Hybrid work environment

Collaborative hybrid work arrangement with flexibility to balance office collaboration and remote work across North American offices

Accessibility accommodations

Committed support for accessibility needs throughout employment, with proactive accommodation planning and ongoing support


Apply for this position

You'll be redirected to the company's application page