Download PDF

Staff Engineer · Platform, Infrastructure & Developer Tooling

Boulder Creek, CA · linkedin.com/in/tedski · github.com/tedski


Summary

Staff engineer specializing in platform, infrastructure, and developer tooling with a decade of SRE experience across large-scale distributed systems. I build things that make engineers more effective, and I treat user feedback, adoption metrics, and documentation as first-class engineering concerns measured with the same rigor as uptime. My leadership instincts were forged in the U.S. Coast Guard and as a Firefighter, where the Incident Command System shaped how I approach reliability: clear ownership, documented procedures, structured on-call culture, and coordination under pressure. Currently building developer experience tooling at Weights & Biases.


Experience

Staff Software Engineer — Developer Experience

Jul 2025 – Present · Weights & Biases (Coreweave) · Remote

Software engineer on the Developer Experience (DevEx) team, building and improving the platform, CI systems, and local development tooling that support engineers throughout the software development lifecycle. Focused on enhancing developer productivity and streamlining engineering workflows.

Career Break — Personal Development & Projects

Apr 2025 – Jul 2025 · Boulder Creek, CA

Deliberate sabbatical focused on agentic AI development, home automation, and personal infrastructure projects. See github.com/tedski.

Staff Site Reliability Engineer → Staff Software Engineer

2019 – Apr 2025 · LinkedIn · Remote

Capacity Engineering (2021 – 2025)

  • Tech lead for Platform Experience within Capacity Engineering, owning all user-facing tooling and feedback pipelines serving 1,000+ engineers across tens of thousands of services; contributed as a peer lead on Forecasting & Measurements, helping shape direction across sub-teams through structured user interviews and roadmap planning.
  • Drove rightsizing and predictive auto-scaling adoption to 85% of services by building web and CLI interfaces shaped by user research — SREs preferred CLI, product engineers preferred web.
  • Evangelized model outputs to service owners, building confidence in automated recommendations and dispelling skepticism to sustain adoption across the org.
  • Trusted by directors and VPs as the go-to analyst for capacity investigations, using Spark, Trino, and Python to validate forecasts, diagnose anomalies, and independently verify performance improvement claims before organizational commitments were made.
  • Designed a standardized metrics framework for Capacity Engineering, classifying signals into user, capacity, and health indicators to improve decision-making and dashboard consistency across the org; began implementing the ingestion pipeline before departing.
  • Authored internal observability guidance adopted across Capacity Engineering, standardizing Grafana-compatible metrics and improving telemetry consistency across the team.
  • Led Code Yellow, a cross-org infrastructure optimization initiative that recovered over 2,000 machines in unused capacity; also led the SRE track onboarding bootcamp for incoming engineers.
  • Maintained an active mentorship practice carrying up to five concurrent mentees, contributing to multiple promotions across the SRE org.

Resilience Engineering — Waterbear (2019 – 2021) · promoted to Staff 2019

  • Founding engineer on the Resilience Engineering team; co-built LinkedOut, LinkedIn’s internal failure injection platform designed to shift resilience testing left and give developers the tooling to validate their own services.
  • Contributed backend and UI features to LinkedOut’s web platform and maintained the Chrome extension, evangelizing the full platform — automated testing, group failure injection, and ad-hoc testing workflows — to drive adoption across engineering teams.
  • Led development of LinkedOut’s automated test platform: scheduled Selenium-based runs across all failure modes at every call graph edge, with a custom image-comparison algorithm to surface visual regressions automatically.
  • Served as tech lead for user research and roadmap prioritization; quarterly ambassador program embedded rotating engineers as co-designers, shifting LinkedOut’s culture from top-down chaos tooling to developer-owned resilience testing.
  • Greenfielded a load balancer tuning tool using PySpark over historical Hive traffic data to recommend optimal tunable values per service, reducing manual tuning toil across infrastructure teams.

Senior Site Reliability Engineer

Jun 2016 – 2019 · LinkedIn · San Francisco Bay Area

  • Bridged SRE and product engineering for the Content Org by shifting operational ownership to developers through tooling training and self-service runbooks, eliminating SRE release gatekeeping and increasing team autonomy and release velocity across the Publishing and Pulse stacks.
  • Transitioned to the founding Resilience Engineering / Waterbear team in 2018, contributing to early LinkedOut platform development ahead of Staff promotion.

Previous Experience — Systems & DevOps Engineering

2011 – 2016 · Lucid Design Group, eBay Inc, Shopping.com

  • Infrastructure automation, configuration management, and DevOps culture across a startup and two large-scale e-commerce environments. Early practitioner of IaC, containerization, and continuous delivery.

Various Roles — Maritime Operations & Systems Administration

2000 – 2009 · U.S. Coast Guard · San Francisco Bay Area / New York

  • Response missions: nine-year veteran conducting search and rescue, law enforcement, and homeland security operations in the San Francisco Bay Area — 800+ SAR missions, small team leadership, and sustained performance under high-pressure conditions.
  • Prevention missions: Waterfront Facility Inspector at the Port of New York, leading a team of 12 managing security and safety audits across 47 facilities; implemented a standardized biometric identification system at 50+ sites while maintaining uninterrupted flow of $132B in annual trade.

Areas of Expertise

Site Reliability Engineering · Developer Tooling & Platform · Resilience & Chaos Engineering · Agentic AI Development · Capacity Engineering · Distributed Systems Observability