Azure archival tier for Zmanda Pro
A customer needed cheaper long-term retention in Azure Blob. I owned the research and delivery: proved that archiving whole repositories was infeasible, then designed selective archival with Cool-tier retention so recovery stays available on demand.
Built StorageManager's token-authenticated provisioning API, the Azure Function App wrapper and a Cool-to-Archive lifecycle ensurer in Go. Merged with zero review changes, backed by a costed savings case.
System sketch
- Zmanda Probackup repositories
- StorageManager APItoken-authenticated provisioning
- Azure Function AppGo wrapper
- Blob Cool tierrecovery on demand
- Lifecycle ensurerCool to Archive
6 figurescustomer deal landed, on a 5x costed savings case
Salesforce licence reconciliation backend
Backend for an enterprise software-asset programme: X.509 and RS256 JWT-bearer OAuth across two Salesforce orgs, a paginated SOQL ingestion pipeline, a config validator, a parse-dispatch whitelist and count-mismatch alerting. Multi-year contract entitlements now drive the cost-savings reporting, where the recurring monthly saving projects into six figures by the next renewal.
Surfaced silent reconciliation defects before they reached dashboards, including identifier mismatches, a seat-collapse coalesce bug and soft-failing pagination, and wrote the clone-safe capture and restore plan that preserved the full build through a prod-to-dev environment refresh.
System sketch
- Two Salesforce orgsentitlement source
- JWT-bearer OAuthX.509 / RS256
- Paginated SOQL ingestconfig validator
- Parse-dispatch whitelistcount-mismatch alerts
- Cost-savings reportingmulti-year entitlements
~1.2Kentitlement rows reconciled across two orgs, defects caught before dashboards
Infrastructure discovery and self-healing agent
An AI-driven agent for 200+ telephony systems and workloads across three departments. It scans subnets, fingerprints services and reads long-running logs through service tracers to predict traffic and detect degradation.
Failing systems are recovered automatically with exponential-backoff retries. Every service is mapped to its owner for instant downtime alerts, and live status is surfaced on a monitoring dashboard.
System sketch
- Subnet scandiscovery
- Service fingerprintwhat is running
- Log tracerstraffic and degradation
- Auto-recoveryexponential backoff
- Owner mapalerts and dashboard
200+systems discovered, monitored and auto-recovered
More shipped
-
Claude Code plugin marketplace
Built the organisation's plugin marketplace and authored smart-pr, a multi-surface unit-test generator and Go best-practice skills.
Demoed live to ~50 engineers
-
Distributed backup manager
Python manager on Linux wrapping Restic, MySQL and Oracle RMAN behind Celery queues for multi-channel point-in-time recovery. The first customer onboarded on it signed a five-figure two-year contract.
Restores 66% faster
MySQL docsOracle RMAN docs
-
Fleet monitoring and cloud automation
WebSocket health checks, custom Prometheus exporters and React dashboards for 300+ Linux workloads, with Azure provisioning and CI/CD automated through Ansible.
80% less manual monitoring, cloud spend down 30%
-
RAG support chatbot
Retrieval-augmented chatbot and Python backend connectors for Zmanda support.
Tickets down 70%, response from 24 hours to under a minute
Try it
-
Distributed ingestion platform
Autonomous rate-limit guards and throughput tuning, designed alongside multi-tenant cloud architecture work.
8,500+ concurrent writes sustained
-
Port monitoring APIs and service manager
Async-safe REST APIs and a Linux system service manager for Zmanda, with API contracts and SLAs agreed with product stakeholders.
Startup latency down 30%
Services docs