As Lead Architect I'm responsible for Whispp's infrastructure and software architecture, with a focus on on-device AI. I designed the fault-tolerant calling stack behind a product with about 9,000 registered users: the AI processing servers operate independently of the orchestration and load balancing layers, so calls in progress survive infrastructure failures. I'm also the on-call engineer for this business-critical infrastructure.
On the device side, I develop Whispp's desktop application in Rust (Tauri), Vue, C++ and ONNX Runtime, with cross-language FFI modules for the performance-critical paths. By profiling execution providers and selecting the right backend, I brought inference latency down from 30 ms to 8 ms (averaged over 100 inferences on Ryzen AI 9 edge hardware). I also explored protecting the model itself, with proofs of concept for watermarking the model and for an obfuscation layer that hardens it against reverse engineering. These stayed proofs of concept and didn't make it into the product.
To keep it all running, I set up the observability stack (Grafana, Prometheus, Loki and Promtail) with custom Golang instrumentation for log tracing and alerting. On top of that runs synthetic monitoring: every 30 minutes, a real call is placed in each region, audio is pushed through the pipeline and recorded at the receiving end, and the result is automatically checked against amplification and latency thresholds.
Team awards
- Software architecture
- On-device AI
- Rust
- Tauri
- C++
- ONNX Runtime
- Vue
- TypeScript
- Python
- Golang
- Grafana
- Prometheus
- Loki