Muhammad Ahmer Saleem — AI Systems Architect · Nixense Vixion

AI systems that survive contact with production.

Eleven years. 80+ production deployments — inside P&G's retail operations, Edwards Lifesciences' clinical research, Canva's creative platform, and Shotdeck's 3M-asset cinematic search. I find the constraint blocking your AI from shipping — then architect, build, and hold the system that clears it.

Every inquiry answered within 24 hours
Muhammad Ahmer Saleem
Muhammad Ahmer Saleem
Staff/Principal AI Engineer
Agentic Systems · RAG & Retrieval Multimodal AI · Computer Vision
Trusted in production by
Shelfr Shotdeck Edwards Synthesys Biscuit AI Trusted by enterprises at scale
01 / The Binding Constraint Method The discipline behind 80+ deployments

Most AI dies somewhere between the demo and the deploy.

Models that score brilliantly in a notebook collapse against the real world — regulators, latency budgets, unit economics, hardware limits. Especially in regulated and high-stakes environments, where I do most of my work. Eleven years of production systems reduce to a single discipline:

Isolate the binding constraint before any model decision.
i.

Diagnose

Find the one non-negotiable — compliance, latency, cost, or hardware — that decides whether the system can exist in production at all.

ii.

Architect backward

Design from the constraint toward the model. Retrieval strategy, data pipeline, and infrastructure are chosen to satisfy it — model choice comes last.

iii.

Ship the whole pipeline

Data strategy through deployment — annotation, training, evaluation, serving. A running system, not a notebook and a wish.

iv.

Hold it in production

Monitoring, retraining, cost control. The constraint gets re-checked as scale grows — because it moves.

The method, in the field — four constraints, four shipped systems
Constraint → ComplianceEdwards Lifesciences

PHI must never leak into training. HIPAA-grade DICOM de-identification built into the clinical pipeline itself — a custom U-Net masking sensitive regions before a single model sees the data.

Constraint → LatencyShotdeck

Sub-second answers across 3M+ assets. Hybrid retrieval architected around the latency budget — pgvector with HNSW indexing, not whichever embedding model benchmarked highest that month.

Constraint → CostSynthesys AI Studio

2M+ generative API calls a year. Unit economics drove the architecture: a strategic shift to serverless GPU execution cut compute from $100K+ annually to a fraction — at 99.9%+ uptime.

Constraint → HardwareIndustrial Edge

Real-time detection on devices, not data centers. Quantization-aware architectures holding low-latency performance through low light, occlusion, and motion blur on Coral TPU and Movidius hardware.

Everything else — model choice, stack, infrastructure — follows from the constraint. Never the other way around. That ordering is why these systems are still running. The full argument, in writing →

02 / The Record Independently verified
11yrs
Production AI — first deployment 2015, systems still running today
80+
Projects across healthcare, retail, media, security, finance & industrial
$400K+
Verified client earnings on Upwork alone, across 10,100+ documented hours
100%
Job Success Score — every engagement, every client
Verification — Upwork: Top 1% globally · A-Team: <2% Acceptance Rate
Hand-picked for Upwork's 2023 Work Without Limits Summit — one of 31 featured (Fortune 500 exclusive)
In their words — clients on the record
Exceptional skills in AI software development… surpassed our expectations… met every project deadline ahead of schedule. 10/10 would recommend.
Braum KatzCEO · Canary
Ahmer's skills were by far exceptional… understanding my requirements, teaching me things along the way, and going above and beyond in delivering the final solution.
Munish SharmaVestas
03 / Case Files Four systems · four constraints · four outcomes

Selected systems, on the record.

CS—012019–2025 · 5 yrs
ShelfrRetail Shelf Intelligence · Founding Sole AI Architect
Binding constraint

Detection accuracy under real-world retail chaos — occlusion, lighting, thousands of SKUs — sustained across five years of continuous deployment.

Joined at founding as the sole AI architect and built the complete shelf-intelligence platform from zero: SKU-level detection, planogram compliance, out-of-stock recognition, share-of-shelf, price-tag and promo-display detection. Owned the full ML lifecycle — annotation pipelines with automated QC, training and evaluation across every product generation, GCP auto-scaling GPU deployment — directing a team of four.

Very professional and managed to pull off quality work in a very short span of time… never hesitated to look for solutions and alternatives. I would highly recommend working with him.
Karim Tribak — CEO & Co-Founder, Shelfr
200+
FMCG brands incl. P&G
+22%
On-shelf availability
+25%
Shelf compliance
2025 P&G External Business Partner Excellence Award
Awarded to the platform
YOLOFaster R-CNNTensorFlow OD APIGCP GPU auto-scalingAnnotation QC pipelines
CS—022019–2020 · Clinical
Edwards LifesciencesClinical Echocardiogram AI · Global Medical Device Leader
Binding constraint

HIPAA & GDPR — patient-identifying data must never reach a training pipeline, without destroying the clinical signal.

Reimplemented the EchoNet-Dynamic pipeline end-to-end in TensorFlow 2.0 — frame-level LV segmentation and spatio-temporal LVEF regression, reproducing published benchmark performance. Engineered beat-to-beat cardiac function assessment via end-systole/end-diastole frame identification, an attention-based MIL framework surfacing diagnostic frames from weakly labeled studies, and a custom U-Net DICOM de-identification pipeline masking PHI regions before training.

~10,000
Echo videos in training pipeline
Per-cycle
Beat-to-beat LVEF via ES/ED detection
0 PHI
Masked before training · HIPAA / GDPR
Benchmark
Published EchoNet performance reproduced
TensorFlow 2.0DeepLabv3ResNet3D / (2+1)DU-NetAttention MILDICOM
CS—032025–Present
ShotdeckMultimodal Retrieval · World's Largest Cinematic Library
Binding constraint

Sub-second natural-language answers across 3M+ cinematic assets — at a latency budget dense single-vector search can't hold.

Architected the core hybrid retrieval system — visual embeddings, textual signals, and structured metadata over pgvector with HNSW indexing — answering queries like "warm backlit interior, 1970s, handheld" in under a second. Built any-modality search (dialogue, reference image, ambient audio, video clip, hex color), a 76-tag cinematographic extraction system across 15 categories, a 20+ class camera-motion classifier, and an automated shot-curation pipeline that replaced manual frame review for a media team of 100+.

3M+
Assets searchable
<1s
Natural-language query latency
100+
GPUs · parallel embedding compute
76
Structured cinematic tags / image
pgvector + HNSWCLIPDINOv2TransNetV2ModalCloudflare R2
CS—042021–2024 · 3 yrs
Synthesys AI StudioGenerative Media Stack · AI Humans Integrated into Canva
Binding constraint

Unit economics at 2M+ API calls a year — fixed GPU fleets were bleeding $100K+ annually against spiky, unpredictable load.

Pioneered and owned the full generative content stack — photo-realistic talking avatars, AI lipsync, hyper-realistic voice cloning, and text-to-image — from raw training data through cloud deployment. Led the strategic migration from fixed AWS/GCP GPU fleets to serverless execution on Modal and Replicate, and engineered the REST/integration layer behind the platform's enterprise partnerships — including the AI Humans product's integration into Canva.

An impressive python developer who is able to work on all tasks we assigned him… honest, dependable, and incredibly hardworking. Without a doubt, we confidently recommend Ahmer.
Nick Koukoulakis — Co-Founder & CEO, Synthesys AI Studio
2M+
API calls / year in production
$100K+→
Annual compute cut to a fraction
99.9%+
Uptime under peak load
Integrated into Canva
AI Humans product
Stable DiffusionWav2LipTacotron2 / CoquiModal · ReplicateAWS / GCP
Agentic AI · Biscuit AI

Autonomous retail sales agents

Primary AI architect leading 5 developers — LangGraph state machines & CrewAI agents that self-configure from enterprise knowledge bases; LangSmith evaluation cutting hallucination risk.

Clinical NLP · HealthTech (Confidential)

HIPAA PHI de-identification at scale

85%+ PHI detection accuracy on real diagnostic reports — context-aware obfuscation preserving clinical utility; ColPali/ColBERT late-interaction retrieval for regulated records.

Biometric AI · FOO Technologies

Face anti-spoofing & liveness

90%+ accuracy against video replay, printed-face and mask attacks; document tampering detection at 91% — deployed as modular REST APIs.

Real-time Avatars · Aphra

Conversational avatar pipeline

Full stack from LLM response through multi-engine TTS, phoneme-level lip-sync and rendering — latency-optimized on serverless GPU infrastructure.

Edge AI · Canaryaware

Industrial safety detection

Real-time person detection on Coral TPU & Movidius NCS2 for forklift and construction-site safety — QAT-optimized through low light, occlusion and motion blur.

Edge-Cloud · Financial Institution

ATM anti-skimming & fraud

Hybrid architecture — live edge inference with centralized continuous learning, encrypted pipelines, and risk-based real-time alerting in a closed microservices system.

04 / Capabilities Eight domains · one discipline

Full-stack ownership, eight production-proven domains.

From data strategy through deployment — and the unglamorous work after launch: monitoring, retraining, cost control.

Agentic Systems

Supervisor · routing · HITL

Production multi-agent systems — supervisor/routing state machines, human-in-the-loop clarification loops, and durable cross-session memory that survives past a single conversation turn.

LangGraphLangChainCrewAIHITL InterruptsDurable Memory

Claude Platform

Claude API · Agent SDK · MCP

Native Claude-platform engineering — prompt-caching architectures, custom MCP servers, and Claude Agent SDK builds for production agentic systems.

Claude APIClaude Agent SDKMCPPrompt Caching

RAG Systems

Citation-strict · audit-grade

Retrieval systems built to be audited, not demoed — late-interaction re-ranking, contextual chunking, and RAGAS-triad evaluation gating every release before it ships. Shipped a citation-strict verification layer enforcing source-and-page-grounded answers across 15M+ tokens.

ColPali · ColBERTHybrid Search (RRF)Contextual ChunkingRAGAS Evaluation

Multimodal AI

Any-to-any · 3M+ asset scale

Cross-modal understanding across image, text, audio, and video — vision-language embeddings and layout-aware document parsing searchable in natural language at multi-million-asset scale.

CLIP · SigLIPDINOv2LLaVA · Qwen-VLDocument AI

Vector Database

HNSW/IVF Indexing · Multi-Tenant

Production vector infrastructure across every major platform — high-throughput ingestion, multi-tenant namespace isolation, and indexing strategies tuned to the query pattern, not the default config.

pgvector · Qdrant · PineconeVespa · Weaviate · FAISSHNSW/IVF IndexingMulti-Tenant Isolation

Computer Vision

11 years in production

Detection, tracking, segmentation, and recognition under field conditions — from 200+ brand retail deployments to biometric anti-spoofing and broadcast sports analytics.

YOLO v5–v12Grounding DINOSAM / SAM 2ViT · SwinByteTrack

Healthcare AI

HIPAA · DICOM · FHIR

Clinical-grade imaging and NLP inside regulated pipelines — echocardiogram analysis, PHI de-identification, and evidence-grounded retrieval that survives an audit.

MONAInnU-NetMedSAMBioBERT · ClinicalBERTEHR integration

Edge AI

Real-time on-device

Inference where the data lives — quantization-aware training, pruning, and edge-cloud hybrid designs holding real-time performance under industrial conditions.

Jetson Nano–OrinCoral TPUTensorRT · OpenVINOINT8 / QAT
05 / Point of View In writing, on the record
Essay — 01 · 6 min read

Why AI pilots die between the demo and the deploy

Eleven years and eighty-plus production systems reduce to one diagnostic question — and most teams ask it too late. The full argument behind the Binding Constraint Method.

Read the essay →
Position — 01Model choice is the last decision in a serious system, not the first.
Position — 02The expensive part of AI isn't training. It's the year after launch.
Position — 03If it can't survive an audit, it isn't a healthcare AI system.
06 / Working Together Three ways in · one process

Three engagement models.

i.

Architecture & DeliveryMulti-month to multi-quarter

End-to-end ownership: data strategy, system design, model development, deployment, and production maintenance. You get a running system — not a handover document.

You walk away with → a production system · documentation · a team that can run it
ii.

Advisory & Fractional AI LeadershipFixed-term or ongoing

Architecture reviews, constraint audits, roadmap and build-vs-buy decisions, hiring support. For teams with strong engineers who need production-AI judgment at the design table.

You walk away with → a constraint audit · an architecture verdict · a defensible roadmap
iii.

Production RescueStarts with a short diagnostic

Stalled pilots, runaway GPU spend, accuracy collapsing in the field. Triage the failure, isolate the actual constraint, re-architect, and stabilize.

You walk away with → a triage report · a re-architecture plan · a stabilized system
How every engagement runs — four stages, no exceptions
Stage 01

Discovery call

Within 48 hours of first contact. Free, direct, and honest about whether this is worth either side's time.

Stage 02

Constraint audit

The binding constraint gets named in writing, with a scoped proposal — fixed-fee for audits and advisory, milestone-based for builds.

Stage 03

Build / advise

Async-first execution with documented progress. Full pipeline ownership or embedded judgment alongside your team.

Stage 04

Production hold

Monitoring, retraining, and cost control after launch — because the constraint moves as you scale.

07 / Questions, Answered What buyers ask before they email

Asked before. Answered plainly.

Why did our AI pilot fail to reach production?+

Most pilots fail because the system was designed around the model instead of around its binding constraint — the one non-negotiable requirement (compliance, latency, cost, or hardware) that decides whether the system can exist in production. The fix is usually re-architecture around that constraint, not more model training. That diagnosis is exactly what a Production Rescue engagement starts with.

Is Nixense Vixion an agency or one person?+

A founder-led consultancy. Every engagement is architected and led by me personally — I'm the one on the calls and in the code. When scope requires hands, delivery teams are assembled: I've directed teams of four and five engineers through multi-year production builds. Everything is documented for handover from day one, so nothing lives only in my head.

How do engagements typically run?+

Four stages: a discovery call within 48 hours of first contact, a constraint audit with a written proposal, the build or advisory phase, and a production-hold phase covering monitoring, retraining, and cost control. Engagements range from short fixed-fee audits to multi-quarter delivery.

How do you handle NDAs, IP, and confidential data?+

Work happens under NDA by default — several past engagements remain confidential, including clinical and financial systems. Deliverables and IP transfer to you. In regulated environments, pipelines are designed so protected data (such as PHI) is de-identified before any model training touches it.

How do you make computer vision and AI HIPAA-compliant?+

Compliance is treated as the binding constraint and built into the architecture, not bolted on: DICOM de-identification before training (custom U-Net masking of PHI regions), context-aware obfuscation that preserves clinical values while removing identifiers, FHIR-compatible pipelines, and evidence-grounded retrieval that produces auditable outputs.

What does an engagement cost?+

Scoped per engagement after the discovery call. Constraint audits and advisory are fixed-fee; builds are milestone-based. The discovery call itself is free and produces a clear read on whether the engagement is worth either side's time.

Do you work with our existing stack, or do we need to adopt yours?+

Your stack, by default — the constraint dictates architecture, and architecture dictates tooling, not the other way around. Systems have shipped on GCP, AWS, and serverless (Modal/Replicate) alike; the default is adapting to what's already in place unless there's a specific technical reason to migrate.

08 / About The person behind the systems

Constraint-first since the first robot.

The first system I shipped was a wall-tracking robot navigating unfamiliar rooms with proximity sensors and a PID controller — 2013, university award, hard physical limits. Every system since has started the same way: find the constraint, then design.

I'm Muhammad Ahmer Saleem. Most of what I've built for clients across the US, UK, and Europe is still running years after the contract ended — that's the number I actually optimize for, not the demo.

I work fully async, remote-first, from Lahore — with deliberate US-timezone overlap and a 24-hour ceiling on response time. Clients notice the same thing across eleven years: nothing waits on a timezone, and nothing ships without the trade-offs spelled out first.

EducationB.Eng Electrical & Electronic — University of Bradford, UK
DistinctionBest Student · Best FYP Award
BaseLahore, Pakistan
Overlap4+ hrs daily · EU & US
ModeDirect access · Nothing lost in handoff
Download full CV (PDF) →
09 / Contact Replies within 24 hours

Describe the system. I'll find the constraint.

Tell me what you're building, where it hurts, and when you need it. Every inquiry gets a reply within 24 hours — a discovery call typically follows within 48.

Goes straight to my inbox — nothing else is stored.

Received. Reply within 24 hours.

Your inquiry is in my inbox. If it's urgent, the direct line on the right reaches me fastest.