Open to Full-Time Roles & High-Impact Engineering

Architecting Production AI Systems & Real-Time Edge Intelligence.

Focusing on: AI Systems Engineer |

Hi, I'm Amanpreet Singh (AmanX). I design and build production-grade Generative AI architectures, enterprise RAG systems with automated hallucination evaluation, and 100% on-device local voice pipelines combining Whisper, quantized LLMs (llama.cpp), and ONNX speech synthesis.

40+ Public Repos & AI Tools
< 1.0s Local Voice Pipeline Latency
4-Vector Automated RAG Hallucination Scoring
100% On-Device Edge Privacy
HR & Hiring Manager Quick-Deck

Candidate Snapshot & Core Superpowers

Production AI & RAG Specialist

Experience implementing multi-agent RAG pipelines with LLM-as-a-judge automated guardrails, preventing hallucinations and measuring groundedness before data reaches executive users.

Edge & Local Voice AI Pioneer

Engineered end-to-end local speech-to-speech pipelines for Python and Android. Combines Faster-Whisper, quantized LFM (350M Q4_0), and ONNX Supertonic TTS for zero-cloud latency and complete privacy.

Full-Stack & Systems Craft

Strong proficiency across Python, FastAPI, Docker, TypeScript, Next.js, and Android Kotlin/JNI. Rapidly converts cutting-edge AI papers into tested, deployable software.

Work Preferences & Availability

Status: Immediate availability for Full-Time roles & High-Impact Contracts.
Location: Open to Remote Worldwide, or Hybrid / Relocation.
Roles: AI Systems Engineer, GenAI Developer, ML Engineer.

Featured Systems & Products

Real-world implementations solving hard engineering problems: from enterprise LLM hallucination prevention to zero-latency on-device voice agents.

On-Device Edge AI Sub-second

Edge Voice Pipeline (Python & Android SDK)

A 100% local, on-device streaming voice pipeline combining Faster-Whisper ASR, quantized LFM 2.5-350M (llama.cpp), and Supertonic 3 ONNX TTS. Operates completely offline with sub-second response times, zero cloud API fees, and absolute user privacy.

Local Architecture Stack
Mic → Faster-Whisper (Q4) → LFM-350M (llama.cpp) → Supertonic TTS
  • Python
  • Android Kotlin / JNI
  • llama.cpp
  • faster-whisper
  • ONNX Runtime
  • CTranslate2
Generative Media AI Full-Stack

Text-to-Videos AI Studio

A full-stack Next.js web application that takes concise prompt inputs and autonomously orchestrates script generation, scene breakdowns, voiceover audio synthesis, and visual rendering to output ready-to-publish 30s+ videos.

Video Synthesis Flow
Prompt → LLM Scene Director → Audio Synthesis → Canvas Video Compositor
  • TypeScript
  • Next.js
  • React
  • Audio Synthesis
  • Tailwind CSS
AI Productivity Suite Multimodal

All-AI Productivity Suite

AI-driven desktop workspace tool featuring intelligent screenshot classification, voice memo transcription, focus Pomodoro timers, and productivity analytics. Engineered to streamline deep-work developer workflows.

Core Capabilities
Voice Memos + Vision OCR + Pomodoro State Machine + Live Analytics
  • TypeScript
  • React
  • Vision Models
  • Whisper Transcription
  • Analytics Engine
Research & Open Source Foundation Models

Alpaca Dataset Generator & LLM Tuning

Custom pipelines for synthetic instruction tuning dataset generation (`Alpaca-style-Dataset-Generator`) and fine-tuning workflows (`LLaMA-Factory`). Active open-source contributor experimenting with DeepSeek-R1 reproductions and neural steering architectures.

Data & Training Pipeline
Seed Tasks → Quality Filtering → Deduplication → LLaMA-Factory LoRA / QLoRA
  • Python
  • PyTorch
  • LLaMA-Factory
  • Synthetic Data
  • HuggingFace

Ask Aman's AI Assistant

Have questions about my technical background, architecture decisions, or role fit? Ask directly or click any pre-configured question below!

AmanX Interactive Candidate Bot ● Online
Model: Local Knowledgebase
Hello! I am Aman's interactive assistant. I can help answer technical inquiries from recruiters, engineering leads, and hiring managers. Feel free to ask a question or select from the options below!

Skills & Engineering Disciplines

Proficiencies honed across production AI platforms, real-time edge pipelines, and full-stack software development.

LLM Systems & RAG

Designing reliable agentic systems and multi-hop retrieval pipelines with strict hallucination guardrails.

Multi-Agent RAG LLM-as-a-Judge Vector DBs (Pinecone/Chroma) Dense Embeddings Fine-Tuning (LoRA) LLaMA-Factory Prompt Engineering OpenRouter Fallback

Edge AI & Audio Pipelines

Deploying quantized models directly on local consumer hardware for sub-second, zero-cloud performance.

faster-whisper llama.cpp ONNX Runtime Supertonic TTS 4-bit Quantization (Q4_0) CTranslate2 Android NDK / JNI Streaming Audio

Backend & Scalable Systems

Robust microservices, async APIs, automated browser scraping, and resilient containerization.

Python (AsyncIO) FastAPI Docker & Compose Playwright Headless RESTful APIs WebSockets Git / GitHub Actions Linux / macOS CLI

Modern Frontend & Mobile

Responsive, accessible user interfaces with clean architecture and reactive state management.

TypeScript Next.js (App Router) React Tailwind CSS Android (Kotlin) State Management Interactive Dashboards Responsive Design

Let's Build Something Extraordinary

I am actively looking for high-impact AI Engineering, Full-Stack GenAI, and Machine Learning roles. Whether you have an immediate hiring need or want to discuss an innovative architecture, I'd love to connect.

Send Email Direct
Copied to clipboard!