Core Capability Matrix
GPU Kernel Engineering
Technical Depth
CUDA / ROCm kernel development, layer fusion, operator optimization
Business Impact
Faster inference, lower hardware cost per query
LLM Power Optimization
Technical Depth
INT8/FP16 quantization, TensorRT inference re-engineering
Business Impact
37%+ power reduction with no accuracy loss
AI Model Fine-Tuning
Technical Depth
LoRA, QLoRA across 7B–70B+ models; cloud infra strategy
Business Impact
Production models in hours, not weeks
Workload Orchestration
Technical Depth
REST API scheduling, VRAM-aware assignment, container isolation
Business Impact
GPU jobs tracked, isolated, and audited end-to-end
Hardware Telemetry
Technical Depth
Redfish / BMC integration, per-GPU power & thermal monitoring
Business Impact
Real-time infrastructure visibility without OS dependency
RAG Infrastructure
Technical Depth
Retrieval-Augmented Generation with distributed data layers
Business Impact
Real-time contextual AI at scale
Healthcare AI
Technical Depth
Dispatch optimization, triage analytics, hospital system integration
Business Impact
Saves critical minutes in emergency response
Case Studies
Real engagements: LLM inference optimization, GPU orchestration, cloud fine-tuning, and hardware telemetry.
LLM Inference Optimization on Constrained GPU Infrastructure
42% higher throughput, 37% lower power, 12 distributed nodes. Full inference path re-engineering with CUDA kernels, TensorRT, and adaptive batching.
GPU Workload Orchestration Framework on Rocky Linux 9.7
Demo-ready in 5 days: REST API, VRAM-aware scheduling, Docker isolation, full audit trail. RTX 3090 + Rocky Linux 9.7.
Cloud GPU Fine-Tuning Strategy for Production LLM Deployment
Tiered strategy 7B–70B+ models: LoRA/QLoRA, Axolotl, DeepSpeed. Provider-agnostic cloud GPU; dataset to production in days.
Real-Time GPU Server Hardware Telemetry via Redfish BMC
Live dashboard every 30s: power, temperature, fan RPM from Lambda Scalar BMCs. HTTPS, Basic Auth, scoped SSL bypass.
Capabilities, Technologies & Engagement Model
What Jashom Has Demonstrated
Custom CUDA kernel engineering for LLMs
Case Study 1: kernel-level optimization of 13B parameter model
INT8/FP16 quantization without accuracy loss
Case Study 1: zero measured accuracy degradation post-quantization
42% throughput improvement on production inference
Case Study 1: measured result on 13B model, 12-node deployment
37% GPU power reduction
Case Study 1: measured against pre-optimization baseline
REST API GPU job scheduling with VRAM awareness
Case Study 2: full FastAPI + SQLite orchestration system
Containerized GPU execution with hard isolation
Case Study 2: NVIDIA_VISIBLE_DEVICES enforced per job
LoRA / QLoRA strategy across 7B–70B models
Case Study 3: tiered fine-tuning framework
Multi-provider cloud GPU management
Case Study 3: AWS, Lambda Labs, CoreWeave, RunPod
Out-of-band BMC hardware telemetry
Case Study 4: Redfish / AST2600 integration
Production platform engineering (TypeScript / Node.js)
Case Study 4: Electron app metric collector extension
Full Technology Stack
GPU & Inference
AI / ML Frameworks
Infrastructure
Monitoring & Telemetry
Cloud Providers
Engagement Model
Fixed-Scope Prototype
Well-defined problem, delivered in 3–5 days. Priced by scope. Examples: GPU orchestration prototype, Redfish telemetry integration, fine-tuning run with evaluation.
Production Engineering
Ongoing GPU engineering, model optimization, or AI system development. Embedded technical partnership with measurable milestones.
Applied Research
Low-power inference architectures, GPU sharing fabric design, model compression and distillation. Research engineering alongside production deliverables.
GPU Audit
Profiling and optimization assessment of your existing GPU infrastructure. Delivered as a prioritized recommendations report with measurable impact projections.