Portfolio

Powering AI. Redefining Efficiency.

Jashom is an applied AI company advancing artificial intelligence while optimizing performance and reducing energy consumption across GPU infrastructure, model deployment, and healthcare AI systems.

Core Capability Matrix

GPU Kernel Engineering

Technical Depth

CUDA / ROCm kernel development, layer fusion, operator optimization

Business Impact

Faster inference, lower hardware cost per query

LLM Power Optimization

Technical Depth

INT8/FP16 quantization, TensorRT inference re-engineering

Business Impact

37%+ power reduction with no accuracy loss

AI Model Fine-Tuning

Technical Depth

LoRA, QLoRA across 7B–70B+ models; cloud infra strategy

Business Impact

Production models in hours, not weeks

Workload Orchestration

Technical Depth

REST API scheduling, VRAM-aware assignment, container isolation

Business Impact

GPU jobs tracked, isolated, and audited end-to-end

Hardware Telemetry

Technical Depth

Redfish / BMC integration, per-GPU power & thermal monitoring

Business Impact

Real-time infrastructure visibility without OS dependency

RAG Infrastructure

Technical Depth

Retrieval-Augmented Generation with distributed data layers

Business Impact

Real-time contextual AI at scale

Healthcare AI

Technical Depth

Dispatch optimization, triage analytics, hospital system integration

Business Impact

Saves critical minutes in emergency response

GPU Portfolio & Case Studies

Case Studies

Real engagements: LLM inference optimization, GPU orchestration, cloud fine-tuning, and hardware telemetry.

Case Study

LLM Inference Optimization on Constrained GPU Infrastructure

42% higher throughput, 37% lower power, 12 distributed nodes. Full inference path re-engineering with CUDA kernels, TensorRT, and adaptive batching.

42% Throughput37% Power ↓12 Nodes
View full case study →
Case Study

GPU Workload Orchestration Framework on Rocky Linux 9.7

Demo-ready in 5 days: REST API, VRAM-aware scheduling, Docker isolation, full audit trail. RTX 3090 + Rocky Linux 9.7.

5 Days4 Endpoints100% Isolation
View full case study →
Case Study

Cloud GPU Fine-Tuning Strategy for Production LLM Deployment

Tiered strategy 7B–70B+ models: LoRA/QLoRA, Axolotl, DeepSpeed. Provider-agnostic cloud GPU; dataset to production in days.

7B–70B+3 TiersDays to Deploy
View full case study →
Case Study

Real-Time GPU Server Hardware Telemetry via Redfish BMC

Live dashboard every 30s: power, temperature, fan RPM from Lambda Scalar BMCs. HTTPS, Basic Auth, scoped SSL bypass.

30s Refresh4 ServersOut-of-band
View full case study →
Portfolio Summary

Capabilities, Technologies & Engagement Model

What Jashom Has Demonstrated

Custom CUDA kernel engineering for LLMs

Case Study 1: kernel-level optimization of 13B parameter model

INT8/FP16 quantization without accuracy loss

Case Study 1: zero measured accuracy degradation post-quantization

42% throughput improvement on production inference

Case Study 1: measured result on 13B model, 12-node deployment

37% GPU power reduction

Case Study 1: measured against pre-optimization baseline

REST API GPU job scheduling with VRAM awareness

Case Study 2: full FastAPI + SQLite orchestration system

Containerized GPU execution with hard isolation

Case Study 2: NVIDIA_VISIBLE_DEVICES enforced per job

LoRA / QLoRA strategy across 7B–70B models

Case Study 3: tiered fine-tuning framework

Multi-provider cloud GPU management

Case Study 3: AWS, Lambda Labs, CoreWeave, RunPod

Out-of-band BMC hardware telemetry

Case Study 4: Redfish / AST2600 integration

Production platform engineering (TypeScript / Node.js)

Case Study 4: Electron app metric collector extension

Technology Stack

Full Technology Stack

GPU & Inference

CUDAROCmTensorRTTriton Inference ServerONNXnvidia-smiNVIDIA NsightINT8/FP16 QuantizationLayer Fusion

AI / ML Frameworks

PyTorchTensorFlowHugging Face TransformersLangChainDeepSpeed (ZeRO-3)UnslothAxolotlTorchTunevLLM

Infrastructure

DockerNVIDIA Container ToolkitFastAPIuvicornSQLiteSQLAlchemysystemdRocky Linux 9.7Ubuntu 22.04Python 3.x

Monitoring & Telemetry

Redfish APISupermicro AST2600 BMCLambda Scalar serversundici (Node.js)TypeScript / Electron

Cloud Providers

AWSGoogle Cloud PlatformMicrosoft AzureLambda LabsCoreWeaveRunPodVast.aiTensorDock
How We Work

Engagement Model

Fixed-Scope Prototype

Well-defined problem, delivered in 3–5 days. Priced by scope. Examples: GPU orchestration prototype, Redfish telemetry integration, fine-tuning run with evaluation.

Production Engineering

Ongoing GPU engineering, model optimization, or AI system development. Embedded technical partnership with measurable milestones.

Applied Research

Low-power inference architectures, GPU sharing fabric design, model compression and distillation. Research engineering alongside production deliverables.

GPU Audit

Profiling and optimization assessment of your existing GPU infrastructure. Delivered as a prioritized recommendations report with measurable impact projections.

Ready to make your GPU infrastructure work harder?