My name is Arup Kumar Sarker, and I am a Research Assistant at the University of Virginia's Biocomplexity Institute. My research lies at the intersection of AI systems, distributed computing, high-performance computing, large language models, agentic AI, multimodal AI, and high-performance data engineering.
A consistent theme throughout my career has been building systems that must operate efficiently under real-world constraints. Whether optimizing a mobile platform with only a few megabytes of memory, developing mission-critical communication systems for first responders, scaling scientific data processing across supercomputers, or reducing the cost of distributed LLM inference, I have been interested in the same fundamental problem: how can we make complex computing systems faster, more scalable, resource-efficient, reliable, and easier to operate?
My current research focuses on the infrastructure required for the next generation of AI applications. Modern AI systems increasingly combine retrieval, large language models, multimodal data, tools, multiple cooperating agents, distributed memory, and heterogeneous CPU/GPU resources. My work explores how these components can be represented as coherent systems in which communication, scheduling, data movement, model state, memory, and computation can be explicitly optimized.
Experience: 16+ Years in the Research and Development.
My recent research spans several interconnected projects:
AAFLOW+ — Distributed LLM State and KV-Cache Orchestration. I developed a stateful operator abstraction that elevates the transformer KV cache into a first-class distributed systems object. AAFLOW+ supports KV materialization, transfer, fork, reuse, composition, and eviction across multi-agent LLM workflows. It uses communication-aware execution and transfer-versus-recomputation cost models to avoid repeatedly processing shared context. Across evaluated configurations, the system demonstrated up to 50.2× lower time-to-first-token, 7.63× lower multi-agent compute cost, 1.72–6.10× lower KV-memory usage, and more than 7.74× higher throughput.
OpRAG, formerly AAFLOW — Scalable Agentic AI and RAG Workflows. OpRAG addresses the fragmentation of modern Agentic AI and retrieval-augmented generation pipelines. It represents preprocessing, embedding, vector retrieval, model execution, and agent coordination as distributed operator graphs. By combining communication-aware scheduling, asynchronous batching, and an Apache Arrow/Cylon zero-copy data plane, the system achieved up to 4.64× end-to-end pipeline acceleration and approximately 2.8× improvement in embedding and vector-upsert stages.
MosaicKV — Multimodal KV-Cache Compression. My ongoing, unpublished research extends model-state optimization to multimodal foundation models. MosaicKV investigates training-free compression of multimodal KV caches using future-query forecasting, cross-modal evidence relationships, value-aware selection, and adaptive memory management. The goal is to preserve information most likely to influence future generation while reducing redundant visual and textual state.
Deep Radical-Cylon, Radical-Cylon, and Cylon — HPC and Scientific AI. My earlier research developed distributed data-engineering and deep-learning systems across heterogeneous CPU/GPU clusters, cloud platforms, and supercomputers. These systems integrate technologies such as Apache Arrow, Cylon, PyTorch, TensorFlow, MPI, UCX, NCCL, CUDA, GLOO, and Radical-Pilot. Deep Radical-Cylon reduced execution time by as much as 75.9 seconds in evaluated scientific AI workflows, while Radical-Cylon operated on workloads ranging from 35 million to 3.5 billion rows and demonstrated performance improvements over comparable batch execution.
Together, these projects form a broader research direction around stateful and high-performance AI infrastructure: treating data, communication, model memory, and execution decisions as first-class components of AI system design.
Before beginning my doctoral research, I spent more than a decade at Samsung R&D, progressing through senior engineering, technical leadership, and staff engineering roles. That experience continues to influence how I approach research today. I care not only about proposing algorithms or abstractions, but also about whether they can be implemented, measured, debugged, scaled, and operated reliably.
One of my most significant industry projects was Mission Critical Services for FirstNet and AT&T, where I led client-side engineering for a team of approximately 33 engineers. The platform supported secure, low-latency voice and video communication for first responders and public-safety personnel. I worked across networking, security, real-time media, and distributed client/server systems using technologies including TLS, SIP, SRTP, TCP/UDP, HTTPS, audio/video codecs, session management, and platform-specific networking.
Mission-critical software fundamentally changed my perspective on system design. For a first responder, latency, reliability, security, and failure recovery are not simply performance metrics—they directly determine whether the system fulfills its purpose. That experience strongly influences my current work on reliable distributed AI systems.
I also led major engineering efforts for Samsung Gear360, including networking, 4K video processing, GPU-accelerated rendering, camera control, and Samsung's proprietary 360-degree stitching technology. My work involved OpenCV, OpenGL, Metal, HEVC/H.264, fisheye calibration, image blending, stereographic projection, network streaming, and GPU optimization. Related image-processing work contributed to a U.S. patent.
Earlier in my Samsung career, I worked on Android audio frameworks, multimedia systems, networking stacks, browser/security protocols, wearable-device communication, and highly resource-constrained feature-phone platforms. Those systems taught me an engineering lesson that remains central to my research: when compute, memory, bandwidth, or latency is constrained, architecture matters even more.
My career has allowed me to work across the complete spectrum from low-level systems programming and networking to distributed HPC runtimes and modern foundation-model infrastructure. I have extensive experience with Python, C/C++, Java, CUDA-oriented systems, distributed communication, deep-learning frameworks, and production software development.
I also consider mentorship an important part of my work. At UVA, I have led and mentored undergraduate researchers contributing to projects including AAFLOW+, OpRAG/AAFLOW, Deep Radical-Cylon, and Radical-Cylon, providing guidance in distributed AI systems, implementation, experimental methodology, performance evaluation, and research communication.
My long-term goal is to build computing systems in which increasingly capable AI models can operate efficiently across distributed resources while preserving performance, scalability, reliability, reproducibility, and resource efficiency. I am particularly interested in the convergence of distributed LLM systems, Agentic AI, Agentic RAG, multimodal memory, HPC, and scientific AI, and in translating research advances in these areas into practical systems capable of supporting the next generation of intelligent applications.
Programming Languages: Python, C, C++, Java, SQL, Rust, Objective-C, Swift, Cython
LLM / Generative AI Systems: Large Language Models, Generative AI, Distributed LLM Inference, KV-Cache Orchestration, KV-Cache Compression, Long-Context Inference, Multimodal AI, Agentic AI, Multi-Agent Systems, Agentic RAG, LLM Performance Modeling, Prefill/Decode Optimization, TTFT/Throughput/Memory Optimization, vLLM, SGLang, Hugging Face Transformers
Machine Learning / Deep Learning: PyTorch, TensorFlow, Transformers, Attention Models, NLP, Computer Vision, 2D/3D Perception, Multimodal Models, Time-Series Models, Embeddings, Representation Learning, Model Training and Evaluation
Agentic AI / RAG: Retrieval-Augmented Generation (RAG), Agentic RAG, Multi-Agent Workflows, Tool-Using Agents, LangChain, LangGraph, LlamaIndex, Model Context Protocol (MCP), Vector Retrieval, Embedding Pipelines
High-Performance Computing / Distributed Systems: Apache Arrow, Cylon, MPI / Open MPI, UCX, NCCL, GLOO, CUDA Runtime, SLURM, Ray, Dask, Apache Spark, cuDF, DDP, FSDP, Radical-Pilot, Heterogeneous CPU/GPU Execution, Multi-Node Computing, Zero-Copy Data Movement, Communication-Aware Scheduling
Data Engineering / Vector Search: Distributed Data Processing, Apache Arrow Data Plane, Cylon Dataframes, ETL/Data Pipelines, FAISS, ChromaDB, Pinecone, Vector Databases, Embedding/Indexing/Upsert Pipelines
Performance Engineering: End-to-End Benchmarking, Strong/Weak Scaling, Workload Characterization, Latency/Throughput/Memory Analysis, GPU/CPU Performance Optimization, Distributed Communication Analysis, Profiling, Cost Modeling, Transfer-vs-Recompute Analysis
Networking / Communication Protocols: TCP/IP, UDP, TLS, HTTPS, SIP, RTP, SRTP, RTSP, SDP, MBCP, SSDP, UPnP, BLE/GATT, Wi-Fi, HTTP Streaming, Secure Real-Time Audio/Video Communication
Multimedia / Computer Vision / Graphics: OpenCV, OpenGL / OpenGL ES, Metal, FFmpeg, HEVC/H.265, H.264, MJPEG, 360-Degree Video, Image Stitching, Fisheye Calibration, Feature Matching, Image Blending, Stereographic Projection, GPU Shader Programming, Video Transcoding and Streaming
Mobile / Embedded Systems: iOS, Android, Tizen, Linux, Samsung MMP, Samsung SUP, Embedded/Mobile Runtime Systems, Low-Memory System Optimization, Device Communication
Software Architecture / Systems Engineering: Distributed Systems, Service-Oriented Architecture, APIs and Web Services, Middleware Frameworks, Concurrent and Multi-Threaded Systems, High-Availability Systems, Resource-Constrained Systems, Python/C++ Integration with Cython
Software Design Patterns: MVC, MVVM, VIPER, Object-Oriented Design, Modular and Layered Architecture
Version Control / Code Review: Git, Perforce, SVN, Gerrit, RBTools
Development & Project Management: Jira, PLM, MPS, Git, Linux Development Environments, CMake, Debugging and Profiling Workflows
Software Development Processes: Agile, Scrum, Kanban, Waterfall, Design Review, Code Review, Performance Validation, Research-to-Production Prototyping
Technical Leadership: System Architecture, Technical Direction, Cross-Functional Collaboration, Research Mentoring, Engineering Team Leadership, Design and Code Review, Experimental Design, Performance Evaluation, Technical Documentation
Department of Computer Science, University of Virginia, VA, USA
PhD in Computer Science;
MSc in Computer Science;
Department of Computer Science & Engineering, University of Dhaka, Bangladesh
MSc in Computer Science and Engineering;
BSc in Computer Science and Engineering;
https://www.linkedin.com/in/arup-sarker-8190212b/
https://scholar.google.com/citations?user=tWBCx3kAAAAJ&hl=en
Gear 360 camera can capture 360 image/video. It is small, sophisticated device, managed by iPhone and Android with Camera, Gallery, Integrated player, Share, Broadcasting, 3D Touch etc.
It is an Audio and Video Communication app where AMR Audio and H265 encoded video will be sent through SRTP to another user by following 3gpp spec of Mission Critical Services, mainly help first responder to handle public safety.
For Android platform and deployed to different Android smart phone
A Manager app combined with service to control Samsung Gear Smart watch
Video Player, Music Player, Frameworks, System, message for E1282T, E2350B, C3312R, S5222R ,E2222,E2220,S3770,C3010S
A mobile and service app for different NGOs to support Maternal health