CS6023 GPU Programming

July 2026

Important Links Course Slides (as google slides)
  1. Intro + Logistics
  2. Computation
  3. Memory
  4. Synchronization
  5. Functions
  6. Support
  7. Streams
  8. Topics
  9. Case Study -- Graphs
  Evaluation
Eval. ItemMarksDeadlines
StudentTAInstructor
A110August 23+10 days--
A212September 13+10--
A312October 4+10--
A4/Project16November 7 (Saturday)+10--
MidSem15+5September 16--+10
EndSem25+5November 19--+10

Attendance
Standard institute rules apply.

Other details

  • Syllabus and structure
  • Prerequisite: CS2710 (PDS Lab) or Equivalent.
  • TAs:
    • Albin James Maliakal cs25s027
    • Shivam cs25m045
    • Bantu Vijayendra Varma cs25m016
    • Shubham jadhav cs24s009
    • Jay Rajesh Rathi cs24s001
    • Prashant Singh cs24d002
  • Instructor: Rupesh Nasre.
  • Venue: SSB 134
  • Slot: G (Monday 12, Wednesday 17, Thursday 10, Friday 9)


Schedule
MonthDatesTopicComments
 July   27, 29, 30, 31  Introduction, Computation  
  • Hello World, One, Two, Three
  • Grid, Blocks, Threads
  • Kernel Launch: 1D, 1D-General, 2D
  •  August  3, 5, 6, 7  Computation  
  • CPU-GPU Communication (cudaMalloc, cudaMemcpy)
  • Global variables
  • Matrix mult.: CPU, Outer parallel, Outer+Inner parallel
  •    10, 12, 13, 14  Computation  
  • Thread Divergence
  • Divergence due to switch
  • Problem Set 1
  •    17, 19, 20, 21  Memory  
  • Memory Coalescing
  • AoS versus SoA
  • Barrier
  •    24, 27, 28  Memory  
  • Linked List Copying
  • Shared Memory
  • Shared Memory with Barrier
  • String Permutation
  • Dynamic Shared Memory
  • Dynamic Shared Memory with Multiple Arrays

    CUDA GDB
  • Error Handling
  • Dangling Pointer
  •    31  Memory  
  • Texture Memory (via CUDA SDK)
  • Constant Memory
  • Bank Conflicts
  • Problem Set 2

    NvProf
  • Original Code
  • Loop Fusion
  • Kernel Fusion
  • Converting Loop to Blocks
  •  September   2, 3, 4  Synchronization  
  • Convolution
  • Worklist Insertion
  • Task Donation
  •     7, 9, 10, 11  Synchronization  
  • Reduction: i + N/2, N - i, i + 1
  • Prefix Sum / Scan
  •     16, 17, 18  Synchronization  MidSem on 16 from 5 PM to 6:30 PM
  • No Global Barrier
  • Global Barrier using Atomics
  • Hierarchical Global Barrier
  •     21, 23, 24, 25  Synchronization  
  • Linked List Insertion
  • CPU-GPU Shared Pinned Memory
  • Persistent Kernel
  • Problem Set 3
  •     28, 30  Synchronization  
  • Array increment: Sequential, Parallel
  • Thrust basics
  • Thrust Reduction
  • Thrust Prefix Sum
  • Thrust-like device vector implementation
  •  October   1, 2  Functions  
  • Basic Stream Program
  • with Asynchronous memcpy
  • with cudaHostAlloc
  • Cooperative Kernels
  •     5, 7, 8, 9  Functions  
  • Dynamic Parallelism
  • Conditional Child Kernels
  • using Global Device Memory
  • with Non-Blocking Streams
  •     12, 14, 15, 16  Functions  
  • MultiGPU: Number of Devices
  • Cross-Device Synchronization
  •    19, 21, 22, 23  Topics  
  • PTX: CUDA Code, Assembly Code
  • Basic Warp Voting
  • Converting Mask to Count (popc)
  • Use of ffs
  • Conditional Participation in ballot
  •     26, 28, 29, 30  Topics  
  • Loop Unrolling, Unrolled Assembly
  • Heterogeneous Computation
  •  November   2, 4, 5, 6  Topics  
  • with Shared Variable
  • Task Distribution
  • OpenMP Reduction
  • with HostAlloc'ed Memory
  • Dynamic Scheduling
  • OpenCL: Driver, Kernel
  •     19    EndSem from 10:00 -- 12:00


    GPU Programming Crossword Puzzle (click and type)

    Courtesy: crosswordlabs.com