Skip to content

Commit 83be559

Browse files
committed
nebula vsearch note 1
Signed-off-by: Zhi Yiliu <2584074296@qq.com>
1 parent 5f77df8 commit 83be559

27 files changed

Lines changed: 3061 additions & 59 deletions

docs/.obsidian/workspace.json

Lines changed: 46 additions & 59 deletions
Original file line numberDiff line numberDiff line change
@@ -7,20 +7,6 @@
77
"id": "07150ca20b5d2386",
88
"type": "tabs",
99
"children": [
10-
{
11-
"id": "87be62f38607b0ae",
12-
"type": "leaf",
13-
"state": {
14-
"type": "markdown",
15-
"state": {
16-
"file": "projects/nebula-vsearch/上篇:初识 Nebula Graph —— 向量类型支持.md",
17-
"mode": "source",
18-
"source": false
19-
},
20-
"icon": "lucide-file",
21-
"title": "上篇:初识 Nebula Graph —— 向量类型支持"
22-
}
23-
},
2410
{
2511
"id": "cfef105082839db4",
2612
"type": "leaf",
@@ -185,53 +171,54 @@
185171
"bases:Create new base": false
186172
}
187173
},
188-
"active": "87be62f38607b0ae",
174+
"active": "cfef105082839db4",
189175
"lastOpenFiles": [
190-
"projects/active slam/kr 3d active slam.md",
191-
"paperreadings/activeslam/3D Active Metric-Semantic SLAM.md",
176+
"notes/cuda/img/vector_add_slo.png",
177+
"notes/cuda/img/slo_v1.png",
178+
"notes/cuda/img/roofline.png",
179+
"notes/cuda/img/memory_chat.png",
180+
"notes/cuda/Vector Add Optimization Example.md",
181+
"notes/cuda/Roofline & Basic Analysis.md",
182+
"notes/cuda/img",
183+
"notes/cuda",
192184
"projects/nebula-vsearch/上篇:初识 Nebula Graph —— 向量类型支持.md",
193-
"projects/active slam/img/data_flow.png",
194-
"projects/active slam/img",
195-
"projects/active slam/新建文件夹",
196-
"paperreadings/activeslam/vins基础.md",
197-
"projects/active slam",
198-
"img.md",
199-
"projects/nebula-vsearch/img/vector_value.png",
200-
"projects/nebula-vsearch/understanding/img/vector.png",
201-
"projects/nebula-vsearch/understanding/8.a kv life.md",
202-
"projects/nebula-vsearch/img/metad.png",
203-
"projects/nebula-vsearch/understanding/7.how to modify sql.md",
204-
"projects/nebula-vsearch/understanding/6.raft-wal.md",
205-
"projects/nebula-vsearch/img/exec.png",
206-
"projects/nebula-vsearch/img/processor.png",
207-
"projects/nebula-vsearch/understanding/img/exec.png",
208-
"projects/nebula-vsearch/understanding/img/LRU.png",
209-
"projects/nebula-vsearch/understanding/img/folly_async.png",
210-
"projects/nebula-vsearch/understanding/img/doPut.png",
211-
"projects/nebula-vsearch/understanding/5.nGQL life.md",
212-
"projects/nebula-vsearch/understanding/4.folly future promise.md",
213-
"projects/nebula-vsearch/understanding/3.concurrent lru cache.md",
214-
"projects/nebula-vsearch/understanding/2.memory management.md",
215-
"projects/nebula-vsearch/understanding/1.visitor pattern.md",
216-
"index.md",
217-
"projects/nebula-vsearch/Untitled",
218-
"projects/nebula-vsearch/summary.md",
219-
"notes/CSAPP/5-优化程序性能.md",
220-
"blogs/posts/img",
221-
"blogs/posts/rocksdb.md",
222-
"blogs/posts/Implement of Concurrent.md",
223-
"blogs/posts/CS144.md",
224-
"blogs/posts/C++异步方案.md",
225-
"blogs/posts/bustub通关指北.md",
226-
"blogs/posts/bision debug.md",
227-
"blogs/posts",
228-
"projects/nebula-vsearch/understanding/img",
229-
"projects/nebula-vsearch/understanding",
230-
"blogs/CS144.md",
231-
"blogs/bision debug.md",
232-
"blogs/img",
233-
"blogs/bustub通关指北.md",
234-
"projects/nebula-vsearch/img",
185+
"notes/nebula-vsearch/understanding/img/vector.png",
186+
"notes/nebula-vsearch/understanding/img/processor.png",
187+
"notes/nebula-vsearch/understanding/img/memory.png",
188+
"notes/nebula-vsearch/understanding/img/LRU.png",
189+
"notes/nebula-vsearch/understanding/img/folly_async.png",
190+
"notes/nebula-vsearch/understanding/img/exec.png",
191+
"notes/nebula-vsearch/understanding/img/doPut.png",
192+
"notes/nebula-vsearch/understanding/img",
193+
"notes/nebula-vsearch/understanding/8.a kv life.md",
194+
"notes/nebula-vsearch/understanding/7.how to modify sql.md",
195+
"notes/nebula-vsearch/understanding/6.raft-wal.md",
196+
"notes/nebula-vsearch/understanding/5.nGQL life.md",
197+
"notes/nebula-vsearch/understanding/4.folly future promise.md",
198+
"notes/nebula-vsearch/understanding/3.concurrent lru cache.md",
199+
"notes/nebula-vsearch/understanding/2.memory management.md",
200+
"notes/nebula-vsearch/understanding/1.visitor pattern.md",
201+
"notes/nebula-vsearch/implements/wal for vector type.md",
202+
"notes/nebula-vsearch/implements/Match for vector property.md",
203+
"notes/nebula-vsearch/implements/DDL for vector type.md",
204+
"notes/nebula-vsearch/implements/DML for vector type.md",
205+
"notes/nebula-vsearch/implements/Create Ann Index.md",
206+
"notes/nebula-vsearch/上篇:初识 Nebula Graph —— 向量类型支持.md",
207+
"notes/nebula-vsearch/understanding",
208+
"notes/nebula-vsearch/summary.md",
209+
"notes/nebula-vsearch/implements",
210+
"notes/nebula-vsearch/img",
211+
"notes/nebula-vsearch",
212+
"notes/tiny-llm/RMSNorm & MLP.md",
213+
"notes/tiny-llm/img",
214+
"notes/tiny-llm/Group Query Attention.md",
215+
"notes/tiny-llm/Batching Inference & KV Cache.md",
216+
"notes/tiny-llm",
217+
"projects/tiny-llm/Group Query Attention.md",
218+
"projects/tiny-llm/RMSNorm & MLP.md",
219+
"projects/tiny-llm/img",
220+
"projects/tiny-llm/Batching Inference & KV Cache.md",
221+
"projects/sglang/SGLang Schedular 技术变迁.md",
235222
"Untitled.canvas",
236223
"Untitled 1.canvas"
237224
]
Lines changed: 61 additions & 0 deletions
Original file line numberDiff line numberDiff line change
@@ -0,0 +1,61 @@
1+
---
2+
title: Roofline & Basic Analysis
3+
date: 2025/10/20 23:15
4+
tags:
5+
- LLMInference
6+
---
7+
8+
# Roofline & Basic Analysis
9+
10+
## AI
11+
12+
## Roofline
13+
14+
Roofline 模型(屋顶线模型)是一种用来**分析程序性能瓶颈**(计算受限还是带宽受限)的方法。
15+
它把**计算性能**(FLOPs/s)和**访存性能**(Bytes/s)联系在一起,以可视化的方式展示性能上限。
16+
17+
$$
18+
Achievable FLOPs=min(AI×Memory BW,Peak FLOPs)
19+
$$
20+
21+
### 以 vector add 为例
22+
23+
- 最简单的累加方式,每个 thread 负责一个线程的计算
24+
25+
```cpp
26+
// FP32
27+
// ElementWise Add grid(N/256),
28+
// block(256) a: Nx1, b: Nx1, c: Nx1, c = elementwise_add(a, b)
29+
__global__ void vector_add_kernel(const float *a, const float *b, float *c,
30+
                                  int n) {
31+
  int idx = blockIdx.x * blockDim.x + threadIdx.x;
32+
  if (idx < n) {
33+
    c[idx] = a[idx] + b[idx];
34+
  }
35+
}
36+
```
37+
38+
| 指标 | 数值 |
39+
| -------------- | ---------------- |
40+
| 每元素 FLOPs | 1 |
41+
| 每元素 Bytes | 12 |
42+
| AI | 0.083 FLOPs/Byte |
43+
| Peak Bandwidth | 1008 GB/s |
44+
| Peak Compute | 82.6 TFLOPs/s |
45+
46+
```shell
47+
性能 (GFLOPs/s)
48+
49+
| ──────────────── ← 82.6 TFLOPs/s (平顶线)
50+
| /
51+
| /
52+
| o /
53+
| (VectorAdd) /
54+
|______________/__________________→ AI (FLOPs/Byte)
55+
0.083
56+
57+
```
58+
59+
- 左下角:访存主导(memory-bound)
60+
- 右上角:计算主导(compute-bound)
61+
- 中间交点:分界点(称为 **ridge point**

0 commit comments

Comments
 (0)