排查内存问题的时候,第一件事就是去 /proc 下面翻文件。Linux 内核在这里暴露了大量内存管理的统计信息。这篇整理了几个最常用的文件,以及每个文件里哪些字段值得关注。
/proc/zoneinfo
root@firefly:~# cat /proc/zoneinfo
Node 0, zone DMA
pages free 808251
min 1912
low 2390
high 2868
scanned 0
spanned 1015296
present 1007616
managed 917418
nr_free_pages 808251
nr_alloc_batch 69
nr_inactive_anon 1878
nr_active_anon 35777
nr_inactive_file 36299
nr_active_file 12876
nr_unevictable 0
nr_mlock 0
nr_anon_pages 35598
nr_mapped 29010
nr_file_pages 51233
nr_dirty 0
nr_writeback 0
nr_slab_reclaimable 8099
nr_slab_unreclaimable 5837
nr_page_table_pages 1105
nr_kernel_stack 326
nr_overhead 0
nr_unstable 0
nr_bounce 0
nr_vmscan_write 0
nr_vmscan_immediate_reclaim 0
nr_writeback_temp 0
nr_isolated_anon 0
nr_isolated_file 0
nr_shmem 2059
nr_dirtied 3391
nr_written 3384
nr_pages_scanned 0
workingset_refault 0
workingset_activate 0
workingset_nodereclaim 0
nr_anon_transparent_hugepages 0
nr_free_cma 0
protection: (0, 0, 0)
pagesets
cpu: 0
count: 33
high: 186
batch: 31
vm stats threshold: 36
cpu: 1
count: 58
high: 186
batch: 31
vm stats threshold: 36
cpu: 2
count: 93
high: 186
batch: 31
vm stats threshold: 36
cpu: 3
count: 139
high: 186
batch: 31
vm stats threshold: 36
cpu: 4
count: 130
high: 186
batch: 31
vm stats threshold: 36
cpu: 5
count: 185
high: 186
batch: 31
vm stats threshold: 36
all_unreclaimable: 0
start_pfn: 512
inactive_ratio: 5
这里显示的是 DMA 这个 zone 的信息。各字段含义:
| 字段 | 说明 |
|---|---|
pages free |
可用空闲页数 |
min / low / high |
空闲页的三条水位线 |
spanned |
zone 总物理页数 |
present |
当前存在的物理页数 |
managed |
内核管理的物理页数 |
nr_free_pages |
空闲页数(和 pages free 一回事) |
nr_inactive_anon / nr_active_anon |
非活动/活动匿名页 |
nr_inactive_file / nr_active_file |
非活动/活动文件页 |
nr_slab_reclaimable / nr_slab_unreclaimable |
可回收和不可回收的 slab |
nr_shmem |
共享内存页 |
nr_dirty / nr_writeback |
脏页和写回中的页 |
nr_anon_pages / nr_file_pages |
匿名页和文件页总数 |
排查时多看一眼的:
pages free
-
- 接近 0 → 内存压力大或泄露
nr_slab_unreclaimable
-
- 持续涨 → 可能是 slab 泄漏
nr_shmem
-
- 过高 → 共享内存占太多,有 OOM 风险
nr_dirty
-
- 一直在高位 → IO 压力大
nr_file_pages / nr_anon_pages
- 比例变化异常 → 定位一下是文件缓存还是进程私有多
/proc/pagetypeinfo
root@firefly:~# cat /proc/pagetypeinfo
Page block order: 9
Pages per block: 512
Free pages count per migrate type at order 0 1 2 3 4 5 6 7 8 9 10
Node 0, zone DMA, type Unmovable 18 10 2 0 1 1 1 0 1 1 0
Node 0, zone DMA, type Movable 0 1 0 0 0 1 1 1 0 0 832
Node 0, zone DMA, type Reclaimable 1 0 1 0 1 1 1 0 1 1 0
Node 0, zone DMA, type HighAtomic 0 0 0 0 0 0 0 0 0 0 0
Number of blocks type Unmovable Movable Reclaimable HighAtomic
Node 0, zone DMA 26 1924 18 0
每个 block 512 页(2^9)。表里按迁移类型和 order 列了空闲页数量。
Unmovable — 不能移动的页
Movable — 可以移动的页
Reclaimable — 可以回收的页
HighAtomic — 高原子分配页
如果 Unmovable 散在各处、高 order 的空闲块凑不出来,基本就是内存碎片了。
/proc/meminfo
root@firefly:~# cat /proc/meminfo
MemTotal: 3669672 kB
MemFree: 3233664 kB
MemAvailable: 3432196 kB
Buffers: 10192 kB
Cached: 194796 kB
SwapCached: 0 kB
Active: 194344 kB
Inactive: 152704 kB
Active(anon): 142788 kB
Inactive(anon): 7508 kB
Active(file): 51556 kB
Inactive(file): 145196 kB
Unevictable: 0 kB
Mlocked: 0 kB
SwapTotal: 0 kB
SwapFree: 0 kB
Dirty: 28 kB
Writeback: 0 kB
AnonPages: 142160 kB
Mapped: 116020 kB
Shmem: 8232 kB
Slab: 55688 kB
SReclaimable: 32372 kB
SUnreclaim: 23316 kB
KernelStack: 5216 kB
PageTables: 4400 kB
NFS_Unstable: 0 kB
Bounce: 0 kB
WritebackTmp: 0 kB
CommitLimit: 1834836 kB
Committed_AS: 1375504 kB
VmallocTotal: 258867136 kB
VmallocUsed: 0 kB
VmallocChunk: 0 kB
HugePages_Total: 0
HugePages_Free: 0
HugePages_Rsvd: 0
HugePages_Surp: 0
Hugepagesize: 2048 kB
这是最常用的,free -m 的数据就来自这里。每个字段:
| 字段 | 含义 |
|---|---|
| MemTotal | 总物理内存 |
| MemFree | 完全空闲的 |
| MemAvailable | 新程序能用的估算值(比 MemFree 更准) |
| Buffers | 块设备缓冲 |
| Cached | 文件缓存(压力大时可回收) |
| SwapCached | 被换出到 swap 后又读回,但还没完全清掉的 |
| Active / Inactive | 活跃/非活跃内存 |
| Slab | 内核 slab 总量,拆 SReclaimable + SUnreclaim |
| CommitLimit | 系统承诺可分配的上限 |
| Committed_AS | 已经承诺出去的量,接近 CommitLimit 时 OOM 风险高 |
排查重点:
MemFree 和 MemAvailable 同时很低 → 内存吃紧
SwapCached 不为 0 → 物理内存不够,在频繁换页
Slab 总量过大,特别是 SUnreclaim 部分 → 内核对象可能泄露
Committed_AS 接近 CommitLimit → 随时可能 OOM
/proc/buddyinfo
root@firefly:~# cat /proc/buddyinfo
Node 0, zone DMA 28 9 9 6 1 3 1 1 2 2 787
每列对应 order 0 ~ 10 的空闲块数量。如果前几列数字大、后面列数字小,说明内存碎片比较严重。
/proc/slabinfo
root@firefly:~# cat /proc/slabinfo
slabinfo - version: 2.1
# name <active_objs> <num_objs> <objsize> <objperslab> <pagesperslab> : tunables ...
dentry 36017 36198 216 18 1 : tunables 0 0 0 : slabdata 2011 2011 0
inode_cache 22152 22152 680 24 4 : tunables 0 0 0 : slabdata 923 923 0
ext4_inode_cache 3425 3425 1296 25 8 : tunables 0 0 0 : slabdata 137 137 0
...
每个缓存的行结构:名称 活跃对象数 总对象数 对象大小 每slab对象数 每slab页数。
排查 slab 泄漏时,盯着某个缓存的 active_objs 看。如果它在系统空闲时还在涨,基本就是问题所在。常见的如 dentry、inode_cache、kmalloc-* 漏起来很要命。
/proc/vmallocinfo
0xffffff8008000000-0xffffff8008011000 69632 of_iomap+0x48/0x5c phys=fee00000 ioremap
0xffffff8008018000-0xffffff800801a000 8192 bpf_prog_alloc+0x48/0xb4 pages=1 vmalloc
0xffffffbdbff72000-0xffffffbdbfff0000 516096 pcpu_get_vm_areas+0x0/0x500 vmalloc
格式:起始地址-结束地址 大小 分配函数 物理地址(ioremap情况) 分配类型。通过这个文件可以看谁用 vmalloc 或 ioremap 占了虚拟内存。
/proc/vmstat
nr_free_pages 807632
nr_alloc_batch 478
nr_inactive_anon 1876
nr_active_anon 35696
...
pgpgin 212188
pgpgout 17004
pswpin 0
pswpout 0
pgmajfault 1418
...
全局内存事件统计。关注点:
pswpin / pswpout — 非零说明物理内存不够了,在大量换页
pgmajfault — 主缺页次数,偏高说明内存紧张
allocstall — 分配停滞次数,非零表示分配压力大
/proc/self/statm
927 111 95 7 0 108 0
当前进程的内存使用(单位 page,通常 4KB):
| 列 | 含义 |
|---|---|
| 1 | 虚拟内存总页数 |
| 2 | 常驻内存(RSS) |
| 3 | 共享页(共享库等) |
| 4 | 代码段 |
| 5 | 数据/堆 |
| 6 | 栈 |
| 7 | swap 出去的页数 |
/proc/self/maps
5578fe8000-5578fef000 r-xp 00000000 b3:07 181803 /bin/cat
558ad0e000-558ad2f000 rw-p 00000000 00:00 0 [heap]
7f9e147000-7f9e286000 r-xp 00000000 b3:07 91231 /lib/aarch64-linux-gnu/libc-2.27.so
...
7fee78b000-7fee7ac000 rw-p 00000000 00:00 0 [stack]
每行结构:起始-结束 权限 偏移 设备:inode 路径。权限字段 r=读 w=写 x=执行 p=私有 s=共享。
排查内存布局问题时,看有没有不合理的巨大映射,或者过多的匿名映射。
/proc/sys/vm
root@firefly:~# ls /proc/sys/vm
dirty_background_ratio dirty_ratio min_free_kbytes
drop_caches max_map_count overcommit_memory
swappiness vfs_cache_pressure ...
这里是可以调整的内存管理参数,几个常用的:
| 参数 | 作用 |
|---|---|
dirty_ratio |
脏页超过内存多少百分比时触发写回 |
dirty_background_ratio |
脏页超过多少时后台进程开始写 |
drop_caches |
写 1/2/3 清除页缓存/dentry/inode 缓存 |
min_free_kbytes |
保留的最小空闲内存 |
swappiness |
越倾向于使用 swap(0~100) |
overcommit_memory |
0=启发式 1=总是允许 2=禁止超额 |
vfs_cache_pressure |
越激进回收 VFS 缓存(值越大) |
/proc/swaps
Filename Type Size Used Priority
/dev/sda2 partition 8388604 0 -1
看交换分区用掉了多少。Used 接近 Size 说明物理内存压力很大。
参考:https://www.cnblogs.com/arnoldlu/p/8568330.html
1246