I got following message during checking FS (fsck.ext4 -fy):
(many thousands FS errors fixed early)
...
Inode 169246461, i_blocks is 211670140568577, should be 0. Fix? yes
Inode 169246423 has illegal block(s). Clear? yes
Illegal block #2 (2914889194) in inode 169246423. CLEARED.
Illegal block #7 (3221425160) in inode 169246423. CLEARED.
Illegal block #10 (3963578733) in inode 169246423. CLEARED.
Illegal block #11 (3770804058) in inode 169246423. CLEARED.
Illegal indirect block (2164754208) in inode 169246423. CLEARED.
Illegal double indirect block (4118543488) in inode 169246423. CLEARED.
Illegal double indirect block (3120627712) in inode 169246423. CLEARED.
Illegal double indirect block (2298555701) in inode 169246423. CLEARED.
Illegal indirect block (3580183415) in inode 169246423. CLEARED.
Illegal indirect block (2538272501) in inode 169246423. CLEARED.
Illegal block #4197388 (3120627712) in inode 169246423. CLEARED.
Error storing directory block information (inode=169246423, block=0, num=3966024): Memory allocation failed
e2fsck: aborted
memory:
total used free shared buffers cached
Mem: 8197852 1827500 6370352 0 42280 514132
-/+ buffers/cache: 1271088 6926764
Swap: 0 0 0
dumpe2fs:
Filesystem volume name: media-data
Last mounted on: /srv/data
Filesystem UUID: 14e2daa4-8fca-472e-a122-06a8fab0a41f
Filesystem magic number: 0xEF53
Filesystem revision #: 1 (dynamic)
Filesystem features: has_journal ext_attr resize_inode dir_index filetype extent flex_bg sparse_super large_file huge_file uninit_bg dir_nlink extra_isize
Filesystem flags: signed_directory_hash
Default mount options: (none)
Filesystem state: clean with errors
Errors behavior: Continue
Filesystem OS type: Linux
Inode count: 469475328
Block count: 1877879808
Reserved block count: 0
Free blocks: 32153724
Free inodes: 469247108
First block: 0
Block size: 4096
Fragment size: 4096
Reserved GDT blocks: 576
Blocks per group: 32768
Fragments per group: 32768
Inodes per group: 8192
Inode blocks per group: 512
RAID stride: 32751
Flex block group size: 16
Filesystem created: Fri Aug 13 23:24:15 2010
Last mount time: Sat Feb 19 16:15:12 2011
Last write time: Sat Feb 19 16:15:56 2011
Mount count: 42
Maximum mount count: 20
Last checked: Wed Aug 18 10:53:33 2010
Check interval: 15552000 (6 months)
Next check after: Mon Feb 14 09:53:33 2011
Lifetime writes: 8104 GB
Reserved blocks uid: 0 (user root)
Reserved blocks gid: 0 (group root)
First inode: 11
Inode size: 256
Required extra isize: 28
Desired extra isize: 28
Journal inode: 8
Default directory hash: half_md4
Directory Hash Seed: f0a9976d-8c6b-462f-9e94-4e7b2631599b
Journal backup: inode blocks
Journal features: journal_incompat_revoke
Journal size: 128M
Journal length: 32768
Journal sequence: 0x00022bf9
Journal start: 0
^C
At second retry I watch free and all time have at least 6GiB memory free but got following message again:
...
Illegal double indirect block (4219387801) in inode 169345606. CLEARED.
Inode 169345606 is too big. Truncate? yes
Block #1049612 (1538191935) causes directory to be too big. CLEARED.
Error storing directory block information (inode=169345606, block=0, num=305115): Memory allocation failed
e2fsck: aborted
And more: Before running tests I was able to mount and use FS, right now I'm getting message during mounting: [ 7855.416191] EXT4-fs (dm-0): ext4_check_descriptors: Checksum for group 768 failed (24562!=0) [ 7855.416198] EXT4-fs (dm-0): group descriptors corrupted! and every fs check One or more block group descriptor checksums are invalid. Fix? yes Group descriptor 768 checksum is invalid. FIXED. Group descriptor 770 checksum is invalid. FIXED. Group descriptor 772 checksum is invalid. FIXED. Group descriptor 774 checksum is invalid. FIXED. Group descriptor 776 checksum is invalid. FIXED. Group descriptor 778 checksum is invalid. FIXED. Group descriptor 780 checksum is invalid. FIXED. Group descriptor 782 checksum is invalid. FIXED. Group descriptor 784 checksum is invalid. FIXED. Group descriptor 786 checksum is invalid. FIXED. Group descriptor 788 checksum is invalid. FIXED. Group descriptor 896 checksum is invalid. FIXED. Group descriptor 15873 checksum is invalid. FIXED. Group descriptor 23040 checksum is invalid. FIXED. Group descriptor 23041 checksum is invalid. FIXED. Group descriptor 23042 checksum is invalid. FIXED. Group descriptor 23043 checksum is invalid. FIXED. Group descriptor 24448 checksum is invalid. FIXED. Group descriptor 25472 checksum is invalid. FIXED. Group descriptor 25477 checksum is invalid. FIXED. Group descriptor 25478 checksum is invalid. FIXED. Group descriptor 25479 checksum is invalid. FIXED. Group descriptor 44290 checksum is invalid. FIXED. Group descriptor 44291 checksum is invalid. FIXED. Group descriptor 44292 checksum is invalid. FIXED. Group descriptor 44293 checksum is invalid. FIXED. media-data contains a file system with errors, check forced. Pass 1: Checking inodes, blocks, and sizes ^Cmedia-data: e2fsck canceled. It repeat every time I do fsck...
retitle 614082 e2fsck dies with an OOM on bad triple-indirect block in a dir inode thanks Apologies for not looking at this sooner. I had assumed this was the standard "user was trying to check a too-big file system with an 32-bit userspace" problem, and then I got busy and it dropped off the end of my inbox. In fact it's a much more interesting problem. What happened was that you got unlucky and trash got written into your inode table. One of the inodes that got overwritten had the directory bit set, and the extents flag was _not_ set. So the i_blocks field was interpreted as a series of blocks. Many of those i_blocks entries were garbage, and were cleared, but the triple indirect block could be interpreted as a valid block number, so we started iterating over all of the double indirection blocks. To make sure that the directory doesn't have "holes" in it, e2fsck then started allocating entries to keep track of all of this via ext2fs_add_dir_block2(), and this is what caused the memory overflow. To fix this what we need to do is to notice when an inode looks insane, and ask the user if they are willing to just nuke it, as opposed to trying to fix all of the individual things that are wrong with it. The way to work around this problem should it occur again is to note the inode number, and then zap it using debugfs: debugfs -w /dev/md1 debugfs: clri <169246423> debugfs: quit Then e2fsck should be able to complete its work. - Ted