system woes likely due to hardware/fs problems

Ian Lance Taylor ian@wasabisystems.com
Tue Apr 6 21:15:00 GMT 2004


Angela Marie Thomas <angela@foam.wonderslug.com> writes:

> Geoff and I were in Australia for 3 weeks so I'm behind in overseers
> mail but did notice the recent htdig discussion and my backups today
> weren't able to complete in a timely fashion either.  I took a look at
> the system and eventually was able to find the following errors:
> 
> Apr  4 22:00:03 sourceware kernel: EXT3-fs error (device lvm(58,7)): ext3_add_entry: bad entry in directory #107864: rec_len %% 4 != 0 - offset=0, inode=1296707338, rec_len=28261, name_len=117

...

> This may indicate more serious hardware issues.  I think it would be 
> prudent to take the system down to verify the consistency of the fileystems
> and take a look at the RAID diagnostics to see whether we have hardware
> issues or just a random glitch causing woe.

I also noticed those errors, but they are all on 58,7 which is
/sourceware/cvs-tmp, which is not very interesting, and is certainly
vulnerable to things like a directory getting rm -rf'ed by two
different processes simultaneously.  But if they could indicate
hardware problems, then more investigation would obviously be
appropriate.

To me it does seem that the system is mainly waiting for disk I/O, and
that the processes waiting are mainly running CVS.  But that is more
of a gut feeling based on top and vmstat, not something I can really
prove.

I looked at some network packets, and on a packet count basis, not a
data size basis, I saw numbers like this:
    SMTP: 32%
    HTTP: 27%
    FTP: 27%
    CVS: 12%

As can be seen, FTP seems to be a pretty big network user, and could
most easily be pushed off to mirror sites.  But I don't know that
network usage is our real problem here.

I note that the bandwidth reports at
    http://sourceware.org/sourceware/bandwidth/
appear to no longer be updated.

Ian



More information about the Overseers mailing list