7.3.1
My site went down recently and I was corresponding with my hosts. These are the last two replies I got. It's all very amiable, but I wonder if someone can comment on these posts...
lrwxrwxrwx 1 root root 0 Aug 24 12:46 cwd ->
/nfs/mysite.com/forums
ls: cannot read symbolic link /proc/293/exe: No such file or directory
lrwxrwxrwx 1 root root 0 Aug 24 12:46 cwd ->
/nfs/mysite.com/forums
lrwxrwxrwx 1 root root 0 Aug 24 12:46 cwd ->
/nfs/mysite.com/forums
lrwxrwxrwx 1 root root 0 Aug 24 12:46 cwd ->
/nfs/mysite.com/forums
lrwxrwxrwx 1 root root 0 Aug 24 12:46 cwd ->
/nfs/mysite.com/forums
lrwxrwxrwx 1 root root 0 Aug 24 12:46 cwd ->
/nfs/mysite.com/forums
.
.
.
.
long list of identical entries
.
.
lrwxrwxrwx 1 root root 0 Aug 24 12:46 cwd ->
/nfs/mysite.com/forums
lrwxrwxrwx 1 root root 0 Aug 24 12:46 cwd ->
/nfs/mysite.com/forums
What was forwarded there shows that the cause for this mornings outage was the site with the most occupancy. While you may have not changed the site, use of, or abuse of the site caused the machine to take itself out of commission.
Please can you review your logs over this period and take steps to throttle or analyse the cause of stuck process/occupancy so we can work together on preventing such an occurrence in future.
I think that the issue we see is that the site is taking about 50% of
all traffic for that server cluster. Whenever the cluster load increases or the servers crash 50% of the time your site is the one that has created 250+ apache processes.
This isn't a 'new' issue or one that has developed recently, it's always been the case that any burst in traffic causes the servers to slow down and if it's a large enough flood of traffic, to crash. Usually it's a spider or search engine at fault but ultimately as things stand we're having to take emergency trouble tickets at all hours of the day and night to fix the servers and typically that means one of us being got out of bed at 3am to reboot one of the cluster boxes.
There are two options that we need to work towards considering. One is a dedicated server of your own to run the site from. The other is a
higher spec/lower density cluster.
Any thoughts?