August 17, 2007

"Kiss the Webmaster" Day

  • Alive Not Dead archive

I was rudely derailed from getting an early and good night’s sleep last night when, around 1:20am, I was unable to connect to alivenotdead. After about fifteen minutes of diagnosing what the problem is (and simultaneous alert messages from Etchy and Jojo – wow, are you guys ALWAYS on AnD?), I figured out that it was very likely a network issue rather than a web site or server issue. The mostly likely cause was a faulty router somewhere within our web host’s facilities.

Let me tell you now: If you’ve never done any tech-based administration like managing servers or desktop computers, than you don’t know just how horrible HORRIBLE this job is. It’s literally the worst part of my job. The only thing you hear about your performance is when something is going wrong – there’s NEVER any GOOD news related to your responsibilities because people EXPECT 100% uptime and quick performance. There are plenty of you out there who have probably done this type of job before so you probably feel my pain. I also had this responsibility when I was at Rotten Tomatoes which was about five times more complex and frustrating (managing 28 servers instead of just one).


Nick Burns, the Computer Guy – with Jackie Chan… not quite server administering, but you get the picture why we turn into such assholes…


Video: http://www.youtube.com/watch?v=R7RuolTf9ho



Well, the next step after diagnosing the problem is to contact the 24-hour support line at our web hosting provider. I call them and, suprisingly, it goes to voicemail. Huh??? I leave a message and follow up with an email to their support email and then try calling again multiple times with the same non-results. At this point, I’m pretty frantic because it’s been an hour with the site being inaccessible so I call the mobile number of our sales representative for our web hosting provider and it ALSO goes to voicemail.

At this point, I’m pretty much out of options but I literally want to do physical violence to someone at our hosting provider.

Incidentally, this is the second time that this exact scenario has happened this week. Some of you might have noticed that the site was down for a period of about an hour three days ago and back then I also didn’t receive any tech support assistance whatsoever. After that fiasco, I wrote an email to our sales rep asking for a standard “SLA” (Service Level Agreement) clause (basically, they would have to refund us money each time they didn’t provided adequate service) added to our contract OR we would seek to terminate early for breach of contract.

Sure enough, 48 hours later and the same exact problem happens except this time the outage is for a full SIX hours. We didn’t come back online until 7am.

I went to bed very VERY grumpy.

Flash forward to 11am and I FINALLY get a call back from the web hosting provider – they say that they got my message (really? because I left about five of them eight hours earlier) and that the cause of the problem was a router issue in the data center that they were in (basically, the company that houses all of THEIR servers). The sales rep seemed surprised that I didn’t get any response to my tech support calls but didn’t apologize or promise any future support.

We selected a web hosting provider here in Hong Kong because of the importance of having adequate performance for our users in the mainland, Hong Kong, and Taiwan but I just don’t see the benefits when the service is so crappy.

I did some research earlier (right after the first outage earlier this week) and found us a new hosting provider in the U.S.. As a consequence, the [alive not dead] site will be physically moving over to the U.S. sometime early next week fingers crossed. It’s going to be more expensive and slightly slower for our Asia users, but hopefully more reliable and lower blood pressure on my part. Not too coincidentally, it’s using the same data center company (Savvis) that Rotten Tomatoes is now currently in although we’ll be using a web hosting provider rather than building/installing the servers ourselves like we did at Rotten Tomatoes (since I’m not physically in the U.S.). The new company comes from a referral from another acquaintance in Hong Kong running an online community who was similarly fed up with the crappy level of support here in HK.

I’ve done some preliminary tests, and while it WILL be slightly slower service in Hong Kong (by about .2 seconds per request, which is pretty noticeable), the difference in performance should actually be neglible for users in mainland China (by about .03 seconds for each request) and MUCH faster for our North American, Japan, and European users. Our long-term solution is to eventually host a separate set of servers within mainland China so that those users (now up to 40%) can access the site quickly without having to pull it throught the Great Firewall which adds on about an additional .2 seconds per request to process.

Also, when we switch over the servers, the site will likely become inaccessible for about three hours around the time we’re switching it as we port the database over and allow the old domain name entry lookups to expire.

In any case, I’m sorry if this post is especially nerdy. I guess I just needed to blow off some steam after being so frustrated with our hosting provider. Hopefully, these infrastructure changes will allow Alive Not Dead to grow quicker…

Incidentally, I want to give a shoutout to our newest artist, Shan Chen, singer, dancer, asset manager, computer science graduate (yay!), and my former East Coast homegirl who I’ve known for several years now. Please “become a fan” of hers on her artist profile page because I expect her to be successful in anything that she aims to do…


Shan get’s checked out by a very naughty monkey…

When I first met her, she ALSO had a thankless job as an IT applications administrator for Citibank (she’s multitalented) so I’m sure she feels my pain. She’s moved on to bigger and brighter things and I’m still here messing with dumb, broken servers…