Story line of the Amazon EC2 and RDS failure and recovery

I wanted to read the timeline of the Amazon “Networking Event”.  So I’ve taken the logs of the EC2 and RDS status updates and put them together for a post.

If you want to learn about how to make a more robust Amazon Web Services (AWS) configuration, read my article on the WebDevStudios blog.

RDS Apr 21, 1:48 AM PDT We are currently investigating connectivity and latency issues with RDS database instances in the US-EAST-1 region.

RDS Apr 21, 2:16 AM PDT We can confirm connectivity issues impacting RDS database instances across multiple availability zones in the US-EAST-1 region.

RDS Apr 21, 3:05 AM PDT We are continuing to see connectivity issues impacting some RDS database instances in multiple availability zones in the US-EAST-1 region. Some Multi AZ failovers are taking longer than expected. We continue to work towards resolution.

RDS Apr 21, 4:03 AM PDT We are making progress on failovers for Multi AZ instances and restore access to them. This event is also impacting RDS instance creation times in a single Availability Zone. We continue to work towards the resolution.

RDS Apr 21, 5:06 AM PDT IO latency issues have recovered in one of the two impacted Availability Zones in US-EAST-1. We continue to make progress on restoring access and resolving IO latency issues for remaining affected RDS database instances.

RDS Apr 21, 6:29 AM PDT We continue to work on restoring access to the affected Multi AZ instances and resolving the IO latency issues impacting RDS instances in the single availability zone.

RDS Apr 21, 8:12 AM PDT Despite the continued effort from the team to resolve the issue we have not made any meaningful progress for the affected database instances since the last update. Create and Restore requests for RDS database instances are not succeeding in US-EAST-1 region.

RDS Apr 21,10:35 AM PDT We are making progress on restoring access and IO latencies for affected RDS instances. We recommend that you do not attempt to recover using Reboot or Restore database instance APIs or try to create a new user snapshot for your RDS instance – currently those requests are not being processed.

RDS Apr 21, 2:35 PM PDT We have restored access to the majority of RDS Multi AZ instances and continue to work on the remaining affected instances. A single Availability Zone in the US-EAST-1 region continues to experience problems for launching new RDS database instances. All other Availability Zones are operating normally. Customers with snapshots/backups of their instances in the affected Availability zone can restore them into another zone. We recommend that customers do not target a specific Availability Zone when creating or restoring new RDS database instances. We have updated our service to avoid placing any RDS instances in the impaired zone for untargeted requests.

RDS Apr 21, 2:41 11:42 PM PDT In line with the most recent Amazon EC2 update, we wanted to let you know that the team continues to be all-hands on deck working on the remaining database instances in the single affected Availability Zone. It’s taking us longer than we anticipated. When we have an updated ETA or meaningful new update, we will make sure to post it here. But, we can assure you that the team is working this hard and will do so as long as it takes to get this resolved.

RDS Apr 22, 2:41 7:08 AM PDT In line with the most recent Amazon EC2 update, we are making steady progress in restoring the remaining affected RDS instances. We expect this progress to continue over the next few hours and we’ll keep folks posted.

RDS Apr 22, 2:41 2:43 PM PDT We are continuing to make progress in restoring access to the remaining affected RDS instances. We expect this progress to continue over the next few hours and we’ll keep folks posted.

EC2 Apr 22, 2:41 AM PDT We continue to make progress in restoring volumes but don’t yet have an estimated time of recovery for the remainder of the affected volumes. We will continue to update this status and provide a time frame when available.
EC2 Apr 22, 6:18 AM PDT We’re starting to see more meaningful progress in restoring volumes (many have been restored in the last few hours) and expect this progress to continue over the next few hours. We expect that well reach a point where a minority of these stuck volumes will need to be restored with a more time consuming process, using backups made to S3 yesterday (these will have longer recovery times for the affected volumes). When we get to that point, we’ll let folks know. As volumes are restored, they become available to running instances, however they will not be able to be detached until we enable the API commands in the affected Availability Zone.

EC2 Apr 22, 8:49 AM PDT We continue to see progress in recovering volumes, and have heard many additional customers confirm that they’re recovering. Our current estimate is that the majority of volumes will be recovered over the next 5 to 6 hours. As we mentioned in our last post, a smaller number of volumes will require a more time consuming process to recover, and we anticipate that those will take longer to recover. We will continue to keep everyone updated as we have additional information.

EC2 Apr 22, 2:15 PM PDT In our last post at 8:49am PDT, we said that we anticipated that the majority of volumes “will be recovered over the next 5 to 6 hours.” These volumes were recovered by ~1:30pm PDT. We mentioned that a “smaller number of volumes will require a more time consuming process to recover, and we anticipate that those will take longer to recover.” We’re now starting to work on those. We’re also now working to enable customers to be able to launch EBS backed instances and create, delete, attach and detach EBS volumes in the affected Availability Zone. Our current estimate is that this will take 3-4 hours until full access is restored. We will continue to keep everyone updated as we have additional information.

EC2 Apr 22, 6:27 PM PDT We’re continuing to work on restoring the remaining affected volumes. The work we’re doing to enable customers to be able to launch EBS backed instances and create, delete, attach and detach EBS volumes in the affected Availability Zone is taking considerably more time than we anticipated. The team is in the midst of troubleshooting a bottleneck in this process and we’ll report back when we have more information to share on the timing of this functionality being fully restored.

EC2 Apr 22, 9:11 PM PDT We wanted to give a more detailed update on the state of our recovery. At this point, we have recovered a large number of the stuck volumes and are in the process of recovering the remainder. We have added significant storage capacity to the cluster, and storage capacity is no longer a bottleneck to recovery. Some portion of these volumes have lost the connection to their instance, and are waiting to be connected before normal operations can resume. In order to re-establish this connection, we need to allow the instances in the affected Availability Zone to access the EC2 control plane service. There are a large number of control plane requests being generated by the system as we re-introduce instances and volumes. The load on our control plane is higher than we anticipated. We are re-introducing these instances slowly in order to moderate the load on the control plane and prevent it from becoming overloaded and affecting other functions. We are currently investigating several avenues EC2 to unblock this bottleneck and significantly increase the rate at which we can restore control plane access to volumes and instances– and move toward a full recovery. The team has been completely focused on restoring access to all customers, and as such has not yet been able to focus on performing a complete post mortem. Once our customers have been taken care of and are fully back up and running, we will post a detailed account of what happened, along with the corrective actions we are undertaking to ensure this doesn’t happen again. Once we have additional information on the progress that is being made, we will post additional updates.

RDS Apr 23, 12:00 AM PDT We are continuing to work on restoring access to the remaining affected RDS instances. We expect the restoration process to continue over the next several hours and we’ll update folks as we have new information.


EC2 Apr 23, 1:55 AM PDT We are continuing to work on unblocking the bottleneck that is limiting the speed with which we can re-establish connections between volumes and their instances. We will continue to keep everyone updated as we have additional information.

RDS Apr 23, 8:45 AM PDT We have made significant progress in resolving stuck IO issues and restoring access to RDS database instances and now have the vast majority of them back operational again. We continue to work on restoring access to the small number of remaining affected instances and we’ll update folks as we have new information.

EC2 Apr 22, 8:54 AM PDT We have made significant progress during the night in manually restoring the remaining stuck volumes, and are continuing to work through the remainder. Additionally we have removed some of the bottlenecks that were preventing us from allowing more instances to re-establish their connection with the stuck volumes, and the majority of those instances and volumes are now connected. We’ve encountered an additional issue that’s preventing the recovery of the remainder of the connections from being established, but are making progress. Once we solve for this bottleneck, we will work on restoring full access for customers to the control plane.

EC2 Apr 22, 11:54 AM PDT Quick update. We’ve tried a couple of ideas to remove the bottleneck in opening up the APIs, each time we’ve learned more but haven’t yet solved the problem. We are making progress, but much more slowly than we’d hoped. Right now we’re setting up more control plane components that should be capable of working through the backlog of attach/detach state changes for EBS volumes. These are coming online, and we’ve been seeing progress on the backlog, but it’s still too early to tell how much this will accelerate the process for us. For customers who are still waiting for restoration of the EBS control plane capability in the impacted AZ, or waiting for recovery of the remaining volumes, we understand that no information for hours at a time is difficult for you. We’ve been operating under the assumption that people prefer us to post only when we have new information. Think enough people have told us that they prefer to hear from us hourly (even if we don’t have meaningful new information) that we’re going to change our cadence and try to update hourly from here on out.

EC2 Apr 22, 12:46 PM PDT We have completed setting up the additional control plane components and we are seeing good scaling of the system. We are now processing through the backlog of state changes and customer requests at a very quick rate. Barring any setbacks, we anticipate getting through the remainder of the backlog in the next hour. We will be in a brief hold after that, assessing whether we can proceed with reactivating the APIs.

RDS Apr 23, 12:54 PM PDT As we mentioned in our last update at 8:45 AM, we now have the vast majority of affected RDS instances back operational again. Since that post, we have continued to work on restoring access to the small number of remaining affected instances. RDS uses EBS, and as such, our pace of recovery is dependent on EBS’s recovery. As mentioned in the most recent EC2 post, EBSs recovery has gone a bit slower than anticipated in the last few hours. This has slowed down RDS recovery as well. We understand how significant this service interruption is for our affected customers and we are working feverishly to address the impact. We’ll update folks as we have new information. Additionally we have heard from customers that you prefer more frequent updates, even if there has been no meaningful progress. We have heard that feedback, and will try to post hourly updates here. Some of these updates will point to EC2’s updates (as they continue to recover the rest of EBS volumes), but we’ll post nonetheless.

HOW TO: How do you redirect all https traffic to http?

There are LOTS of ways to do this.

 

This is what I do so that I can use the same code in any site’s .htaccess

RewriteEngine On
RewriteBase /
Options +FollowSymlinks
RewriteCond %{HTTPS} =on
RewriteRule ^(.*)$ http://%{SERVER_NAME}/$1 [R=301,L]

 

I like this better than checking for port 443 because sometimes load balancers handle the certificate and encryption then send the information to port 80.   This works in that situation. Of course you can adjust this code to redirect to www.%{SERVER_NAME} if that is your preference.

 

Of course redirecting all https traffic to http is equally simple by adding a ! or not to the https check and adjusting your target.

RewriteEngine On
RewriteBase /
Options +FollowSymlinks
RewriteCond %{HTTPS} !=on
RewriteRule ^(.*)$ https://%{SERVER_NAME}/$1 [R=301,L]

 

 

Getting Started With BackPress

This summer I took the opportunity to write my first BackPress application.  It was an educational experience that came with a number of surprises. I thought I would share my experience with you, along with some tips to get you out of the gates with a BackPress project of your own.

My Project

In 2008, the comedic duo Rhett and Link created a YouTube contest called SuperNote. I created a PHP app to support this project. It created its own user database, collected videos from YouTube via the API and allowed users to login to and review the video and record, among other things, how long the video’s creator was able to hold a vocal note without interruption. It also allowed different levels of user interaction with different fields visible for each user role.  Everything about this project was written from scratch, with the exception of using WordPress 2.6’s wp-db.php for database access.  It worked alright. I used sessions for logging in and that was always somewhat flakey.

In 2010, the guys decided this project would be a bi-annual event. Since we would roll out this code every two years, it was worth an investment in its stability.  Enter BackPress,  stage right.  Because the code in BackPress makes up the core of projects such as BuddyPress and in some ways WordPress itself, I knew it would have undergone much more rigorous testing than my own home brewed code had.

Additionally the project was expanding in scope.  We wanted to have the project integrated the WordPress blog, which was in turn integrated with Ning and the same password could be used on each.  A requirement was that only active members of the community could have access to the tool and that this access was granted or denied via the WordPress user interface.

 

What BackPress IS and what it ISN’T.

When I started this project, I was thinking of BackPress as a very small white labeled WordPress.  I expected that I could put BackPress on a website and it would do SOMETHING. I wasn’t expecting anything pretty, but I was expecting to at least be able to visit the site and get a mostly blank “Hello World” screen created by BackPress.  I expected to be able to login and maybe see a profile page with my login info and nothing more.

I was expecting WAAAAAAAY too much.

Here is what the BackPress.org website says:

BackPress is a PHP library of core functionality for web applications. It grew out of the immensely popular WordPress project, and is also the core of the bbPress and GlotPress sister-projects.

It includes a variety of the foundations you need to build robust, scalable web applications including (amongst other things):

  • Logging,
  • User Roles and Capabilities (Permission systems),
  • Database connections (across multiple servers and multiple datacenters),
  • HTTP Transactions,
  • XML-RPC Server and Client,
  • Object caching,
  • Formatting,
  • XSS and SQL injection protection, including a variety of powerful escaping functions,
  • Taxonomies and
  • Options management

And the best part? It’s all licensed under a generous GPL2 license, so you can use it in your own free or commercial projects.

That is all completely true, but the word that should be underscored, bolded and blink tagged is “LIBRARY”. (Ok so I couldn’t bring myself to <blink> it…).

You need to think of BackPress like a C library.  You won’t have a working program of any kind until you write it.  In BackPress, there is no UI for any of the code.  If you want to have your users login, you need to write that.  If you want to actually display something to the screen via index.php, you have to build it and integrate the parts of BackPress you need.  BackPress does EVERYTHING listed on the .org site, but only provides the frameworks.

One of the things that surprised me the most was the basic nature of both the Option and Role management APIs.  For example, you can set options, retrieve options and all of that across a single WordPress “session” or visit.  It all works great – but if you want the values to be persistent and the options to be available the next time the site is visited, you will need to include a storage mechanism.  You could use a database or memcache or flat files, but without connecting your options and roles system to SOMETHING, everything you set goes away.

Plan Ahead

Because I’d written a large amount of my project in 2008, I was able to do a conversion to BackPress in one all night coding session and other than the surprise of no persistent storage of options, it all worked.  However, you should plan on doing a full UX design and data structure design for your project.

You should have answers to the following – not necessarily in any order:

  • Are you creating a single use project or do you need to consider other installations?  This will help you decide whether or not you need routines to create the database tables and default options, or if you want to hand create them and just have your BackPress site use them.
  • Do you need to store options? How will you store them?
  • Do you need to check user roles and if so what do you want them to be? Where/how will you store them?
  • If you are connecting with WordPress, will your tables, used to store options, users and etc, be located in the same WordPress database? If so, will your tables have a unique prefix or will you use pre-existing WordPress tables? Or if not will you access your tables from your own database?  Where will the database access information be stored?
  • Do you need to provide profile editing?
  • Do you need a link to login from the home page? Or do you want to only display a login prompt only when the user is not logged in?
  • Do you need a log out link?
  • Do you need a UI to allow administration of the site itself?  Do you have a role for that?
  • etc etc etc

How to get started

The best thing that you can do to get started is to look at an existing project.  There aren’t a lot of them out there.  Because so much of the code is custom purposed, not everyone thinks their code is ready to be shared.  Additionally, you want to protect the client’s IP and still release a working project. Since I’ve not yet shared the SuperNote code, I’m going to give you the link that I found that provided the most useful launch point.  

SupportPress: https://supportpress.svn.wordpress.org/trunk/

Browsing the code of SupportPress, I was finally able to realize what BackPress was and what it wasn’t.  SupportPress is a neat project that has only been released via SVN (and maybe even that much of a release is Google’s fault as I only found reference to it in search results.).  I don’t think it is yet code complete, but it has all the basic functionality you would need and can help guide you as to what you will need to create.  It also contains code from BuddyPress (which also can be examined for how BuddyPress used BackPress) and you may not want to use the same solutions that SupportPress chose.  So I would advise using it as a guide only.

When I release the source code for the SuperNote Review Tool (Hmmm “SNoRT” I like it!!!!) I will include a link here:

I hope this helps.  I’ll do my best to answer any questions raised in the comments.  Also if there is interest I’ll create some posts with code examples from my project.  It’s been a few months since it was written and it is all getting to be a bit fuzzy in my memory. So, if you want more articles on the subject, let me know soon!

Adding a line to /etc/hosts via Perl

Having appended a line to /etc/hosts through BASH, what I really wanted to do, was add it via PERL.

The basic operation would be the same: load a text file into a single variable, search a string for text, and if not found append text to a file in Perl.

So, here is the conversion of that earlier bash script:

#!/usr/bin/perl
# Configure environment to promote good programming practices
use strict;
use warnings;

# Define initial variables & constants
use constant NOT_FOUND => -1;        # All good programmers use Constants
my $domain="example.com";            # Change it to meet your needs
my $needle="www.$domain";            # I search with www. prefix
my $hostline="127.0.0.1 $domain www.$domain";
my $filename="/etc/hosts";           # Standard host file location

# Read the file into a scalar var in the most backwards compatible way
local( $/, *FILEHANDLE );
open(FILEHANDLE, $filename) or die("Cannot access $filename");
my $haystack = ;
close(FILEHANDLE);

# Look for the domain in the file already
my $result = index($haystack, $needle);

# Index returns 
if ( $result == NOT_FOUND )
{
  print "$needle NOT found in $filename\n";
  # If the line wasn't found, add it using an echo append >>
  open(FILEHANDLE, ">>" . $filename);
  print FILEHANDLE $hostline . "\n"; #write newline
  close(FILEHANDLE);
  print "$hostline added to $filename\n";
} 
else
{
  print "$needle already exists in $filename\n";
}

As a refresher on PERL, implied lessons in this are:
To concatenate strings in PERL you can either embed them within the double quoted string or concat them with the period like PHP does.
Lines end in semicolons of course
If causes change from brackets and fi in BASH to parents and braces in PERL
Constants are done with “use constant”
vars are usually defined with a my like basics var – there are of course variations on this.
Print vs echo
comparisons get the double = as in php

Installing SVN (Server & Client) on CentOS

The process is EXTREMELY simple. One line to install SVN, two to create the repository, and one to run the daemon:

yum install subversion.i386 mkdir /svn svnadmin create /svn/my-repo/ svnserve -d

Here’s what my results produced (The first line confirms I have subversion available in yum): [root@hosting ~]# yum list | grep ‘subversion’ subversion.i386 1.4.2-4.el5_3.1 base subversion-devel.i386 1.4.2-4.el5_3.1 base subversion-javahl.i386 1.4.2-4.el5_3.1 base subversion-perl.i386 1.4.2-4.el5_3.1 base subversion-ruby.i386 1.4.2-4.el5_3.1 base [root@hosting ~]# yum install subversion.i386 Loaded plugins: fastestmirror Determining fastest mirrors addons | 951 B 00:00 base | 2.1 kB 00:00 extras | 2.1 kB 00:00 updates | 1.9 kB 00:00 wiredtree | 951 B 00:00 Excluding Packages in global exclude list Finished Setting up Install Process Resolving Dependencies –> Running transaction check —> Package subversion.i386 0:1.4.2-4.el5_3.1 set to be updated –> Processing Dependency: perl(URI) >= 1.17 for package: subversion –> Processing Dependency: neon >= 0.25.5-6.el5 for package: subversion –> Processing Dependency: libneon.so.25 for package: subversion –> Processing Dependency: libapr-1.so.0 for package: subversion –> Processing Dependency: libaprutil-1.so.0 for package: subversion –> Running transaction check —> Package apr.i386 0:1.2.7-11.el5_3.1 set to be updated —> Package apr-util.i386 0:1.2.7-11.el5 set to be updated –> Processing Dependency: libpq.so.4 for package: apr-util —> Package neon.i386 0:0.25.5-10.el5_4.1 set to be updated —> Package wt-URI.noarch 0:1.35-1 set to be updated –> Processing Dependency: perl(Business::ISBN) for package: wt-URI –> Running transaction check —> Package postgresql-libs.i386 0:8.1.21-1.el5_5.1 set to be updated —> Package wt-Business-ISBN.noarch 0:2.00_01-1 set to be updated –> Processing Dependency: perl(Business::ISBN::Data) >= 1.09 for package: wt-Business-ISBN –> Running transaction check —> Package wt-Business-ISBN-Data.noarch 0:1.13-1 set to be updated –> Finished Dependency Resolution Dependencies Resolved ============================================================================================================================================ Package Arch Version Repository Size ============================================================================================================================================ Installing: subversion i386 1.4.2-4.el5_3.1 base 2.3 M Installing for dependencies: apr i386 1.2.7-11.el5_3.1 base 123 k apr-util i386 1.2.7-11.el5 base 80 k neon i386 0.25.5-10.el5_4.1 base 101 k postgresql-libs i386 8.1.21-1.el5_5.1 updates 196 k wt-Business-ISBN noarch 2.00_01-1 wiredtree 353 k wt-Business-ISBN-Data noarch 1.13-1 wiredtree 12 k wt-URI noarch 1.35-1 wiredtree 146 k Transaction Summary ============================================================================================================================================ Install 8 Package(s) Upgrade 0 Package(s) Total download size: 3.3 M Is this ok [y/N]: y Downloading Packages: (1/8): wt-Business-ISBN-Data-1.13-1.noarch.rpm | 12 kB 00:00 (2/8): apr-util-1.2.7-11.el5.i386.rpm | 80 kB 00:00 (3/8): neon-0.25.5-10.el5_4.1.i386.rpm | 101 kB 00:00 (4/8): apr-1.2.7-11.el5_3.1.i386.rpm | 123 kB 00:00 (5/8): wt-URI-1.35-1.noarch.rpm | 146 kB 00:00 (6/8): postgresql-libs-8.1.21-1.el5_5.1.i386.rpm | 196 kB 00:00 (7/8): wt-Business-ISBN-2.00_01-1.noarch.rpm | 353 kB 00:00 (8/8): subversion-1.4.2-4.el5_3.1.i386.rpm | 2.3 MB 00:00 ——————————————————————————————————————————————– Total 1.6 MB/s | 3.3 MB 00:02 Running rpm_check_debug Running Transaction Test Finished Transaction Test Transaction Test Succeeded Running Transaction Installing : apr 1/8 Installing : neon 2/8 Installing : postgresql-libs 3/8 Installing : wt-Business-ISBN-Data 4/8 Installing : apr-util 5/8 Installing : wt-URI 6/8 Installing : subversion 7/8 Installing : wt-Business-ISBN 8/8 Installed: subversion.i386 0:1.4.2-4.el5_3.1 Dependency Installed: apr.i386 0:1.2.7-11.el5_3.1 apr-util.i386 0:1.2.7-11.el5 neon.i386 0:0.25.5-10.el5_4.1 postgresql-libs.i386 0:8.1.21-1.el5_5.1 wt-Business-ISBN.noarch 0:2.00_01-1 wt-Business-ISBN-Data.noarch 0:1.13-1 wt-URI.noarch 0:1.35-1 Complete! [root@hosting ~]# mkdir /svn [root@hosting ~]# svnadmin create /svn/my-repo/ [root@hosting ~]# svnserve -d

The last thing to do is to configure the password if you want one.

vi /svn/my-repo/conf/svnserve.conf

vi /svn/my-repo/conf/passwd

The details of those two files are beyond the scope of this post. Besides I’m sure you’ll want to triple check, as every one else does, that the svn password file is in plain text. Yes, that’s correct. Plain text.  That makes you think about all the svn repositories that you used secure passwords to access now doesn’t it?

Building Up Reporting Tools

While working with Lee Newton over at b5media I was able to watch him build up some server tools over time that were invaluable to diagnosing exactly what was going on on the server.

Now I find I need to make some of my own. Here’s how I am doing it.

For now I am going to concentrate on access logs.

Where these logs are varies server by server, but if you are running a standard cPanel setup, chances are you can find a directory named /usr/local/apache/domlogs with files in it named after your domain name. In this case I picked one of the sites I host: nakedpastor.com

So if I do a:

cd /usr/local/apache/domlogs
tail -10 nakedpastor.com

I will get the last 10 lines of the access log file

Here’s one example:

24.555.555.27 – – [09/Aug/2010:20:13:45 -0400] “GET /wp-content/uploads/2010/08/IMG_0001.jpg HTTP/1.1” 304 – “http://www.nakedpastor.com/” “Mozilla/5.0 (iPhone; U; CPU iPhone OS 3_1_3 like Mac OS X; en-us) AppleWebKit/528.18 (KHTML, like Gecko) Version/4.0 Mobile/7E18 Safari/528.16”

Someone on an iPhone is looking at a picture on the site (which David would get more Google Juice from if he had named it better).

The trick is going to be to break down that line into important bits of information that will help me diagnose how my server is being used. For example, I might want to know if one IP address is flooding me. I might want to know if I am getting a HUGE number of requests for one particular file or if I am serving a large number of errors. If I ran this on a combined log file I, I would want to know if one domain was getting all of the traffic. Top referrers might be a fun thing to look at too. There are lots of little bits of info in there that could be helpful.

My tool chest includes:

tail – request a certain number of lines from the END of the file as shown above. During testing I will use -15 to get the last 15 lines but when I make this live, I’ll want to look at the last several thousand at least.
head – requests the number of lines at the top of a file. In this case it gives me the most pertinent results
sort – Will put the most important results at the top
cut – probably not helpful initially as the fields are not fixed width
awk – Used to parse the lines into chunks so I can see what is important. See also here
grep – Used to search for text
uniq – Used with -c uniq counts the number of occurrences of each variance
| – The piping symbol used to send the results of one command right into the next.

Let’s go after something simple first. The IP address. I want to take the last 100 lines of the error log, get the the ip address which will be the first word in the line, count how many times each ip address is used, sort it numerically by count and return the top 10. In bash, that is pronounced as:

tail -10000 /usr/local/apache/domlogs/nakedpastor.com | awk ‘{print $1}’ | sort | uniq -c | sort -nr | head -10

You can find examples of that line lots of places out there. In fact I copied and pasted that from another site. Their line used a “tail -n” instead of “head”, but it did the same thing You may want to note that you have to call sort before you call uniq in order for unique to work right..

For the rest of the examples I’m going to use the awk command to break down the line into separate fields separated by quotes or spaces. this line prints the number 7 because there are 7 fields separated by quotes:

tail -5 nakedpastor.com|awk -F ‘”‘ ‘{c=NF; print c}’

If I search for quotes and then search for space, I can get the result code

tail -15 nakedpastor.com|awk -F ‘”‘ ‘{print $3}’|awk ‘{print $1}’

Or the requested page:

tail -15 nakedpastor.com|awk -F ‘”‘ ‘{print $2}’|awk ‘{print $2}’

Or the referrer:

tail -15 nakedpastor.com|awk -F ‘”‘ ‘{print $4}’|awk ‘{print $1}’

Or the agent:

tail -15 nakedpastor.com|awk -F ‘”‘ ‘{print $6}’|awk ‘{print $1}’

So using these examples I can get the 20 most popular agent/OS combos:

tail -1000 /usr/local/apache/domlogs/nakedpastor.com | awk -F ‘”‘ ‘{print $6}’| sort | uniq -c | sort -nr | head -20

So, there you have some tools to use the next time you want to see what is going on on your server. Along with free -m and top you can get some neat info.

Just remember that you are getting snapshots. When I looked a little bit ago. Knowmore.com’s bot took up 1390 of the 10000 lines I was looking at. That’s over 10% of the traffic going to just that bot. HOWEVER if I looked at 5000 lines, they didn’t appear at all. Looking at 100000 lines they only appeared 1394 times. So, don’t just use one sample size. It can be misleading.

I will probably take these lines and combine them into some .bachrc functions along the line of what I’d discussed in my Three helpful additions to your .bashrc post.

Brain Storming on Blocking Bad-ads

I’m just jotting down some notes about using the Google Safe Browsing API to prevent a site from serving malicious/bad ads.

Problem Defined

  • Ads are put on a site via javascript by calls as simple as “getad(‘adposition1’)”. JavaScript is executed via the client’s browser after the page is served.
  • Those calls don’t touch any of our servers, they go from the client to the Google/Glam/Whatever Ad Server. So we don’t see the ads before they appear on the customer screens.
  • The ads being served may be malicious
    • Any ad that is served can link to a site that has been infected. We will want to block this.
    • Any ad that is served can “take over” the page and redirect the page to a site that may or may not have malware. We want to block ALL take over attempts.
    • There may be other types of ads that we wish to block.  Potentially we might wish to block specific ads on specific sites (i.e. a sexual connotations in ads on pre-teen audience sites). This may be beyond the initial scope and/or incur unwanted execution expenses.
  • Serving a malicious ad can get a site listed as “infected” even though your server has had nothing to do with ANY of the ad content.

Obstacles

  • Any extra calls WILL slow the page load process.
  • Each page load MUST call the ad serving script again
  • If an ad can be identified as bad, some other type of content must be served in that position to ensure page integrity.
  • The request for ad content HAS to come from the customer side because many ads are geo-specific and the customer’s IP determines what ad shows at what time.
  • You don’t want to set up a system where the site itself can submit a site as “bad” as anyone could sniff that info and seed our black list with bad data.
  • The results of the first getad() call could result in more javascript which must, in turn, be processed by the browser to produce the final ad. Potentially, several layers of JS could exist before the real ad is served. (e.g. 2 layers of indirection before ad: Google Ad Manager JS —serves—> Glam Ad embeded JS call —serves—> JS call to 3rd Party Ad Server —serves—> Ad). This pattern is real and happens often.

Possible solutions

  • Status Quo: As problem sites are reported to us, determine which ad is bad, report it to the ad server & hope they fix it before google sees it and lists the site as a dangerous site in it’s tool bar and in chrome.
    • Unless you are “lucky” you don’t get the badad.
    • Once you get the badad, it is hard to determine the initial JS that caused the problem
  • Embed everything JS with its own iframe
    • Will block take overs
    • May or may not prevent Google from listing the site, probably not.
    • Will break ads that are contextual based
  • Check the ad entirely on the client side via a black list: GSB API (http://code.google.com/apis/safebrowsing/) or PhishTank (http://data.phishtank.com/data/online-valid.xml)
    • This Good/Bad check could be done with a single call with the API call
    • Calls to external servers are dependent upon the health/bandwidth of that server
    • This could also be done via downloading the black list and checking off of that: http://code.google.com/p/jgooglesafebrowsing/wiki/Quick_Start_Guide
    • Blacklist downloading would cost time and would have to be updated periodically.
  • Implement a hybrid solution where a call is done to our servers to see if the an ad is good or bad.  (Server side base code: http://lampsecurity.org/php-google-safe-browsing-api )
    • Ad call is processed in JS eval (Will have to be checked for nested JS calls)
    • MD5 of ad is sent to the server. The results are Good/Bad/Unknown.  (Pass the url?)
    • If the result is Good, ad is served and process exits
    • If the result is Bad, either go to step 1, or serve place holder/known good ad & exit.
    • If the result is Unknown, send the JS to the server for verification. The server processes the code and returns a Good/Bad result.
    • If the result is Good, ad is served and process exits
    • If the result is Bad, either go to step 1, or serve place holder/known good ad & exit
  • Other solutions?

Reading

Anyway, I had this going through my head and wanted to get this all written out. So I can have a place to check back on this tomorrow…

WordPress/WordPress mu Merge Definitively Confirmed

There’s been rumor and confusion over the last week about whether WordPress and WordPress mu were merging as Matt seemed to imply at WordCamp SF. The announcement was so shocking that the true meaning was uncertain. For example, the avid WordPress evangelist Lorelle was left with the impression that WordPress.org would become a community site. Thankfully, Donncha, WordPress mu’s lead, gave the conclusive word on the subject this morning:

Basically, the thin layer of code that allows WordPress MU to host multiple WordPress blogs will be merged into WordPress. I expect the WordPress MU project itself will come to an end because it won’t be needed any more (which saddens me), but on the other hand many more people will be working on that very same MU code which means more features and more bugfixes and faster too.

Donncha, I would view this with the honor it does you. It is not much of a stretch to say that with your work on mu, you’ve made a lasting contribution to the shape of world and how people get information and will relate to each other over the upcoming years. More and more and more sites are run on mu, while the whole buddy press/bbpress/mu paradigm is taking off and will change the shape of the web. The adoption of the mu’s features into the WP core is a signal of what is to come and it will be an exciting ride!

Congrats guy!