Tuesday, December 6, 2011

Running a first Hadoop job with Cloudera's Tutorial VM

After watching the O'Reilly video on Map Reduce I decided that I'd like to know more about Hadoop. After doing some Googling I found that a firm by the name of Cloudera has pre-populated VMs available for playing around with located here.

The tutorial located here looks very much like the Apache Hadoop tutorial.

I was running through the tutorial and ran into a number of road blocks. I eventually got around the road blocks but thought it would be handy to document the issues that I had in case I ever revisit this tutorial in the future.

First issue that I ran into was that I couldn't get the Cloudera VM up and running with VirtualBox. After some trial and error I found that I could get the CentOS based VM up and running after selecting the IO APIC checkbox in the configuration for the VM:



The second issue was that I couldn't install the VirtualBox Guest Additions since CentOS didn't have all the kernel source installed. What a bummer.

I was able to install the necessary kernel files by issuing the following command:

yum install kernel-devel-2.6.18-274.7.1.el5

Then I found that the VirtualBox Guest Additions still couldn't be installed as there was no GCC compiler. That was easily fixed with:

yum install gcc

After that, the VirtualBox Guest Additions installed just fine and with a reboot I know have much better control over the VM. The default 640x480 is a bit stiffling. Now I have it running full screen at 1920x1600 and it is much, much nicer.

Next step was working on the WordCount example.

As a side note: In the O'Reilly video we did a very similar word counting map reduce job and I have to say that I much prefer the terse Python code over the Java solution. But that is just personal preference.

The tutorial provides the source code for WordCount.java and I ran into some issues where with some deviation from the tutorial. The tutorial gives a tip on the proper environment variables for HADOOP_HOME and HADOOP_VERSION but the tutorial is out of sync with the Cloudera VM.

The tutorial states that the proper version information is "0.20.2-cdh3u1" when it is actually "0.20.2-cdh3u2". Not really a big deal but when following a tutorial on a subject that is brand new, this can be frustrating.

The next issue that I ran into was due to my forgetting most of my Java development skills. Java development is not something that I do day in and out so some of that information had been garbage collected off of my mental heap to make way for other information (probably due to memorizing useless movie quotes).

I created a sub-dir for the WordCount.java code and compiled and created a .jar as provided by the instructions but my first attempts at executing a Hadoop job failed with Hadoop complaining that it couldn't find "org.myorg.WordClass" as seen below.

Compiling the WordClass.java code:

[root@localhost wordcount_classes]# javac -classpath ${HADOOP_HOME}/hadoop-${HADOOP_VERSION}-core.jar WordCount.java
[root@localhost wordcount_classes]# ls -al
total 24
drwxr-xr-x 2 root root 4096 Dec 6 18:34 .
drwxr-xr-x 3 root root 4096 Dec 6 18:32 ..
-rw-r--r-- 1 root root 1546 Dec 6 18:34 WordCount.class
-rw-r--r-- 1 root root 1869 Dec 6 18:33 WordCount.java
-rw-r--r-- 1 root root 1938 Dec 6 18:34 WordCount$Map.class
-rw-r--r-- 1 root root 1611 Dec 6 18:34 WordCount$Reduce.class
[root@localhost wordcount_classes]# cd ..
[root@localhost wordcount]# jar -cvf wordcount.jar -C wordcount_classes/ .
added manifest
adding: WordCount$Map.class(in = 1938) (out= 798)(deflated 58%)
adding: WordCount.java(in = 1869) (out= 644)(deflated 65%)
adding: WordCount.class(in = 1546) (out= 749)(deflated 51%)
adding: WordCount$Reduce.class(in = 1611) (out= 649)(deflated 59%)
[root@localhost wordcount]# jar tf wordcount.jar
META-INF/
META-INF/MANIFEST.MF
WordCount$Map.class
WordCount.java
WordCount.class
WordCount$Reduce.class
[root@localhost wordcount]#


When I attempted to run my first Hadoop job I got this output:


[root@localhost bad.wordcount]# /usr/bin/hadoop jar wordcount.jar org.myorg.WordCount /usr/joe/wordcount/input /usr/joe/wordcount/output_1
Exception in thread "main" java.lang.ClassNotFoundException: org.myorg.WordCount
at java.net.URLClassLoader$1.run(URLClassLoader.java:202)
at java.security.AccessController.doPrivileged(Native Method)
at java.net.URLClassLoader.findClass(URLClassLoader.java:190)
at java.lang.ClassLoader.loadClass(ClassLoader.java:307)
at java.lang.ClassLoader.loadClass(ClassLoader.java:248)
at java.lang.Class.forName0(Native Method)
at java.lang.Class.forName(Class.java:247)
at org.apache.hadoop.util.RunJar.main(RunJar.java:179)


What got me was the "Exception in thread "main" java.lang.ClassNotFoundException: org.myorg.WordCount."

I look over the code and as far as I can tell, everything is just fine and there is not reason why the WordClass shouldn't be found. After mulling the problem for a while my brain went into brain persistence layer and pulled out the proper way of taking care of the issue.

The source code defines the package name as org.myorg but I hadn't created the proper sub-dirs to reflect the org.myorg package.

I created the org subdir, then myorg in the org subdir and then compiled the code again:



I re-created the .jar pulling in all the files under the org subdir the jar file I found that Hadoop would be much happier and would find my org.myorg.WordCount class finally.

The next issue that I ran into was due not understanding where Hadoop would full the input files for the word count example. In the O'Reilly Map Reduce video STDIN and STDOUT were used and I just figured that I'd be able to specifiy the input and output subdirs from the local file system. I was incorrect.

I attempted to execute the Hadoop job with the following parameters referencing the input and output subdirs:


[root@localhost wordcount]# /usr/bin/hadoop jar wordcount.jar org.myorg.WordCount /home/cloudera/Desktop/wordcount/input /home/cloudera/Desktop/wordcount/output/
11/12/06 18:57:50 WARN mapred.JobClient: Use GenericOptionsParser for parsing the arguments. Applications should implement Tool for the same.
11/12/06 18:57:51 WARN snappy.LoadSnappy: Snappy native library is available
11/12/06 18:57:51 INFO util.NativeCodeLoader: Loaded the native-hadoop library
11/12/06 18:57:51 INFO snappy.LoadSnappy: Snappy native library loaded
11/12/06 18:57:51 INFO mapred.JobClient: Cleaning up the staging area hdfs://0.0.0.0/var/lib/hadoop-0.20/cache/mapred/mapred/staging/root/.staging/job_201112061431_0006
Exception in thread "main" org.apache.hadoop.mapred.InvalidInputException: Input path does not exist: hdfs://0.0.0.0/home/cloudera/Desktop/wordcount/input
at org.apache.hadoop.mapred.FileInputFormat.listStatus(FileInputFormat.java:194)
at org.apache.hadoop.mapred.FileInputFormat.getSplits(FileInputFormat.java:205)
at org.apache.hadoop.mapred.JobClient.writeOldSplits(JobClient.java:971)
at org.apache.hadoop.mapred.JobClient.writeSplits(JobClient.java:963)
at org.apache.hadoop.mapred.JobClient.access$500(JobClient.java:170)
at org.apache.hadoop.mapred.JobClient$2.run(JobClient.java:880)
at org.apache.hadoop.mapred.JobClient$2.run(JobClient.java:833)
at java.security.AccessController.doPrivileged(Native Method)
at javax.security.auth.Subject.doAs(Subject.java:396)
at org.apache.hadoop.security.UserGroupInformation.doAs(UserGroupInformation.java:1127)
at org.apache.hadoop.mapred.JobClient.submitJobInternal(JobClient.java:833)
at org.apache.hadoop.mapred.JobClient.submitJob(JobClient.java:807)
at org.apache.hadoop.mapred.JobClient.runJob(JobClient.java:1242)
at org.myorg.WordCount.main(WordCount.java:55)
at sun.reflect.NativeMethodAccessorImpl.invoke0(Native Method)
at sun.reflect.NativeMethodAccessorImpl.invoke(NativeMethodAccessorImpl.java:39)
at sun.reflect.DelegatingMethodAccessorImpl.invoke(DelegatingMethodAccessorImpl.java:25)
at java.lang.reflect.Method.invoke(Method.java:597)
at org.apache.hadoop.util.RunJar.main(RunJar.java:186)


Notice that I used the path to my Desktop (yeah, I know, I shouldn't be putting the files under the Desktop subdir but it makes it easy to acces the files via the GUI. Obviously I would never do this on a real development system) subdirs referencing /input I created.

What I didn't realize at the time was that he paths given to Hadoop are relative to the HDFS file system and not the local ext3 file system. After reading the Cloudera Quick Start Guide PDF it all started to make sense.

I needed to populate HDFS with the input and output subdirs along with the input files noted in the tutorial. The paths from the tutorial reference the path of /usr/joe/wordcount/input and /usr/joe/wordcount/output.

I created the input subdir using the proper DFS command:


/usr/bin/hadoop dfs -mkdir /usr/joe/wordcount/input


And I copied the previously created input files, file01 and file02:


/usr/bin/hadoop dfs -put file01 /usr/joe/wordcount/input
/usr/bin/hadoop dfs -put file02 /usr/joe/wordcount/input


Now it was time for the big event, now that I put everything into place I tried the tutorial command over again:


[root@localhost wordcount]# /usr/bin/hadoop jar wordcount.jar org.myorg.WordCount /usr/joe/wordcount/input /usr/joe/wordcount/output
11/12/06 19:13:32 WARN mapred.JobClient: Use GenericOptionsParser for parsing the arguments. Applications should implement Tool for the same.
11/12/06 19:13:33 WARN snappy.LoadSnappy: Snappy native library is available
11/12/06 19:13:33 INFO util.NativeCodeLoader: Loaded the native-hadoop library
11/12/06 19:13:33 INFO snappy.LoadSnappy: Snappy native library loaded
11/12/06 19:13:33 INFO mapred.FileInputFormat: Total input paths to process : 2
11/12/06 19:13:34 INFO mapred.JobClient: Running job: job_201112061431_0007
11/12/06 19:13:35 INFO mapred.JobClient: map 0% reduce 0%
11/12/06 19:13:46 INFO mapred.JobClient: map 33% reduce 0%
11/12/06 19:13:47 INFO mapred.JobClient: map 66% reduce 0%
11/12/06 19:13:52 INFO mapred.JobClient: map 100% reduce 0%
11/12/06 19:14:06 INFO mapred.JobClient: map 100% reduce 100%
11/12/06 19:14:09 INFO mapred.JobClient: Job complete: job_201112061431_0007
11/12/06 19:14:09 INFO mapred.JobClient: Counters: 23
11/12/06 19:14:09 INFO mapred.JobClient: Job Counters
11/12/06 19:14:09 INFO mapred.JobClient: Launched reduce tasks=1
11/12/06 19:14:09 INFO mapred.JobClient: SLOTS_MILLIS_MAPS=25035
11/12/06 19:14:09 INFO mapred.JobClient: Total time spent by all reduces waiting after reserving slots (ms)=0
11/12/06 19:14:09 INFO mapred.JobClient: Total time spent by all maps waiting after reserving slots (ms)=0
11/12/06 19:14:09 INFO mapred.JobClient: Launched map tasks=3
11/12/06 19:14:09 INFO mapred.JobClient: Data-local map tasks=3
11/12/06 19:14:09 INFO mapred.JobClient: SLOTS_MILLIS_REDUCES=19860
11/12/06 19:14:09 INFO mapred.JobClient: FileSystemCounters
11/12/06 19:14:09 INFO mapred.JobClient: FILE_BYTES_READ=79
11/12/06 19:14:09 INFO mapred.JobClient: HDFS_BYTES_READ=348
11/12/06 19:14:09 INFO mapred.JobClient: FILE_BYTES_WRITTEN=215844
11/12/06 19:14:09 INFO mapred.JobClient: HDFS_BYTES_WRITTEN=41
11/12/06 19:14:09 INFO mapred.JobClient: Map-Reduce Framework
11/12/06 19:14:09 INFO mapred.JobClient: Reduce input groups=5
11/12/06 19:14:09 INFO mapred.JobClient: Combine output records=6
11/12/06 19:14:09 INFO mapred.JobClient: Map input records=2
11/12/06 19:14:09 INFO mapred.JobClient: Reduce shuffle bytes=91
11/12/06 19:14:09 INFO mapred.JobClient: Reduce output records=5
11/12/06 19:14:09 INFO mapred.JobClient: Spilled Records=12
11/12/06 19:14:09 INFO mapred.JobClient: Map output bytes=82
11/12/06 19:14:09 INFO mapred.JobClient: Map input bytes=50
11/12/06 19:14:09 INFO mapred.JobClient: Combine input records=8
11/12/06 19:14:09 INFO mapred.JobClient: Map output records=8
11/12/06 19:14:09 INFO mapred.JobClient: SPLIT_RAW_BYTES=294
11/12/06 19:14:09 INFO mapred.JobClient: Reduce input records=6


Wait? What is this? Could it be? Yes! It worked! YES!YES!YES!

My dancing around the room was enough to wake up my Basset Hound. He looked up at me, cocked his head as to say, "Hey, why are you being so goofy? Make yourself useful and get me another doggie snack."

Checking the results of the job I see the following:


[root@localhost wordcount]# /usr/bin/hadoop dfs -ls /usr/joe/wordcount/output
Found 3 items
-rw-r--r-- 1 root supergroup 0 2011-12-06 19:14 /usr/joe/wordcount/output/_SUCCESS
drwxr-xr-x - root supergroup 0 2011-12-06 19:13 /usr/joe/wordcount/output/_logs
-rw-r--r-- 1 root supergroup 41 2011-12-06 19:14 /usr/joe/wordcount/output/part-00000


Look at that! _SUCCESS!

Checking the part-00000 file:


[root@localhost wordcount]# /usr/bin/hadoop dfs -cat /usr/joe/wordcount/output/part-00000
Bye 1
Goodbye 1
Hadoop 2
Hello 2
World 2


Sure. It was a lot of work to just count a few words in a text file, but it was a real good learning experience this afternoon. After banging my head against the brick wall for a long enough period I got around the potholes that I ran into and feel that I'll be able to continue the tutorial and learn the basics of Hadoop.

Not a bad bit of afternoon vacation learning.

Saturday, December 3, 2011

I'm dreaming of a geeky Xmas

Black Friday and Cyber Monday has come and gone and the team and I kicked ass... We kicked major ass. All that work over the past 11 months paid off over that single week of retail Hell and our systems just simply kicked ass...

But now, it is time for all the vacation that I didn't take during the year while prepping for the few busiest days of the year. I am taking off the entire month of December (and still rolling 40 hours of vacation into next year!)

What to do with all this time?

On CyberMonday O'Reilly had a 60% off sale on books and videos so I decided to purchase a few videos to get up to speed in some technologies that I haven't fiddled around with.

Over December I hope to go over the following videos:

An Introduction to MapReduce with Pete Warden
Erlang by Example with Cesarini and Thompson
Hilary Mason: An Introduction to Machine Learning with Web Data

I've had a very brief introduction to Machine Learning during a GDAT class with Dr. Gunther and found it very interesting and something that could be cool for automated monitoring of systems to determine when something is going astray.

Also I plan on covering some Haskell learnings with the Channel9 lecture series on the subject of Haskell and functional programming. I've done a little bit of FP with R and it is time to wrap my head around FP more with Haskell and Erlang.

The current plan on the Haskell front is to gather with fellow Geeks at the Addison, TX location of Thoughtworks for Geeknight to watch the Channel9 videos and learn the ways of FP first hand.

I'm sure that good times will ensue!

Wednesday, November 23, 2011

Wrapping perl around tqharvc for generating scatter plots

TeamQuest View is a handy utility for looking at metrics collected on servers by TeamQuest. But one thing that I haven't liked about TeamQuest View is creating graphs from collected metrics. Fortunately, with the installation of TeamQuest View a command line utility by the name of tqharvc is also installed. It is possible to execute tqharvc.exe from the command line to collect metrics and create plots from the output.

I got the idea of writing some perl code that slices and dices the TeamQuest .RPT files that can be generated by TeamQuest View and then passing on the queries to tqharvc to a graphing routine to generate both plots and .CSV files for later processing if need be.

For the graphing of data I utilize GD::Graph for generating a scatter plot.

For the example in this blog entry I have a very simple .RPT file by the name of physc.RPT and as the name implies, it simply reports physc usage from a LPAR:


[General]
Report Name = "physc"
View = Line
Format = Date+Time,
Parameter

[dParm1]
System = "[?]"
Category Group = "CPU"
Category = "by LPAR"
Subcategory = ""
Statistic = "physc"
Value Types = "Average"


Simple enough, isn't it?

My perl routine builds up a query for tqharvc by slicing and dicing the .RPT file and iterates through all the sliced and diced metrics, collects the output data and generates the .CSV output and plot image.

The tqharvc command line query is formed as:

tqharvc.exe -m [hostName] -s [startDate]:[startTime] -e [endDate]:[endTime] -a 10-minute -k [metric]

Say that I want to query "serverA" for the metric "CPU:by LPAR:physc" between 11/1/2011 midnight to 11/07/2011 11:59:59 PM with metrics from every 10 minutes?

I call tqharvc with the following command line:

tqharvc.exe -m serverA -s 11012011:000000 -e 11072011:235959 -a 10-minute -k "CPU:by LPAR:physc"

My routine takes several command line parameters:

tqGatherMetrics.pl server.list physc.rpt 11/01/2011 00:00:00 11/06/2011 23:59:59

server.list is simply a text file with a list of servers to query:


serverA
serverB
serverC
serverD
serverE
serverF


phsyc.rpt is the aforementioned .RPT file. 11/01/2011 is the start date and 00:0:000 is the start time. 11/07/2011 is the end date and 23:59:59 is the end time of the queries.

Below is the output to STDERR from the perl routine as it crunches through the data from the servers:


Running query against serverA.
Executing TQ Query for serverA:CPU:by LPAR.
Processing query results.
Extracted 864 lines of data from query.
Processing results for serverA : CPU:by LPAR : 7
Generating graph for "serverA_CPU_by_LPAR_physc.gif"
Running query against serverB.
Executing TQ Query for serverB:CPU:by LPAR.
Processing query results.
Extracted 864 lines of data from query.
Processing results for serverB : CPU:by LPAR : 7
Generating graph for "serverB_CPU_by_LPAR_physc.gif"
Running query against serverC.
Executing TQ Query for serverC:CPU:by LPAR.
Processing query results.
Extracted 864 lines of data from query.
Processing results for serverC : CPU:by LPAR : 7
Generating graph for "serverC_CPU_by_LPAR_physc.gif"
Running query against serverD.
Executing TQ Query for serverD:CPU:by LPAR.
Processing query results.
Extracted 864 lines of data from query.
Processing results for serverD : CPU:by LPAR : 7
Generating graph for "serverD_CPU_by_LPAR_physc.gif"
Running query against serverE.
Executing TQ Query for serverE:CPU:by LPAR.
Processing query results.
Extracted 864 lines of data from query.
Processing results for serverE : CPU:by LPAR : 7
Generating graph for "serverE_CPU_by_LPAR_physc.gif"
Running query against serverF.
Executing TQ Query for serverF:CPU:by LPAR.
Processing query results.
Extracted 864 lines of data from query.
Processing results for serverF : CPU:by LPAR : 7
Generating graph for "serverF_CPU_by_LPAR_physc.gif"


The output of the plot appears as:



With the above command line I was able to generate plots for six servers in the server.list file.

If the .RPT file had listed multiple metrics, there would have been one plot generated per server for each metric. I chose the single metric physc as a simple example.

The perl code to generate the plot follows:

1:  use GD::Graph::points;  
2: use Statistics::Descriptive;
3: use Time::Local;
4: use POSIX qw/ceil/;
5:
6: use strict;
7:
8: my $sourceList = shift;
9: my $tqReport = shift;
10: my $startDate = shift;
11: my $startTime = shift;
12: my $endDate = shift;
13: my $endTime = shift;
14:
15: $startDate =~ s/\///g;
16: $endDate =~ s/\///g;
17:
18: $startTime =~ s/\://g;
19: $endTime =~ s/\://g;
20:
21: if (length($sourceList) == 0) {
22: die "Usage:tqGatherMerics.pl <sourceList> <tqReport> <startDate> <stareTime> <endDate> <endTime>\n";
23: };
24:
25: my $tqHarvestBinary = "C:\\Program Files\\TeamQuest\\manager\\bin\\tqharvc.exe";
26:
27: if ( -f "$tqHarvestBinary" ) {
28: if ( -f "$sourceList" ) {
29: if ( -f "$tqReport" ) {
30:
31: my @hostNames = ();
32: my %metricHash = ();
33:
34: open(HOSTNAMES, "$sourceList") || die "$!";
35: while (<HOSTNAMES>) {
36: chomp($_);
37: if ($_ !~ /^#/) {
38: push(@hostNames, $_);
39: };
40: };
41: close(HOSTNAMES);
42:
43: open(REPORT, "$tqReport") || die "$!";
44:
45: my $catGroup = "::";
46: my $catName = "::";
47: my $subCat = "::";
48:
49: while (<REPORT>) {
50:
51: my @statArray = ();
52:
53: chomp($_);
54:
55: if ($_ =~ /Category Group = \"(.+?)\"/) {
56: $catGroup = $1;
57: };
58:
59: if ($_ =~ /Category = \"(.+?)\"/) {
60: $catName = $1;
61: };
62:
63: if ($_ =~ /Subcategory = \"(.+?)\"/) {
64: $subCat = $1;
65: };
66:
67: if ($_ =~ /Statistic =/) {
68: my $tmpString = "";
69: $_ =~ s/Statistic =//g;
70: $tmpString = $_;
71: do {
72: $_ = <REPORT>;
73: chomp($_);
74: if ($_ !~ /^Resource|^Value Types/) {
75: $tmpString .= $_;
76: };
77: } until ($_ =~ /^Resource|^Value Types/);
78: my @statArray = split(/\,/, $tmpString);
79: $metricHash{"${catGroup}:${catName}"} = \@statArray;
80: };
81:
82: };
83:
84: close(REPORT);
85:
86: foreach my $hostName (@hostNames) {
87: print STDERR "Running query against $hostName.\n";
88: my %metricData = ();
89: foreach my $paramName (sort(keys(%metricHash))) {
90: my %columnHash = ();
91: my $linesExtracted = 0;
92: my $shellCmd = "\"$tqHarvestBinary\" -m $hostName -s $startDate:$startTime -e $endDate:$endTime -a 10-minute -k \"$paramName\"";
93: # my $shellCmd = "\"$tqHarvestBinary\" -m $hostName -s $startDate:$startTime -e $endDate:$endTime -a 1-minute -k \"$paramName\"";
94: print STDERR "\t\tExecuting TQ Query for $hostName:$paramName.\n";
95:
96: open(OUTPUT, "$shellCmd |") || die "$!";
97:
98: print STDERR "\t\tProcessing query results.\n";
99:
100: my $totalColumns = 0;
101:
102: while (<OUTPUT>) {
103: chomp($_);
104: if ($_ =~ /^Time:/) {
105: my @columns = split(/\,/, $_);
106: my $statName = "";
107: for (my $index = 0; $index < $#columns; $index++) {
108: foreach $statName (@{$metricHash{$paramName}}) {
109: $statName =~ s/^\s+//g; # ltrim
110: $statName =~ s/\s+$//g; # rtrim
111: $statName =~ s/\"//g;
112:
113: my $columnName = $columns[$index];
114: if (index($columnName, $statName, 0) >= 0) {
115: $columnHash{$index} = $columns[$index];
116: $totalColumns++;
117: };
118: };
119: };
120: } else {
121: if ($_ =~ /^[0-9]/) {
122: chomp($_);
123: my @columns = split(/\,/, $_);
124: foreach my $index (sort(keys(%columnHash))) {
125: $metricData{"$columns[0] $columns[1]"}{$columnHash{$index}} = $columns[$index];
126: };
127: $linesExtracted++;
128: };
129: };
130: };
131:
132: close(OUTPUT);
133:
134: if (($linesExtracted > 0) && ($totalColumns > 0)) {
135: print STDERR "\tExtracted $linesExtracted lines of data from query.\n";
136: my @domainData = ();
137:
138: foreach my $timeStamp (sort dateSort keys(%metricData)) {
139: push(@domainData, $timeStamp);
140: };
141:
142: foreach my $metricIndex (sort(keys(%columnHash))) {
143: print STDERR "\t\tProcessing results for $hostName : $paramName : $metricIndex\n";
144:
145: my $metricName = $columnHash{$metricIndex};
146: my @rangeData = ();
147: my $stat = Statistics::Descriptive::Full->new();
148:
149: foreach my $timeStamp (@domainData) {
150: push(@rangeData, $metricData{$timeStamp}{$columnHash{$metricIndex}});
151: $stat->add_data($metricData{$timeStamp}{$columnHash{$metricIndex}});
152: };
153:
154: my $graphName = "${hostName}_${paramName}_${metricName}";
155: my $csvName = "";
156:
157: $graphName =~ s/\\/_/g;
158: $graphName =~ s/\//_/g;
159: $graphName =~ s/\%/_/g;
160: $graphName =~ s/\:/_/g;
161: $graphName =~ s/\s/_/g;
162:
163: $csvName = $graphName;
164: $graphName .= ".gif";
165: $csvName .= ".csv";
166:
167: print STDERR "\t\tGenerating graph for \"$graphName\"\n";
168:
169: open(CSVOUTPUT, ">$csvName");
170: print CSVOUTPUT "Timestamp,$paramName:$metricName\n";
171:
172: my $i = 0;
173: foreach my $timeStamp (@domainData) {
174: print CSVOUTPUT "$domainData[$i],$rangeData[$i]\n";
175: $i++;
176: };
177:
178: close(CSVOUTPUT);
179:
180: my $dataMax = $stat->max();
181: my $dataMin = $stat->min();
182:
183: if ($dataMax < 1) {
184: $dataMax = 0.5;
185: };
186:
187: if ($dataMin > 0) {
188: $dataMin = 0;
189: };
190:
191: for (my $rangeIndex = 0; $rangeIndex < $#rangeData; $rangeIndex++) {
192: if ($rangeData[$rangeIndex] == 0) {
193: $rangeData[$rangeIndex] = $dataMax * 5;
194: };
195: };
196:
197: my @data = (\@domainData, \@rangeData);
198: my $tqGraph = GD::Graph::points->new(1024, int(768/2));
199: my $totalMeasurements = $#{$data[0]} + 1;
200:
201: $tqGraph->set(x_label_skip => int($#domainData/40),
202: x_labels_vertical => 1,
203: markers => [6],
204: marker_size => 2,
205: y_label => "$metricName",
206: y_min_value => $dataMin,
207: y_max_value => ceil($dataMax * 1.1),
208: title => "${hostName}:${paramName}:${metricName}, N = " . $stat->count(). "; avg = " . $stat->mean() . "; SD = " . $stat->standard_deviation() . "; 90th = " . $stat->percentile(90) . ".",
209: line_types => [1],
210: line_width => 1,
211: dclrs => ['red'],
212: ) or warn $tqGraph->error;
213:
214: $tqGraph->set_legend("TQ Measurement");
215: $tqGraph->set_legend_font(GD::gdMediumBoldFont);
216: $tqGraph->set_x_axis_font(GD::gdMediumBoldFont);
217: $tqGraph->set_y_axis_font(GD::gdMediumBoldFont);
218:
219: my $tqImage = $tqGraph->plot(\@data) or die $tqGraph->error;
220:
221: open(PICTURE, ">$graphName");
222: binmode PICTURE;
223: print PICTURE $tqImage->gif;
224: close(PICTURE);
225:
226: };
227:
228: } else {
229: print STDERR "#################### Nothing extracted for Hostname \"$hostName\" and metric \"$paramName\" ####################\n";
230: };
231:
232: };
233: };
234:
235: } else {
236: print STDERR "Could not find the TeamQuest Report file.\n";
237: };
238: } else {
239: print STDERR "Could not find the list of hostnames to run against.\n";
240: };
241: } else {
242: print STDERR "Could not find the TeamQuest Manager TQHarvest binary at \"$tqHarvestBinary\". Cannot continue.\n";
243: };
244:
245: sub dateSort {
246: my $a_value = dateToEpoch($a);
247: my $b_value = dateToEpoch($b);
248:
249: if ($a_value > $b_value) {
250: return 1;
251: } else {
252: if ($b_value > $a_value) {
253: return -1;
254: } else {
255: return 0;
256: };
257: };
258:
259: };
260:
261: sub dateToEpoch {
262: my ($timeStamp) = @_;
263: my ($dateString, $timeString) = split(/ /, $timeStamp);
264: my ($month, $day, $year) = split(/\//, $dateString);
265: my ($hour, $min, $sec) = split(/:/, $timeString);
266:
267: $year += 2000;
268: $month -= 1;
269:
270: return timegm($sec,$min,$hour,$day,$month,$year);
271:
272: };
273: