Hadoop MapReduce low latency matters! BGI Shenzen wins IDC Innovation Award


Last week IDC released the winners for the HPC Innovation Excellence Awards.

As per IDC’s announcement, BGI Shenzen saved millions of dollars while enabling faster processing of their large genome data sets. The IBM Platform Symphony product was used for the Hadoop MapReduce applications. Platform Symphony provides a low latency Hadoop MapReduce implementation. In conjunction with its unique shared memory management and the data aware low latency scheduler, it accelerates many life sciences applications by as much as 10 times over open source Apache Hadoop

Below is an excerpt from the release:

BGI Shenzhen (China). BGI has developed a set of distributed computing applications using a MapReduce framework to process large genome data sets on clusters. By applying advanced software technologies including HDFS, Lustre, GlusterFS, and the Platform Symphony MapReduce framework, the institute has saved more than $20 million to date. For some application workloads, BGI achieved a significant improvement in processing capabilities while enabling the reuse of storage, resulting in reduced infrastructure costs while delivering results in less than 10 minutes, versus the prior time of 2.5 hours. Some of the applications enabled through the MapReduce framework included: sequencing of 1% of the Human Genome for the International Human Genome Project; contributing 10% to the International Human HapMap Project; conducting research in combating SARS, and a German variant of the E. coli virus; and completely sequencing the rice genome, the silkworm potato genome, and the human gut metagenome. Project leader: Lin Fang

Congratulations to the BGI team for their well deserved recognition.

Rohit Valia
Director, Marketing
HPC Cloud and Analytics

The win in Monaco and a hot season continues



Mark Webber's winning dive into the pool, he's mid air!

As Red Bull Racing prepares for their upcoming race events, a recent highlight was at the Monaco Grand Prix where the team won the race for the third consecutive time! It was a nail biting race and that also resulted in a second consecutive win for Mark Webber in Monaco.

Platform Compuitng  logo on the car.  Very cool.
The partnership between Red Bull Racing and Platform Computing, an IBM company is an exciting one, seeing the team rise to the top and win the most prestigious race on the calendar… again was truly amazing! It was nothing but smiles on the Energy Station on race day. A big CONGRATULATIONS to the team and the best of luck as the season pushes ahead!

For more information on the partnership between Platform Computing and Red Bull Racing see
www.ibm.com/platformcomputing or check out the video or read the case study.

Next up the Grand Prix of Europe.

Infrastructure Cloud Strategies

San Francisco, Monday Feb 20th, evening. I'm wondering what path most Platform customers will take to the cloud over the next decade, and perhaps more interesting, what will the intermediate steps be as they prepare for that journey. I'm here for the molecular tri-con, and had the pleasure, or more properly, the learning experience of listening to people like Deepak Singh of Amazon, Chris Dagdigian of Bioteam, and Jason Stowe of Cycle Computing.

For a while, global organizations have recognized that its much easier to move computing resources and demands to the location where data is housed rather than the reverse. As such, its now quite common to hear that regions within globally distributed companies have specialties, and take on projects of their own. Naturally, this isn’t as efficient as it could be, and perhaps the promise of 3D compression services for remote visualization will eventually unlock the user from only working on data close to him. But that’s another story and another blog post.

More and more, in these cloud computing conferences, the need for data is emphasized, or seen as a barrier for widespread use of cloud computing. Simply put, the time to transfer data up to the cloud process it, and bring it back often negates the value of using a metered pay per use infinite compute service. For me this is like not being able to drive a new Nissan GTR out of the garage simply because it doesn’t have any gas in the tank.

Amazon EC2 has recognized the import of data and made uploads into the cloud a free process. Of course, storing that data in the cloud isn't free, but then again, neither is buying redundant filers, and locating them in geographically dispersed data centers. But more cleverly, once data is sitting in the cloud, not only is access to it nearly guaranteed by amazon, but that data is easily and cost effectively manipulated in any number of ways using the EC2 instances as operators. There is no secret sauce here, no super-sophisticated technology which Amazon has developed with scores of software developers with Mensa membership cards in their wallets.

Cloud vendors take note (IBM people included) Once an organization starts locating corporate data in a particular cloud as a business policy, that cloud provider has won the battle for where that data will be processed. And processing demand only grows with time.

And the winner is... Jaguar Land Rover for business IT project of the year!

Held at a glittering event in London, Platform customer Jaguar Land Rover (JLR) has been awarded the prestigious business project of the year award at the annual UK IT Industry Awards hosted by Computing Magazine in the UK.

Platform Computing worked with Jaguar Land Rover to create an advanced IT environment to underpin the organisation’s virtual car product development, complying with strict safety and environment regulations. JLR has deployed a state-of-the-art system consisting of scalable compute clusters, engineering workstations all built from commodity technologies. The project was recognised for its complexity, ability to reduce the time to market, engineering costs and environmental impact of product development for JLR.

Big congratulations to the team for this significant achievement!

Big Data report from SC’11

In my previous blog I expressed high expectations for the Big Data-related activities at this year’s Supercomputing conference. Coming back from the show, I’d say the enthusiasm and knowledge on Big Data within the HPC community truly surprised me. Here are the major highlights from the show:

  • Good flow of traffic at the Platform booth for Platform MapReduce. Many visitors stopped by our booth to learn more about Platform MapReduce – a distributed runtime engine for MapReduce applications. I found it easy to talk to the HPC crowd because many folks in this  community are already familiar with Platform LSF and Platform Symphony; both are flagship products from Platform that have been deployed and tested in large-scale distributed computing environment for many years. Since Platform MapReduce is built on similar core technology as what’s in those mature products, the HPC community quickly understood the key features and functions it brings to Big Data environments. Even though many users are still at  early stage of either developing MapReduce applications or looking into new programming models, they understand that a sophisticated workload scheduling engine and resource management tool will become critically important once they are ready to deploy their applications into production. Many HPC sites were also interested in exploring the potential of leveraging their existing infrastructure for processing data-intensive applications. For instance, questions on how a MPI and MapRedcue jobs can coexist on the same cluster were frequently asked at the show. The good news, Platform MapReduce is the only solution that can provide capability of supporting mixed workloads.
  •  “Hadoop for Adults” -- This was a quote from one of the attendees after sitting through our breakfast briefing on overcoming MapReduce barriers. We LOVE it! The briefing lured over 130 people and well exceeded our expectations! Our presentation on how to overcome the major challenges in current Hadoop MapReduce implementations drew great interest. “Hadoop for Adults” sums up the distinct benefits Platform MapReduce brings. Platform Computing knows how to manage large-scale distributed computing environments. Bringing that same technology into Big Data environments is a natural extension for us. The reaction at SC’11 for Platform MapReduce was encouraging and a validation on our expertise in scheduling and managing workloads and overall infrastructure in a center.
  • Growing momentum on application development. As sophisticated as always, the HPC community is at the forefront of developing applications to solve data-intensive problems across various industries and disciplines: cyber security, bioinformatics, electronic industry and financial services are just a few examples. Many Big Data related projects are being funded at HPC data centers and we are expecting a proliferation of applications coming out of those projects very soon.

The show is officially over but the excitement around Big Data will continue. For me, not only have I gained tremendous insights on the Big Data momentum in HPC, but I’m also pleased to see the overwhelming reaction for Platform MapReduce within the HPC community. Nothing beats pitching the right product to the right audience!  

Linux Interacting with Windows HPC Server

There are many interesting technology showcases at SC’11 this week. One of the Birds of a Feather sessions this week will discuss a solution implemented at Virginia Tech Transportation Institute (VTTI) that mixes Linux with Windows HPC Server in a cluster for processing large amount of data.


Without a proper tool or a lot of practice, getting Linux and Windows to work together seamlessly to provide a unified interface for end users is a very challenging task. Having both systems coexist in an HPC cluster environment adds an order of magnitude of additional complexity compared to an already complex enough HPC Linux cluster.


This is because Windows and Linux “speak very different languages” in many areas such as user account management, file path and directory structure, cluster management practice, application integrations etc.


The good news is the Platform Computing engineering team did some heavy lifting in product development for this project. Platform HPC integrates with the full software stack required to run an HPC Linux cluster. Its major differentiator compared to alternative solutions is that Platform HPC is application aware. When adding Windows HPC Server into the HPC cluster, the solution delivered by Platform HPC ensures it provides a unified user experiences across Linux and Windows, and hides the difference and complexity between the two OSs.


Platform HPC team has developed a step by step guide for implementing an end-to-end solution with provisioning both Windows and Linux, unified user authentication, unified job scheduling, automated workload driven OS switch, application integrations, and unified end-user interfaces.



This solution significantly reduces the complexity of a mixed Windows and Linux cluster, so users can focus on their applications and their productive work, as opposed to managing the complexity of the mixed Windows and Linux cluster.

Wading throughthe Pre-SC’11 HPC News

The pre-SC11 dawn has already started to heat up with announcements being made by new vendors who see the promise of cloud computing aiding the needs of HPC users everywhere.


Just recently, a competitor of Platform Computing announced their entry into the “cloud bursting” space . They claim that their technology functions with all of the common workload management systems. Without much detail, the only conclusion that can be drawn from the announcement is that they have built a system to “poll – act – sleep” in a loop. While possibly illustrative of the promise of cloud computing, this simplistic view ignores what we believe is the most fundamental challenge to cloud bursting – data locality.

By “data locality” we refer to the fact that compute resources must have some access to data for them to be useful. When datasets get large (input or output), getting access to globally distributed compute resources may have dubious value until workload schedulers understand what data exists where and can affect transfers of data between localities to take better advantage of those remote resources.


Separately, we are very pleased to see others in the HPC industry focusing on the ease of use / ease of build idea. In fact just recently, Rightscale made an announcement to this effect for building clusters in the cloud Platform has been talking about the importance of this for some time, namely with our Platform HPC product. Stay tuned for more announcements as we take this same story into the cloud.


Finally, we were also very happy to see that HPC as a service continues to gain momentum. Just recently there was a story from the Netherlands where a small cluster of very “thick” servers running KVM have been used to create a serviced based HPC infrastructure to be rented to various researchers in the academic and government communities.


No doubt these announcements will be just the beginning of a coming onslaught from many vendors waiting to announce during the week. We expect that cloud and HPC will take one step closer together than they were last year in New Orleans. Stay tuned to the Platform blogs, we’ll provide a summary of the show and any key items we hear about.

Big Data’s Big Show at SC’11

It’s less than a week away, everyone in the HPC community are drumming up for SC’11. As someone who has been to SC for the past seven years, I was pleased to see that Big Data appears to be the  new kid on the block this year. Roughly 20 sessions will be dedicated to Big Data related topics at this year’s Supercomputing show. From basic training on Hadoop and MapReduce to the challenges and opportunities for exascale data analytics, we will hear wide range of discussions on  Big Data.  Platform Computing will also be hosting a breakfast briefing (free!) on how to overcome your MapReduce barriers in the morning of Wed., Nov 16 at the show. Registration details can be found here.

Traditionally, hot topics in HPC are often around performance, scaling, latency and bandwidth. It’s only been in the last couple of years that data intensive computing has become an area of interest in HPC, and it is heating up quickly! The reality is, now that hardware is getting faster and cheaper - thanks mainly to the advancement in processor technologies -- users can run more problems faster. As a result, more data is being generated and a lot of that data contains tremendous insights we could utilize to make better products and decisions. Sure, Big Data exists in Web 2.0, retail, telco companies as well as many other verticals in the enterprise space, yet there is no shortage of use cases in HPC. Areas such as cybersecurity, fraud detection, next-gen sequencing analysis are just a few such examples that fall into HPC arena – often applications in these space are both computationally and data intensive. 

HPC has long been considered the incubator for many bleeding edge innovations that will later trickle down and benefit mainstream applications. We believe Big Data is no exception. The annual Supercomputing show is always about showcasing leading edge science and technologies, Big Data certainly fits the bill.  I am looking forward to learning about new solutions for Big Data problems developed in HPC, as well as getting a better understanding of the specific requirements for this particular market. It will be an exciting week ahead, and we expect to hear some big buzz around Big Data at SC’11!

 

Stop by and visit Platform Computing at SC11 in booth #1117!

Blog Series – Five Challenges for Hadoop MapReduce in the Enterprise, Part 5


Challenge #5: Lack of Multiple Data Source Support

With this blog entry, we have reached the final chapter of the Hadoop challenge series. In this blog, I am going to discuss the fifth challenge for current Hadoop implementations – the lack of multiple data source support.
The current Hadoop implementation does not support multiple data sources; instead, it supports only one distributed file system at a time, the default being the Hadoop Distributed File System (HDFS). This restriction creates barriers for organizations whose data is stored in different file systems other than HDFS while implementing MapReduce applications in Hadoop.  For non-HDFS users, enabling MapReduce applications running in Hadoop environment means they have to first move their data from the current file system into HDFS, which can turn into a timing consuming and very expensive operation.  This limited capability for heterogeneous data support also leads to inferior performance and poor resource utilization due to an inflexible infrastructure.

The reality is, users want a platform that 1) supports various types of input data at the same time and outputs to their desired data sources (which could be different from the input type); 2) completes the   data conversion at runtime so no extra ETL steps are needed after the run.  (See the chart below for a high-level architecture layout for heterogeneous data support.) For instance, a user can send his input data stored in HDFS and output to a relational database such as Oracle and MySQL upon the completion of the run.  Such capabilities eliminate the data movement at both the beginning and the final stages of a MapReduce run, therefore, dramatically reducing the cost while improving the operation efficiency and driving   faster time to results. 

A high level diagram on heterogeneous data support

This capability of  heterogeneous data support in runtime can be considered as an alternative approach to traditional ETL function. The advantage of the former, while compared with the existing ETL tools, is that it provides a faster, cheaper and integrated new platform for users running Big Data applications.   

Having identified the 5 major challenges in the current Hadoop MapReduce implementation, we at Platform Computing has developed a solution – Platform MapReduce, an enterprise-class distributed runtime engine to not only address those barriers mentioned in this blog series, but also bring additional capabilities requested by users wanting to move their Big Data applications into production.  For detailed features and benefits delivered by Platform MapReduce, please visit: http://www.platform.com/mapreduce


Please join us at SC11 for a free breakfast briefing: “Overcoming Your MapReduce Barriers”.  Register today to secure your spot!

Blog Series – Five Challenges for Hadoop MapReduce in the Enterprise, Part 4


Challenge #4: Lack of Quality of Service

We are back after a short break.  The challenge with the current implementation of Hadoop MapReduce continues.  In this blog let’s take a look at the fourth challenge in the existing Hadoop stack – the lack of quality of service.

By high quality of service, we are referring to the capability of dynamically allocating available IT infrastructure based on workloads requirements, maximizing resource utilization and preventing silos. Those capabilities lead to better application performance and faster time to results, and therefore, provide high return on investment for the IT organization.  The current architecture design of the existing open source Hadoop stack puts limitations on the above capabilities. As mentioned in part 2 of this blog series, the single job tracker in the current Hadoop implementation is not separated from the resource manager, so as a result, the job tracker does not provide sufficient resource management functionalities to allow dynamic lending and borrowing of available IT resources. This creates a static IT environment in which each Hadoop application can only run on a pre-assigned set of resources at a given time and no exceptions are allowed. As the requirements of the application changes, resources will need to be re-configured manually to meet new demand. Such a static IT infrastructure brings the following issues for an IT organization:

·         Unable to provide the necessary and guaranteed services to multiple lines of businesses
·         Unable to manage real-time workload requirements
·         Slower performance and time to discovery
·         Increased management complexity

The result?  Underutilized resources and a higher total cost of ownership for IT.

In contrast to a static IT infrastructure, a sophisticated runtime built on a service oriented architecture (SOA) evolution.  brings quality of service to IT organizations committed to providing high quality services to their internal and/or external clients. Such a runtime solution will help transform IT into a true service provider and help meet demanding requirements (-high availability, dynamic resource allocations, ease of management, etc.) in the production environment. As new technologies such as Hadoop and MapReduce continue their penetration into the mainstream market, more applications will be developed and moved into production. Quality of service will undoubtedly become a critical consideration for IT in the next wave of the Big Data


Please join us at SC11 for a free breakfast briefing: “Overcoming Your MapReduce Barriers”.  Register today to secure your spot!




Why Combine Platform Computing with IBM?

You may have read the news that Platform Computing has signed a definitive agreement to be acquired by IBM and you may wonder why. I’d like to share with you our thinking at Platform and what our relevance is to you and the dramatic evolution of enterprise computing. Even though not an old man yet and usually too busy doing stuff, for once I will try to look at the past, present and future.

After the first two generations of IT architecture, centralized mainframe and networked client/server, IT has finally advanced to its third generation architecture of (true) distributed computing. An unlimited number of resource components, such as servers, storage devices and interconnects, are glued together by a layer of management software to form a logically centralized system – call it virtual mainframe, cluster, grid, cloud, or whatever you want. The users don’t really care where the “server” is, as long as they can access application services – probably over a wireless connection. Oh, they also want those services to be inexpensive and preferably on a pay-for-use basis. Like a master athlete making the most challenging routines look so easy, such a simple computing model actually calls for a sophisticated distributed computing architecture. Users’ priorities and budgets differ, apps’ resource demands fluctuate, and the types of hardware they need vary. So, the management software for such a system needs to be able to integrate whatever resources, morph them dynamically to fit each app’s needs, arbitrate amongst competing apps’ demands, and deliver IT as a service as cost effectively as possible. This idea gave birth to commodity clusters, enterprise grids, and now cloud. This has been the vision of Platform Computing since we began 19 years ago.



Just as client/server took 20 years to mature into the mainstream, clusters and grids have taken 20 years, and cloud for general business apps is still just emerging. Two areas have been leading the way: HPC/technical computing followed by Internet services. It’s no accident that Platform Computing was founded in 1992 by Jingwen Wang and I, two renegade Computer Science professors with no business experience or even interest. We wanted to translate ‘80s distributed operating systems research into cluster and grid management products. That’s when the lowly x86 servers were becoming powerful enough to do the big proprietary servers’ job, especially if a bunch of them banded together to form a cluster, and later on an enterprise grid with multiple clusters. One system for all apps, shared with order. Initially, we talked to all the major systems vendors to transfer university research results to them, but there was no taker. So, we decided to practice what we preached. We have been growing and profitable all these 19 years with no external funding. Using clusters and grids, we replaced a supercomputer at Pratt & Whitney to run 10 times more Boeing 777 jet engine simulations, and we supported CERN to process insane amounts of data looking for God’s Particle. While the propeller heads were having fun, enterprises in financial services, manufacturing, pharmaceuticals, oil & gas, electronics, and the entertainment industries turned to these low cost, massively parallel systems to design better products and devise more clever services. To make money, they compute. To out-compete, they out-compute.

The second area adopting clusters, grids and cloud, following HPC/technical computing, is Internet services. By the early 2000s, business at Amazon and Google was going gangbusters, yet they wanted a more cost effective and infinitely scalable system versus buying expensive proprietary systems. They developed their own management software to lash together x86 servers running Linux. They even developed their own middleware, such as MapReduce for processing massive amounts of “unstructured” data.



This brings us to the present day and the pending acquisition of Platform by IBM. Over the last 19 years, Platform has developed a set of distributed middleware and workload and resource management software to run apps on clusters and grids. To keep leading our customers forward, we have extended our software to private cloud management for more types of apps, including Web services, MapReduce and all kinds of analytics. We want to do for enterprises what Google did for themselves, by delivering the management software that glues together whatever hardware resources and applications these enterprises use for production. In other words, Google computing for the enterprise. Platform Computing.



So, it’s all about apps (or IaaS J). Old apps going distributed, new apps built as distributed. Platform’s 19 years of profitable growth has been fueled by delivering value to more and more types of apps for more and more customers. Platform has continued to invest in product innovation and customer services.

The foundation of this acquisition is the ever expanding technical computing market going mainstream. IDC has been tracking this technical computing systems market segment at $14B, or 20% of the overall systems market. It is growing at 8%/year, or twice the growth rate of servers overall. Both IDC and users also point out that the biggest bottleneck to wider adoption is the complexity of clusters and grids, and thus the escalating needs for middleware and management software to hide all the moving parts and just deliver IT as a service. You see, it’s well worth paying a little for management software to get the most out of your hardware. Platform has a single mission: to rapidly deliver effective distributed computing management software to the enterprise. On our own, especially in the early days when going was tough, we have been doing a pretty good job for some enterprises in some parts of the world. But, we are only 536 heroes. Combined with IBM, we can get to all the enterprises worldwide. We have helped our customers to run their businesses better, faster, cheaper. After 19 years, IBM convinced us that there can also be a “better, faster, cheaper” way to help more customers and to grow our business. As they say, it’s all about leverage and scale.

We all have to grow up, including the propeller heads. Some visionary users will continue to buy the pieces of hardware and software to lash together their own systems. Most enterprises expect to get whole systems ready to run their apps, but they don’t want to be tied down to proprietary systems and vendors. They want choices. Heterogeneity is the norm rather than exception. Platform’s management software delivers the capabilities they want while enabling their choices of hardware, OS and apps. 


IBM’s Systems and Technology Group wants to remain a systems business, not a hardware business nor a parts business. Therefore, IBM’s renewed emphasis is on systems software in its own right. IBM and Platform, the two complementary market leaders in technical computing systems and management software respectively, are coming together to provide overall market leadership and help customers to do more cost effective computing. In IBM speak, it’s smarter systems for smarter computing enabling a Smarter Planet. Not smarter people. Just normal people doing smarter things supported by smarter systems.

Now that I hopefully have you convinced that we at Platform are not nuts coming together with IBM, we hope to show you that Platform’s products and technologies have legs to go beyond clusters and grids. After all, HPC/technical computing has always been a fountainhead of new technology innovation feeding into the mainstream. Distributed computing as a new IT architecture is one such example. Our newer products for private cloud management, Platform ISF, and for unstructured data analytics, Platform MapReduce, are showing some early promise, even awards, followed by revenue. 

IBM expects Platform to operate as a coherent business unit within its Systems and Technology Group. We got some promises from folks at IBM. We will accelerate our investments and growth. We will deliver on our product roadmaps. We will continue to provide our industry-best support and services. We will work even harder to add value to our partners, including IBM’s competitors. We want to make new friends while keeping the old, for one is silver while the other is gold. We might even get to keep our brand name. After all, distributed computing needs a platform, and there is only one Platform Computing. We are an optimistic bunch. We want to deliver to you the best of both worlds – you know what I mean. Give us a chance to show you what we can do for you tomorrow. Our customers and partners have journeyed with Platform all these years and have not regretted it. We are grateful to them eternally.

So, with a pile of approvals, Platform Computing as a standalone company may come to an end, but the journey continues to clusters, grids, clouds, or whatever you want to call the future. The prelude is drawing to a close, and the symphony is about to start. We want you to join us at this show.

Thank you for listening.

Taming Big Data – A Recap of the O’Reilly Strata Conference

O’Reilly Strata, a conference dedicated to data science, held its second meeting of the year from September 22 – 23 in New York City.  The conference drew close to 500 attendees, including all the major technology providers in the space as well as user organizations from various industries that are dealing with Big Data problems. Platform Computing was one of the sponsors of the conference, and our introduction of the newly released Platform MapReduce 1.5 generated wide interest among conference participants.

Platform team greets booth visitors 

My takeaways from the conference are following:
1)   Data science is hot. The Strata conference lured participants from various industries. We met startups working on web analytics, energy companies trying to analyze their log files, banks of looking for use cases and solutions to help analyze their large data sets fast, and of course, Web 2.0 companies who are already at the forefront of tackling Big Data problems exploring  better solutions than what they are currently using. The enthusiasm around cracking Big Data   has never been higher. 
2)   Big Data market is still young. Many discussions we had at the show revealed that majority of the main stream organizations (excluding Web2.0 companies like Google, Yahoo, Facebook) are still at the early stages of the adoption where they are either exploring various technologies on the market, figuring out the proper applications built upon new technologies such as Hadoop and MapReduce, or in the midst of building small pilot projects. Comments such as “We are thinking about moving certain applications to Hadoop,” or “We only have one Hadoop project running at this time” were often heard at the conference.  A majority of the organizations that have large amounts of data are just beginning to tap into Big Data and looking into proper use cases.. The FAQ these days is “What to do with my data?” and there is no simple answer to that. 
3)   There’s a shortage of skilled resources. While Hadoop and MapReduce appear to be promising approaches to access and analyze Big Data, they are also new to developers, and the learning curve is rather steep. Adding to the length of the learning cycle is the fact that development tools in the ecosystem are still yet to mature. The reality is that there is lack of production quality applications for main stream user, most codes today are developed in-house and still being tested in R&D labs.  
4)   The market is fragmented. Various tools have been developed to fulfill the ecosystem while the mainstream market is still catching up on the basics.. We believe the gap will close and the market will eventually hit the tipping point as more applications become available. But for now, the path for Big Data remains wildly unpredictable. 
Big Data is here to stay. However, “Making Data Work”, the slogan at O’Reilley Strata conference, is no easy task. Companies dealing with  large amounts of data have a lot on their plates today --  disruptive technologies,  new application development,  understanding the meaning of the  new discoveries and their business impact, just to name a few. Needless to say, a well built-out ecosystem is critical to support all the efforts taking place in the market.    At Platform Computing, we are committed to not only providing the best solution in the ecosystem through leveraging our proven technology, but also working toward  bringing a viable,  end-to-end solution to the market.      



ISC cloud 2011

The European conference concerning the intersection between cloud computing and HPC has just finished, and it's very pleasing to report that this conference delivered considerable helpings of useful and exciting information on the topic.

Though cloud computing and HPC have tended to stay separated, the HPC community starting with the sc2010 conference, interest has been gaining primarily because access to additional temporary resource is very temping. However, other reasons for HPC users and architects to evaluate the viability of cloud include total cost of ownership comparisons, and startup businesses which may need temporary access to HPC but do not have the capital to purchase dedicated infrastructure.

Conclusions varied from presenter to presenter, tough some things were generally agreed upon:

  1. if using Amazon EC2, HPC applications must use the cluster compute instance to achieve comparable performance to local clusters.
  2. fine grained MPI applications are not well suited to the cloud simply because none of the major vendors offer infiniband or other low latency interconnect on the back end
  3. running long term in the cloud, even with favorable pricing agreements is much more expensive than running in local data centers, as long as those data centers already exist. (no one presented a cost analysis which included the datacenter build costs as an amortized cost of doing HPC.)


Another interesting trend was the different points of view depending on where the presenter came from. Generally, researchers from national labs had the point of view that cloud computing was not comparable to their in-house supercomputers and was not a viable alternative for them. Also, compared to the scale of their in-house systems, the resources available from Amazon or others were seen as quite limited.

Conversely, presenters from industry had the opposite point of view (notably a presentation given by Guillaume Alleon from EADS). Their much more modest requirements seemed to map much better into the available cloud infrastructure and the conclusion was positive for cloud being a capable alternative to in-house HPC.

Perhaps this is another aspect of the disparity between capability and capacity HPC computing. One maps well into the cloud, the other doesn't.

Overall it was a very useful two days. My only regret was not being able to present Platform's view on HPC cloud. See my next blog for some technologies to keep an eye on for overcoming cloud adoption barriers. Also, if anyone is interested in HPC and the cloud, this was the best and richest content event I've ever attended. Highly recommended.

Congratulations to Rick Parker – a Platform Customer, Partner and SuperNova Semifinalist!


Over the past year, I’ve had the pleasure of working closely with someone I regard as a true visionary when it comes to IT management and unlocking the power of private cloud computing. I am referring to Rick Parker who, until recently, has been the IT Director at Fetch Technologies, a longtime customer of Platform Computing, which leverages Platform ISF for their private cloud management. Now I am proud to now call Rick by a new name: Protostar! From a pool of more than 70 applicants, Rick was recognized in Constellation Research’s SuperNova Awards among an elite group of semifinalists that have overcome the odds in successfully applying cloud computing technologies within their organizations.

Most award programs recognize technology suppliers for advancements in the market.  Few programs recognize individuals for their courage in battling the odds to affect change in their organizations.  The Constellation SuperNova Awards celebrate the explorers and pioneers who successfully put new technologies to work and the leaders that have created disruptions in their market. Rick fits the bill perfectly, and an all-star cast of judges (including Larry Dignan at ZDNET and Frank Scavo at Constellation - to name a couple) have agreed. The award recognized applicants who embody the human spirit to innovate, overcome adversity, and successfully deliver market changing approaches.  

Here is an excerpt from the award nomination to give you a peek into Rick’s story of implementing cloud computing at Fetch (with a little help from Platform Computing solutions!).
Since joining Fetch, Rick has successfully implemented a dynamic, nearly completely virtualized data center leveraging private clouds and Platform Computing’s ISF cloud management solution. Rick’s journey originally began with a simple idea: to build the perfect data center that moved IT away from the business of server management and toward true data center management. As part of this vision, Rick founded Bedouin Networks to create one of the first, if not the first, public cloud services in 2006 and deliver radically improved cost effectiveness and reliability in data center design.
In 2007, Rick carried his passion for disruptive data centers with him when he joined Fetch Technologies. Fetch Technologies is a Software-as-a-Service (SaaS) provider that enables organizations to extract, aggregate and use real-time information from websites and, as such, depends on its ability to maximize data center resources, efficiently and effectively. At any given moment, Rick and his IT team can get a call to provision more compute resources for Fetch’s fast-growing customer base. Before using Platform ISF, Fetch provisioned resources manually to increase SaaS capacity, which usually took several hours of personnel time per server. The cost effectiveness of Rick’s innovative design not only made Fetch’s products and services more profitable, it made them possible. The expenditure that would have been required for the hundreds of physical servers, networking, and data center costs to enable the additional capacity required could have exceeded the potential revenue.

A Protostar
Rick has been a vital asset to the Fetch team and an incredible partner to the Platform Computing team over the years. We congratulate him on this latest recognition and wish him the best of luck in his newest adventure – in search of more disruption – in his new position as Cloud Architect at Activision. We look forward to working with Rick on his next private cloud adventure. Congratulations, Protostar!

One small step for man, one enormous leap for science

News from CERN last week that E may not, in fact, equal mc2 was earth shattering. As the news broke, physicists everywhere quivered in the knowledge that everything they once thought true may no longer hold and the debate that continues to encircle the announcement this week is fascinating. Commentary ranges from those excited by the ongoing uncertainties of the modern world to those who are adamant mistakes have been made.

This comment from Matt Alden-Farrow on the BBC sums up the situation nicely:

“This discovery could one day change our understanding of the universe and the way in which things work. Doesn’t mean previous scientists were wrong; all science is built on the foundation of others work.”

From our perspective, this comment not only sums up the debate, but also the reality of the situation. Scientific discoveries are always built on the findings of those that went before and the ability to advance knowledge often depends on the tools available.

Isaac Newton developed his theory of gravity when an Apple fell on his head – the sophisticated technology we rely on today just didn’t exist. His ‘technology’ was logic. Einstein used chemicals and mathematical formulae which had been discovered and proven. CERN used the large hadron collider and high performance computing.

The reality is that scientific knowledge is built in baby steps, and the time these take is often determined by the time it takes for the available technology to catch up with the existing level of knowledge. If we had never known Einstein’s theory of relativity, who’s to say that CERN would have even attempted to measure the speed of particle movement?

IDC HPC User Forum – San Diego

Platform just returned from attending the IDC HPC User Forum held from Sept. 6-8 in San Diego.


As opposed to previous years, this year’s event seemed to have drawn fewer people from the second and third tiersof the HPC industry. Overall attendance for this event also appeared to be about half of that compared to the event in April in Texas.


This time the IDC HPC User Forum was dominated by a focus on software and the need for recasting programming models,. There was also a renewed focus on getting ISVs and open source development teams to adopt programming models that can scale far beyond the limits they currently have. Two factors are driving this emphasis
  • Extremely parallel internals for compute nodes (from both a multi-core and an accelerator [CUDA, Intel, AMD] points of view).
  • The focus on “exa” scale, which by all counts will be achieved by putting together ever increasing numbers of commodity servers

Typically there is a theme to the presentations for the multi-day event, and this forum was no different. Industry presentations were very focused on the material science being performed primarily by the US national labs and future applications of the results being obtained. The product horizon on the technologies presented was estimated at approximately 10 years.


In contrast to the rest of the industry which is very cloud focused right now, cloud computing was presented or mentioned only three times by various vendors and also mentioned by by the Lawrence Berkeley National Laboratory (LBNL) at the forum. When it comes to cloud, there seems to be a split between what the vendors are focusing on and what the attendees believe. Specifically, attendees from national laboratories tend to be focused on “capability” computing (e.g. large massively parallel jobs running on thousands of processors). Jeff Broughton from LBNL presented some data from a paper that showed how, for the most part, cloud computing instances are not ready for the challenge of doing HPC.


Though we can’t refute any of the data or claims made by Mr. Broughton, the conclusions that may be drawn from his data might extend beyond what the facts support. For instance, in our experience here at Platform, we’ve found that most HPC requirements in industry do not span more than 128 cores in a single parallel job nor do they require more than 256 GB of memory. The requirements for most companies doing HPC are significantly more modest and are therefore much more viable to be addressed by a cloud computing model.


We at Platform have long been fighting the “all or nothing” notion of HPC employing cloud technology. Rather, we believe that industry – especially in Tiers 2 and 3, to a lesser or greater degree, will be able to make extremely beneficial use of cloud computing to address their more modest HPC requirements. Platform is focused on developing products to help these customers easily realize this benefit. Stay tuned for more on Platform’s cloud family of products for HPC—there will be more on this in the coming months…