HPC from A - Z (Part 26) - Z

Z is for… Zodiac
The wonders of the universe will remain of interest to the human race until the end of time or at least until we think we know everything about everything. Whichever comes first…

With the developments in technology of late, the amount of information we have about constellations and other universes is growing exponentially. For all we know, in fifty years we may be taking holidays in space or have discovered a form of life just like us in a far away galaxy.

But understanding what is really out there is far more complex than books and films make it seem. The pretty pictures of constellations don’t do astronomy justice! The amount detail needed to track star and planet movements and understand which direction constellations are moving in requires some seriously high resolution telescopes. Just think about the amount of ‘zoom’ required to detect traces of flowing liquid on Mars. This is well beyond the capabilities of your standard Canon or Nikon – that’s for sure!

With high resolution come high data volumes. So, like all the posts before this, HPC is crucial for cosmology, astrophysics, and high energy physics research. Without it, results could take years to find instead of months of minutes. By the time the path of the celestial sphere is mapped it could easily be into its second or even third cycle.

HPC can also be used in more theoretical contexts. For example, researches at Ohio State University required the compute power provided by the Ohio Supercomputing Center Glenn Cluster to run simulations and modeling required for their study on the effects of star formation and growth of black holes.

As we finally reach the end of our ABC series, there’s no denying the critical role that compute power plays in our day-to-day lives. Technology is developing at a startling pace, and with each and every new development comes more data and a consequent need to process and make sense of it. Without HPC our technological advancements would not be nearly as fast, and we as a society would not have the insight and capabilities that we do today.

Blog Series – Five Challenges for Hadoop MapReduce in the Enterprise, Part 2

Challenge #2: Current Hadoop MapReduce implementations lack flexibility and reliable resource management

As outlined in Part 1 of this series, here at Platform, we’ve identified five significant challenges that we believe are currently hindering Hadoop MapReduce adoption in enterprise environments.  The second challenge, addressed here, is a lack of flexibility and resource management provided by the open source solutions currently on the market.

Current Hadoop MapReduce implementations derived from open source are not equipped to address the dynamic resource allocation required by various applications.  They are also susceptible to single points of failure on HDFS NameNode, as well as on JobTracker. As mentioned in Part 1,  these shortcomings are due to the fundamental architectural design in the open source implementation in which the job tracker is not separated from the resource manager.  As IT continues its transformation from a cost center to a service-oriented organization, the need for an enterprise–class platform capable of providing services for multiple lines of business will rise.  In order to support MapReduce applications running in a robust production environment, a runtime engine offering dynamic resource management (such as borrowing and lending capabilities) is critical for helping IT deliver its services to multiple business units while meeting their service level agreements.  Dynamic resource allocation capabilities promise to not only yield extremely high resource utilization but also eliminate IT silos, therefore bringing tangible ROI to enterprise IT data centers.

Equally important is high reliability. An enterprise-class MapReduce implementation must be highly reliable so there are no single points of failure. Some may argue that the existing solution in Hadoop MapReduce has shown very low rates of failure and therefore reliability is not of high importance.  However, our experience and long history of working with enterprise-class customers has proved that in mission critical environments, the cost of one failure is measured in millions of dollars and is in no way justifiable for the organization. Eliminating single points of failure could significantly minimize the downtime risk for IT. For many organizations, that translates to faster time to results and higher profits.

HPC from A-Z (part 25) - Y

Y is for… yield strength

Think how many children play with plasticine and Play-Doh and day dream about being an artist or a sculptor when they grow up. The majority of adults actually follow different career paths but the true passions of a child often resonate with their adult selves in a different and more advanced form. Just look at the number of engineers in the world!

‘Yield strength’ or ‘yield point’ is in many ways, basic knowledge for ‘grown up’ sculptors. It is the stress at which a material begins to deform and can no longer return to its original shape. This is vital information for design as it represents an upper limit of a load that can be applied on a surface and similarly is an important consideration in materials production. How would we create new objects and materials without this knowledge?

Times have moved on since my Play-Doh days and the engineers of today no longer need to determine the yield strength of a material by stacking weights on top of materials one by one. Instead, this type of testing is nearly all done through simulations so that the extensive (and expensive) real world testing is only conducted on a select few prototypes. Think about the materials used to create space shuttles. Not only do they cost a small fortune but making a mistake and using a material with the wrong yield strength could actually impact human life. Simulation is crucial to avoid these types of issues.

The role of HPC? It’s all about the high performing cluster technology that is required for the engineers at the heart of material development to vet prototypes without ever having to develop scale models. It’s a profit enabler, enabling faster product development and time to market for materials and the designers that make use of them.

Blog Series – Five Challenges for Hadoop MapReduce in the Enterprise


With the emergence of “big data” has come a number of new programming methodologies for collecting, processing and analyzing the large volume and often unstructured data. Although Hadoop MapReduce is one of the promising approaches for processing and organizing results from unstructured data, the engine running underneath MapReduce applications  is not yet enterprise ready. At Platform Computing, we have identified five major challenges in the current Hadoop MapReduce implementation:

·         Lack of performance and scalability
·         Lack of flexible and reliable resource management
·         Lack of application deployment support
·         Lack of quality of service
·         Lack of multiple data source support

I will be taking an in-depth look at each of the above challenges in this blog series. To finish, I will share our  vision on what an enterprise–class solution should be  that will not only address the five challenges customers are currently facing,  but also expand beyond those boundaries to explore the capabilities  of the next generation Hadoop MapReduce runtime engine.

 Challenge #1:  Lack of performance and scalability

Currently the open source Hadoop MapReduce programming model does not provide the performance and scalability needed for production environment, this is mainly due to its fundamental architectural design.   On the performance measure, to be most useful in a robust enterprise environment a MapReduce job should take  sub-millisecond to start,  but the job startup time in the current open source MapReduce implementation is measured in seconds. This high latency at the beginning can lead to subsequent delays in getting to the final results and cause significant financial loss to an organization. For instance, in capital markets of the financial service sector, a millisecond of delay can cost a firm millions of dollars.  On the scalability front, customers are looking for a runtime solution that is not only capable of  scaling one MapReduce application as the problem size grows,  but one that can also support multiple applications of different  kinds running across thousands of cores and servers at the same time.  The current Hadoop MapReduce implementation does not provide such capabilities. As a result, for each MapReduce job, a customer has to assign a dedicated cluster to run that particular application, one at a time.  This lack of scalability will not only introduce additional complexity into an already complex IT data center and make it hard to manage, but it also creates a siloed IT environment in which resources are poorly utilized.

A lack of guaranteed performance and scalability is just one of the roadblocks preventing enterprise customers from running MapReduce applications at production scale.  In the next blog,   we will discuss the shortcomings in resource management in the current Hadoop MapReduce offering and examine the impact it brings to organizations tackling “Big Data” problems.

HPC from A-Z (part 24) - X

X is for X-ray analysis

There’s no denying that science is significantly more advanced than ever before. In fact, I would argue that today’s doctors and scientists would be lost if asked to work in the conditions experienced by their counterparts in times gone by.

Modern technology has allowed medicine and associated scientific disciplines to go beyond the simple diagnosis of coughs and colds or a broken leg. Instead, researchers can now extract huge amounts of detail about atom arrangement and electron density using techniques such as x-ray crystallography modelling and gather physical proof for their theories.

For example, a deeper understanding of biological molecules means bioscientific researchers can examine molecular dynamics, look further into protein analysis and identification, sequence alignment and annotation, and structure determination analysis. The results are fascinating and they are really driving the growth of knowledge in the scientific profession. What many don’t realize, however, is the sheer volume of data behind discoveries and how much information advanced research techniques actually generates.

Another example is medical imaging studies which use data from digitized X-ray images. In this case HPC is used to glean information from numerous medical databases. In fact, supercomputing facilities, along with industry, are exploring ways today that will ensure access to massive amounts of stored digital data for future research.

With HPC able to provide such insight into scientific happenings here on earth, just think what insights could come from Chandra about the wider Universe (and ET!) in years to come…