Showing posts with label cluster. Show all posts
Showing posts with label cluster. Show all posts

Is Platform HPC 3 Just a GUI around Linux?

Platform HPC 3 has a new skin, a brand new web interface that is modern and future looking. While we are emphasizing the usability of Platform HPC with this new product release, some people think Platform HPC is just a GUI around Linux. They couldn’t be more mistaken.


GNOME and KDE are graphical desktop interfaces for Linux, which are included in most Linux distributions. They make Linux easier to use. However, making Linux easier to use does not turn a bunch of servers into a cluster. A cluster that acts as a single system for users and administrators requires more capabilities from management software.


First, it requires a cluster management system. This allows administrators to easily manage a large number of nodes within a cluster without perming tasks node by node. Management tasks include installing the OS and software, patching the OS and upgrading applications, etc. on nodes. Cluster monitoring is also an important component of the cluster management. Without cluster monitoring capabilities, administrators would have to login to individual nodes to get cluster health and performance information, which is impractical for a cluster with more than 10 nodes. Using consolidated alerts can release administrators from having to closely watch the cluster all day long. Cluster management is the foundation that allows all nodes within a cluster to act like a single system from an administration perspective.


Second, a cluster requires workload management. This helps make a cluster into a single system for multiple users. Without workload management, a cluster is many discrete servers from user’s perspective. Instead, workload management helps run each user’s application instance as a “job”. It schedules jobs according user specified requirements and available resources. Workload management software automates fault tolerance and load balancing. It also turns multiple discrete commodity servers into a single reliable system.


Finally, a reporting system is essential for management so administrators can understand how cluster capabilities have been used and whether or not additional capabilities are required. It also shows how well users are being served using performance indicators like average job wait time, job run time, etc. This is to ensure a good return on investment for the valuable computer cluster.


There exist open source software solutions for cluster management, workload management, and reporting on the market today. However, nothing replaces a solution that integrates them together with an easy-to-use web interface while the command line interface is still there. This dramatically reduces system set-up time and the learning curve for users. It also helps users focus on their work rather dealing with complex cluster issues. For power users, the flexible command lines are still there for customization and extension. This is what Platform HPC 3 delivers. It is far more than just a nice GUI around Linux.

Performance and Productivity of an HPC Cluster (3)

My last blog described the results of implementing a refreshed 32-node CFD (computational fluid dynamics) cluster by using a completely different solution from the one implemented in the previous cluster. Using Platform HPC for the refresh, the new cluster significantly reduced its job failure rate and had a 15% usage increase and 25% throughput increase.


After the HPC cluster was in production with happy users, the management team began making plans to further improve application performance and user productivity. They were able to start to investigate and plan for the future. GPU adoption and a new ways to speed up the design process are two areas identified.


GPU Use


Using GPU to accelerate applications is what this organization is looking at next. Some commercial applications ISVs have already released applications that support GPU acceleration, for example, ANSYS. Adding GPU to the cluster could dramatically reduce the application run time. However, it will obviously increase the cluster complexity. Fortunately Platform HPC can help reduce such complexity. With Platform HPC, it’s much easier for administrators to deploy GPU required software. Platform Computing has repackaged NVIDIA’s CUDA software into a format that can be automatically deployed across all compute nodes in a cluster by leveraging the cluster management capability of Platform HPC. The workload scheduler in Platform HPC is also GPU aware so it can schedule GPU resources such that users running GPU jobs don’t need to worry about which node has GPUs or if the GPUs are used by other users. These capabilities would go a long way to helping this organization overcome the hurdles of adopting GPU technology.


Design Process Acceleration


After users transferred from using the command line interface to the web interface, they still had some complex scripts that they needed to automate job flows. These scripts were originally built by a few power users. New users typically just copied them and wrapped the scripts with a few lines of code as required. But when there was a problem, it was very difficult to debug. One technology they are evaluating to help alleviate this problem is Platform Process Manager. Platform Process Manager allows users to program a flow in a graphical way without writing scripts. It can be integrated with Platform HPC’s web interface. Once a flow is developed, users can click a button to launch the flow. Job dependencies are taken care of by the Platform Process Manager. This company is exploring the possibility of using Platform Process Manager as a technology to automate their job flow to further increase user productivity. The graphical view of the flow also makes the job flows much easier to maintain and improve, which can also be used to document best practices in their engineering process.


The partnership with Platform Computing allows this design firm to advance digital design efficiency with increased user productivity. By doing more simulations to improve the quality of the design, high performance computing gives them the competitive advantage to be leading edge.

Performance and Productivity of an HPC Cluster (2)

In my last blog, I described an audit we recently did for a small manufacturing company on their 32-node CFD (computational fluid dynamics) cluster. The objective of the audit was to both evaluate the efficiency of the cluster and develop an upgrade plan. In this blog, I will describe their experience and the results they had after deploying Platform HPC.


As a refresher, through the auditing process, we identified that the company had less than optimal cluster usage, with 40% of the cluster value wasted and an 18% loss of user productivity.


The solution for making the refreshed cluster more efficient was geared towards shorter set-up time, higher cluster utilization and creating quick troubleshooting processes. Platform HPC Enterprise was selected for the company’s upgrade. After the cluster hardware was installed and connected, it took less than one day to install all the software components on the head node, together with the operating systems and management software components provisioned on all compute nodes. It took another day to get ANSYS Fluent and ANSYS Mechanical integrated with the MPI library (Platform MPI) and Platform HPC web portal. After these steps, the cluster went in production almost immediately. Platform HPC Enterprise did the heavy lifting for the management software installation and application integrations. Similar tasks that previously took two months to complete were now completed within three days.


One of the improvements we made was in the application license configuration process, where we removed group specific application license reservations. This was one of the reasons contributing to the low cluster utilization. The license reservation system was supposed to guarantee application licenses for mission critical user groups. Rather than using a reservation system, Platform HPC allows dynamic license preemption. This means jobs with high priority can stop low priority jobs, pre-empting application licenses if necessary. The low priority jobs do not die. Instead, they resume when high priority jobs are completed and licenses are freed up. We saw a 15% increase in cluster usage using this configuration.
Another performance boost came as a result of the web-based interface in Platform HPC, because users do not need to deal with scripts. Using the web interface, they can manage their CFD jobs, as well as the data. This dramatically reduced job failure rates for the simulation. Administrators also got relief from dealing with user problems, thus resulting in a 10% time savings. This allows the administrator to now spend time on other duties.


The new cluster also uses Voltaire Fabric Collective Accelerator (FCA), which speeds up communication among all tasks for an MPI job. With the integration between Platform MPI and Voltaire FCA, ANSYS Fluent performance increased by 10-20%.


After the initial deployment and performance tuning, we saw a 25% job throughput improvement, a 15% license usage increase, and a 10% reduction in administration effort overall. This is a significant improvement over the previous cluster.


In the next blog, I will discuss the next steps planned for this site to further improve user productivity.

What a Part Time HPC Cluster Admin Needs

In most organizations that use small and medium HPC clusters (smaller than 200 nodes), a HPC cluster is treated as a separate system from an IT management perspective. As a result, the amount of IT administration effort allocated to HPC clusters is very limited. Often, a part-time Linux administrator is tasked with taking care of an HPC cluster that is used by tens and even hundreds of users. Because most IT administrators are not HPC experts, they usually rely on the management tool included with the HPC cluster package to perform their daily work. Those packages are usually a stack of open source software with very limited support. Because these software packages are assembled from functionality perspective, they are not integrated. But what IT administrator really need is a robust, easy-to-use management tool to keep the cluster up and running rather spending time integrating the software stack provided with the cluster. This is a very different use case than large HPC shops that have lots of HPC expertise.


We’ve designed and built Platform HPC specifically for organizations that need small or medium-sized clusters. Starting with an easy one-step installation, Platform HPC automates many of the complex tasks of managing a cluster, including provision and patch nodes, integrate applications, and troubleshooting user problems, saving time and hassles for part-time IT administrators. A web interface also makes remote administration a reality for users. As long as administrator can access a web browser that can reach the head node of the HPC cluster, he can easily monitor, manage and troubleshoot problems.


Then again, the tools and technologies included in Platform HPC were originally designed for large scale environments. So under the hood, those tools are powerful, scalable, and highly customizable. For those running a small cluster, whether or not they are rapidly growing, Platform HPC would be their best choice now and for the future.

Santa and HPC

“So just how does Santa manage to travel the earth in just one evening and deliver all those presents” is a question you might get asked this Christmas. Sure, you can opt for the “magic” answer but perhaps Santa has a really awesome HPC environment up there at the North Pole?

Red Bull Racing is using HPC software to significantly accelerate its computer-aided design and engineering processes for its winning Formula One Cars. If Red Bull can use HPC to optimise designs that maximise the downforce and reduce the drag cars create on the track, then why can’t Santa use HPC for his sleigh design?

As Santa is going to experience some turbulent and icy conditions on his route this year, I’d like to think his elves have been running simulations via HPC to ensure his sleigh can cope with these adverse conditions, without sacrificing on speed. Throw into the mix the rising cost of gasoline, in monetary and environmental terms, Santa’s going to have to ensure his sleigh is energy efficient too. Of course the rise in population is something he needs to keep in mind, Santa now needs to make far more stops that he did many years ago, and can take advantage of HPC to analyze all those naughty and nice lists to take the most efficient route possible.

The design conundrum outlined above really does point toward the need for an HPC environment at the North Pole. Santa just can’t afford to take a risk with the design of his sleigh. There would just be too many disappointed children to imagine.

So when you’re faced with the “just how does Santa do it” question, why not take this as an opportunity to introduce the child to the infinite possibilities of High Performance Computing!



Note: All opinions in this blog are my own and not officially endorsed by other people named Saint Nick.

Overheard at SC’10…

It was a busy but exciting week for me last week in New Orleans during Supercomputing 2010 (SC’10). I’m really pleased to report that our complete cluster management solution, Platform HPC, which was announced one week prior to the conference, resonated very well with many of the show attendees I talked to.

As our VP of Product Management and Marketing Ken Hertzler noted in his blog announcing the product release a couple of weeks ago, Platform HPC was designed to make cluster building an easier, less painful process for people who need clusters, but may not be cluster experts. It’s a complete solution that boasts built-in workload management, application integration, and leading MPI, each of which give Platform HPC a leg up on our competition. In addition to the completeness, ease of use and commercial support are other two aspects that set Platform HPC apart. In packaging this product, we found that support was a huge differentiator with our channel IHV and ISV partners—they told us they always have happier customers when they can sell a complete cluster management solution where commercial support is delivered together with their hardware or applications. They spoke, so we listened!

So you can imagine how pleased we were to also be listening in on some of the comments overheard in our booth at SC’10 last week…

“I thought the best thing about your tool was the completely integrated solution stack it provides. This provides a strong competitive advantage over other advanced scheduling/provisioning/and monitoring tools in that you are able to offer an entire solution stack and have control over it to meet customers’ needs. I don't think Adaptive Computing has this, as Moab requires plugins to xCAT, CMU, or ROCKS.”

Supercomputing goes commercial in China

Just having come back from the HPC China Conference in Beijing two weeks ago where TianHe 1A was announced, I am now packing for SC’10 in New Orleans.


For those of you unfamiliar with TianHe 1A, it is currently the world’s fastest supercomputer capable of computing 2.507 petaflops, or 2,507 trillion calculations per second. TianHe 1A, together with Nebulae (the word’s second fastest supercomputer according to the July 2010 TOP500) will be hosted in government funded supercomputing centers in China. However, unlike many other countries where the top systems are almost exclusively running code for research, supercomputing centers in China will be running like commercial companies to sell computing capacity to industries and research firms. Their model center, the Shanghai Supercomputing Center (#24 on July 2010 TOP500) has about 300 customers and keeps full capacity operations 24x7. As the result, those big systems will run more commercial code than the systems in national labs. Such commercialized operations will demand high software quality and strong local support.


The recently launched Platform HPC matches needs like these very well. With a complete management solution with a leading MPI implementation, Platform HPC brings performance and productivity to production environments. It provides all the necessary components to manage large systems including user web portal, application integrations, resource allocation and job scheduling, cluster management and monitoring, and job usage accounting. Most importantly, it is supported by a single vendor to dramatically decrease troubleshooting time.


I’m looking forward to seeing how such HPC cloud model on top systems shakes the future of HPC community.

Introducing Platform HPC, the most complete cluster solution available

The past few weeks have been really busy here at Platform leading up to today, but I’m proud to say that we have just released the latest product in our flagship HPC product family, Platform HPC. As the most complete cluster software solution currently on the market, Platform HPC makes cluster management and workload scheduling easy for users who are not experts in clusters, but need HPC clusters to help improve their application performance. One main goal with Platform HPC is to allow technical application users across a variety of vertical markets—from engineering and scientific research to oil and gas and digital media—to spend their time and energy concentrating on their research and designs rather than setting up and managing their computing environments. Until now, clusters have been a daunting prospect for most workstation and technical application users. While engineers, scientists and graphic artists may all be expert users of the apps they work with on a daily basis, most may not be IT experts—and within most organizations, cluster building has been left to the IT department because clusters can be challenging to set up. To address this, we specifically designed Platform HPC with technical app users in mind so that both experienced or novice HPC users can quickly and easily deploy, run and manage their own clusters and still meet the application performance and workload management demands they require.

Another primary goal of Platform HPC is to make is easier for user to use their cluster from their application. Based on our Platform LSF workload scheduler, Platform HPC’s interface is accessible through a web portal interface that is integrated through application templates with some leading technical applications already on the market. We’re supporting apps by Abaqus, ANSYS, Blast, Fluent, LS-DYNA, MSC Natran and Schlumberger among others, with plans to add more in the next few months. Platform HPC uses a single installer and is pre-certified by leading hardware and software vendors and available through our certified channel partners. The new product also allows multiple apps to run on the cluster regardless of OS—it supports both Linux and Windows environments. Talk about maximizing your existing infrastructure and compute resources! We also fully support the product and provide one point of contact for support, despite the product being sold through the channel - just another way to make clusters easier and more approachable for users who aren’t cluster experts!

Platform support for GPU Environments

In addition to our Platform HPC product launch, we’re also pleased to announce our continued support for GPU-aware application workload scheduling and monitoring. Both Platform LSF and Platform HPC now incorporate GPU-aware scheduling features for better and more efficient workload management in GPU environments. For more on our support of GPUs, see today’s release.

Finally, we’ll be at Supercomputing ’10 (SC’10) in New Orleans next week, so please stop by the Platform booth (#2739) if you’d be interested in a demo. And stay tuned for more exciting Platform product news at SC’10, as well!

Microsoft and Platform supercharge the HPC cluster!

It’s no secret that time is money, and designers, engineers and product analysts need to design, prototype and deliver products to market faster in order to meet shrinking product lifecycles and be competitive. To date, the obvious solution has been to do more and larger simulations that can create and even test parts prior to any metal being machined. Increased model sizes and the need for reduced run times has led many design teams to purchase and deploy several high performance computing (HPC) clusters for the muscle required to power their much needed solutions.

The challenge of this obvious solution is that the applications selected by design teams for complex simulations--which can include mechanical analysis, thermal modeling, rendering and fluid flow--are not always optimized for the same operating systems. This has led many enterprises to deploy multiple HPC clusters—one for each operating system—to allow users to run different applications, resulting in cluster silos and underutilized capacity. Design teams need a cluster that can dynamically switch between Windows and Linux, depending upon workload, and Platform Adaptive Cluster is what design teams can use to deploy a hybrid Windows/Linux HPC cluster.

A hybrid Windows/Linux cluster also cuts cost. Since the cluster can support both Linux and Windows less hardware needs to be purchased, and Platform Adaptive Cluster also increases effective application performance by allowing applications to access the capacity of the entire cluster. Cluster silos are eliminated and cluster resources are maximized.

Earlier this week, Microsoft made some waves in the HPC market by with the release of its much-awaited Windows HPC Server 2008 R2. When Windows HPC Server 2008 R2 release is used in combination with Platform Adaptive Cluster, HPC designers, engineers and product analysts can optimally coordinate the management of their HPC environment and workloads.

What’s great about the new Windows HPC Server 2008 R2 is that it integrates well with all the leading HPC solutions such as Dell, Cray, and Platform, allowing lots of collaboration and performance improvements across an enterprise’s HPC solution portfolio. With the Platform-Microsoft solution, enterprises no longer need to deploy separate HPC clusters to accommodate different operating systems, making it easier and more cost effective to deploy and use hybrid Windows and Linux clusters. The net result of the Platform-Microsoft solution is that design teams can better leverage technology in their quest to deliver better products to market faster.

Platform HPC Enterprise Edition Launched

Application performance is no longer relying on the computer clock speed – parallelization is the only way to dramatically scale application performance. Referred to as a cluster, it is the foundation to scale an application, but unlike an operating system on single machine (i.e., Linux), setting up and using a cluster remains a daunting task for many users. Why? First, there are multiple software components that are required to run the applications in a cluster environment. This includes parallelization libraries (MPI), job distribution workload management, and administrative cluster management. Second, all these components must be integrated to work together. Though they seem to install fine, they often break in production. So, can these software components be packaged, integrated, and tested together up-front to behave more reliably like an operating system?


Pre-assembling cluster software with hardware has been tried in the past by many different user organizations and usually resulted in complicated, time-consuming systems to manage. Platform Computing has taken on this challenge to solve this problem. The launch of Platform HPC Enterprise Edition is a more complete offering, as compared to their previous launch of Platform HPC Workgroup Edition last year. What I like is their unique seamless integration between cluster management and workload management along with their MPI library and system monitoring. In my experience, now both users and administrators can get their job done by interfacing through a single unified web portal. This makes a cluster more like a single operating system. The product’s single install and single user interface makes a highly technical and complex high performance Linux cluster easy to use for the average cluster user.


Join Platform on May 26, 2010, at 11:00 a.m. ET for a webinar to discuss the product features, capabilities and benefits of the new Platform HPC Enterprise Edition.

Making the HPC lifecycle easier

Last week, I attended the IDC HPC User Forum in the Detroit area. Among the many interesting topics discussed at the Forum was how HPC helps manufacturing firms remain competitive. Large organizations are definitely in leading positions when it comes to HPC. After many years of investment, there is a lot of well established HPC expertise within large companies. IDC data shows the main hurdle for HPC getting penetration beyond those large
organizations is the cost of dealing with complexity of HPC systems. This includes cluster setup and getting applications optimized to take advantage of the latest hardware.


Many HPC vendors provide open source software stacks so that users can avoid hunting down their own components. This helps simplify HPC system setup to a certain degree. However, because many HPC software components are not designed to work together by their nature, users still face the challenge of integrating them to get them work together seamlessly. I talked to many of the attendees of IDC HPC User Forum, and they’ve all spent a tremendous amount of time and effort setting up their HPC environments. Many of them admitted this would be an impossible mission for average manufacturing firms.


With the goal of simplifying HPC cluster life cycle management, Platform’s HPC Suite, which includes HPC Enterprise Edition and HPC Workgroup Edition, provides one-stop shopping for those who don’t have the expertise to deal with the complexity of the HPC system software stack and application integrations. Offered through Platform Computing’s hardware partners, the HPC Suite (or HPC Workgroup Edition) helps average users get applications up and running with optimized performance. The key differentiator of this offering is that it provides the necessary software components to connect the OS to applications in an HPC cluster environment. By working closely with our hardware partners, we’re in the process of certifying he solution and and optimizing it for their specific hardware. With the application integrations, users can start to run parallel applications right away.

For more on the HPC Suite, check our products page here.

Wag the Dog: Private Clouds are Key to Regaining Supply Chain Control

We’ve been talking a lot about control issues here at Platform. Not in the Janet Jackson, Type-A sense of the word, but rather as it relates to private cloud computing and who’s in control of IT resources at most organizations. The most common association as it relates to cloud is the perceived delegation of control to business users. I say "perceived" because with private clouds IT will still own the service catalog to ensure compliance, budget adherence, etc. Digging a bit deeper, however, you find the Fortune 500 using private clouds in an entirely different control angle: vendor leverage.

Now, as a vendor providing a private cloud solution, you might find this a bit odd for us to talk about vendor leverage, perhaps even blasphemous. My perspective is coming not as an observer of the cloud vendor war going on in the industry right now, however, but rather from speaking with our customers and prospects that are tired of the big-vendor lock-in that they have inherited from architectures of years gone by, and see a private cloud as an opportunity to shift some things around and get the ball back in their court.

Interested? I recently wrote a piece for SandHill.com on the topic-
Wag the Dog: Private Clouds are Key to Regaining Supply Chain Control. Take a look and let me know what you think.

SC09 Wrap-up

Economic difficulties aside, the annual SuperComputing conference, now in its 22nd year, somehow defies the odds! SC09 in Portland, Oregon set an attendance record this year – roughly 10,000 attendees and 318 exhibitors hailing from 71 countries and all 50 U.S. states, the District of Columbia and Puerto Rico. In addition to physical size, the SC conference continues to grow steadily in terms of market impact each year as well. SC09 was no exception. The conference featured some of the most interesting and innovative HPC scientific and technical breakthroughs from around the world.

 

For many of us working at Platform, the SC conference always has the familiar feel of a second home. The show formed the perfect backdrop for Platform Computing to showcase its new corporate brand identity. The new look and feel also gave a tangible boost to the level interest in the company's offerings and booth traffic from show attendees remained heavy throughout the event. No SC event can be complete without a party, and this year was no exception, our 5th Annual party was held at Rock Bottom Brewery, with over 600 attending and enjoying, great beer, food, prizes and live entertainment.

 

Just a few days before SC09, we announced the general availability of Platform ISF, our dynamic IT solution for enterprises to build and run their private clouds. Try it free for 30 days http://my.platform.com/public/eval-input.action. This is the solution Forrester Principal Analyst James Staten called: "the most complete internal cloud software solution we’ve seen so far.” Despite all the cloud hype, there's little doubt the market for building and managing private or internal clouds will grow to be quite large, eventually exceeding the market for HPC. Also announced was our new cluster management solution, Platform ISF Adaptive Cluster a product that dynamically changes the operating systems and personalities of compute nodes managed by Platform LSF through our new HPC portal, http://www.platform.com/private-cloud-computing/platform-isf-adaptive-cluster. And finally our industry-leading workload management solution for HPC environments Platform LSF 7 Update 6 is now available http://www.platform.com/grids/platform-lsf.

 

SC09 conference is all about high-performance computing -- and Platform Computing is the recognized leader in HPC management solutions. We were reminded of this fact at SC09 by our friends at Intel, who recognized Platform with an Intel Cluster Ready Pioneer Award. The name of the award is both a nod the Oregon Trail and SC09's return to Portland as well as formal recognition of Platform's early commitment to removing the risk and complexity out of buying HPC clusters. Thank you, Intel. We're honored to be a part of such a forward-thinking, customer-centric program as Intel Cluster Ready.

Clusters, Grids, Clouds … Whatever!

No, I’m not trying to be heretical. The point is that I find it more useful to get beyond the buzz to the real details and lessons behind cloud, utility, grid, or whatever we call the latest computing model.

 

There are a few high performance computing (HPC) type scientific applications that are being talked about living on the public cloud infrastructure, such as the European Space Agency project Gaia which will attempt to map 1% of the galaxy, and the backend map/reduce functionality will reside on a cloud, or the fact that the US DOE is funding Argonne National Lab’s Magellan project to do general scientific computing problems as a test case for HPC in the cloud.

 

Some cloud concepts are worthwhile in a private HPC setting, especially if your setting is multiple grids or one big grid but divided into many logical grids via queues or workload management. Elasticity (the ability to grow and shrink as needs dictate), self service provisioning, tracking and billing for chargeback, and flexibility/agility to handle multiple application requirements (such as OS and patches, reliability/availability, amount of CPU or memory, data locality awareness, etc) can all improve data center utilization and responsiveness.

 

Most HPC grids or clusters already have some of these features, such as self-service and metering, but the ideas of flexibility and elasticity have not been a realizable goal until recently.

 

Today the primary driver for cloud computing is to maximize the utilization of resources. Typically that is a goal for a workload management (WLM) as well, but too often the HPC landscape is turned into silos (either grid based or queue based), which can mean that overall utilization is only in the 30-50% range much of the time. And even when it is in the 60-70% range, there is quite a bit of either backlogged demand that cannot get scheduled or systems running at low loads consuming expensive power and cooling resources, meaning that there is definite room for improvement.

 

Some of the primary causes of lower than ideal utilization within many current grid implementations include the budgeting process (typically done project by project, with each one buying new equipment for their own grid, or part of the existing grid with dedicated queues for them), service level requirements (such as immediate access to critical resources without pre-emption, which means that certain machines are left idle much of the time to allow for these jobs), and a fixed, rather than dynamically changeable, OS on each node (meaning that applications requiring a different stack cannot reuse the same equipment).

 

There are also a number of potential roadblocks to getting many HPC applications into the cloud. Let’s first investigate one of the core assumptions when the term cloud comes up, and that is one of virtualization. Indeed, there are many out there that state that this is a foundational and required stepping stone to cloud computing. I would argue, rather, that concepts that are generally embodied into the technology of virtualization such as agility and flexibility, as well as non direct ownership of resources, would be the foundation, rather than any one tool or type of tool. This is one of the key concepts to a couple of the biggest features of cloud computing, flexibility and elasticity. It is certainly possible to achieve these goals using physical systems instead of virtual machines, just not as easily. But that brings up the question of why … why don’t people just put their HPC applications into VMs and be done with it? The primary reasons are centered around performance … most of these applications were created to scale out across hundreds or even thousands of systems in order to achieve one primary goal: Get the highest possible performance at the lowest possible price. And quite often that also means very specialized networks (such as InifiniBand (IB) using RDMA instead of TCP/IP), specialized global file systems (such as Lustre), specialized memory mapping and cache locality, etc all of which get somewhat or completely disrupted in a VM environment.

 

There are several companies addressing the problem, and one example is that Platform Computing has recently announced a new capability called HPC Adaptive Clusters, which takes these concepts and applies them equally to physical machines or to VMs. The physical instances would be multi-boot capable allowing for smart workload scheduler to dynamically change the landscape as needed and by policy in order to handle various job types, whether they need a flavor of Linux, Windows or whatever (thus, our tagline, “Clusters, Grids, Clouds, Whatever”). Additionally, as technology advances such as Intel’s new Nehalem processors with tools and APIs for power capping, socket and eventually individual core control, these physical boxes can even be setup appropriately for the application load as well as saving power and cooling costs whenever possible.

 

Platform has been a leader in WLM for over a decade, and now they are adding in the ability to efficiently combine resources, with dynamic control and distribution, as well as ever smarter workload management … Thus, HPC Adaptive Cluster. Check it out at http://www.platform.com/Products/platform-isf/platform-isf-adaptive-cluster

 

Phil Morris

CTO, HPC BU