Showing posts with label GPU HPC workstation cluster. Show all posts
Showing posts with label GPU HPC workstation cluster. Show all posts

Performance and Productivity of an HPC Cluster (3)

My last blog described the results of implementing a refreshed 32-node CFD (computational fluid dynamics) cluster by using a completely different solution from the one implemented in the previous cluster. Using Platform HPC for the refresh, the new cluster significantly reduced its job failure rate and had a 15% usage increase and 25% throughput increase.


After the HPC cluster was in production with happy users, the management team began making plans to further improve application performance and user productivity. They were able to start to investigate and plan for the future. GPU adoption and a new ways to speed up the design process are two areas identified.


GPU Use


Using GPU to accelerate applications is what this organization is looking at next. Some commercial applications ISVs have already released applications that support GPU acceleration, for example, ANSYS. Adding GPU to the cluster could dramatically reduce the application run time. However, it will obviously increase the cluster complexity. Fortunately Platform HPC can help reduce such complexity. With Platform HPC, it’s much easier for administrators to deploy GPU required software. Platform Computing has repackaged NVIDIA’s CUDA software into a format that can be automatically deployed across all compute nodes in a cluster by leveraging the cluster management capability of Platform HPC. The workload scheduler in Platform HPC is also GPU aware so it can schedule GPU resources such that users running GPU jobs don’t need to worry about which node has GPUs or if the GPUs are used by other users. These capabilities would go a long way to helping this organization overcome the hurdles of adopting GPU technology.


Design Process Acceleration


After users transferred from using the command line interface to the web interface, they still had some complex scripts that they needed to automate job flows. These scripts were originally built by a few power users. New users typically just copied them and wrapped the scripts with a few lines of code as required. But when there was a problem, it was very difficult to debug. One technology they are evaluating to help alleviate this problem is Platform Process Manager. Platform Process Manager allows users to program a flow in a graphical way without writing scripts. It can be integrated with Platform HPC’s web interface. Once a flow is developed, users can click a button to launch the flow. Job dependencies are taken care of by the Platform Process Manager. This company is exploring the possibility of using Platform Process Manager as a technology to automate their job flow to further increase user productivity. The graphical view of the flow also makes the job flows much easier to maintain and improve, which can also be used to document best practices in their engineering process.


The partnership with Platform Computing allows this design firm to advance digital design efficiency with increased user productivity. By doing more simulations to improve the quality of the design, high performance computing gives them the competitive advantage to be leading edge.

HPC shifting sands: Distributed or centralized infrastructure

As NVIDIA takes the HPC world by storm by lending the power of their GPUs (graphic processing units) to HPC applications, ISVs, system architects, and users are all trying to find their footing when it comes to planning how they’ll be using HPC in the next year or two. To make a historical comparison, the core counts, memory footprints, and I/O bandwidth achievable in a single workstation today would have enough to make HPC aficionados salivate only two years ago.


Indeed, “deskside” clusters were the prediction in 2007, expected to bring about the demise of the centralized cluster—with more computing power than any one user could ever need. Well, the HPC community are a greedy bunch. And in some ways the prediction has become half true. “Deskside clusters” have actually developed into “desktop SMPs”.


It seems that the line between when a user needs a cluster and when he can live with his personal workstation will be shifting. I believe that’s due to the enormous computing power offered by GPU accelerators. Yes, most applications aren’t ready yet, and, yes, the Tesla hardware is expensive. But those two barriers are true only today and are likely to quickly disappear.


Time will bring new and more mainstream applications into GPU enablement and force their now square peg into NVIDIA’s round hole. The elite prices of GPGPUs are nothing more than market perception. I just purchased an NVIDIA card for less than $200 for a media center PC that has 96 cores in it.


This trend adds up to a temporary shift in the centralized vs. distributed computing tug-of-war. Every one of the workstations on a corporate refresh cycle will have graphics capability. Servers, on the other hand, are often on a slower refresh cycle and are scrutinized more carefully by corporate bean counters who won’t buy GPUs servers until CIO/CTOs tell them to.

Desktop compute harvesting, once a fringe HPC endeavor, may suddenly have more ROI than it ever has as a technique for providing HPC scale for minimal incremental cost. The question is: Will the window of opportunity for such techniques close before a complete solution can be developed and delivered?

Of course, the inclusion of GPUs in servers will happen. When that occurs, the balance will shift back. Until then, though, the horsepower of cloud or even local clusters has become a little less attractive when there’s a processing stallion underneath the desk waiting to be bridled.