Showing posts with label Computing. Show all posts
Showing posts with label Computing. Show all posts

Wednesday, May 24, 2017

Thoughts on HyperConverged, and the Future of HyperConverged (Part 2)

So how did we get here? Where did HCI come from?

If we look back at the history of HCI, it seems to have evolved from the idea of using clustered, "whitebox" x86 servers to create a clustered storage system. There were a number of early entrants in the space, some dating back to 2006. Another vector was the idea of a "Virtual Storage Appliance" or VSA, software which ran in a VM, connected to local server hard disk drives, and presented that internal storage to the guest VMs over the internal IP network. The first VSA was from Lefthand in 2007. But the real hyper-converged push started around 2009 with the founding of integrated HCI players Nutanix and SimpliVity.

We also have to look at where the HCI market is today. It is arguably dominated by three primary players: Nutanix; SimpliVity (now part of HPE); and VMware VSAN. They represent the lion's share of the HCI market, and we will come back to them.

If you look at the earlier clustered storage companies, they either offered a scale-out NAS, a kind of commodity alternative to Isilon, a scale-out block storage solution, or a scale-out unified storage solution. These early players came into existence when "grid computing" was the buzzterm of the day, and these architectures were also called "grid storage".

In 2009 Nutanix was founded. There were other virtual storage appliance start-ups, such as Virsto Software (which eventually became VMware VSAN), but it is fair to define the official beginning of the hyper-converged era as August 2011, when Nutanix emerged from stealth. The same month, VMware released vSphere 5 which included its first implementation of a VSA (vSphere Storage Appliance). SimpliVity would emerge from stealth one year later in August 2012. VMware's VSA did not gain traction, and VMware announced its intent to acquire Virsto six months later in February 2013 which represented VMware's serious interest in HCI.

As Nutanix and SimpliVity started to grow, and with VMware's very public acquisition of Virsto, and obvious plans to enter the HCI market, many of the earlier clustered storage vendors and virtual storage appliance vendors redefined themselves as hyper-converged players. Several new industry buzzterms were developed: "Server SAN"; "Virtual SAN"; and "Software Defined Storage", or "SDS".

Many of the early clustered storage system vendors redefined themselves as SDS or HCI players, moving their clustered storage software from bare-metal to run in VMs, and allowing their clustered storage software to run alongside guest VMs on the same server. VSA vendors added more sophisticated clustering, replication, and scalability to their products.

From this, it is fair to say modern HCI owes itself to three parents: Commodity clustered storage systems; virtual storage appliances; and purpose built integrated HCI systems.

To me, the most interesting thing is many of the earlier clustered storage or "grid storage" players had little to no success, but the HCI players saw significant early success. Part of this may have been how each targeted the market. Clustered/grid storage historically had been seen as targeting the high-performance and academic community for technical computing use cases. HCI targeted business organizations and VMware virtualization workloads.

But what cannot be dismissed is the reality the early clustered storage ystems did not provide the level of performance and reliability required for enterprise workloads. The early clustered storage systems were not designed for transactional, random I/O workloads. They were better suited for sequential I/O. The early HCI players focused on addressing write latency and random I/O with aggressive write and read caching. The also focused on ease of use and eliminating the need for storage administrators to provision storage to VMware administrators.

At this point it is interesting to note, there were other players aggressively targeting VMware virtualized workloads. Tintri had come out of stealth five months before Nutanix with its VMware optimized storage platform. It too targeted the VMware admin and sought to used its product to bypass the traditional storage management team in an organization.


So that is the history lesson and the end of Part 2.

Sunday, May 21, 2017

Thoughts on HyperConverged, and the Future of HyperConverged (Part 1)

Almost two years ago I made some observations on HyperConverged Infrastructure, and where I think it needed to go to be successful. I posted these to Twitter at the time. I still stand by some of those observations, for others I am not as sure. But I have done a lot more thinking about the HCI phenomena, and believe change is coming to HCI.

To this point, I recently saw an update of the Gartner Hype Cycle, which showed HCI at the zenith of the "Peak of Inflated Expectations". I agree with this. The question is what comes next? Probably a vendor shake-out.

 
But another question to ask is "What comes after HCI?" The idea HCI is the end-game for IT infrastructure is a naive assumption. There may be better architectures being worked on by start-ups as I write this.

These were my original observations on HCI:

HCI must support multiple hypervisors, and no hypervisor (i.e., Containers, Hadoop, Oracle RAC, etc.).

At the time, Microsoft was pushing Hyper-V very hard, and I thought Hyper-V was going to make significant penetration into the enterprise. At the same time, some organizations were experimenting with OpenStack and KVM. Today, looking back, VMware still dominates. Hyper-V exists mainly in on-prem Azure Stack deployments, and KVM struggles without a single brand behind it.

As for no-hypervisor HCI (my idea being a combination of OpenStack with Containers and an HCI filesytem embedded in Linux for something like Oracle RAC), this has yet to take off. There is a chance we could see something like it for OpenStack.

HCI must become all-flash for virtualized workloads.

For the most part, this has become true. And the reality is, All-Flash saved HCI, which probably would not have been able to keep up with the performance requirements of virtualized workloads in its hybrid form.

HCI filesystems must be or become flash aware (WAF, etc.).

HCI filesystems have been adapted for flash, but I do not believe they have reached a point to make them comparable to All Flash Arrays in reducing flash wear. They have been able to avoid this by using high Drive Write Per Day (DWPD) SSDs in their caching tier to coalesce writes to low DWPD SSDs in their capacity tier. I see two problems with this approach. The first is the use of a high DWPD SSD as a cache is a carry-over from the hybrid HCI filesystem architecture. There it provided a significant performance boost. When combined with an SSD capacity tier, it provides no performance boost, and only a write wear mitigation benefit. The second issue is high DWPD SSDs are not a high volume part for SSD manufacturers, who would rather manufacture lower DWPD, higher capacity, higher revenue SSDs. Ultimately, high DWPD SSDs may fade away like SLC and eMLC SSDs did. If that happens, what will HCI vendors do?

HCI must move to parity/erasure coding data protection and move away from mirroring/replication based data protection (RF2/RF3).

I believed this was necessary for All-Flash HCI due to the cost of flash, and the capacity of SSDs at the time. I am less sure of this now, at least as a $/GB requirement. I think parity/erasure coding will only be driven by availability requirements, and not $/GB requirements.

HCI must support storage only nodes and compute only nodes for asymmetric scaling.

I believe this even more today. With All-Flash HCI, storage efficiencies (a.k.a., Data Reduction technologies) became critical. When you look at the Virtual Desktop (VDI) use case for HCI, deduplication means storage capacity does not grow linearly with VDI instances. In fact, it hardly grows at all. But what does grow is a need for write caching. If I invested in HCI for VDI, and deployed 200 VDI instances across 4 HCI nodes, and later decided to grow my VDI to 400 instances, I might need 4 more nodes of compute, but deduplication might mean I need only 10% more storage capacity, which I might already have on my existing nodes. I might need a caching SSD on each new node, but not 5 to 11 data drives.

The reverse holds true as well. If I assume a certain storage efficiency ratio, but due to adding workloads with different data types (say pre-compressed image files) my storage efficiency drops, today I have to add compute and hypervisor instances (and associated licenses) just to gain access to more storage capacity. If I could add a storage only node or two, it would provide flexibility. Also, it might offer the ability to introduce tiering between an all-SSD production tier, and a NL-SAS capacity tier.


This is the end of Part 1. Over the next several parts, I will dig much deeper into these thoughts, including thoughts on what comes after HCI.

Thursday, March 16, 2017

Everything I need to know about NetApp’s All-Flash Storage Portfolio I learned from watching College Football

Okay, silly title. I got the idea when Andy Grimes referred to NetApp’s all-flash storage portfolio as a “Triple Option”. To me, when I hear triple option, I think of the famous Wishbone triple option offense popular in college football in the 1970s and 1980s. And that got me to thinking of how NetApp’s flash portfolio had similarities to the old Wishbone offense.

The Wishbone triple-option is basically three running plays in one. The first option is the fullback dive play. This is an up the middle run with no lead blocker. It is up to the fullback to use his strength and power to make yardage. The second option is the quarterback running the ball. While most quarterbacks are not great runners, the real threat of the quarterback in running offenses is the play action pass, where a running play is faked, but the quarterback instead passes the ball. In today’s college football, while the Wishbone may have faded, option football remains popular, and many of the most exciting players are “dual-threat” quarterbacks who can both run well and pass well. But, back to the Wishbone. The third option is the halfback, an agile, quick running back who often depends more on his ability to cut, make moves, and change direction to make the play successful.

In considering this analogy, I wanted to find the right pictures or videos of Wishbone football to make the comparisons to NetApp’s flash portfolio, but found the older pictures and videos from the 1980s to not be that great. So I decided to take the three basic concepts: The powerful fullback, the dual-threat quarterback, and the agile halfback and look at more recent examples. I just happen to use examples from my alma mater, Auburn University, because I knew of a few plays that visually represent the comparisons I am about to make.

So first up is the fullback. The fullback is all about power. It is not about finesse. The fullback position is not glamorous. The fullback had to have the strength to face the defense head-on. To me, the obvious comparison in the NetApp flash portfolio is the EF-Series. The EF is all about performance: Low latency, high bandwidth, without extra bells and whistles which can slow other platforms down.

While I don’t have a good fullback example, I have a similar powerful running back demonstrating the comparison I am trying to make. Here we see Rudi Johnson on a power play break eight tackles and dragging defenders 70 yards to a touchdown from the 2000 Auburn-Wyoming game.

Rudi Johnson great 70 yard TD against Wyoming 2000



The next comparison is to the dual-threat quarterback. The dual-threat quarterback can run or pass with equal effectiveness. In NetApp’s flash portfolio, the obvious comparison is the All-Flash FAS (AFF), the only multi-protocol (SAN and NAS) all-flash storage array from a leading vendor. The multi-protocol capability of AFF (Fibre Channel, iSCSI, and FCoE SAN; NFS and SMB NAS) allows storage consolidation, and truly brings the all-flash data center to reality.

The play which best demonstrates the dual-threat quarterback’s potential is the run-pass option (RPO), where a quarterback rolls out and can either keep the ball and run with it, or pass it to a receiver if the receiver is open. Here we see Nick Marshall on an RPO play which tied the 2013 Iron Bowl with 33 seconds left in the game. The reason the play worked is Nick Marshall, a gifted runner, had already run for 99 yards including a touchdown.

2013 Iron Bowl: Marshall to Coates



That brings us to the halfback, also known as the tailback, or just the running back. For the sake of this discussion, and keeping with the original Wishbone concept, I will use the term halfback. The handful of teams who still run a variation of the Wishbone (Georgia Tech, Navy, Army, Air Force, and a few others), tend to use smaller, more agile athletes as halfbacks. These running backs usually get the ball on the outside, and leverage their agility to make the defenders miss. When I think of agility in flash storage, I think of SolidFire. Agility is a key feature of SolidFire. It scales with agility, provisions with agility, adapts with agility, and is the best storage for agile infrastructures like private clouds, especially private clouds using OpenStack. The best recent example I have seen of a running back leveraging agility to make a play is this run by Kerryon Johnson against Arkansas State.

Watch Kerryon Johnson's incredible touchdown against Arkansas State





So enough fun for now. But if you have a dedicated application needing performance acceleration, such as a performance critical database, NetApp’s EF-Series might be your tackle-breaking fullback powering through spaghetti code and getting the job completed despite the challenge. If you are looking to move to an all-flash data center and need consolidated flash storage to accelerate iSCSI MS-SQL databases and NFS VMware datastores on the same infrastructure, AFF is your dual-threat quarterback. And if you are looking to deploy a private cloud with the agility to grow with your workload, SolidFire is your agile halfback.

Sunday, January 17, 2016

"True" Private Clouds

Wikibon is talking about "True" Private Clouds. I think their definition is too narrow, and gets into the weeds. It misses the true customer of a "true" private cloud. And there are two customers. The first is the organizational customer that purchases a private cloud. The second is the internal end-consumer of cloud services.

To Wikibon's credit, the definition of "Private Cloud" is an issue that needs to be addressed. In my career I have seen too many organizations overuse the term "Private Cloud". I have seen a VMware cluster deployed on disparate hardware with no upper level cloud management platform called a private cloud. I have seen converged infrastructure, acquired but managed identically to non-converged infrastructure (as discrete components each managed by their functional staff) called private clouds.

Converged infrastructure plays a role in a private cloud, be even that term is challenged. I have seen disparate servers and storage, purchased separately at different times, cobbled together and called converged infrastructure after the fact. I have also seen single-SKU converged infrastructure broken apart, support for component infrastructure separated, and individual components upgraded on different life-cycles.

From an operations perspective, I have seen mature IT organizations in large enterprises provide similar levels of managed services as traditional managed service providers. I have also seen the converged infrastructure single-support model dramatically fail organizational customers, and provide no better single support that that provided by an reseller or managed service provider.

If the goal of a "true" private cloud is to provide a similar level of service offering to internal end-consumers they receive from a public cloud, but with higher levels of compliance and data sovereignty, then much of the detailed requirements Wikibon mentions are not necessary. As long as the organization can provide an offering to internal end-consumers which is competitive (on cost,  ease of consumption, and reliability), it should meet the definition.

Here are what I believe are required of a "True" Private Cloud:
  • Acquired in consolidated units of management, virtualization, compute, network, and storage with common amortization, and common life-cycle management.
  • Components supported as an integrated whole, with a single number, first-call support model, and escalated support abstracted from the internal end-consumer.
  • Compute, storage, network, and virtualization managed as a single entity by a single, cross-functional team.
  • Provisioned and managed via a cloud management platform (CMP).
  • Consumed by internal end-consumer as a shared resource in logical, not physical increments, i.e., VMs and GBs.
  • End-consumer offerings include multiple performance and data protection SLAs.
  • Provides charge-back to internal end-consumers.
  • Provides the Private Cloud operator performance, capacity, and licensing budgeting of the infrastructure; performance metering and capacity measurement to manage over-subcription, prevent over-consumption (especially of performance), and allow for elastic performance and capacity scaling; and provide built-in performance and capacity planning for predictable infrastructure growth.
  • Managed by high IT maturity organizational customer IT staff, or optionally part of a managed services offering  that does not require organizational customer IT staff to manage.
  • Financed to organizational customer either through capital purchase, capital lease, operational lease, capacity lease, or pay-per-use offering.

Some organizational customers will want to capitalize the "True" Private Cloud and manage it themselves. Others will want to basically rent the whole stack to include the software, and have it managed for them. But the common denominator should be how the internal end-consumer consumes the offering. It should look, feel, and cost as much like the public cloud as possible.

Wednesday, December 23, 2009

x86 Rises, Part 4: The emergence of Linux as a viable datacenter OS

Several years ago I drafted a white paper I called "x86 Everywhere". I started it in the fall of 2004, let it sit, and updated it in April 2005. It remains unfinished, but with the release today of Intel's Nehalem processor, I took a look at it again. Here it is:

Three trends could allow what I call "x86 Everywhere" to happen.

The third trend necessary for "x86 Everywhere" is the possibility of the emergence of Linux as a viable datacenter OS.

This seems less likely than high-end x86 servers at this point, but it is certainly possible in several years time, if the efforts of the Datacenter Linux project bear fruit. Windows on 32-bit x86 systems did not penetrate the datacenter, in part because the hardware was not 64-bit, the hardware was not scalable, and customers did not trust Windows with their critical data.

Today, the hardware is 64-bit, AMD Opteron is scalable to eight-sockets today, Intel is pursuing efforts that will likely address the scalability limitations of Xeon, both AMD and Intel are aggressively pursuing multicore chip strategies, and customers trust Linux in places they formerly only trusted UNIX. The result is a very real, industry standard ABI/ISA platform combination that scales from embedded systems, to an inexpensive developer platform (the PC), to midrange enterprise datacenter computers. This could be enough to cause a tipping point, creating a fundamental driver for the Datacenter Linux initiative. Such a change in the primary enterprise compute platform from RISC/UNIX to x86/Linux would likely be highly disruptive to the industry, and would rival the move of commercial computing in the early 1990's from proprietary minicomputers to SMP RISC/UNIX servers. Once established in the datacenter as a viable midrange enterprise platform, like SPARC/Solaris it becomes a straightforward scaling exercise for x86/Linux to establish itself as a high-end platform.

Finally, while not a trend driving large scale x86 adoption, there are other developments to consider. Intel has a virtualization technology, called Vanderpool on desktops and Silvervale on servers, that will help provide partitioning on its systems. AMD has also stated it intends to offer a virtualization layer, called Pacifica. AMD has also stated it plans to improve RAS features of its Opteron, and it is likely Intel will do the same with Xeon, using features it already offers on Itanium. Both of these key technology areas will improve adoption of x86 servers in the enterprise market.

How will this play out?

First, Dell's strategy is to only enter established markets, and to do so with a superior fulfillment system. For markets that are not at that point, Dell has used partnerships, such as its existing partnership with EMC. Dell also partners with Unisys to resell Unisys' 8-way Intel Xeon systems. Therefore the most likely path for Dell is to primarily continue the status quo, assuming four socket x86 systems and below represent the lion's share of the server market. If there is a need to address the greater than eight-socket x86 server market, Dell could expand the Unisys agreement beyond 8-way. If Dell expands into the Opteron market, and needs to address the greater than eight-socket x86 server market, it could partner with Newisys (also an Austin TX company).

IBM already is a player with its Enterprise X Architecture (EXA) for Intel systems. However, IBM has close ties to Newisys (the founder is ex-IBM, and the Horus chipset is based on similar principals to EXA), IBM sold its North Carolina based PC Server manufacturing plants to SCI-Samna, IBM has a strong presence in Austin TX, Newisys' home, and IBM has strategic agreements with AMD around CPU fabrication technology. It is possible IBM could offer the Newisys system in addition to its own EXA systems.

HP is committed to x86 in the four-socket and below space, and is a strong backer of Linux. If the x86/Linux platform gains momentum, it would simultaneously weaken Itanium sales. This would require a strategy change for HP, but such a change would be necessary to remain a viable datacenter systems vendor. To address this, HP could OEM a solution if needed to address a short term requirement. HP did this with NEC's high-end Itanium system before HP adapted its Superdome system to accept Itanium processors. Here the most likely partner would be Newisys, with similar Texas roots to the Compaq, whose former Texas offices server as headquarters for HP's x86 division in the post-merger HP. Longer term, HP's relationship with Intel could produce a high-end x86 system, especially given the common chipset Intel promises for Itanium and Xeon. In fact, HP's “Arches” system, the follow-on to Superdome, could easily accept future Xeon processors, given the common Itanium chipset. HP could also acquire a solution, but the most likely acquisition in this case would be Unisys. A Unisys acquisition would be defensive as well if Unisys had or was considering a significant Dell agreement.

Sun has some of the closest ties to AMD, and Sun has the technology to build large systems. Sun already plans eight-socket Opteron systems. If a significant market for larger than eight socket x86 servers emerges, Sun will have to decide how to address that market. However, balancing the high-end SPARC and x86 business would be a challenge for Sun. If the scalable x86 market shows great promise, the best technical solution for Sun could be an even tighter AMD partnership with technology sharing to allow common systems to be built with either AMD or SPARC processors. The potential for Sun to leverage common technologies such as coherent Hypertransport for SPARC systems as well as Opteron could offer considerable economies of scale. This could make the most sense in the post APL timeframe. A secondary solution, which also offers a near term solution, would be an OEM deal with Newisys. Sun has relationships with SCI-Samna, OEMing Newisys' two socket and four socket Opteron servers as the V20z and V40z, and Sun contracts with SCI-Samna to manufacture low-end UltraSPARC servers. A deal with Newisys around higher-end systems would also server to more strongly establish Sun in the Texas information technology community, clearly one of the top IT centers in the world, and the most important in the x86 business.

AMD's best interests are served if it does not depend on other vendor's chipsets for scalability. Therefore, offering a higher-end Opteron processor with more coherent Hypertransport links allowing greater glueless SMP scalability is the most likely path for AMD.

Similarly, Intel's best interests are served if it can offer everything needed to build a scalable server directly to the distributor. This is the shift needed to move high-end servers into the commodity space, and allow Dell to enter the market with superior logistics.

Based on all of this, a two-phased industry approach is likely. The first being server-vendor based proprietary scalable solutions (such as IBM's EXA, Unisys' CMP, and Newisys' Horus), followed by processor vendor solutions based on in-chip features.

Who is threatened most by x86 Everywhere? One could say Sun, who relies on SPARC systems for the vast majority of its revenues. However if x86 Everywhere happens, SPARC's installed base is still very large, and will not be replaced overnight. The bigger victim is likely IBM, who is trying to repeat Sun's SPARC success with its POWER architecture. In fact, assuming a Sun/AMD partnership could allow Sun to build SPARC or Opteron systems from common technology (i.e., memory controllers and memory subsystems, coherent Hypertransport MP interconnects, and common Hypertransport I/O bridges), SPARC systems could be continued as long as customer demand supported the design of SPARC processors.

The big loser in this appears to be Newisys. SCI-Samna's business model is two-fold: Contract manufacturing and OEM manufacturing. Newisys' low-end systems fit well in the OEM model, and SCI-Samna has had success selling these systems to its OEM partners. However, the high-end Horus systems do not fit the OEM model. Several have tried OEMing datacenter servers, and few have succeeded. In the late 1990s, Unisys OEMed its x86 CMP system to both Dell and Compaq. The Dell OEM lasted only months. Dell realized a 32-way datacenter server did not fit its direct business model. Compaq's deal lasted a little longer, but it too abandoned the OEM arrangement. Other OEM deals include HP's OEMing of NEC's first generation Itanium system, which delivered few sales. The most successful OEM deal of datacenter servers appears to be Bull Worldwide's OEMing of IBM's pSeries servers, but this arrangement created significant channel conflict for IBM in europe, and seems to always be in danger whenever IBM announced a new generation of RISC/UNIX servers. Fujitsu's deal with Siemens is not considered as an OEM deal here because it is really more of a partnership. The Fujitsu-Siemens model is worth considering by Newisys, as it is a successful model of a business relationship between a high-end server manufacturer and a IT solutions provider. The most likely target customers for Newisys' Horus system are IT integrators such as EDS. IBM has a high-end x86 server in its product portfolio. EDS does not. IT integrators can provide the professional services required in selling such systems. Also, because this would be an OEM arrangement, there is the opportunity for greater margins and services to the IT integrator, compared to deals which involve simply reselling an server vendor's product.

x86 Rises, Part 3: x86 Grows in Performance and Scalability

x86 Rises, Part 2: Decreasing Value of Big UNIX

x86 Rises, Part 1: The Background

Friday, October 02, 2009

x86 Rises, Part 3: x86 Grows in Performance and Scalability

Several years ago I drafted a white paper I called "x86 Everywhere". I started it in the fall of 2004, let it sit, and updated it in April 2005. It remains unfinished, but with the release today of Intel's Nehalem processor, I took a look at it again. Here it is:

Three trends could allow what I call "x86 Everywhere" to happen.

The second trend is the prospect of several vendors offering scalable 64-bit x86 systems large enough to meet most customer's workloads.

The desktop megahertz wars of the late 1990s and early 2000s between Intel and AMD drove x86 performance at a rate exceeding Moore's law. This directly benefited Intel x86 server performance, making x86 servers available for larger workloads. At the same time, enterprise applications were being rearchitected to multi-tier web-based applications, requiring deployment of additional web and application servers. RISC still had advantages over x86 in this environment, as running Microsoft Windows on x86 servers required the purchase of client access licenses (CALs) for each discreet user. This was extremely expensive for emerging self-service web-based ERP and CRM applications, but it was impossible for B2C ecommerce applications. Enter Linux. In the late 1990s, Linux became established as an entry server operating system, which unlike Microsoft Windows, did not require the purchase of client access licenses (CALs) for each user. Linux quickly became established as the web server OS of choice. The result was a positive feeback loop. Application server ISVs aggressively ported their J2EE appservers to Linux, and improved their clustering so their appservers would work well on clusters of low-cost entry x86 servers. ERP vendors quickly followed porting their application tier to Linux on x86. The low purchase cost of the Linux/x86 architecture was driven home by the dot-com bust and worldwide recession of the early 2000s.

At the same time as the desktop megahertz war, the smaller x86 chip manufacturers each tried to establish their products into a niche area. Via acquired Cyrix and focused in the “system on a chip” market for very low-cost desktops. Transmeta focused on very low power consumption chips for low-end laptops and embedded markets. AMD, long a player in the budget desktop market, decided to focus on the server market by designing an x86 architecture, called “Hammer” which addressed the weaknesses of Intel's existing Xeon x86 server processor, primarily the latter's lack of 64-bit memory addressing. The release of Hammer, branded as Opteron, forced Intel to follow suit with its 64-bit x86 technology, long rumored under the codename “Yamhill”, and branded as EM64T technology.

The emergence of a truly competitive x86 server processor marketplace is driving new innovation in x86 processors, as AMD tries to stay one step ahead of Intel, and as Intel tries to leapfrog AMD. Dual-core processors, improved power management, virtualization technologies, and other improvements are announced on a regular basis.

After the emergence of 64-bit x86 technology in 2004, in 2005 dual-core x86 processors were released. These two technologies have strong synergies. 64-bit addressing increases the size of the workload which can run on an x86 server, and dual-core processors increases the size of server which can be built with x86 processors.

With dual-core 64-bit x86 processors now shipping, and four-core 64-bit x86 processors possible in two to three years, four to eight socket servers may provide the capacity required for most customers' workloads. Beyond that, workloads requiring large, single system image servers (HPTC, large data warehouses, etc.), may be relegated to a niche market. Ordinarily, such a niche market could still justify large, scalable RISC/UNIX systems. But the market for large, single system image servers is not limited to RISC/UNIX. For some time, the scalable x86 market has been a targeted by some system vendors.

In the mid-1990s, Sequent, with its NUMA-Q system, was one of the first vendors of large, scalable x86 systems. Data General offered a very similar NUMA system during the same time period. Both of these systems provided very limited performance because of their architecture. Data General's system failed to gain significant market share, and was end of lifed not long after EMC acquired Data General. Sequent targeted decision support and data warehouse workloads with its NUMA-Q system and had some success. Sequent was acquired by IBM, and IBM released a more advanced x86 NUMA system which offered greater node to node bandwidth and large L4 caches to better manage inter-node latencies. In 2005 IBM released its third generation of x86 NUMA systems.

In the late 1990s, Unisys built a large, scalable SMP x86 system using a technology it calls cellular multiprocessing, or CMP. This technology was derived from Unisys' Clearpath mainframe systems. In fact, Unisys offers a version of its x86 CMP system which runs the Clearpath mainframe OS ported to the x86 architecture. Despite the mainframe heritage and mainframe variant of Unisys' x86 CMP systems, sales have not been strong. These systems were limited by the lack of scalability of Intel's x86 architecture, as well as the x86's lack of 64-bit memory addressing. Unisys now offers a second-generation CMP design, with simpler eight socket entry systems as well as large 32 socket systems.

Both IBM and Unisys offer 32-socket Intel Xeon systems, but both of these systems continue to be limited by the inherent lack of scalability in Intel's Xeon architecture.

The limits of x86 scalability changed with AMD's Opteron. Opteron is the first scalable x86 processor architecture. By virtue of its high-performance, coherent Hypertransport MP interconnect, Opteron is scalable in SMP design. Because of its 64-bit memory addressing, Opteron is scalable in memory capacity, with memory addressing balanced with processor performance. Four to eight socket x86 servers are no longer crippled with saturated SMP busses or inadequate memory capacity. Intel has followed suit with 64-bit memory addressing for Xeon, and a unique dual front side bus (FSB). But the dual FSB, while providing temporary relief to Xeon's saturated SMP bus, is actually designed for the soon to be released dual-core Xeon processors. Dual-core Xeons will likely once again saturate the SMP busses. Better SMP interconnects will be required for efficient scaling of Xeon systems to four sockets and above.

Over the next several years, x86 systems with eight-sockets and greater will become more prevalent. Newisys, a division of SCI-Samna, a major OEM manufacturer of AMD Opteron systems, is planning a 32-way Opteron chipset called Horus. Intel has promised future Itanium and Xeon processors will support a common chipset, allowing a next generation scalable Itanium server architecture to also serve as a scalable Xeon platform. This means traditional large scalable Itanium system vendors, HP, SGI, and NEC could enter the large scalable Xeon system market. The other possibilities are a higher-end AMD Opteron chip with more Hypertransport links allowing more scalable glueless MP topologies, similar to Compaq Alpha EV7's architecture, or the possibility of Intel introducing a scalable glueless chip to chip interconnect. It is important to note, Intel has access to the design of the EV7 interconnect and now employees the developers of the EV7's interconnect through an agreement with Compaq before Compaq was acquired by HP. Regardless, increased SMP scalability of x86 servers seems likely in the next few years.

Related Posts:

x86 Rises, Part 2: Decreasing Value of Big UNIX

x86 Rises, Part 1: The Background

Tuesday, June 16, 2009

x86 Rises, Part 2: Decreasing Value of Big UNIX

Several years ago I drafted a white paper I called "x86 Everywhere". I started it in the fall of 2004, let it sit, and updated it in April 2005. It remains unfinished, but with the release today of Intel's Nehalem processor, I took a look at it again. Here is Part 2:

Three trends could allow what I call "x86 Everywhere" to happen.

The first trend is the decrease in value of large, partitionable, RISC/UNIX systems.

All major commercial RISC/UNIX systems vendors offer large systems that can support large workloads, or can be partitioned to support many medium-sized workloads. The primary reasons for deploying a medium-sized workload in a partition on a large server are expected growth beyond the capacity of typical midrange servers, higher system resource utilization, system management efficiencies of server consolidation, and customer politics and preferences. Each of these reasons is coming under assault by the advancement of Moore's law, and as a result, the value proposition of large, partitionable datacenter servers is declining.

The performance improvements brought about by Moore's law over the last several years have outpaced customer workload growth, allowing midrange systems to handle the expected growth of most customer workloads. In addition, the price of midrange RISC/UNIX system has declined significantly over the last several years, starting with Sun's UltraSPARC III based V880, whose price point was then met by IBM with the POWER4-based p650, and HP's strategy of offering standard configurations of PA-RISC and Itanium midrange systems at very aggressive prices. Moore's law has caused system utilization to drop, as processors are now very powerful.

Traditional physical based partitioning, such as Sun's Dynamic System Domains and HP's Node Partitions (nPars) do not provide adequate granularity given the performance of today's processors. The result is the rise of software-based partitioning, logical partitioning, and virtual machine technology, which are portable to smaller, less expensive RISC/UNIX systems. In the case of purely software based partitioning technology, it is portable to other ISAs such as x86 platforms. For example, virtual machine technology is primarily being used on x86 systems via VMware's products. These shift in server partitioning technology are also decreasing the value proposition of large and midrange RISC/UNIX servers.

The recent emphasis in the industry for provisioning and system management solutions, along with policy-based computing solutions to manage large numbers of discreet servers has yet to significantly change the industry, however improvements in this area have improved the system management efficiencies of distributed servers. This, along with some of the inherent provisioning and management efficiencies of software-based partitioning technologies (shared network and disk resources) have resulted in a decrease in the relative value of large partitionable systems.

One should note, this decrease in value is real. It is not simply a customer perception. First, physical partitioning is simply too expensive a method to achieve partitioning in a server. Markets define prices, not vendors. Costs define margins, not prices. In a scenario with two otherwise equivalent servers, one using physical partitioning, the other using logical partitioning, the logical partitionable server will offer the vendor greater margins. Similarly, designing a server with physical partitioning which offers the same granularity as logical partitioning would likely be abandoned for having too high a cost. Second, customers really are moving workloads from previous generation large servers to smaller servers of the current generation, rather than partitions on larger current generation servers. In 1998 a Sun customer might consider paying the 50% price premium of an E10K over multiple E4500s. The value the 50% premium represented, primarily in growth capacity, justified the premium. Today the premium an E20K has over multiple V890s or V490s is so much higher (around 150% more), few customers can justify the value the E20K price premium provides.

The effect of this is a leveling of playing field between RISC/UNIX servers and x86 servers. Midrange RISC/UNIX servers are becoming simpler and cheaper. Midrange x86 servers have become more robust. RISC ISAs versus the x86 ISA is become a "Coke versus Pepsi" decision: a flavor choice.

Related Post:

x86 Rises, Part 1: The Background

Tuesday, March 31, 2009

x86 Rises, Part 1, The Background

Several years ago I drafted a white paper I called "x86 Everywhere". I started it in the fall of 2004, let it sit, and updated it in April 2005. It remains unfinished, but with the release today of Intel's Nehalem processor, I took a look at it again. Here it is:

What is “x86 Everywhere”? x86 Everywhere is a concept that the dominance the x86 instruction set architecture (ISA) currently has on the desktop and entry server markets will expand into the midrange and high-end datacenter server markets, eventually reaching a tipping point, and displacing most RISC/UNIX platforms. Over time the x86 ISA establishes a monopoly in the datacenter similar to its current monopoly on the desktop.

The drivers for such a scenario are purely economic, but this does not refer to server acquisition costs. Instead it refers to the economic advantages a single, dominant ISA would bring to system vendors and independent software vendors. This is not the first time such a scenario has been speculated. In the early and mid 1990s, when Microsoft announced Window NT as a portable, multiplatform operating system for both RISC and x86, many speculated Windows would become the dominant operating system and programming application binary interface (ABI) from the desktop to large datacenter servers. A few years later, many speculated Intel's IA-64 “Merced” (later branded Itanium) ISA would dominate all computers, displacing RISC from the datacenter. Desktop PCs, entry and midrange servers running Microsoft Windows and Novell Netware, and high-end datacenter servers running UNIX would all use the IA-64 architecture. Despite this speculation, few put two and two together and speculated a Windows/IA-64 monopoly platform combination. The latest domination scenario proposed a few years ago was Linux would displace all UNIX variants. In this scenario, system vendors with their own UNIX variants would simply abandon their UNIX distributions and instead port Linux to their RISC architectures. This scenario is amazingly similar to the speculation about Microsoft Windows in the mid 1990s. Then experts suggested RISC vendors would abandon their UNIX variants to instead embrace Windows.

There is a huge difference with x86 Everywhere. The difference is the current installed base of x86 systems, and the current willingness of customers to use x86 systems for critical tasks. This is not to say other ISAs will exist. While RISC/UNIX established dominance in the datacenter in the 1990s, mainframes still exist, and while x86 is dominant on the desktop, the Apple Macintosh continues to be successful as an alternative platform. However, in this scenario, traditional RISC/UNIX systems are rendered to a smaller, niche market.

Three trends could allow what I call "x86 Everywhere" to happen.

I will cover those three trends in my next post.

Wednesday, August 08, 2007

I Told You So

Eight months ago I postulated a crazy idea. It is one of those things I predicted would happen within the next three to five years. The idea was to embed a Xen hypervisor into the BIOS of an x86 server.

Today I read Dell (yes Dell) is planning to embed a hypervisor into an x86 computer. The hypervisor is likely to be VMware ESX based. This makes sense, because Dell already has a relationship with EMC, the owner of VMware. It also makes sense because Dell knows it must differentiate its products rather than simply being a low-cost provider.

More at:

The Register: Dell to stuff hypervisors in flash memory

The Inquirer: Dell plans to embrace virtualisation

ZDNet Blogs: Speculation about embedded hypervisors

SearchServerVirtualization.com: VMware prepping embedded 'ESX Lite' hypervisor

Special hat tip to Timothy Prickett Morgan:

The UNIX Guardian: The X Factor: Virtualization Belongs in the System, Not in the Software

Related Post:

Is this a crazy idea?

Tuesday, May 29, 2007

The Return of In-Flight Broadband

I stumbled onto this purely by accident.

Like the mythical bird, Phoenix, in-flight broadband Internet access may soon rise from the ashes of Connexion by Boeing. Panasonic Avionics Corporation has improved upon Connexion's original concept in its eXconnect offering.

eXconnect improves on Connexion by leasing satellite Internet connectivity, reducing the breakeven point for the service. It also improves speed using newer technology. The airplane antenna is more compact and lighter, saving fuel costs. And Panasonic is looking to partner with existing airport WiFi vendors, so you can pay for your in-flight connection, and use the airport departure lounge WiFi prior to boarding under a single package. Finally, the goal is to start the service at a price similar to Connexion by Boeing, and get the price down to about $20 for the duration of long-distance international flight within a year of launching.

It makes total sense the in-flight entertainment companies are bringing back in-flight broadband, and as any second attempt, it should be faster, smaller, and cheaper.

Now if they would just bring back Concorde.

Links:

Panasonic May Relaunch Connexion

Aircraft Interiors: Panasonic plans broadband launch in fourth quarter

Tuesday, April 03, 2007

Rhetorically Brilliant

So Sun Microsystems is bringing back a dedicated microelectronics division.

Why? Or more specifically, why now?

Here are my thoughts as to why.

First, the obvious. Sun in the past had an solid OEM SPARC business, primarily in the low-end embedded and SPARC clone workstation market, as well as low-end servers, and specialized systems such as telco market products and hardened systems for the military. These systems included workstations and servers where customers built their own systems based on UltraSPARC processors OEMed from Sun, as well as systems where the entire motherboard was OEMed from Sun.

Back when the techical desktop was ruled by UNIX/RISC, and Sun UltraSPARC II processors were solid desktop performers, this made a lot of sense. It also made sense with the initial UltraSPARC IIIi systems. However, the MHz race between Intel and AMD, along with the rise of Linux caused this specialized OEM business to shift towards x86, and Sun's OEM business shrank accordingly.

However, it never entirely went away, as Tadpole still sells UltraSPARC IIi and UltraSPARC IIIi based laptop computers, and the mil spec SPARC systems business still exists.

So why bring a dedicated microelectronics OEM business back? One reason is Sun has a very good OEM server chip with the UltraSPARC T1 "Niagara" processor. This processor has a system on a chip architecture which makes it easy for companies to build compute solutions around. And there are companies focused on markets where US-T1 fits well, such as telco, security, networking, etc. The other is Sun has Niagara 2 in the works, which could allow Sun to reposition the original Niagara as more of an embedded play, but also offer Niagara 2 to OEM customers where it may fit. Also, Sun now has a 10Gb Ethernet ASIC they wish to offer to the OEM market. And as you know if you follow Sun, volume matters. Sun knows it cannot drive its network ASIC into the larger market by itself.

But I think that is only part of the story. It explains the "Why". It does not explain the "When", or more precisely, the "Why Now".

As you may know, Sun will soon announce the servers which are part of the Sun-Fujitsu "Advanced Product Line" (APL) project. These systems will use Fujitsu's SPARC64-VI processors. Sun has had to face a lot of FUD around the future of its SPARC processors due to its decision to partner on these midrange and high-end systems.

What better preemptive action to take against he obvious FUD and annoying reporter questions than to create an Executive Vice President of SPARC? Who is the obvious leader of the SPARC processor business? People in the industry know who David Yen is. Does anyone know who run's Fujitsu's SPARC64 business?

Sun has grabbed the leadership of the SPARC industry in a very visible way just before the announcement of a product line which uses third-part SPARC processors.

All I can say is, this is rhetorically brilliant.

It has Jonathan Schwartz's fingerprints all over it.

Thursday, March 29, 2007

10 Gigabit Ethernet "Crossover"

I just saw this article comparing Fibre Channel, 10Gb Ethernet, and InfiniBand, and I thought it was interesting.

It points out information from the Dell'Oro consulting firm showing Gigabit Ethernet ports first outshipped Fast Ethernet ports in 2004, some seven years after GigE was introduced, and five years after the 1000BASE-T spec was introduced in 1999. (There is a good article here on the 1000BASE-T PHY.)

So if it took five years from the 1000Base-T release until 1000BASE-T ports exceeded 100BASE-T ports, despite backwards compatibility.

"Crossover", as this point is known, is very important in for a new replacement technology. Once a new standard or product reaches crossover, the second half of market penetration typically occurs quickly. Backwards compatibility helps, but realize a compatible 100/1000BASE-T port does not mean a switch blade with 48 of those ports will be compatible with an older blade chassis. Also, the increased cost of the new technology causes some resistance.

Today, there is a much bigger challenge for 10GBASE-T: Power consumption. From the article on the 1000BASE-T PHY I linked to earlier:

"Because of the complexity of the signal-processing task, a 10/100/1000Base-T copper PHY is the dominant consumer of power in essentially all gigabit switch designs supporting copper media. First-generation 1000Base-T copper PHYs introduced in 1999 in 0.35-micron CMOS consumed well over 5 watts of power-too high for widespread use in high-density Gigabit Ethernet switch form factors."
With 10GBASE-T, the power required is much higher. Chelsio's new 10GBASE-T NIC requires 24 watts of power, and can only drive a signal over 50 meters of the 100 meter distance of the 10GBASE-T spec. Now some of the NICs power is the supporting TCP Offload Engine (TOE) and other circuitry. But it is probably safe to say 15 watts for each 10GBASE-T switch or NIC port is currently required.

So while some say the promise of consolidated I/O will drive the transition from Gigabit Ethernet to 10 Gigabit Ethernet faster than the transition from Fast Ethernet to Gigabit Ethernet, power consumption will likely slow this transition significantly.

So indeed the transition from Gigabit Ethernet to 10 Gigabit Ethernet may follow the transition from Fast Ethernet to Gigabit Ethernet. In that case, if 1000BASE-T is a gauge, and the 10GBASE-T spec was just approved in 2006, it could take until 2011 before 10GBASE-T ports outnumber 1000BASE-T ports.

What is good about this is the software (iWARP, iSER, NFSoRDMA, pNFSoRDMA, etc.) and supporting networking protocols for this new generation of Ethernet will have plenty of time to catch up. This means once crossover does happen, it should have a very strong impact on the market.

Meanwhile, it appears there continues to be plenty of opportunity for Fibre Channel for storage and InfiniBand for low-latency IPC.

Related posts:
Predictions for the future of low-latency computing, it's not where you think it is
Part 1 | Part 2 | Part 3

Tuesday, March 27, 2007

I wonder why Microsoft has not bought Adobe

Adobe, with its new Creative Suite 3, is all the buzz now. But PhotoShop is not what makes Adobe interesting. It is Adobe's ability to establish two of its formats (PDF and Flash) as defacto standards.

Microsoft loves defacto standards. And Microsoft hates the fact that Adobe owns THE web animation standard (which is becoming the streaming media standard), and Adobe owns THE print-formatted document standard.

Microsoft tried to create an alternative to PDF, but has had zero success, despite the fact most PDF's original documents are created in Microsoft Office applications.

Which begs the question. Microsoft has a market cap of about $275 billion, compared to Adobe's $25 billion. Why Microsoft has not attempted a takeover, hostile or otherwise, of Adobe, is beyond me.

Perhaps it is all the bad blood between Microsoft and Adobe in the past. Maybe it is Microsoft's failed acquisition of SoftImage in the 1990s. But we are in a new decade, make that a new millennium.

Monday, February 05, 2007

Predictions for the future of low-latency computing, it's not where you think it is (Part 3)

Where is the future of general-purpose computer networks going? What could 10Gb Ethernet to the desktop enable? Good questions. For the last three years, I have carried a laptop with a 1Gb connection. Only twice in that time has it connected at above 100Mb. The truth is most LAN switches are still 100Mb, and 100Mb is more than enough bandwidth to support both computing and VoIP phones. A MPEG4 HDTV signal only requires about 4Mb/sec. You can put a lot on a 100Mb connection.

So back to 10Gb to the desktop. I recently built a new PC. I chose an Nvidia GeForce 7600 GS-based fanless video card. This is based on an older video chip, but leverages semiconductor process shrinks to deliver a much lower power solution. However, this card easily surpasses the performance of a state of the art, high-end UNIX workstation graphics card of four years ago. The cards for RISC/UNIX workstations of that era were PCI based, as the RISC/UNIX workstations did not have the Intel AGP slot. A 64-bit, 66MHz PCI slot provided about 500MB/sec of bandwidth. The idea behind these cards was all graphics processing was done on the card, which had very high bandwidth memory, and only instructions and small amounts of data were passed thorough the 500MB/sec PCI bus. Hold that thought.

At about the same time, there was an interest in creating high-end visualization solutions using these high-performance PCI graphics cards in small servers interconnected with a low-latency network such as Myrinet. The idea is these “graphics grids” could replace high-end visualization solutions from SGI and Evans and Sutherland.

So, what would happen if the 1000MB/sec 10Gb Ethernet replaced the PCI bus? Now, take the low-power graphics card, and add a TOE enabled NIC, and I have a high-performance networked display, without the need for an entire computer to support it. But what about Ethernet's latency? Those previous visualization clusters used Myrinet for latency as well as bandwidth. That is where iWARP comes in. One can run a protocol over a low-latency RDMA connection. This could fundamentally change computing, as 10Gb Ethernet replaces the system bus, and the graphics card becomes an add-on device to the display. Add a keyboard and mouse interface, and you now have an engineering thin-client, perfectly suited for a virtualized desktop running in a VM on a larger server.

Speaking of iWARP, the reduction of latency iWARP offers will be more important than the increase in bandwidth 10Gb Ethernet offers. You read it here first: Low-latency Ethernet will have a far greater impact than higher-performance Ethernet. Why? Simple. Lower-latency offers more potential for innovation than more bandwidth. There are plenty of options for bandwidth today, such as EtherChannel for IP and 4Gb Fibre-Channel for storage.

There is a key trend happening in computing over the last few years. Some call it commoditization. Some call it the “trend to free”. Operating systems “went to free” with the emergence of Linux. Some have said all software is “going to free”, with the emergence of open source. Others have spoken of “free bandwidth”. Certainly, one can look at Web 2.0 as an example of what happens to Internet sites when it is assumed everyone has a broadband connection. The emergence of AMD's Opteron and Intel's EM64T 64-bit extensions to the x86 architecture means 64-bit memory addressing is now “free” with the purchase of an x86 system, and is no longer requires an expensive RISC/UNIX platform. And with the emergence of Xen as a standard option for major Linux distributions, and multiple free options from VMware (VMware Player, VMware Server), virtualization is “going to free”.

What happens when something like this becomes “free”? New innovation is enabled at a level above the free layer. And that is why low-latency Ethernet will be so empowering to innovation.

Low-latency computing has always been a very high cost technology. For decades, low-latency computing has been limited to the realm of supercomputers, mainframes, and their logical follow-ons, HPC clusters and high-end RISC/UNIX systems. As a result, most of the battle against latency has occurred in software. The application clustering enabled by 1Gb Ethernet required proprietary software to manage state among many cluster nodes. Replication, caching, and specialized protocols were required to make it all work. In fact, in the clustered Java appserver space, the clustering technology became a key differentiator. But the truth is, if there was a ubiquitous low-latency interconnect available at the time, and intelligent operating system clustering, much less work would have been required on the part of the ISV. Simply put, if BEA was building a clustered Java appserver for a DEC VAXcluster, it would have been much easier, and they would have come to market much faster.

One can look at ISVs currently developing on InfiniBand as early adopters of low-latency computing. The basic “80/20” rule would suggest for every player investing in IB, there are four who could benefit but are not. Or it could be a 90/10 rule. It is not hard to believe for every ISV currently trying to gain competitive advantage with IB, there are nine others who don't feel it is currently worth the effort to pursue high-performance, low-latency networking as an enabler. This is much like early ISV support of Linux. Some felt it was a good fit for their product, others waited for more market acceptance and maturity.

So when low-latency networking becomes free, that is when all x86 servers come with iWARP ready on-board 10Gb Ethernet with TOEs, operating system and application developers will have new assumptions about cluster latency. It could open up initiatives for true single-system image clustering in Linux, and perhaps even Windows. Applications which previously were not clusterable, may be made so, which may disrupt existing applications. The promise of grid/utility computing becomes much more viable with a unified fabric. Blade server backplanes will probably be RDMA Ethernet based. Perhaps a shared storage clustered database alternative to Oracle RAC will emerge. Fundamental changes in the world of real-time computing, such as electronics, data capture, etc. are very likely. Radical changes to client computing are certainly possible, with thin clients offering far more potential than before. Basically, every form of computing which was weird or expensive because it required highly specialized, high-performance interconnects, will be commoditized.

This is my prediction of what 10Gb iWARP Ethernet will enable: Shared system image clustering will emerge as the defacto form of clustering, global filesystems will emerge as the defacto server filesystems, and "grid computing" (shared resource clustering) will become the normal method of deploying multiple servers. My guess for a target date for this becoming the norm in computing will be around 2015.

Fortunately, we do have an opportunity to examine in real-time what happens when a high-end computing technology becomes commoditized. The technology which offers this opportunity to observe is cheap, high-performance 3d graphics card technology. Once an industry unto itself, then an optional feature of a high-performance, expensive workstation, and now standard equipment of an ordinary desktop PC, only now is an x86 desktop operating system being released (Windows Vista), which requires a 3d graphics card. 3d displays are officially commoditized, and they are assumed to be there. Watch what happens in the graphical user interface space over the next few years. It will be a good benchmark for the innovation which occurs around a technology which has been commoditized.

Part 1 | Part 2

Thursday, January 25, 2007

Is this a crazy idea?

The Linux BIOS projects seeks to put a small Linux image into a ROM to manage PC type hardware.

Many have developed "boot from thumbdrive" operating systems, which put the whole OS into a few hundred megabytes, similar to a "Live CD".

Xensource has a developed bare metal hypervisors, including one which targets Windows-only environments. I assume these use a locked-down Linux or BSD kernel to provide the Xen "Domain 0" function. Xensource also provides a Xen Live CD for evaluation.

Meanwhile, as virtualization like Xen and VMware continue to mature, as CPUs evolve to support virtualization (Intel VT, AMD-V), and PCI Express I/O virtualization coming soon, x86 virtualization will become more robust and approach native performance. It is likely virtualized x86 servers will become the norm for production environments.

What would happen if the Linux BIOS, boot from flash memory, and a bare-metal Xen Live CD ideas merged? Imagine a "XenBIOS" project, using a few hundred megabytes of flash memory to hold a live image on board. It would mean virtualization sedimenting into hardware, not into operating systems, as most are currently predicting.

The effect of free, hardware based virtualization which is automatically there would make for very interesting x86 servers. Even more so with a few on-board, fully virtualized, multi-fabric I/O channels. Kind a baby mainframe.

Maybe it's a crazy idea. Maybe its a vision of the future of computing.

Wednesday, January 24, 2007

Predictions for the future of low-latency computing, it's not where you think it is (Part 2)

Many people make Mistaken Assumptions when speaking on the history of computer networking.

One assumption made is that 100Mb, then 1Gb Ethernet replaced many other protocols. Certainly, in some cases, this is true. But people claiming this often overstate the facts.

Ethernet is first and foremost a LAN protocol, not a specialized, high-performance cluster interconnect for connecting multiple, large shared memory systems or vector supercomputers. To put it simply, I doubt anyone can point to any case of Ethernet replacing HIPPI. HIPPI was used both a storage connection or a cluster interconnect for supercomputers. Clearly what killed HIPPI in storage interconnects was fibre-channel. In cluster interconnects, one of the only vendors using HIPPI was SGI, for clustering multiple multi-hundred CPU Origin systems together. SGI used InfiniBand for clustering its Altix follow-ons to the Orgin.

Certainly 10Mb shared Ethernet killed Token Ring. But think about it, how prevalent were Token ring networks? For how many people was Ethernet the first LAN technology they experienced?

What is important is not that Ethernet killed Token Ring, but that Ethernet, by being standardized and multivendor, drove local area networking prices down low enough to become ubiquitous, which enabled the emergence of LAN email and the client server revolution in the early to mid 1990s.

Certainly 100Mb switched Ethernet with QoS killed ATM to the desktop (very small, niche market). MPLS was probably the key technology replacing ATM in the Wide Area Network, and FDDI in the Metro Area Network, but now Metro Ethernet is often used over MPLS connections.

What is important is not that 100Mb Ethernet killed ATM's promise in the LAN (primarily of network video), but that 100Mb Ethernet, by being standardized and multivendor, drove high-performance local area networking prices down low enough to become ubiquitous, and along with inexpensive Ethernet routers, the emergence of campus wide LANs, which enabled the emergence web-based computing using Java and other technologies in the late 1990s. As for network video, that came in a highly compressed form, primarily over the Internet via 1.5Mb down/256Kb up DSL and cable modem connections, not via bidirectional 100Mb Fast Ethernet connections.

Now what was 1Gb Ethernet supposed to kill? Answer: Fibre-Channel. Many predicted back in 1997-1998 GigE would kill Fibre Channel. I remember first it was going to be NFS over GigE, then it was DAFS over GigE, then it was iSCSI over GigE, as if to blame the protocol for the failure, instead of the true reason: TCP/IP overhead in the pre-TOE, pre-GHz class CPU era. Instead, GigE enabled the easy clustering of applications between servers, such as Java application servers like BEA WebLogic and IBM WebSphere, and databases such as Oracle 9i RAC, rather than servers to storage. Oddly enough, iSCSI has reemerged in the last couple of years not as a replacement for Fibre-Channel, but for storage replication and as a remote boot technology for centrally managed client PCs.

Do you see a trend here? It is not the technology which is superseded which determines the success of the new technology, it is the new innovation the new technology enables. Those who predicted uses of GigE by looking at other 1Gb networks (i.e., Fibre-Channel), instead of a faster Ethernet, were wrong. Just as those before them.

So the truth is while it is wrong to bet against Ethernet, it is also wrong to assume Ethernet has killed every other networking protocol before it. It is also the antithesis of innovative thought to look at a technology by what old things it can kill, rather than what new things it can enable.

But often few people are experiencing the latest speed on their desktop and laptop computers, and this may be part of the reason some people naturally look at current high-performance networking to estimate where 10Gb Ethernet will make its impact. The problem is, as we have see, higher-performance Ethernet's impact is always somewhere other than where the preceding high performance networking technology was. Often, higher-performance Ethernet solves a different, unforseen problem than the preceding high-performance networking technology of similar bandwidth.

My next post will look at 10 gigabit Ethernet and the some predictions on the future of computer networking.

Part 1 | Part 3

Wednesday, January 17, 2007

Predictions for the future of low-latency computing, it's not where you think it is (Part 1)

Low-latency cluster interconnects have always been esoteric. Digital Equipment Corporation (DEC) used reflective memory channel (along with many other interconnects) in its VAX clustering. Sequent used Scalable Coherent Interconnect (SCI) for its NUMA-Q. SGI used HIPPI to connect its Power Challenge, and later Origin servers. IBM had its SP interconnect for its RS/6000 clusters. Sun Microsystems flirted with its Fire Link technology for its high-end SPARC servers. In the late 1990s, Myrinet ruled the day for x86 high-performance computing clustering.

Today, the low-latency interconnect of choice is InfiniBand (IB). Unlike the other technologies, IB is both an industry standard, and offered by multiple vendors. SCI, while standards-based, was a single-vendor implementation. Similarly, Myricom submitted Myrinet as a standard, but remains single vendor. A multi-vendor environment drives down prices while forcing increased innovation.

Another key aspect of IB is the software is also proceeding down a standards path. The OpenIB Alliance was created to drive a standardized, open source set of APIs and drivers for InfiniBand.
An interesting thing happened on the path to OpenIB. The emergence of 10 gigabit Ethernet, and its required TCP/IP Offload Engine (TOE) Network Interface Cards (NICs), offered another standards-based, high performance interconnect. As a result, OpenIB was rechristened as the OpenFabrics Alliance, and became “fabric neutral”.

My next post will look at some of the mistaken assumptions about the progress of computer networking.

Part 2 | Part 3