Showing posts with label information technology. Show all posts
Showing posts with label information technology. Show all posts

Monday, December 9, 2013

Failure guru Amy Edmondson deconstructs the Healthcare.gov fiasco

Amid all the breathless news coverage of the failed rollout of the Obamacare Healthcare.gov website, we now have some genuine analysis, courtesy of one of my heroes, Amy Edmondson of Harvard Business School ("The Mistakes Behind Healthcare.gov Are Probably Lurking In Your Company, Too"). She may be more qualified than anyone to weigh in, given her deep research experience in learning from mistakes and failure in very complex situations (including healthcare). A couple of potent excerpts:

Healthcare.gov is a good example of the importance of learning small and fast, rather than rolling out a risky new product or service launch all at once. Cycling out in phases includes the expectation of early failures – and demands all hands on deck to learn from them along the way. A roll-out, in contrast, implies that something is all set, ready to go — like a carpet. All it needs is a bit of momentum to propel it forward. For complex initiatives, of course, this is simply not the case. Getting people motivated enough to change is not the real challenge; it’s getting them engaged enough to learn — to become part of a discovery process.

and...

Managers must make it clear that they understand that excellent performance does not mean not making mistakes — it means learning quickly from mistakes and sharing the lessons widely.

Thursday, September 5, 2013

Improving large, distributed information systems by inducing failures

I've been thinking about this powerhouse paper from the Association for Computing Machinery's acmqueue site for the last week. It's called "The Antifragile Organization: Embracing Failure to Improve Resilience and Maximize Availability" by Ariel Tseitlin.

Taking its starting point from Nassim Taleb's arguments about antifragility (see references here and here), Tseitlin discusses ways of making large distributed information services antifragile - i.e., able to capitalize and improve themselves in the wake of disruption. His focus is on testing and simulation, the question of how to exercise highly complex systems to ensure they will not collapse as a result of stressors.

As Tseitlin observes, traditional scripted testing is utterly unsuited to this task - in systems of any size, it's impossible to build (or even imagine) the total number of test cases required to prove a system's robustness. Moreover, even the largest test system is a fraction of the size and complexity of the production environment.

Taking a radically different approach, companies like Amazon and Netflix are increasingly causing intentional disruption within their production systems to assess whether their resilience mechanisms are working as expected - and whether new vulnerabilities have emerged. Tseitlin describes several ways that this is done:

Once you have accepted the idea of inducing failure regularly, there are a few choices on how to proceed. One option is GameDays, a set of scheduled exercises where failure is manually introduced or simulated to mirror real-world failure, with the goal of both identifying the results and practicing the response—a fire drill of sorts. Used by the likes of Amazon and Google, GameDays are a great way to induce failure on a regular basis, validate assumptions about system behavior, and improve organizational response.

But what if you want a solution that is more scalable and automated—one that doesn't run once per quarter but rather once per week or even per day? You don't want failure to be a fire drill. You want it to be a nonevent—something that happens all the time in the background so that when a real failure occurs, it will simply blend in without any impact.

One way of achieving this is to engineer failure to occur in the live environment. This is how the idea for "monkeys" (autonomous agents really, but monkeys inspire the imagination) came to Netflix to wreak havoc and induce failure. Later the monkeys were grouped together and labeled the Simian Army.

Netflix's "monkeys" include, among others, a Chaos Monkey, which randomly terminates virtual instances in the production environment; and a Latency Monkey, which inserts delays into various components of the network.

These are run regularly and the Netflix team carefully measures to see whether the system adapts suitably to the disruption. The random and low-level nature of these tests helps avoid limits of human-scripted testing, and the fact they are running in the production environment means they are not limited by the size of a test environment.

Is it risky? Not if the system has been engineered not to collapse under stressors. Top quality distributed systems are built to isolate failures and degrade gracefully rather than result in catastrophic downtime. As such, it does require a system that has been in production long enough to develop stability - Twitter in its first two years would not be a good candidate for this, but Twitter at present would be.

Tseitlin concludes the paper by discussing how these exercises in resilience are building toward true antifragility, by using tools such as post-exercise blameless postmortems and requiring that developers also be operators, the better to anticipate code that might cause operational issues down the line.

Thursday, August 22, 2013

"You can't build a web system that will never break"

From a post on Fred Wilson's AVC blog:

Once you have a successful product in the market, you need to turn your attention to scaling it. The system you and your team built will break if you don't keep tweaking it as demand grows. Greg Pass, who was VP Engineering at Twitter during the period where Twitter really scaled, talks about instrumenting your service so you can see when its reaching a breaking point, and then fixing the bottleneck before the system breaks. He taught me that you can't build something that will never break. You have to constantly be rebuilding parts of the system and you need to have the data and processes to know which parts to focus on at what time.

Thursday, March 28, 2013

Enterprise startups who don't build a professional services business making a mistake

Mark Suster on his Both Sides of the Table blog writes this: "Many young startups are being advised not to have a professional services business and in my opinion this is a big mistake."

This is a good lesson for anyone starting up a company selling to the business market. Ignore VC calls to go lean and not invest in professional services - rather, invest wisely to ensure you have successful product implementations and the resulting referenceable customers. Wise advice. Here's a brief excerpt:

The most important way to sell a product for an early-stage business (or frankly any stage) is to have strong referenceable customers. These are the lifeblood of your sales organization. Referenceable means they are willing to be part of your sales collateral, willing to take calls from key leads, willing to speak at your conferences, etc.

How do you get referenceable customers? You build a great product and make sure it is used in such a way as to deliver real benefit to your customers versus just the promise of a benefit outlined in your marketing materials.

As much as many non-experienced investors might like to believe, even great products don’t just roll themselves out. You need to implement them. This often means getting the product to talk with other existing products, implementing the product to match the specific needs of a customer’s internal processes, training, monitoring usage and encouraging adoption
.

It also reminded me of this advice from Scott Weiss; his company decided to invest in customer-facing resources to give them big-company customer service, so the large enterprises they were targeting would not be afraid to buy from them.

Wednesday, March 20, 2013

A cool and candid post-mortem from the founder of Gowalla

This piece by Gowalla founder Josh Williams in Medium.com is a great recapitulation of the "check-in wars" from 2008-2011, when location based services like Gowalls and Foursquare emerged (to be joined by Facebook) and fought to establish this new turf in the online world. Gowalla was a casualty of this war, and Williams' candid recounting of the story is a service to other tech entrepreneurs. While he doesn't directly deal with mistakes that contributed to Gowalla's downfall, the story helpfully lays out the extreme uncertainty that comes with any new venture (especially online ventures). Here's a taste:

In one of the many highlights of Gowalla, we crafted an amazing tie-in with Disney that was loved by both our community and theirs. Each of these endeavors cost us time and money. Unfortunately the relative payoff for us was simply less due to network effect.

While Gowalla continued to grow, the trajectory was not what it needed to be. At least not in terms of winning the game we had chosen to play.

We were the younger, prettier, but less popular sister of Foursquare. And even that had changed. In time, Foursquare had dramatically improved the design and experience of its service. This was no longer a defensible platform for us as a company.

Around this time we knew that our path was in trouble.

Monday, July 2, 2012

Zappos site messes up pricing, then shows sense of humor in dealing with it

This is a posting from Zappos' site 6pm.com dated May 2010.

6pm.com
Hey everyone – As many of you may know (and I’m sure a lot of you do not), 6pm.com is our sister site.  6pm.com is where brandaholics go for their guilt free daily fix of the brands they crave.  Every day, the site highlights discounts on products ranging up to 70% off.  Well, this morning, we made a big mistake in our pricing engine that capped everything on the site at $49.95.  The mistake started at midnight and went until around 6:00am pst.  When we figured out the mistake was happening, we had to shut down the site for a bit until we got the pricing problem fixed. 
While we’re sure this was a great deal for customers, it was inadvertent, and we took a big loss (over $1.6 million - ouch) selling so many items so far under cost.  However, it was our mistake.  We will be honoring all purchases that took place on 6pm.com during our mess up.  We apologize to anyone that was confused and/or frustrated during out little hiccup and thank you all for being such great customers.  We hope you continue to Shop. Save. Smile. at6pm.com
Cheers!
Aaron Magness
Director of Brand Marketing & Business Development
Zappos Development, Inc.
Twitter: @macknuttie
Update: Upon further investigation and clarification with our merchandising team, I realized that the statement about "capping everything on the site at $49.95" was not 100% accurate. There are some items that are sold on both 6pm.com and Zappos.com, and those items were not affected by the pricing mistake. The pricing mistake applied to items sold on 6pm.com but not Zappos.com (the vast majority of inventory available on 6pm.com). The actual dollar figure of our loss is accurate - over $1.6 million. Let's just say this was not a boring weekend for us.
Update 2: We've received a number of inquiries asking for more details as to what happened, so here are more details from Tony Hsieh (CEO, Zappos.com, Inc.):
We have a pricing engine that runs and sets prices according to the rules it is given by business owners. Unfortunately, the way to input new rules into the current version of our pricing engine requires near-programmer skills to manipulate, and a few symbols were missed in the coding of a new rule, which resulted in items that were sold exclusively on 6pm.com to have a maximum price of $49.95. (Items that are sold on both 6pm.com and Zappos.com were not affected.)
We already had planned on improving our internal pricing engine so that it will have a much easier-to-use interface for our business owners. We are also planning on adding additional checks and balances to hopefully prevent this type of thing from happening again.
To those of you asking if anybody was fired, the answer is no, nobody was fired - this was a learning experience for all of us. Even though our terms and conditions state that we do not need to fulfill orders that are placed due to pricing mistakes, and even though this mistake cost us over
$1.6 million, we felt that the right thing to do for our customers was to eat the loss and fulfill all the orders that had been placed before we discovered the problem.
PS: To put an end to any further speculation about my tweet (
http://twitter.com/zappos/status/14576863056 ), I will also confirm that I did not, in fact, eat any ice cream on Sunday night.
Tony Hsieh

Hat tip to Positive Sharing.

Monday, May 14, 2012

Renny Gleeson: the 404 error page is an opportunity to engage with users

In this TED talk, Renny Gleeson discusses the HTML 404 error (page not found), and how companies have creatively designed their version of this error page to build connection with their users. Says Gleeson, "A simple mistake should remind me of why I love you."


Friday, April 27, 2012

Urban Outfitters apologizes for web outage, with kittens!

Urban Outfitters ran an online sale this week which they promoted with an email blast (see below).


During the sale, their website crashed. And to their credit, the company sent another email out apologizing for the outage, and adding free shipping. If I had tried to order something that day and was unable to, this email would have made me feel better:


This is evidence that you can turn a negative into a positive with some humility and a sense of humor. A deal-sweetener doesn't hurt either.

Thanks to the Listrak Email Marketing blog, which posted on this (and from which I got the screenshots).

Friday, April 20, 2012

TerraCycle's Tom Szaky makes a tough decision

An undertone of this project is decision-making. Each decision point is an opportunity to make a mistake. Avoiding decisions are mistakes in themselves. Yet in a complex environment there is no playbook for assuring a decision leads to a good outcome.

It's rare to get a peek into someone else's significant decision-making process, so kudos to TerraCycle CEO Tom Szaky, who discussed a key business decision in the New York Times "You're The Boss" blog this week.

TerraCycle has been trying to add a new type of business to its existing one - collecting juice pouches and remanufacturing them into consumer goods such as purses. In this case, the juice companies (such as Capri Sun) subsidize TerraCycle's costs for collection. The new business (called World of Waste, or W.O.W.) involves engaging consumers to cover the costs to collect and recycle other products for which the manufacturers aren't willing to pay TerraCycle to collect.

Enhancing TerraCycle's information systems to support this new model has caused delays in the program's introduction. Those delays are the subject of Szaky's post. I was most interested in the section where he talked about the management debate between investing in enhancing the existing juice pouch business model versus adding this new, quite different model:

I have made a firm decision that W.O.W. will not be delayed again. Not everyone on my senior management team agreed with this decision. The essence of our debate was whether we should spend all of our resources building our existing infrastructure (which is proven but dependent on one source of funding, our brand partners) or whether we should take three or four months and improve our I.T. infrastructure so that it can handle the new demands of W.O.W. (with a new system, W.O.W. would allow us to receive funds directly from consumers and thereby diversify our revenue). Of course, W.O.W., if it works, would not start generating real revenue until 2013, because we would not be able to introduce it, at best, until late in the third quarter of this year.

The arguments on both sides are compelling and logical. On the one had, it makes sense to strengthen a proven business model and have it immediately generate more revenue, which can in turn help drive our ability to invest further in the company. On the other hand, adding a new revenue generator and a major extension of our offerings could open us up to new customers and bring a big new revenue opportunity.

So far, my experience in business (which admittedly is limited, given that I’m 30) has taught me that taking such gambles is a worthwhile endeavor — even if the odds are against you. I have also found that in such moments it is difficult to make decisions that everyone on the team will support. Perhaps this is why most organizations have a leader and are not run by committee.

Dialogue, debate, decide, fall in line. This is a classic decision process. As Szaky says, there is a leader so these hard decisions can be made. And when he says "taking such gambles is a worthwhile endeavor - even if the odds are against you," he is saying that the bold decision energizes the team and makes it work like crazy to make the decision successful. (For more on this point, check out my article on The 99 Percent.)