Showing posts with label aviation safety. Show all posts
Showing posts with label aviation safety. Show all posts

Wednesday, November 12, 2008

What about GM?

Toyota has a track record of taking over totally dysfunctional GM plants and making them functional by just changing management, with same unions, same facilities, same labor, same equipment. See history of NUMMI in Fremont, CA, going from GM's worst plant to the best in one year.

references:

Becoming Lean , Jeffrey Liker, page 62-63
http://books.google.com...

Stop Rising Healthcare Costs Using Toyota Lean Production Methods

By Robert Chalice (page 53)


http://books.google.com...

There may be no reason to lose the jobs, plants, or contracts.

Actually, what changed was not just management, but the whole underlying philosophy on which the plant was managed, which is the crucial change.

Chalice cites these factors as the new "five core values: teamwork, equity, involvement, mutual trust and respect, and safety."

In short, workers were treated as first class partners in the plant, not as some kind of "asset" to be "managed" and "controlled." They were listened to. They were respected.

Yes, that does make all the difference, in automobiles, as Liker points out, or in hospitals, as Chalice points out, supported by the Keystone study of John's Hopkins Dr. Peter Pronovost in Michigan, showing that when nurses were actually listened to by doctors, patients were significantly better off and had better outcomes.

Gasp. I took Dr. Pronovost's class in Patient Safety last year, and, yes, it really is that "simple." Culture drives safety and productivity. Culture drives the bottom line, not technology.

If you want things to work, you have to learn about human beings, and culture, and work within the constraints that puts on you. Humans are not machines and work way better than machines if allowed to (McGreggor's Theory Y), or way worse than machines if forced to (Theory X).

It's the job of the stockholders and stakeholders to realize that, and put management in place that will support the work force instead of trying to exploit it.

Period.


Wednesday, May 30, 2007

The road to error - illustrated

There are many different kinds of errors that organizational systems of humans can make, but one of the trickiest is directly related to the questions of "integrity", "transparency", and "prejudice." I want to relate these to the classic "swiss cheese" multi-layered defense system that James Reason made famous:


[ source of that slide: ..? ]

Instead of looking at the layers the way he does, let's just use one slice of cheese as a model, and examine what can happen when an organization, initially one person, has a base fully covered but then the organization starts to grow and add people.



The problem is that, as the organization spreads out one conceptual task over more and more people, gaps start to occur in the coverage. They occur particularly in the area where it's a little fuzzy which person or team's job it is to handle that task.





This seems to me to be an intrinsic failure mode for organizations. It turns out, that regardless how good a job anyone in a company can do, if they don't actually do it, their skill level doesn't matter. Furthermore, a very common way for people not to do a job is for them not to realize that it's their job to do. In some organizations this might be accompanied by a twinge of remorse, but then a resigned "It's not my job!" and forgetting about the task.

So, when a task that used to be something one person does get's divided up among many people, there is a risk that none of those people will decide the task is their to do, regardless how well intentioned or skilled they are. This effect can completely neutralize years of effort getting skilled at a task. Things, almost literally, "fall through the cracks."

And the cracks almost always appear, if the task and organization keep growing and growing and adding more and more people to distribute a single conceptual task among. Soon, the organization looks like the following, with entire "silos" of separate groups, and each silo broken into a pecking order of elites, middle class, and bottom rung workers of some kind. Now there are a lot of gaps, but still, the gaps are fairly small.



But, as the organization continues to grow and evolve more specialized skills in each local area, the people in each box start to spend more time talking to each other than they do talking to people outside their own little box. It's more convenient, and the language is more directly relevant. We all speak the same language. It begins to become "us" here in this box, versus "them" out there in other boxes.



Still, the teams may be cooperating, but that won't last. Sooner or later, messages are missed, or silence itself becomes interpreted as a hostile message. Something falls through the cracks, there is a storm of blame and recimination, and a deadly spiral sets in of becoming more and more convinced that all problems are due to the people in other boxes, who are surely idiots or else have evil intent. The boxes draw away from each other, in a mild form of disgust. The "us" becomes fractured into many different kinds of "us".



As the communication between teams becomes more hostile, "management" may decide to simlify the problems by having all communications go through them. The number of connections going into one box is now at most two, one from above, and one going to a box below in the pecking order. This allows the fabric of the cheese to twist around the thin connecting segments, as if around an axle. Within each section of cheese, this is unnoticed, because their world is still fine, locally.



Then, the layer of cheese may start to warp and become a curved surface, not a flat surface. Again, seen from within that section, everything is fine, because the observers in that "flat land" are measuring a curved surface with curved rulers, and it looks just fine. Even simple facts and reasoning from other sections, however, don't seem to make sense anymore, because they don't line up correctly. This is attributed to the other group losing touch with reality.




Finally, the fabric of the organization is so frayed and fragmented that whole pieces fall off, unnoticed from within. Now you can "drive a small truck" through the gaps and holes, but again this is not visible from inside each segment, because it spends zero time pondering the middle territory or white space. That space is "not our job" but is "someone else's job".

This condition of an organization is now somewhat stable. Life goes on, and a number of errors come and go, with everyone attributing the errors to everyone else, and shaking their heads at how those "others" aren't doing their jobs. Other groups are seen as actively hostile enemies, blaming us for things we didn't do. Relations deteriorate. Errors abound.

Now the amazing thing is that this can occur even though each team is doing an almost perfect job of managing what they see as their own turf.

The error occurs in a place we are so unfamiliar with we don't even have a name for it. I call it the M.C. Esher Waterfall Error, after this work of Escher. At first glance and even close inspection, the image seems a little strange, but harmless.


A closer inspection reveals that the water, however, is following an impossible path.

It flows down a waterfall, then flows down a zigzag of channels, and finds itself back at the top of the waterfall, so it falls down the waterfall, ...
etc. forever. It's a perpetual motion machine.

The vertical columns in the middle tier in front have something terribly wrong with them too.

And yet, if you look at any small part of this lithograph, nothing seems wrong.

This is a problem we are simply not used to encountering - the detail level is correct, but the larger global level is clearly absurd and wrong.

We have "emergent error", sort of the opposite of synergy.

The swiss cheese and waterfall pictures are meant to illustrate that organizations break down in a funny way, where all the pieces continue to work, but the overall integrity falls apart, in a very subtle and unnoticed way. In fact, it is generally hard to get anyone to pay attention to the fact that something serious is wrong, because anyone can see, from inside, that everything (that you see from inside) is correct. (We have run into Godel's Theorem as a problem.)

Conclusions:
1) Just because everything locally measures as fine does not mean things are fine.
2) Even if everyone can do a perfect job, that won't matter if they don't do it.
3) They won't do it if it's not perceived as "their job".
4) This mode of breakdown is very insidious, but I think it is also very common.

This kind of expansion and condensation and specialization needs to be balanced with a corresponding effort at reintegration, although it may seem a minor and non-urgent task.

Then, something huge comes through the gap, and everyone is astounded that such a thing could happen.

Another post will deal with ways to address it. This post is just to document that there is a type of problem that organizations can suffer, a malady or disorder or disease, that is very difficult to trace locally. It always seems to be coming from "over there", but if you go "over there" you see that it isn't coming from "over there" either. It locks itself down with blame, stereotyping, and sullen bitterness about having to put up with "those idiots" in the other departments who keep messing things up. It is hard to decipher because the simplest messages from other departments don't even make sense and you have to wonder if they've remembered to take their medications lately. The more errors go through the hole, the more people lock into blaming each other, and the more the subsections curl up to avoid touching the other sections and withdraw into their own comfortable world where people talk sense and behave rationally.

No one is doing anything wrong, and everyone is doing something wrong, but the wrongness is subtle. It has something to do with whether everyone is OK with not being clear whose job a task might be, and not being able to find out whose job it is. If people are "responsibility seeking", this may be less likely than if they are "responsibility avoiding" as an ethic. If people feel an error is "not my problem" or "someone else's problem" this can worsen.

If the world is divided into "us" and "them", there is always a middle ground that is very confusing and not clearly us and not clearly them. Errors flow to that ground, like pressurized gas trying to escape. If there are cracks between teams, errors seem eerily capable of finding them. The errors are remarkably resilient to efforts to track them down and fix them, and seem to keep happening, as if those idiots over there have no learning curve at all.

But, it is a very dangerous wrongness, if this problem occurs on a global scale, and teams don't just get annoyed at each other and fight figurative wars, but actually start dropping explosive devices on each other in order to stop the continual assault they feel they are under.

It may also be that the efforts to reduce animosity by controlling all communications between hostile teams by routing them through management is well intended, but is based on a model of communication that is single-channel, explicit, context-independent, and rooted deeply in processing linear strings of symbols - where one mistake can throw off everything. The communication that takes place before the body is fractured and fragmented howerver is more like image-processing: it is multi-channel, implicit, context dependent, and not based on symbol processing, so it is robust and fairly immune to point-noise. In fact, generally, changing a single pixel in an image has zero effect on the contained communication.

It may be that what is needed is a lot more socializing, and sloppy, many-to-many uncontrolled interactions, as a kind of glue to keep the pieces from falling apart. As Daniel Goleman notes in his book Social Intelligence, humans have a great many different ways to synchronize and synch up and coordinate with each other, most of which are non-verbal, very fast, and intrinsically sloppy and prone to pointwise error. Those errors are made up by having massive parallel communications, not by reducing communications to a single channel that is very tightly regulated. There is not enough bandwidth in a single channel to synchronize two disparate groups at all points. The groups can "twist" and "rotate" around that channel, and move out of synch. Best efforts mysteriously fail.

References
Human Error - Models and Management, James Reason. BMJ 2000;320:768-770 ( 18 March )

Monday, May 14, 2007

My Comair 5191 crash analysis now available

Because "systems thinking" is a difficult concept to describe, I wrote and just posted a paper analyzing a commercial aircraft disaster - the crash of Comair 5191 - in Lexington Kentucky, August, 2006. This is a full-length (30 page) analysis with pictures and diagrams and source materials, aside from the cockpit voice recorder transcripts, which are linked below. The final NTSB findings on the case are not yet out, to my knowledge.

It's a little rough around the edges, but it starts with the basic astounded question of how two, fully trained pilots, not under pressure, could taxi to and attempt to take off from the wrong runway, resulting in the death of all on-board except the one who was flying the plane, who was pulled from the flaming wreckage by a first responder. The runway was a few hundred yards too short for the plane to have made it off the ground safely.

So, it goes from "How on earth could this have happened!?" to "Oh... There but for the grace of God go I." Only the new commercial pilots on the pilot chat blogs couldn't imagine how such a thing could ever happen to them. It brought to mind the old saying "There are bold pilots, and there are old pilots." In this case, however, the rest of the world conspired to set the stage.

As with "errors" in hospitals, it typically takes a whole team of people to align their actions in the wrong way (the "swiss cheese model"), for someone to buy the gun, someone to load the gun, someone to cock the hammer, someone to hand it to the poor last guy in the chain, and that guy to pull the trigger. For legal purposes, blame is assessed one way, a way this paper does not assess. For purposes of safety engineering, and seeing where interventions might help to avoid ever having this happen again, we need to look at a whole different set of factors that set the stage for this "accident".

Please contact me if you'd like to use this paper (or a newer, better version) for instructional material. Thanks!

( Note: I am a private pilot, but I'm not a member of the NTSB or any official agency, and this analysis is a personal analysis for instructional purposes in safety engineering, not intended for legal purposes. I have no relationship that I know of to anyone involved in this case. These are all real, living people and my reconstruction may be entirely wrong. The point is to honor those who died by learning everything we can from their deaths so this won't happen again.)

Prior Posts:
Comair 5191 - Confirmation Bias and Framing (1/20/07)
Cockpit voice recorder transcripts
Washington DC Crash of Air Florida was 25 years ago - remembered
(with links to BMJ, High-reliability engineering, TEM, etc.)

Sunday, May 13, 2007

Powerpoint on why too much quality doesn't work



It seems to me that there can be such a thing as too many procedures, to the point where, as John Gall would say, the component that will fail is sthe one that you put in to make the system "fail-safe". That is, the thing that will kill you is the thing that you put in place to save you.

This may have been true of Comair Flight 5191, that crashed while attempting a mistaken takeoff from the wrong runway in Lexington, Kentucky last August.

Some pilots have wryly compared the pre-takeoff check list on a 727 to a "Michner novel", in terms of its length. In the case of Comair 5191, judging from the transcripts of the cockpit voice reocorder, the right-hand seat co-pilot apparently spent the entire taxi-time with head down, going though the checklist, while the left-hand seat pilot taxied the plane to the wrong runway, opened the throttles and told the copilot "you've got it."

We have some mixed signals on how to deal with this. In a perfect case, even a very "lightweight" solution is more than adequate. The picture illustrates a single sheet of typing paper rolled up and taped into that shape, holding up 3 books.

However, when it comes to checklists, or standard operating procedures (SOP's), sometimes there are simply too many. Cultures in the red-quadrant of the "competing values" diagram think that the problem is always too few procedures, and want to add more. They think the graph of reliability versus standardization goes up forever. In practice, the graph seems to be more hill-shaped, going up to a point, then somewhat down as more and more procedures start getting in the way, and finally result in catastrophic failure as the system crashes under the weight of it's own safety system. One can think, perhaps, of the "right number of laws" to have optimal regulation of an industry, and which of those two curves applies in whatever is your own case.

I put up a powerpoint slide presentation (no audio) considering this issue.
you can get it here, titled "spectacular.ppt".