Showing posts with label algorithms. Show all posts
Showing posts with label algorithms. Show all posts

Sunday, June 22, 2014

Decision-by-Algorithm: No Silver Bullet

This post contains excerpts from "The Scored Society: Due Process for Automated Predictions," a 2014 law review article authored by Danielle Keats Citron and Frank A. Pasquale III. 


* * * * * * *

Big Data is increasingly mined to rank and rate individuals. Predictive algorithms assess individuals as good credit risks, desirable employees, reliable tenants, and valuable customers. People’s crucial life opportunities are on the line, including their ability to obtain loans, work, housing, and insurance.

The scoring trend is often touted as good news. Advocates applaud the removal of human beings and their flaws from the assessment process. Automated systems are claimed to rate individuals all in the same way, thus averting discrimination. But this account is misleading. Human beings program predictive algorithms. Their biases and values are embedded into the software’s instructions, known as the source code, and predictive algorithms.  Please see What Gets Lost? Risks of Translating Psychological Models and Legal Requirements to Computer Code.

Credit scoring has been lauded as shifting decision-makers’ attention from troubling stereotypes to bias-free assessments of would-be borrowers’ actual records of handling credit. The notion is that the more objective data at a lender’s disposal, the less likely a decision will be based on protected characteristics like race or gender. But far from eliminating existing discriminatory practices, credit-scoring algorithms are instead granting them an imprimatur, systematizing them in hidden ways.

A credit card company uses behavioral-scoring algorithms to rate consumers a worse credit risk because they used their cards to pay for marriage counseling, therapy, or tire-repair services. Online evaluation systems score interviewees, with color-coded rating of red signaling a “poor candidate,” yellow as middling, and green as “hire away.”

Beyond biases embedded into code, some automated correlations and inferences may appear objective, but may in reality reflect bias. Algorithms may place a low score on occupations like migratory work or low paying service jobs. This correlation may have no discriminatory intent, but if a majority of those workers are racial minorities, such variables can unfairly impact consumers’ loan application decisions.

Credit scores are only as free from bias as the software and data behind them. Software engineers construct the datasets mined by scoring systems; they define the parameters of data-mining analyses; they create the clusters, links, and decision trees applied. They generate the predictive models applied. The biases and values of system developers and software programmers are embedded into each and every step of development.

Just as concerns about scoring systems are heightened, their human element is diminishing. Although software engineers initially identify the correlations and inferences programmed into algorithms, Big Data promises to eliminate the human “middleman” at some point in the process.

According to a January 9, 2014 article in CIO.com, IBM says cognitive computing systems like Watson are capable of understanding the subtleties, idiosyncrasies, idioms and nuance of human language by mimicking how humans reason and process information.

Whereas traditional computing systems are programmed to calculate rapidly and perform deterministic tasks, IBM says cognitive systems analyze information and draw insights from the analysis using probabilistic analytics. And they effectively continuously reprogram themselves based on what they learn from their interactions with data.

Said IBM CEO Ginni Rometty, "In 2011, we introduced a new era [of computing] to you. It is cognitive. It was a new species, if I could call it that. It is taught, not programmed. It gets smarter over time. It makes better judgments over time." "It is not a super search engine," she adds. "It can find a needle in a haystack, but it also understands the haystack."

This "new species" of computing has its challenges. According to "IBM Struggles to Turn Watson Computer Into Big Business," a recent Wall Street Journal article:
Watson is having more trouble solving real-life problems than "Jeopardy" questions, according to a review of internal IBM documents and interviews with Watson's first customers. 
For example, Watson's basic learning process requires IBM engineers to master the technicalities of a customer's business—and translate those requirements into usable software. The process has been arduous.
Klaus-Peter Adlassnig is a computer scientist at the Medical University of Vienna and the editor-in-chief of the journal Artificial Intelligence in Medicine. The problem with Watson, as he sees it, is that it’s essentially a really good search engine that can answer questions posed in natural language. Over time, Watson does learn from its mistakes, but Adlassnig suspects that the sort of knowledge Watson acquires from medical texts and case studies is “very flat and very broad.” In a clinical setting, the computer would make for a very thorough but cripplingly literal-minded doctor—not necessarily the most valuable addition to a medical staff.

As Hector J. Levesque, a professor at the University of Toronto and a founding member of the American Association of Artificial Intelligence, wrote:

 "As a field, I believe that we tend to suffer from what might be called serial silver bulletism, defined as follows:
the tendency to believe in a silver bullet for AI, coupled with the belief that previous beliefs about silver bullets were hopelessly naıve. 
We see this in the fads and fashions of AI research over the years: first, automated theorem proving is going to solve it all; then, the methods appear too weak, and we favour expert systems; then the programs are not situated enough, and we move to behaviour-based robotics; then we come to believe that learning from big data is the answer; and on it goes."

Similarly, employment assessment companies have marketed the benefits of "science, precision and data" over the past fifteen years under the guise of neural networks, artificial intelligence, big data and deep learning, yet what has changed? Employee engagement levels have hardly budged and employee turnover remains a continuing and expensive challenge for employers. Please see Gut Check: How Intelligent is Artificial Intelligence?









Saturday, June 14, 2014

Algorithms: Deeply Human Choices Behind Cold Mechanisms

This post is comprised of excerpts and a graphic from Rethinking Personal Data: A New Lens for Strengthening Trust, a document published by the World Economic Forum and  prepared in collaboration with A.T. Kearney. The document addresses the key trust challenges facing the personal data economy, and offers a set of near-term and long-term insights for addressing these issues.  

* * * * * * *

Complex and opaque, algorithms generate the predictions, recommendations and inferences for decision-making in a data-driven society. While easily dismissed as abstract empirical processes, algorithms are deeply human. They reflect the intentions and values of the individuals and institutions which design and deploy them. The ability for algorithms to augment existing power asymmetries gives rise to debates on their influence over data-driven policy-making.

The nature of these debates are complex, value-laden and give rise to some fundamental societal choices. Questions of individual autonomy, the sovereignty of individuals, digital human rights, equitable value distribution and free will are all a part of these conversations. There are no easy answers. Through this long-term lens on the impact of proactive computing, the focal point for discussion begins to shift away from personal data, per se, to computer-based profiles of individuals and groups of individuals. These profiles — fueled by fine-grained behavioral and sensor data — make it possible to monitor, predict and instrument social phenomena at the micro and macro levels.

The world of “smart” environments, where cars, eyeglasses and just about everything else coalesce into the Internet of Things, creates a sea change in how data will be processed. Rather than being based on “interactive” human-machine computing, smart environments rely upon “proactive computing”. By design, these proactive environments are one step ahead of individuals. Connected cars need to anticipate accidents before they happen. Evacuating flood prone areas needs to occur before major storms hit.

The emphasis on proactive computing will change the role of human intervention from a governance perspective. Lacking a full understanding of how complex systems work, the ability of humans to understand, make decisions and adapt can be too slow, incomplete and unreliable. In this brave new world, building trust from the “principles up” will be essential and require new forms of governance that are open, inclusive, self-healing and generative.

From a community and societal perspective, as civil “regulation-by-algorithm” begins to scale, incumbent interests and power asymmetries will play an increasing role in establishing who gets
access to an array of commercial and governmental services. As such, there is a need to ensure that the algorithms driving proactive and anticipatory decisions will be lawful, fair and can be explained intelligibly. Meaningful responses must be given “when individuals are singled out to receive differentiated treatment by an automated recommendation system”.


One emerging set of concerns is the institutional ability “to discover and exploit the limits of an individual’s ability to pursue their own self-interest.” Given that a majority of consumer interactions in the future will be mediated via devices and commercially oriented communications platforms, data-centric institutions will have the means and incentives to trigger “predictable irrationality”
from individuals.

With a vast trail of “digital breadcrumbs” accessible for companies to mine and tailor highly personalized experiences, a growing set of concerns is arising on how individuals could be profiled and targeted at moments of key vulnerability (decision fatigue, information overload, etc.) and limit their ability to act with agency and in their own self-interest. With the lives of individuals becoming increasingly mediated by algorithms, a richer understanding is needed for how people adapt their behaviors to empower themselves and gain more control over the manner of how profiles and algorithms shape their lives in areas such as credit scores, retail experiences, differential pricing, reputational currencies, insurance rates, etc.

One of the most strategic insights on strengthening trust is the concept of exploring ways to share intended consequences of data usage to individuals. For example, the 2012 Draft European Data Protection Act (section 20), calls for “the obligation for data controllers to provide information about the envisaged effects of such processing on the data subject”.

To address this emerging set of concerns, establishing a cross-disciplinary community of forward-looking experts, complexity scientists, biologists, policy-makers and business leaders with an appreciation of the long-term societal impact was identified as a priority. This group would proactively help design and test systems that balanced the commercial, legal, civil and technological
incentives shaping outcomes at the individual and social level. They would need to develop some form of legal protection to limit liabilities and provide a safe space to explore complex issues in a
real-world setting. One attribute of this safe space would be for it to be governed by an institutional review board where ethics and the interests of individuals could have a meaningful and relevant voice (similar to how they are used by the biomedical and behavioural science sectors). Institutions concerned about legal uncertainties, regulatory action or civil lawsuits could have a richer means for assessing ethical concerns using these approaches.

Thursday, May 15, 2014

White House: Big Data's Role in Employment Discrimination

In“Big Data: Seizing Opportunities, Preserving Values," a White House review of how the government and private sector use large sets of data found that such information could be used to discriminate against Americans on issues such as employment. As noted in the review, "while big data can be used for great social good, it can also be used in ways that perpetrate social harms or render outcomes that have inequitable impacts, even when discrimination is not intended."

Clear Windshield or Rearview Mirror?


Algorithms embody a profound deference to precedent; they draw on the past to act on (and enact) the future. The apparent omniscience of big data may in truth be nothing more than misdirection. Instead of offering a clear windshield, the big data phenomenon may be more like a big rear-view mirror telling us nothing about the future.

Does this deference to precedent result in a self-reinforcing and self-perpetuating system, where individuals are burdened by a history that they are encouraged to repeat and from which they are unable to escape?

Already burdened segments of the population can become further victimized through the use of sophisticated algorithms in support of the identification, classification, segmentation, and targeting of individuals as members of analytically constructed groups. In creating these groups, the algorithms rely upon correlations that lead to viewing people as members of populations, or categories, or groups, rather than as individuals (i.e., persons who live more than X miles from an employer's location). Please see From What Distance is Discrimination Acceptable?

Just as neighborhoods can serve as a proxy for racial or ethnic identity, there are new worries that big data technologies could be used to “digitally redline” unwanted groups, either as customers, employees, tenants, or recipients of credit. A significant finding of the White House report is that big data could enable new forms of discrimination.

Correlation Does Not Equal Causation

Decisions made or affected by correlation are inherently flawed. Correlation does not equal causation. This point is made vividly by Tyler Vigen, a law student at Harvard who, in his spare time, put together a website that finds very, very high correlations - as shown below - between things that are absolutely not related.

Screenshot_2014-05-12_12.46.40
Screenshot_2014-05-12_12.46.30
Screenshot_2014-05-12_12.45.09
Each of these have correlation coefficents in excess of 0.99, serving to demonstrate the point that a strong correlation isn't nearly enough to make strong conclusions about how two phenomena are related to each other.

Shrouding Opacity In The Guise of Legitimacy

Some of the most profound challenges revealed by the White House Report concern how big data analytics may lead to disparate inequitable treatment, particularly of disadvantaged groups, or create such an opaque decision-making environment that individual autonomy is lost in an impenetrable set of algorithms.

Workforce analytic systems, designed in part to mitigate risks for employers, have become sources of material risk, both to job applicants and employers. The systems create the perception of stability through probabilistic reasoning and the experience of accuracy, reliability, and comprehensiveness through automation and presentation. But in so doing, technology systems draw  attention away from uncertainty and partiality. Please see Workforce Science: A Critical Look at Big Data and the Selection and Management of Employees.

Moreover, they shroud opacity—and the challenges for oversight that opacity presents—in the guise of legitimacy, providing the allure of shortcuts and safe harbors for actors both challenged by resource constraints and desperate for acceptable means to demonstrate compliance with legal mandates and market expectations.

Programming and mathematical idiom (e.g., correlations) can shield layers of embedded assumptions from higher level decisionmakers at an employer who are charged with meaningful oversight and can mask important concerns with a veneer of transparency.

This problem is compounded in the case of regulators outside the firm, who frequently lack the resources or vantage to peer inside buried decision processes. In recognition of this problem, the White House Report states that "[t]he federal government must pay attention to the potential for big data technologies to facilitate discrimination inconsistent with the country’s laws and values" and, as one of the six policy recommendations in the report, :
The federal government’s lead civil rights and consumer protection agencies, including the Department of Justice, the Federal Trade Commission, the Consumer Financial Protection Bureau, and the Equal Employment Opportunity Commission, should expand their technical expertise to be able to identify practices and outcomes facilitated by big data analytics that have a discriminatory impact on protected classes, and develop a plan for investigating and resolving violations of law in such cases. In assessing the potential concerns to address, the agencies may consider the classes of data, contexts of collection, and segments of the population that warrant particular attention, including for example genomic information or information about people with disabilities. 

Monday, November 18, 2013

Do We Regulate Algorithms, or Do Algorithms Regulate Us?

The genesis for this posting comes from the following articles:
This posting includes portions of the articles and modifies them to address issues relating to big data and the use of algorithmic decisionmaking in the area of pre-employment assessments and workforce optimization.

Embedding Bias

Every step in the big data pipeline raises concerns: the privacy implications of amassing, connecting, and using personal information, the implicit and explicit biases embedded in both datasets and algorithms, and the individual and societal consequences of the resulting classifications and segmentation.

While many companies and government agencies foster an illusion that classification is (or should be) an area of absolute algorithmic rule—that decisions are neutral, organic, and even automatically rendered without human intervention—reality is a far messier mix of technical and human curating. Data isn't something that's abstract and value-neutral. Data only exists when it's collected, and collecting data is a human activity. And in turn, the act of collecting and analyzing data changes (one could even say "interprets") us. 

Both the datasets and the algorithms reflect choices, among others, about data, connections, inferences, interpretation, and thresholds for inclusion that advance a specific purpose. Like maps that represent the physical environment in varied ways to serve different needs—mountaineering, sightseeing, or shopping—classification systems are neither neutral nor objective, but are biased toward their purposes. They reflect the explicit and implicit values of their designers.Assumptions are embedded in a data model upon its creation. Data sources are shaped through ‘washing’, integration, and algorithmic calculations in order to be commensurate to an acceptable level that allows a data set to be created.

Errors are not only possible, but they are likely to occur at each stage in the process of assessment that proceeds from identification to its conclusion in a discriminatory act. Error is inherent in the nature of the processes through which reality is represented as digitally encoded data. Some of these errors will be random, but most will reflect the biases inherent in the theories, and the goals, the instruments and the institutions that govern the collections of data in the first place.

Clear Windshield or Rearview Mirror?

The decisions made by the users of sophisticated analytics determine the provision, denial, enhancement, or restriction of the opportunities that citizens and consumers face both inside and outside formal markets.


Algorithms embody a profound deference to precedent; they draw on the past to act on (and enact) the future. The apparent omniscience of big data may in truth be nothing more than misdirection. Instead of offering a clear windshield, the big data phenomenon may be more like a big rear-view mirror telling us nothing about the future.

Does this deference to precedent result in a self-reinforcing and self-perpetuating system, where individuals are forever burdened by a history that they are encouraged to repeat and from which they are unable to escape? Does deference to past patterns augment path dependence, reduce individual choice, and result in cumulative disadvantage?

Already burdened segments of the population can become further victimized through the use of sophisticated algorithms in support of the identification, classification, segmentation, and targeting of individuals as members of analytically constructed groups. In creating these groups, the algorithms rely upon generalizations that lead to viewing people as members of populations, or categories, or groups, rather than as individuals (i.e., persons who live more than X miles from a jobsite).

Shrouding Opacity In The Guise of Legitimacy

Workforce analytic systems, designed in part to mitigate risks for employers, have now become sources of material risk. The systems create the perception of stability through probabilistic reasoning and the experience of accuracy, reliability, and comprehensiveness through automation and presentation. But in so doing, technology systems draw organizational attention away from uncertainty and partiality. They embed, and then justify, self-interested assumptions and hypotheses.

Moreover, they shroud opacity—and the challenges for oversight that opacity presents—in the guise of legitimacy, providing the allure of shortcuts and safe harbors for actors both challenged by resource constraints and desperate for acceptable means to demonstrate compliance with legal mandates and market expectations.

The technical language of workforce analytic systems obscures the accountability of the decisions they channel. Programming and mathematical idiom can shield layers of embedded assumptions from high-level firm decisionmakers charged with meaningful oversight and can mask important concerns with a veneer of transparency. This problem is compounded in the case of regulators outside the firm, who frequently lack the resources or vantage to peer inside buried decision processes and must instead rely on the resulting conclusions about risks and safeguards offered them by the parties they regulate.

Do We Regulate Algorithms, or Do Algorithms Regulate Us?

Can an algorithm be agnostic? Algorithms may be rule-based mechanisms that fulfill requests, but they are also governing agents that are choosing between competing, and sometimes conflicting, data objects.

The potential and pitfalls of an increasingly algorithmic world beg the question of whether legal and policy changes are needed to regulate our changing environment. Should we regulate, or further regulate, algorithms in certain contexts? What would such regulation look like? Is it even possible? What ill effects might regulation itself cause? Given the ubiquity of algorithms, do they, in a sense, regulate us?

We regulate markets, and market behavior, out of concerns for equity, as well as out of concern for efficiency. The fact that the impacts of design flaws are inequitably distributed is at least one basis for justifying regulatory intervention.

The regulatory challenge is to find ways to internalize the many external costs generated by the rapidly expanding use of analytics. That is, to find ways to force the providers and users of discriminatory technologies to pay the full social costs of their use. Requirements to warn, or otherwise inform users and their customers about the risks associated with the use of these systems should not absolve system producers of their own responsibility for reducing or mitigating the harms. This is part of imposing economic burdens or using incentives as tools to shape behavior most efficiently and effectively.







Friday, September 20, 2013

What Gets Lost? Risks of Translating Psychological Models and Legal Requirements to Computer Code

The genesis for this posting is the article "Technologies of Compliance: Risk and Regulation in a Digital Age" authored by Kenneth A. Bamberger and found at 88 Texas L. Rev. 669 (2010). This posting takes portions of the article, modified to address the issue of job applicant assessments, and intersperses information on elements of workforce analytics to provide examples of the risks and challenges raised in the Bamberger article.

Workforce analytic systems are powerful tools, but they pose real perils. They force computer programmers to attempt to interpret psychological models, legal requirements and managerial logic; they mask the uncertainty of the very hazards with which lawmakers and regulators are concerned; they skew decisionmaking through an “automation bias” as a substitute for sound judgment; and their lack of transparency thwarts oversight and accountability.

Lost In Translation

The hiring assessment functionality of workforce analytics contains three divergent logic systems, legal, psychological and managerial. The legal logic system derives, in part, from the Americans with Disabilities Act (ADA) and its accompanying regulations and related caselaw. The psychological logic system derives primarily from the five-factor model of personality, or Big Five, as it has evolved over the past 20-25 years. The managerial logic derives from the implementation of the human resource function of the employer. Technology in the form of workforce analytics then attempts to tie these three logic systems together in order to create an automated assessment program that attempts to determine applicant "suitability" or "fit."

Information technology is not value-neutral, but embodies bias inherent in both its social and organizational context and its form. It is not infinitely plastic, but, through its systematization, trends towards inflexibility. It is not merely a transparent tool of intentional organizational control, but in turn shapes organizational definitions, perceptions, and decision structures. In addition to controlling the primary risks it seeks to address, then, it can raise—and then mask—different sorts of risk in its implementation.

For example, many workforce analytic companies utilize the Big Five model in creating their personality assessments. As its name implies, the Big Five looks at five traits: openness, conscientiousness, extraversion, agreeableness, and neuroticism, with each trait conceptualized on an axis from low to high (e.g., low neuroticism, high neuroticism). The Big Five, operating under various names, existed for a number of decades prior to its "rebirth" in the early 1990s, where it was embraced by organizational psychologists.

Since the late 1990s workforce analytics companies like Unicru (now owned by Kronos) have adopted the Big Five for use in their job applicant assessment program. The workforce analytic companies have created "model" psychological profiles and tested applicants against those profiles. In general, applicants receive either green, yellow or red scores on the basis of a 50/25/25 cutoff. Applicants scoring red are generally not interviewed, let alone hired.

The use of technology systems to hardwire workforce analytics raises a number of fundamental issues regarding the translation of legal mandates, psychological models and business practices into computer code and the resulting distortions. These translation distortions arise from the organizational and social context in which translation occurs; choices “embody biases that exist independently, and usually prior to the creation of the system.” And they arise as well from the nature of the technology itself “and the attempt to make human constructs amenable to computers.”

These distortions are compounded when psychological models, legal standards and managerial processes are turned over to programmers for translation into predictive algorithms and computer code. These programmers may know nothing of the psychological models, legal standards and management processes. Some are employees of separate IT divisions within firms; many are employees of third-party systems vendors. Wherever they work, their translation efforts are colored by their own disciplinary assumptions, the technical constraints of requirements engineering, and limits arising from the cost and capacity of computing.

Managerial processes may be poor vehicles for capturing nuance in legal policy and psychological models, especially in a context like employment discrimination where regulators have eschewed rules for standards and where the interpretation of those standards by regulators and psychological professionals may change over time. For example, the ADA prohibits pre-employment medical examinations and psychological tests used by workforce analytic companies may be considered medical examinations (please see ADA, FFM and DSM). Employers and workforce analytic companies have interpreted the medical examination requirement as prohibiting the use of tests that are designed to diagnose mental illnesses. 

This interpretation creates two significant risks for employers and workforce analytic companies. First, the legal standard does not speak to "tests designed to diagnose mental illnesses;" rather, it is "whether the test is designed to reveal an impairment of physical or mental health such as those listed in the Diagnostic and Statistical Manual of Mental Disorders." "Designed to reveal" is semantically and substantively different from "designed to diagnose" and, as set out in Employment Tests are Designed to Reveal an Impairment, Big Five-based tests are designed to reveal impairments by their screening out process. As depicted in the graphic, tests designed to diagnose are a subset of the overall category of tests designed to reveal an impairment. Using the "designed to diagnose" category as the proxy for medical examinations puts employers and workforce analytic companies at significant risk of violating the ADA medical examination prohibition and the confidential medical information safeguards under the ADA. It may also result in other claims against the workforce analytic companies by job applicants, employers and insurers. The mistaken use of the "designed to diagnose" category as a proxy for medical examinations could be considered a design defect in the product liability arena or as negligent in a tort claim..


The second significant risk for employers and workforce analytic companies arises from their failure to account for the evolution of the Big Five model from non-clinical model to clinical model.  In her seminal review of the personality disorder literature published in 2007, Dr. Lee Anna Clark stated that “the five-factor model of personality is widely accepted as representing the higher-order structure of both normal and abnormal personality traits.” A clear sign of this evolution comes with the publication of the most current volume of the Diagnostic and Statistical Manual for Mental Disorders (DSM-5), published in May 2013, where the model used to diagnose many personality disorders is based on the Big Five. 

Consequently, even if the standard for defining a medical examination was focused solely on the use of a test that diagnosed a mental illness, the five-factor model has now evolved into a diagnostic tool used by the psychiatric community to define mental impairments, including personality disorders, of the kind set out in the DSM-5. The failure of employers and workforce analytic companies to account for the evolutionary development of the five-factor model puts them at significant risk due to their belief that the five-factor model is not a diagnostic tool - a belief that time and scientific advances have now overturned. 

Automation Bias

While computer code and predictive-analytic methods might be accessible to programmers, they remain opaque to users —for whom, often, only the outcomes remain visible. In the case of job applicants, even this information (assessment outcome) is not visible to them - results are not disclosed by employers or workforce analytic companies. Programmers “code[] layer after layer of policies and other types of rules” that managers and directors cannot hope to understand or unwind.

Human judgment is subject to an automation bias, which fosters a tendency to “disregard or not search for contradictory information insight of a computer-generated solution that is accepted as correct.” Such bias has been found to be most pronounced when computer technology fails to flag a problem.

In a recent study from the medical context, researchers compared the diagnostic accuracy of two groups of experienced mammogram readers (radiologists, radiographers, and breast clinicians)—one aided by a Computer Aided Detection (CAD) program and the other lacking access to the technology. The study revealed that the first group was almost twice as likely to miss signs of cancer if the CAD did not flag the concerning presentation than the second group that did not rely on the program.

Automation bias may be found in the algorithms created and used by workforce analytic companies to provide "insights" to their employer customers. For example, Kenexa, an IBM company, has determined that distance from work, commute time and frequency of household moves all have a correlation with attrition in call-center and fast-food jobs. Applicants who live more than five miles from work, have a lengthy commute or have moved more frequently are scored down by the algorithms, making them less desirable candidates.

Painting with the broad brush of distance from work, commute time and moving frequency may result in well-qualified applicants being excluded. The Kenexa insights are generalized correlations; they say nothing about any particular applicant.

What are the risks of employers slavishly adhering to the results of the algorithm? Part of the answer comes from identifying groups of people who have longer commutes and move more frequently than others, lower-income persons who, according to the U.S. Census, are disproportionately African-American and Hispanic.

Through the application of these “insights,” many low-income persons are electronically redlined, meaning employers will pass over qualified applicants because they live (or don’t live) in certain areas, or because they have moved. The reasons for moving do not matter — whether it is to find a better school for their children, to escape domestic violence, the elimination of mass transit in their community, or as a consequence of job loss due to a company shutdown (please see From What Distance is Discrimination Acceptable?)

An employer who does not look past the simple results of the assessment algorithms not only harms itself by failing to consider well-qualified employees, the employer puts itself at risk for employment discrimination claims by classes of persons (e.g., African-American, Hispanic) protected by federal and state employment laws.

Institutionalized (Mis)Understanding

Institutionalization of workforce management practices might permit evolutionary improvements in existing measurements, but it masks areas where risk types are ignored or analysis is insufficient and where more revolutionary, paradigm-shifting advances might be warranted.

These understandings (or misunderstandings) can be institutionalized across the field of workforce analytics. As workforce analytic practices are disseminated through the industry by professional groups, workforce analytic practitioners, management scholars, and third-party technology vendors and consultants, they standardize an approach that other firms adopt, seeking legitimacy.

Workforce analytic systems, designed in part to mitigate risks, have now become sources of risk themselves. They create the perception of stability through probabilistic reasoning and the experience of accuracy, reliability, and comprehensiveness through automation and presentation. But in so doing, technology systems draw organizational attention away from uncertainty and partiality. They can embed, and then justify, self-interested assumptions and hypotheses.

Moreover, they shroud opacity—and the challenges for oversight that opacity presents—in the guise of legitimacy, providing the allure of shortcuts and safe harbors for actors both challenged by resource constraints and desperate for acceptable means to demonstrate compliance with legal mandates and market expectations.

The technical language of workforce analytic systems obscures the accountability of the decisions they channel. Programming and mathematical idiom can shield layers of embedded assumptions from high-level firm decisionmakers charged with meaningful oversight and can mask important concerns with a veneer of transparency. This problem is compounded in the case of regulators outside the firm, who frequently lack the resources or vantage to peer inside buried decision processes and must instead rely on the resulting conclusions about risks and safeguards offered them by the parties they regulate.

Risks of a Technological Monoculture

Technology-based workforce analytic systems proliferate, in part, because policy makers have rejected rule-based mandates in favor of regulatory principles that rely on the exercise of context-specific judgment by regulated entities for their implementation . Yet workforce analytic technology can turn each of these regulatory choices on its head. The need to translate psychological, legal and managerial logic into a fourth distinct logic of computer code and quantitative analytics creates the possibility that legal choices will be skewed by the biases inherent in that process.

Such biases introduce several risks: that choices will be shaped both by assumptions divorced from sound management and incentives unrelated to public ends (e.g., hiring discrimination leading to larger income support payments - SSDI, SSI); that the rule-bound nature of code will substitute one-time technological “fixes” for ongoing human oversight and assessment (e.g., failure to recognize the evolution of the Big Five becoming a diagnostic tool); and that the standardization of risk-assessment approaches will eliminate variety—and therefore robustness in workforce analytic efforts, developing systemic risks of which individual actors may not be aware.

Systemic risks have developed because there is a technological "monoculture" in the workforce analytic industry. The problems are analogous to those of the biological domain. A deeply entrenched standard prevents the introduction of technological ideas that deviate too far for accepted norms. This means that the industry may languish with inefficient or non-optimal solutions to problems, even though efficient, optimal, and technically feasible solutions exist. The technical feasibility of these superior solutions is not important; they are excluded because they are incompatible with the status quo technology.

In a diverse population, any particular weakness or vulnerability is likely confined to only a small segment of the whole population, making population-wide catastrophes extremely unlikely. In a homogeneous population, any vulnerability is manifested by everyone, creating a risk of total extinction. In the case of workforce analytics, one successful challenge the Big Five-based model either being an illegal medical examination or screening out persons with disabilities introduces systemic risk to all customers of that workforce analytics company - one loss will lead to multiple challenges, along the lines of the asbestos litigation (please see The Next Asbestos? The Next FLSA?).Systemic risk is not limited to the ecosystem of the workforce analytic company being challenged, but extends to all companies that market or utilize Big Five-based assessments.

The potential costs are enormous. If the assessment is an illegal medical examination, then each applicant has a claim based on the use of an illegal medical examination. Some employers use the assessments to screen millions of applicants each year; each applicant is a potential plaintiff. Further, if the test is a medical examination, then each applicant has a claim for the misuse of confidential medical information (if the test is a medical examination, the applicant responses are confidential medical information). Not only does that lead to claims based on privacy violations, but all systems, solutions and databases that incorporate the information obtained from the assessments will need to be "sanitized." In a very real sense, the data may be the virus and the costs of "cleansing" those systems may well dwarf the very significant damages payable to applicants (please see When the First Domino Falls: Consequences to Employers of Embracing Workforce Assessment Solutions).

Wednesday, August 7, 2013

From What Distance is Discrimination Acceptable?

Some 60% of American workers earn hourly wages. Of these, about half change jobs each year, so firms that employ lots of entry-level workers, such as call centers, supermarkets, home improvement stores and fast-food chains, have to vet million of applications every year.

Xerox Evolv(ing)

A recent article in MIT Technology Review reports that Xerox is screening tens of thousands of applicants for low-wage jobs in its call centers using software from a startup company called EvolvAccording to its website, Evolv “is a workforce science software company that harnesses big data, predictive analytics and cloud computing to help businesses improve workplace productivity and profitability” and its customers include 20 of the Fortune 100.


Working with Xerox, Evolv found that one of the best predictors that a customer-service employee will stick with a job is that he lives nearby and can get to work easily.  As Evolv states in its Q3 2013 Workforce Performance Report:
The distance that employees live from work affects how long they choose to stay at a job. Unsurprisingly, employees that live 0-5 miles from their place of work have the longest median tenure. They remain at their jobs 20% longer than employees with the shortest median tenure.
Although it still sells photocopiers, Xerox has also become one of the world’s largest outsourcing companies. It provides services like running customer service centers, handling health claims, and processing credit-card applications that brought in $11.5 billion in revenue last year.

That business relies on a huge workforce of 54,000 customer service agents, and because of high attrition in hourly jobs, Xerox will have to replace 20,000 of them this year, says Teri Morse, vice president for recruiting at Xerox Services.

The Dictatorship of Data

Morse says Xerox today won’t even look at resumes of those who score in the “red” category of Evolv’s initial behavioral assessment, a 30-minute online exam that workers fill out at home. Early on, while piloting the system, Morse says Xerox still hired against the advice of the data. Now, she says, “people who do poorly we no longer hire.”

We are more susceptible than we may think to the “dictatorship of data” — that is, to letting the data govern us in ways that may do as much harm as good. The threat is that we will let ourselves be mindlessly bound by the output of our analyses even when we have reasonable grounds for suspecting something is amiss. Or that we will become obsessed with collecting facts and figures for data’s sake. Or that we will attribute a degree of truth to the data which it does not deserve.
For more and more companies, like Xerox, the hiring boss is an algorithm. Jobs that were once filled on the basis of work history and interviews are left to personality tests and data analysis. The new hiring tools are part of a broader effort to gather and analyze employee data.

The risks to employers of utilizing online personality tests in their employment application process have been set out in a number of prior posts, including What Are the Issues, Courts Find Tests to be Illegal, The Next Asbestor? The Next FLSA?, Damages and Indemnification Challenges to Employers, and Welcomed as Customers; Rejected as Employers. The remainder of this post sets out the employment discrimination litigation risks to employers (and society) of blindly following the "insights" of Kenexa and Evolv in distance from job site and housing mobility.

Interestingly, while Evolv now touts the distance from job insight for use by its clients, the company expressed a different view in a 2012 Wall Street Journal article, which reads;
Evolv is cautious about exploiting some of the relationships it turns up for fear of violating equal opportunity laws. While it has found employees who live farther from call-center jobs are more likely to quit, it doesn't use that information in its scoring in the U.S. because it could be linked to race.
From What Distance is Discrimination Acceptable?

Kenexa, purchased by IBM in December 2012, will test approximately 40 million applicants this year for thousands of clients. Kenexa believes that a lengthy commute raises the risk of attrition in call-center and fast-food jobs. It asks applicants for call-center and fast-food jobs to describe their commute by picking options ranging from "less than 10 minutes" to "more than 45 minutes."

The longer the commute, the lower their recommendation score for these jobs, says Jeff Weekley, who oversees the assessments.Applicants also can be asked how long they have been at their current address and how many times they have moved. People who move more frequently "have a higher likelihood of leaving," Mr. Weekley said.

Painting with the broad brush of distance from job site, commute time and moving frequency results in well-qualified applicants being excluded, applicants who might have ended up being among the longest tenured of employees. The Kenexa and Evolv findings are generalized correlations (i.e., persons living closer to the job site tend to have longer tenure than persons living farther from the job site). The insights say nothing about any particular applicant.

As a consequence, employers will pass over qualified applicants solely because they live (or don't live) in certain areas. Not only does the employer do a disservice to itself and the applicant, they increase the risk of employment litigation, with its consequent costs. 

Distance From Jobsite

A recent New York Time article, "In Climbing Income Ladder, Location Matters," reads, in part:
Stacey Calvin spends almost as much time commuting to her job — on a bus, two trains and another bus — as she does working part-time at a day care center.  ...
Her nearly four-hour round-trip [job commute] stems largely from the economic geography of Atlanta, which is one of America’s most affluent metropolitan areas yet also one of the most physically divided by income. The low-income neighborhoods here often stretch for miles, with rows of houses and low-slung apartments, interrupted by the occasional strip mall, and lacking much in the way of good-paying jobs
This geography appears to play a major role in making Atlanta one of the metropolitan areas where it is most difficult for lower-income households to rise into the middle class and beyond, according to a new study that other researchers are calling the most detailed portrait yet of income mobility in the United States.
The dearth of good-paying jobs in low-income neighborhoods means that residents of those neighborhoods have a longer commute. The 2010 Census showed that poverty rates are significantly higher for blacks and Hispanics. Consequently, hiring decisions predicated on distance from job site, intentionally or not, discriminate against certain races.

Housing Mobility

As shown in the table below, poor and near-poor families tend to move much more frequently than their higher income neighbors and the general population.


According to a 2011 study by the Center for Public Housing, entitled "Should I Stay or Should I Go? Exploring the Effects of Housing Instability and Mobility on Children," a wide range of often complex forces appears to drive frequent mobility, and residential instability in general — the formation and dissolution of households, an inability to afford one’s housing costs, the loss of employment, the lack of a safety net, lack of quality housing or a safer neighborhood.

Correlation Is Not Causation

When two variables, A and B, are found to be correlated, there are several possibilities:

1. A causes B
2. B causes A
3. A causes B at the same time as B causes A (a self-reinforcing       system)
4. Some third factor causes both A and B

The correlation is simple coincidence. It is wrong to assume any of these possibilities.

Kenexa, Evolv and their clients, however, assume that A (proximity to job site) causes B (reduced attrition and better performance). That assumption leads them to disfavor otherwise qualified applicants who do not live within a five-mile radius.

The correlation could also demonstrate B (reduced attrition and better performance) is caused by C (proximity of job site to applicants homes). Instead of being a hiring insight, the correlation might function better as being a job site location insight. Given the relative immobility of persons and companies, locating a job site (call center, etc.) close to communities with high numbers of lower-income persons could lead to a more sustainable competitive advantage.

Per Kenexa, the correlation between (A) persons who move and (B) shorter job tenure is that A causes B. However, per the 2011 study by the Center of Public Housing, it may well be that (B) shorter job tenure causes (B) persons to move. As shown in the table below, taken from the 2011 study, more than 76% of involuntary moves were a result of job loss.


Although data does give rise to information and insight, they are not the same. Data's value to business relies on human intelligence, on how well managers and leaders formulate questions and interpret results. More data doesn't mean you will get "proportionately" more information. In fact, the more data you have, the less information you gain as a proportion of the data (concepts of marginal utility, signal to noise and diminishing returns).

Maybe It's the Work, Not the Workforce, That Need Analysis ...

Some 60% of American workers earn hourly wages. Of these, about half change jobs each year. The cost to U.S. businesses of worker attrition and lost productivity is $350 billion annually. How much is that?


What should an employer do? Increase starting pay? Since most people work, at least in part, for the money - giving them more money might encourage them to stay longer and work harder. It seems to work reasonably well with corporate executives.  Top executive compensation averaged $9.4 million last year at the 50 largest employers of low-wage workers.

What do many employers do? Retain "workforce science" companies like Evolv and Kenexa to administer personality tests and analyze applicant data in order to address the employee turnover issue.

As noted above, working with Xerox, Evolv found that one predictor that a customer-service employee will stick with a job is that s/he lives nearby and can get to work easily. Kenexa had a similar "insight," and added that people who move more frequently have a higher likelihood of leaving.

Are there any groups of people who might live farther from the work site and may move more frequently than others? Yes, lower-income persons, disproportionately women, black, Hispanic and the mentally ill. They can't afford to live where the jobs are and move more frequently because of an inability to afford housing or the loss of employment.

So, not only are low-income persons poorly paid, many are electronically redlined from hiring consideration. What type of “workforce science” fails to take into account the most important variable (pay) and yet offers “solutions” based on this flawed science?

Are Employer Referral Programs Encouraging Discriminatory Hiring Practices?

According to Evolv, it recently “used rich data on hundreds of thousands of employees to demonstrate that referred workers show measurably better productivity and retention.” The research showed that 65 percent of companies have a referral program and 36 percent filled their last opening through an employee referral. A primary insight from Evolv is that, when it comes to retention, referred workers were around 13 percent less likely to quit.

As noted above, previous insights of Evolv and Kenexa – distance from work and housing mobility – lead to workforce selection processes that discriminate against blacks and Hispanics. Combining the employer referral program insight with the distance from work and housing mobility insights likely exacerbates the discriminatory impact of workforce science and its use in the hiring process.

About 40 percent of white Americans and about 25 percent of non-white Americans are surrounded exclusively by friends of their own race, according to an ongoing Reuters/Ipsos poll. Even looking at a broader circle of acquaintances to include coworkers as well as friends and relatives, 30 percent of Americans are not mixing with others of a different race, the poll showed.

The workforce insights regarding distance from job and employee mobility results in fewer blacks and Hispanics being hired. Consequently, if 36% of job openings are filled by referrals from employees and 30% of those employees do not have friends or relatives of another race, blacks and Hispanics will be underrepresented in the workforce hired as a result of referrals.





Sunday, August 4, 2013

The Dictatorship of Data or Fooled by Randomness

Big Data has spawned a cult of infallibility — a vision of prediction obviating explanation and math trumping science. In The End of Theory: The Data Deluge Makes the Scientific Method ObsoleteChris Anderson wrote, "With enough data, the numbers speak for themselves."

The trouble is that you don't always know when to believe them. When you've got algorithms weighing hundreds of factors over a huge data set, you can't really know why they come to a particular decision or whether it really makes sense

As Geoff Nunberg, who teaches at the School of Information at the University of California Berkeley stated in an NPR interview, big data is no more exact a notion than big hair. Nothing magic happens when you get to the 18th or 19th zero. After all, digital data has been accumulating for decades in quantities that always seemed unimaginably vast at the time.

An exponential curve looks just as overwhelming wherever you get onboard. And anyway, nobody really knows how to quantify this stuff precisely. Whatever the sticklers say, data isn't a plural noun like pebbles. It's a mass noun like dust.

What's new is the way data is generated and processed. It's like dust in that regard, too. We kick up clouds of it wherever we go. Cell phones and cable boxes; Google and Amazon, Facebook and Twitter; the bar codes on milk cartons; and the RFID chip that whips you through the toll plaza - each of them captures a sliver of what we're doing, and nowadays they're all calling home.

It's only when all those little chunks are aggregated that they turn into big data, then the software called analytics can scour it for patterns. Epidemiologists watch for blips in Google queries to localize flu outbreaks. Economists use them to spot shifts in consumer confidence. Police analytics comb over crime data looking for hot zones. 

Big Data and Google Flu Trends

In 2008, Google launched Google Flu Trends, which used big data and search algorithms to estimate the prevalence of flu outbreaks. For the first few years, Google Flu Trends metrics tracked closely with Center for Disease Control (CDC) data - and they were delivered several days before the CDC data.

But for the 2012-13 flu season, a comparison with traditional surveillance data showed that Google Flu Trends, which estimates prevalence from flu-related Internet searches, had drastically overestimated peak flu levels

It is not the first time that a flu season has tripped Google up. In 2009, Flu Trends had to tweak its algorithms after its models badly underestimated ILI in the United States at the start of the H1N1 (swine flu) pandemic

The glitch may be no more than a temporary setback for a promising strategy and Google is sure to refine its algorithms. But as flu-tracking techniques based on mining of web data and on social media proliferate, the episode is a reminder that they will complement, but not substitute for, traditional epidemiological surveillance networks.

This stumble doesn't necessarily make Google Flu Trends irrelevant. But it does mean that Google needs to recalibrate the way they mine big data to track the spread of disease, accounting for searches that may not be linked with infections. "You need to be constantly adapting these models, they don’t work in a vacuum," says Harvard Medical School epidemiologist John Brownstein.

Twitter Flu, Too?

Other researchers are turning to what is probably the largest publicly accessible alternative trove of social-media data: Twitter. Several groups have published work suggesting that models of flu-related tweets can be closely fitted to past official ILI data, and various services, such as MappyHealth and Sickweather, are testing whether real-time analyses of tweets can reliably assess levels of flu.


But Lyn Finelli, head of the CDC’s Influenza Surveillance and Outbreak Response Team, is skeptical. “The Twitter analyses have much less promise” than Google Flu or Flu Near You, she says, arguing that Twitter’s signal-to-noise ratio is very low, and that the most active Twitter users are young adults and so are not representative of the general public.

Further, we know many Twitter accounts are automated response programs called "bots," fake accounts, or "cyborgs" -- human controlled accounts assisted by bots. Recent estimates suggest there could be as many as 20 million fake accounts. So even before we get into the methodological minefield of how you assess health issues on Twitter, let's ask whether those issues are expressed by people or just automated algorithms.

Fooled by Randomness - The Dictatorship of Data

As stated by the authors in Big Data: A Revolution That Will Transform How We Live, Work, and Think:
Correlations let us analyze a phenomenon not by shedding light on its inner workings but by identifying a useful proxy for it. Of course, even strong correlations are never perfect. It is quite possible that two things may behave similarly just by coincidence. We may simply be “fooled by randomness” to borrow a phrase from the empiricist Nassim Nicholas Taleb. With correlations, there is no certainty, only probability.  
Moreover, because big-data analysis is based on theories, we can’t escape them. They shape both our methods and our results. It begins with how we select the data. Our decisions may be driven by convenience: Is the data readily available? Or by economics: Can the data be captured cheaply? Our choices are influenced by theories. What we choose influences what we find, as the digital-technology researchers danah boyd and Kate Crawford have argued. After all, Google used search terms as a proxy for the flu, not the length of people’s hair. Similarly, when we analyze the data, we choose tools that rest on theories. And as we interpret the results we again apply theories. The age of big data clearly is not without theories – they are present throughout, with all that this entails. 
We are more susceptible than we may think to the “dictatorship of data” — that is, to letting the data govern us in ways that may do as much harm as good. The threat is that we will let ourselves be mindlessly bound by the output of our analyses even when we have reasonable grounds for suspecting something is amiss. Or that we will become obsessed with collecting facts and figures for data’s sake. Or that we will attribute a degree of truth to the data which it does not deserve.