Tracking, Valuing and Evaluating the Diversity of HSS Research Outputs

This was originally a script for a video I recorded for a meeting of the Australian Deans and Directors of Creative Arts (DDCA) group. It is addressed to those who are in institutional positions who need to make judgements about the qualities of research outputs in Humanities and Social Sciences with a particular focus on outputs that are “non-traditional” (actually often very “traditional” in their field!) or otherwise “not easily measurable by a STEM audience”.

In my view there are two key roles for those specifically engaged in supporting those doing HSS research. The first is to advocate for financial systems that can support small and numerous research projects and the funders who want to support them. The second key role is to defend the diversity of research practices, outputs and impacts in HSS and to create a meaningful narrative around their qualities for institutional audiences.

I’m going to focus on the second here, but I wanted to link both to underline the point that what is of value in HSS research is precisely its diversity, and its difference from the homogeneity and scale of conventional STEM research and its outputs. There is not infrequently a greater diversity of research outputs from single HSS groups, centres and departments than there is from entire faculties of science, engineering and medicine.

That diversity is at the core of the value that HSS research and scholarship brings to the table. 

Every social epistemology of knowledge from Latour to Longino via Merton, Ravetz, Kuhn and Fleck have at their heart the idea that the core is the negotiation of knowledge across boundaries, or between groups. This places diversity, of all kinds, at the heart of this process, and linked to it the institutional systems that keep it alive. Arguably it is the purpose of a university to support and bring into contact diverse approaches, experiences and ways of knowing.

The issue of course is that managing that diversity is a challenge. At the scale of a university, systems are a necessity. Consistency is needed and economies of scale are real. Particularly in HSS but across other disciplines as well, it is impossible for any leader to be across the detail of how good every specific body of work is. If you believe in the argument for diversity then this as it should be – anything else would be a sign of insufficient diversity. That doesn’t mean that leaders can’t be effective at telling the stories about why diverse research matters. Nor does it mean that they can’t make judgement calls on the distribution of resources. But it means those decisions are subjective and contextual. And that the evidence that supports them will be specific to each case.

But that doesn’t solve the problem either. So what are the pragmatic and principled approaches available to addressing this complexity? In my view the key lies in separating out the challenge into two parts, one relatively straightforward and immediately useful, the other more complex but I would argue ultimately tractable.

An important, but oddly overlooked distinction in research evaluation and particularly technology discussions is the distinction between tracking outputs and evaluating them. There is often a tendency to lump these together and to imagine for example, that if scholarly books were better tracked then that would necessarily mean there would be improved citation data for them. The homogeneity of STEM research outputs means that issues with data coverage and completeness are often conflated to the question “is it indexed” using the presence of a specific journal in a specific database both as a reliable marker of quality itself but also of confidence in the quality and relevance of generally accepted performance data.

So, although this is an oversimplification, in STEM the tracking of outputs is standardised and homogenous. You look it up in an index like Web of Science.

In HSS, and particularly for creative works, the situation is far more complex. Institutions generally have very poor data on the scope and volume of creative outputs, or other forms of (so-called) “non-traditional” outputs. Work done by Niamh Quigley, within the COKI team (Quigley, 2022) showed generally used recording systems (systems like Symplectic, Pure and the like) tended to be poor at capturing aspects of research outputs important to creative researchers. Further, there was a tendency for the information that was manually inputted by researchers (that they care enough to type in!) to be stripped away as it made its way through institutional systems to the data warehouses used for high level analysis.

From a systems perspective what is needed is systems that make it easy for creative practice researchers to register outputs, to flexibly record the evidence that they consider important about the qualities of these works, and to also make it easy for institutions in turn to gather that information up. Recent technical developments provide improved paths to making this happen.

The Open Research Contributor and ID (ORCID) system provides a mechanism for researchers to authoritatively identify themselves, and to record outputs they have contributed to. In 2025 ORCID have increased the number of work types (Petro, 2025) to include a range relevant in humanities (including images, video, designs, sound and other creative works) as well as improving the mapping of some other relevant work types including reports. A researcher can choose what works they make visible, and thereby designate as “research”. This is an important distinction for some creative workers. Institutions can then harvest those outputs without requiring further work from researchers. 

ORCID doesn’t solve all the problems though. Badly implemented as an institutional system it is yet another place creative researchers have to battle through manually adding their work while colleagues in STEM subjects have it all done for them by journal publishers. Critical to making this work is a commitment from you and your teams that ORCID be the only place researchers have to make sure their work is available and pulling that information from ORCID is an institutional responsibility.

But that still leaves the manual input issue. There isn’t – currently – a complete solution for that (although one – or several – could be built) but there is at least a solution that offers some additional value for creative researchers. Zenodo is a large scale repository, managed by CERN (yes the particle accelerator in Switzerland CERN) that offers a place for anyone to deposit research outputs. It provides a safe, publicly accessible platform from which researchers can showcase their work. It even allows for dark deposit for those creative researchers who need to restrict access to the core of their work for artistic, ethical or commercial reasons.

Most importantly Zenodo provides DataCite DOIs and this means that – while you do have to manually put information into Zenodo – that information can automatically flow through to ORCID. One place to input the information and then it can flow to the other places it needs to be.

This flow of information should be the case for any repository that integrates with DataCite (or Crossref) to provide DOIs. Zenodo is large scale and flexible, but there can be an argument for specialist repositories. I am generally sceptical of the value of creating new local repositories, whether at institutions or nationally, but the key questions to ask of any repository are “does it integrate with DOI providers” and “does it provide the recording functionality creative research practitioners need”? That flexibility is key and Zenodo has some strengths here that can also help us to address, if not yet fully solve, the second problem. 

Creative practice researchers know the reasons that their work is important, and can help provide evidence of that. But that evidence is highly diverse. 

I know the pressures that institutional leaders face to articulate a comparative quantitative value of research outputs to the rest of the university. And I would argue that we nonetheless need to collectively hold the line that this makes no sense. If you value a diversity of research and researchers, you can’t protect that diversity by homogenising the way that value is measured. 

But that doesn’t mean we can’t articulate value. It doesn’t mean we can’t compare how disparate research projects and outputs achieve specific goals. It just means, in the best tradition of HSS scholarship, that the evidence that we use is contextual. That it is diverse.

And here is how repositories like Zenodo can help. Where they support the creation of packages of work, researchers can bundle up the work itself with evidence of its value and impact. This will start unstructured, but over time, and in collaboration with research communities it will be possible to develop some standards. For specific cases. Where appropriate. This might mean giving some background on the gallery that invited an exhibition, the footfall or an audience survey. It might be details of the film festival that selected a documentary for screening, or audience numbers when it was syndicated or broadcast. It might be structured in a particular way to make it easier to process (or even generate). But the key is to build on systems that always allow for the flexibility of a narrative statement that allows for the flexibility that research statements already provide. 

This workflow places researchers in control, gives them flexibility they need to provide access and to evidence the value of their work. It provides a single point of entry from which communities, institutions, and others can draw in the data to help track those outputs and collate the evidence of their value. It is not controlled by corporate interests but by academic communities. It is not subject to the funding whims of one national government but an international set of consortia.

Perhaps best of all, this is an area where the humanities and social sciences can lead. It solves a problem by putting researchers in control while supporting institutional needs. It helps and simplifies monitoring while offering a path to gathering more sophisticated evidence to tell better stories. These are not problems that are restricted to HSS, but you could argue that much of STEM evaluation has lost its way, focusing on numbers instead of substance, journal lists instead of journal content, impact factors instead of actual impact. This recontextualisation, and re-situation is where humanities and social sciences scholars, and leadership can show a way to better forms of evaluation.

References

(S)low impact research and the importance of open in maximising re-use

Open
Image by tribalicious via Flickr

This is an edited version of the text that I spoke from at the Altmetrics Workshop in Koblenz in June. There is also an audio recording of the talk I gave available as well as the submitted abstract for the workshop.

I developed an interest in research evaluation as an advocate of open research process. It is clear that researchers are not going to change themselves so someone is going to have to change them and it is funders who wield the biggest stick. The only question, I thought,  was how to persuade them to use it

Of course it’s not that simple. It turns out that funders are highly constrained as well. They can lead from the front but not too far out in front if they want to retain the confidence of their community. And the actual decision making processes remain dominated by senior researchers. Successful senior researchers with little interest in rocking the boat too much.

The thing you realize as you dig deeper into this as that the key lies in finding motivations that work across the interests of different stakeholders. The challenge lies in finding the shared objectives. What it is that unites both researchers and funders, as well as government and the wider community. So what can we find that is shared?

I’d like to suggest that one answer to that is Impact. The research community as a whole has stake in convincing government that research funding is well invested. Government also has a stake in understanding how to maximize the return on its investment. Researchers do want to make a difference, even if that difference is a long way off. You need a scattergun approach to get the big results, but that means supporting a diverse range of research in the knowledge that some of it will go nowhere but some of it will pay off.

Impact has a bad name but if we step aside from the gut reactions and look at what we actually want out of research then we start to see a need to raise some challenging questions. What is research for?  What is its role in our society really? What outcomes would we like to see from it, and over what timeframes? What would we want to evaluate those outcomes against? Economic impact yes, as well as social, health, policy, and environmental impact. This is called the ‘triple bottom line’ in Australia. But alongside these there is also research impact.

All these have something in common. Re-use. What we mean by impact is re-use. Re-use in industry, re-use in public health and education, re-use in policy development and enactment, and re-use in research.

And this frame brings some interesting possibilities. We can measure some types of re-use. Citation, retweets, re-use of data or materials, or methods or software. We can think about gathering evidence of other types of re-use, and of improving the systems that acknowledge re-use. If we can expand the culture of citation and linking to new objects and new forms of re-use, particularly for objects on the web, where there is some good low hanging fruit, then we can gather a much stronger and more comprehensive evidence base to support all sorts of decision making.

There are also problems and challenges. The same ones that any social metrics bring. Concentration and community effects, the Matthew effect of the rich getting richer. We need to understand these feedback effects much better and I am very glad there are significant projects addressing this.

But there is also something more compelling for me in this view. It let’s us reframe the debate around basic research. The argument goes we need basic research to support future breakthroughs. We know neither what we will need nor where it will come from. But we know that its very hard to predict – that’s why we support curiosity driven research as an important part of the portfolio of projects. Yet the dissemination of this investment in the future is amongst the weakest in our research portfolio. At best a few papers are released then hidden in journals that most of the world has no access to and in many cases without the data, or other products either being indexed or even made available. And this lack of effective dissemination is often because the work is perceived as low, or perhaps better, slow impact.

We may not be able to demonstrate or to measure significant re-use of the outputs of this research for many years. But what we can do is focus on optimizing the capacity, the potential, for future exploitation. Where we can’t demonstrate re-use and impact we should demand that researchers demonstrate that they have optimized their outputs to enable future re-use and impact.

And this brings me full circle. My belief is that the way to ensure the best opportunities for downstream re-use, over all timeframes, is that the research outputs are open, in the Budapest Declaration sense. But we don’t have to take my word for it, we can gather evidence. Making everything naively open will not always be the best answer, but we need to understand where that is and how best to deal with it. We need to gather evidence of re-use over time to understand how to optimize our outputs to maximize their impact.

But if we choose to value re-use, to value the downstream impact that our research or have, or could have, then we can make this debate not about politics or ideology but how about how best to take the public investment in research and to invest it for the outcomes that we need as a society.

 

 

 

 

Enhanced by Zemanta

Evidence to the European Commission Hearing on Access to Scientific Information

European Commission
Image by tiseb via Flickr

On Monday 30 May I gave evidence at a European Commission hearing on Access to Scientific Information. This is the text that I spoke from. Just to re-inforce my usual disclaimer I was not speaking on behalf of my employer but as an independent researcher.

We live in a world where there is more information available at the tips of our fingers than even existed 10 or 20 years ago. Much of what we use to evaluate research today was built in a world where the underlying data was difficult and expensive to collect. Companies were built, massive data sets collected and curated and our whole edifice of reputation building and assessment grew up based on what was available. As the systems became more sophisticated new measures became incorporated but the fundamental basis of our systems weren’t questioned. Somewhere along the line we forgot that we had never actually been measuring what mattered, just what we could.

Today we can track, measure, and aggregate much more, and much more detailed information. It’s not just that we can ask how much a dataset is being downloaded but that we can ask who is downloading it, academics or school children, and more, we can ask who was the person who wrote the blog post or posted it to Facebook that led to that spike in downloads.

This is technically feasible today. And make no mistake it will happen. And this provides enormous potential benefits. But in my view it should also give us pause. It gives us a real opportunity to ask why it is that we are measuring these things. The richness of the answers available to us means we should spend some time working out what the right questions are.

There are many reasons for evaluating research and researchers. I want to touch on just three. The first is researchers evaluating themselves against their peers. While this is informed by data it will always be highly subjective and vary discipline by discipline. It is worthy of study but not I think something that is subject to policy interventions.

The second area is in attempting to make objective decisions about the distribution of research resources. This is clearly a contentious issue. Formulaic approaches can be made more transparent and less easy to legal attack but are relatively easy to game. A deeper challenge is that by their nature all metrics are backwards looking. They can only report on things that have happened. Indicators are generally lagging (true of most of the measures in wide current use) but what we need are leading indicators. It is likely that human opinion will continue to beat naive metrics in this area for some time.

Finally there is the question of using evidence to design the optimal architecture for the whole research enterprise. Evidence based policy making in research policy has historically been sadly lacking. We have an opportunity to change that through building a strong, transparent, and useful evidence base but only if we simultaneously work to understand the social context of that evidence. How does collecting information change researcher behavior? How are these measures gamed? What outcomes are important? How does all of this differ cross national and disciplinary boundaries, or amongst age groups?

It is my belief, shared with many that will speak today, that open approaches will lead to faster, more efficient, and more cost effective research. Other groups and organizations have concerns around business models, quality assurance, and sustainability of these newer approaches. We don’t need to argue about this in a vacuum. We can collect evidence, debate what the most important measures are, and come to an informed and nuanced inclusion based on real data and real understanding.

To do this we need to take action in a number areas:

1. We need data on evaluation and we need to able to share it.

Research organizations must be encouraged to maintain records of the downstream usage of their published artifacts. Where there is a mandate for data availability this should include mandated public access to data on usage.

The commission and national funders should clearly articulate that that provision of usage data is a key service for publishers of articles, data, and software to provide, and that where a direct payment is made for publication provision for such data should be included. Such data must be technically and legally reusable.

The commission and national funders should support work towards standardizing vocabularies and formats for this data as well critiquing it’s quality and usefulness. This work will necessarily be diverse with disciplinary, national, and object type differences but there is value in coordinating actions. At a recent workshop where funders, service providers, developers and researchers convened we made significant progress towards agreeing routes towards standardization of the vocabularies to describe research outputs.

2. We need to integrate our systems of recognition and attribution into the way the web works through identifying research objects and linking them together in standard ways.

The effectiveness of the web lies in its framework of addressable items connected by links. Researchers have a strong culture of making links and recognizing contributions through attribution and citation of scholarly articles and books but this has only recently being surfaced in a way that consumer web tools can view and use. And practice is patchy and inconsistent for new forms of scholarly output such as data, software and online writing.

The commission should support efforts to open up scholarly bibliography to the mechanics of the web through policy and technical actions. The recent Hargreaves report explicitly notes limitations on text mining and information retrieval as an area where the EU should act to modernize copyright law.

The commission should act to support efforts to develop and gain wide community support for unique identifiers for research outputs, and for researchers. Again these efforts are diverse and it will be community adoption which determines their usefulness but coordination and communication actions will be useful here. Where there is critical mass, such as may be the case for ORCID and DataCite, this crucial cultural infrastructure should merit direct support.

Similarly the commission should support actions to develop standardized expressions of links, through developing citation and linking standards for scholarly material. Again the work of DataCite, CoData, Dryad and other initiatives as well as technical standards development is crucial here.

3. Finally we must closely study the context in which our data collection and indicator assessment develops. Social systems cannot be measured without perturbing them and we can do no good with data or evidence if we do not understand and respect both the systems being measured and the effects of implementing any policy decision.

We need to understand the measures we might develop, what forms of evaluation they are useful for and how change can be effected where appropriate. This will require significant work as well as an appreciation of the close coupling of the whole system.
We have a generational opportunity to make our research infrastructure better through effective evaluation and evidence based policy making and architecture development. But we will squander this opportunity if we either take a utopian view of what might technically feasible, or fail to act for a fear of a dystopian future. The way to approach this is through a careful, timely, transparent and thoughtful approach to understanding ourselves and the system we work within.

The commission should act to ensure that current nascent efforts work efficiently towards delivering the technical, cultural, and legal infrastructure that will support an informed debate through a combination of communication, coordination, and policy actions.

Enhanced by Zemanta