Introducing the BHL Platform Modernization Project — Survey Results & What's Next

Hi everyone,

My name is Rahul Sharma, and I’m writing to formally introduce myself to the community as the Project Manager & Cloud Engineer for the BHL Platform Modernization initiative, authorized by the BHL Executive Committee. Many of you already know BHL well as users, contributors, or partners. I wanted to use this post to explain what this project is, why we’re doing it now, and to share the first results from our recent user survey, which will directly shape the work ahead.

Why we’re modernizing the platform

BHL is the largest open-access provider of biodiversity literature and archival material in the world serving users in over 195 countries, connecting 660+ contributing institutions, and hosting more than 64 million pages spanning five centuries of biodiversity knowledge. Our public APIs alone receive over 30 million calls a year.

That scale is something we’re proud of, but it also means our legacy infrastructure is under real strain. Over the years, the community has submitted more than 10,300 issues and feature requests and 140+ published user stories, consistently asking for a platform that’s more accessible, more discoverable, and more interoperable with the wider biodiversity data ecosystem.

This project is our response to that feedback. At a high level, we’re aiming to:

  • Modernize BHL’s infrastructure and reduce our dependence on legacy systems
  • Build a standards-aligned, interoperable data infrastructure that connects cleanly with platforms like GBIF, Catalogue of Life, and Wikimedia
  • Maintain full, uninterrupted access to BHL content throughout the transition.
  • Improve discoverability and reuse of BHL’s content, both within BHL and across the open web
  • Strengthen platform performance, security, and reliability
  • Keep BHL responsive to community needs and open to community contribution, now and going forward.

This project is meant to position BHL as a stronger, more interoperable digital commons for the whole biodiversity informatics community.

Why we ran the survey?

Before locking in our technical direction, we wanted to hear directly from you, the people who actually use BHL day to day rather than relying only on historical tickets and internal input. The survey opened on July 16, 2026, and ran through July 27, 2026, and is a key input into the Project Plan we’re now developing.

What we heard (as of July 27, 2026)

We had 414 responses, and a few patterns came through clearly:

  • Who responded: Our respondents are overwhelmingly long-time users, 64.5% have used BHL for 5+ years, and another 18.8% for 2–5 years. Nearly half (48.1%) use BHL primarily for scientific research or taxonomy, with library/archive/museum work (18.8%) and art, writing, and humanities research (14%) also well represented.
  • Engagement: Almost half of respondents visit BHL weekly or more often (31.4% weekly, 15.2% daily or almost daily), with 38.6% visiting monthly.
  • Top priority, by far was “search”. When asked how important various potential improvements were, Enhanced Text Search came out clearly on top, with about 79% of respondents calling it “Very Important.” Image Search and an Improved Page Viewer followed as strong secondary priorities. This held true across both our “Scientific” and “Library” user segments with one notable exception: Image Search saw a real split, rated “Very Important” by nearly 70% of Library users but only about 40% of Scientific users. Every other priority area was closely aligned across both groups.

Open-ended feedback: When we asked about challenges, the single largest category was search functionality and relevance consistent with the ratings data above. On the flip side, when asked for general comments, the largest category by far was general praise and gratitude, which was genuinely great to see. We also picked up a smaller but notable thread of mixed opinions specifically about AI/ML tools some respondents are eager to see them, others are more cautious which tells us we’ll need to communicate carefully as that work develops.

How this fits into the bigger picture?

These results reinforce most of what’s already reflected in the Project Charter’s high-level requirements particularly around enhanced search, improved discoverability, and OCR/HTR quality and give us real, current data to prioritize against as we move into detailed planning. They also surface a few things worth watching closely as we go: making sure mobile and API work stay properly resourced even though they didn’t top the list, and building a clear, honest communications plan around AI/ML given the range of opinions in the community.

We’re currently in the early planning phase (Months 0–4 per our project timeline) and this planning is very crucial because this plans willcarry this project through development, testing, and eventual migration all while keeping the existing platform fully accessible throughout.

Thank you, and over to you

Thank you to everyone who took the time to respond to the survey. Your input is genuinely shaping the direction of this project. I also want to thank the wider BHL community for your patience and continued engagement as we are taking on a project of this scale.

I’d love to open the floor to this community now: What questions do you have about the modernization project? And if you weren’t able to take the survey, I’d still welcome your thoughts here in the thread.

Looking forward to the discussion.

Best,
Rahul Sharma
Project Manager, BHL Platform Modernization

If you’d like to explore the data yourself, the raw survey responses, along with a detailed PDF summary and a high-level PowerPoint overview, are linked below.

(Please note the PPT contains data until Aug 19, 2026) after which we have officially stopped accepting any further responses.)

Raw User Survey Responses: https://shorturl.at/B36jc
PDF Summary: https://shorturl.at/vgpIB
PPT Summary: https://shorturl.at/JnlIh

4 Likes

Thanks for posting this @rahul, nice to see the results of the survey and also the raw data. Lots to explore here, but yes, the main signal is make search better. Of course, defining “better” is a challenge, but I think there are demonstrable problems with the current search, and there are also some cool new apporaches to search (e.g., vector-based search for images and text) that open up all sorts of possibilities for BHL.

I took the survey results and ran them through Claude (as did Mike). Here’s a sumamry of what users want improved:

Claude also took a stab at aassigning reponses to user-communities:

Everyone wants better text search, image search was a bigger priority if you weren’t a scientist (art and humanities ranked image search higher than text search).

The most arm-wavvy chart is “intensity” versus “reach”, which is basically how many people care about something versus how much they care. High and to the left is basically a small group of people really upset about something, especially those who can’t actually get into BHL. Search has big reach (everybody cares) but nobody is screaming about it, which I would interpret as everybody thinks it’s a no-brainer (hence they don’t need to scream about it).

I also asked Claude for a text-based summary, which it delivered below, with a fauir bit of “AI speak” (e.g., “this is the arbitrage”):

Where the biggest win is in the survey

“Better search” was the top request — but on its own it isn’t actionable. Reading what people actually described, it splits into seven distinct asks with very different price tags:

Expensive, but highest leverage

• Article-level indexing — find the paper inside the volume (~25 mentions)
• OCR quality — the substrate under every text search (~20)

Medium

• Scope a search to one journal / title / item; search within results (~20)
• Author disambiguation + duplicate creators (~15)
• Fuzzy matching for title and spelling variants (~10)

Cheap — and this is the arbitrage

• Highlight search terms on the page, and carry the query into “search inside” (~8)
• Combinable filters (date + language + subject) and sort by relevance (~8)

The cheap ones are worth doing first. Multiple people described the same experience: search, click a result, land on a page with nothing highlighted — and several named Internet Archive and HathiTrust as getting it right. That’s a viewer change, not a search-engine rebuild, and it turns “returns a page” into “answers the question.”

The biggest structural win is still article-level metadata. It’s the root of “having to go through whole journals to find an article”, the duplicate-author noise, and the DOI requests from ZooBank, WoRMS, the Reptile Database and Kew/IPNI — all at once.