Back in 2015, Elsevier Labs adopted Databricks to build NLP algorithms against their content archive. Shared notebooks replaced fragmented individual workflows, letting multiple teams process terabytes of text and petabytes of supplementary data.
Benefits & outcomes
Project timelines dropped from weeks to days.
Team involvement grew from 2–3 specialists to 15+ contributors.
CoreNLP integrated in a day; a concordancer for large-scale regex querying built in under an hour.
Code and data centralised, making R&D outputs visible and reusable internally.
Notes & quotes
“We used to have to do a lot of manual data transfer… Databricks eliminated that.” — Ron Daniel, Labs Director
“I can take one of Elsevier Labs’ existing Databricks notebooks and immediately demonstrate what they were looking for.” — Brad Allen, Chief Architect
Notebooks doubled as demo environments for senior stakeholders — possibly the first time an R&D notebook was boardroom-ready.
Publisher / company: Elsevier / Databricks
Date of mention: 11/12/2015
Source: https://www.databricks.com/blog/2015/11/12/elsevier-labs-deploys-databricks-for-unified-content-analysis.html
