NVDA 217.50 ▼0.02%GOOGL 343.80 ▼3.84%MSFT 503.81 ▼0.44%AMD 474.32 ▲1.01%INTC 97.71 ▲0.19%TSMC 422.06 ▲0.86%AMZN 272.27 ▼2.09%META 599.12 ▲0.71%AAPL 304.91 ▼1.09%PLTR 174.94 ▼0.17%
Markets at last close

Research

Cornell datasets use clicks to forecast research impact

·1 min read

Cornell researchers curated new arXiv and GitHub datasets to test whether early digital engagement can forecast future scientific and technical impact. The work focuses on lead-lag forecasting, using signals such as downloads, views, likes, pushes and stars to identify contributions that may become important before conventional measures catch up.

The arXiv dataset covers downloads for 2.3 million papers uploaded to the repository. Using established machine learning models, the team found that early downloads could predict popularity at five years with as little as a month of data, offering a faster alternative to citation counts, which can take years to accumulate.

A parallel GitHub dataset compiled information from about 938,000 repositories and showed a similar pattern: early pushes and stars were associated with more forks five years later. The researchers say the datasets could help improve forecasting methods for scientific breakthroughs, software trends and other domains where early attention may signal long-term value.

The team is also exploring whether early downloads can help identify paper quality on arXiv as AI contributes to a flood of low-quality submissions, and whether concurrent reading patterns can reveal emerging fields.

Originally reported by news.cornell.eduRead the source →
Related coverage