NVDA 224.09 ▲3.03%GOOGL 343.54 ▼0.08%MSFT 492.43 ▼2.26%AMD 482.93 ▲1.82%INTC 100.95 ▲3.32%TSMC 429.15 ▲1.68%AMZN 267.28 ▼1.83%META 578.85 ▼3.38%AAPL 302.25 ▼0.87%PLTR 171.04 ▼2.23%
Markets at last close

Research

Cornell datasets use clicks to forecast research impact

·1 min read

Cornell researchers curated new arXiv and GitHub datasets to test whether early digital engagement can forecast future scientific and technical impact. The work focuses on lead-lag forecasting, using signals such as downloads, views, likes, pushes and stars to identify contributions that may become important before conventional measures catch up.

The arXiv dataset covers downloads for 2.3 million papers uploaded to the repository. Using established machine learning models, the team found that early downloads could predict popularity at five years with as little as a month of data, offering a faster alternative to citation counts, which can take years to accumulate.

A parallel GitHub dataset compiled information from about 938,000 repositories and showed a similar pattern: early pushes and stars were associated with more forks five years later. The researchers say the datasets could help improve forecasting methods for scientific breakthroughs, software trends and other domains where early attention may signal long-term value.

The team is also exploring whether early downloads can help identify paper quality on arXiv as AI contributes to a flood of low-quality submissions, and whether concurrent reading patterns can reveal emerging fields.

Originally reported by news.cornell.eduRead the source →
Related coverage