You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
{{ message }}
This repository was archived by the owner on Jan 15, 2024. It is now read-only.
Repository navigation
This repository was archived by the owner on Jan 15, 2024. It is now read-only.
[Sparse Attention][Performance] Accelerate the performance of sparse attention + Benchmark #1397
We are having ongoing efforts about supporting sparse attention in GluonNLP: #1395. To better accelerate related kernels, we can compare the performance of these potential solutions, including:
Use BlockSparse kernel to implement the operator
We may try out these implementations
We are having ongoing efforts about supporting sparse attention in GluonNLP: #1395. To better accelerate related kernels, we can compare the performance of these potential solutions, including:
We may try out these implementations