Meng, K., Bau, D., Andonian, A. & Belinkov, Y. Locating and editing factual associations in GPT.
Adv. Neural Inf. Process. Syst. (2022). at
<https://proceedings.neurips.cc/paper_files/paper/2022/hash/6f1d43d5a82a37e89b0665b33bf3a182-
Abstract-Conference.html>
26. Meng, K., Sharma, A. S., Andonian, A., Belinkov, Y. & Bau, D. Mass-editing memory in a
transformer. in International Conference on Learning Representations (arxiv.org, 2023). at
<https://arxiv.org/abs/2210.07229>
27. Mitchell, E., Lin, C., Bosselut, A., Manning, C. D. & Finn, C. Memory-Based Model Editing at
Scale. in Proceedings of the 39th International Conference on Machine Learning (eds. Chaudhuri,
K., Jegelka, S., Song, L., Szepesvari, C., Niu, G. & Sabato, S.) 162, 15817–15831 (PMLR, 17--23
Jul 2022).
28. Hartvigsen, T., Sankaranarayanan, S., Palangi, H., Kim, Y. & Ghassemi, M. Aging with GRACE:
Lifelong Model Editing with Discrete Key-Value Adaptors. in Advances in Neural Information
Processing Systems (2023). at <https://arxiv.org/abs/2211.11031>
29. Mitchell, E., Lin, C., Bosselut, A., Finn, C. & Manning, C. Fast model editing at scale. in
International Conference on Learning Representations (arxiv.org, 2022). at
<https://arxiv.org/abs/2110.11309>
30. Sinitsin, A., Plokhotnyuk, V., Pyrkin, D., Popov, S. & Babenko, A. Editable Neural Networks. in
International Conference on Learning Representations (2020). at <http://arxiv.org/abs/2004.00345>
31. De Cao, N., Aziz, W. & Titov, I. Editing Factual Knowledge in Language Models. in Proceedings of
the 2021 Conference on Empirical Methods in Natural Language Processing 6491–6506
(Association for Computational Linguistics, 2021).
32. Zhong, Z., Wu, Z., Manning, C. D., Potts, C. & Chen, D. MQuAKE: Assessing Knowledge Editing
in Language Models via Multi-Hop Questions. arXiv [cs.CL] (2023). at
<http://arxiv.org/abs/2305.14795>
33. Cohen, R., Biran, E., Yoran, O., Globerson, A. & Geva, M. Evaluating the ripple effects of
knowledge editing in language models. Trans. Assoc. Comput. Linguist. 12, 283–298 (2023).
De Cao, N., Aziz, W. & Titov, I. Editing Factual Knowledge in Language Models. arXiv [cs.CL]
(2021). at <http://arxiv.org/abs/2104.08164>
35. Meng, K., Sharma, A. S., Andonian, A., Belinkov, Y. & Bau, D. Mass-Editing Memory in a
Transformer. arXiv [cs.CL] (2022). at <http://arxiv.org/abs/2210.07229>
36. Mitchell, E., Lin, C., Bosselut, A., Finn, C. & Manning, C. D. Fast Model Editing at Scale. arXiv
[cs.LG] (2021). at <http://arxiv.org/abs/2110.11309>
37. Hartvigsen, T., Sankaranarayanan, S., Palangi, H., Kim, Y. & Ghassemi, M. Aging with GRACE:
Lifelong Model Editing with Key-Value Adaptors. (2022). at
<https://openreview.net/pdf?id=ngCT1EelZk>
Language Models: Problems, Methods, and Opportunities. arXiv [cs.CL] (2023). at
<http://arxiv.org/abs/2305.13172>
41. Hase, P., Hofweber, T., Zhou, X., Stengel-Eskin, E. & Bansal, M. Fundamental problems with model
editing: How should rational belief revision work in LLMs? arXiv [cs.CL] (2024). at
<https://scholar.google.com/citations?view_op=view_citation&hl=en&citation_for_view=FO90FgM
AAAAJ:M3ejUd6NZC8C>
42. Cheng, S., Tian, B., Liu, Q., Chen, X., Wang, Y., Chen, H. & Zhang, N. Can We Edit Multimodal
Large Language Models? in Proceedings of the 2023 Conference on Empirical Methods in Natural
Language Processing (eds. Bouamor, H., Pino, J. & Bali, K.) 13877–13888 (Association for
Computational Linguistics, 2023).
Here are the URLs for the specified papers:
1. **Locating and editing factual associations in GPT**
Meng, K., Bau, D., Andonian, A. & Belinkov, Y. (2022).
[Link to Paper](https://proceedings.neurips.cc/paper_files/paper/2022/hash/6f1d43d5a82a37e89b0665b33bf3a182-Abstract-Conference.html)
2. **Mass-editing memory in a transformer**
Meng, K., Sharma, A. S., Andonian, A., Belinkov, Y. & Bau, D. (2023).
[Link to Paper](https://arxiv.org/abs/2210.07229)
3. **Memory-Based Model Editing at Scale**
Mitchell, E., Lin, C., Bosselut, A., Manning, C. D. & Finn, C. (2022).
[Link to Paper](https://proceedings.mlr.press/v162/mitchell22a.html)
4. **Aging with GRACE: Lifelong Model Editing with Discrete Key-Value Adaptors**
Hartvigsen, T., Sankaranarayanan, S., Palangi, H., Kim, Y. & Ghassemi, M. (2023).
[Link to Paper](https://arxiv.org/abs/2211.11031)
5. **Fast model editing at scale**
Mitchell, E., Lin, C., Bosselut, A., Finn, C. & Manning, C. D. (2022).
[Link to Paper](https://arxiv.org/abs/2110.11309)
6. **Editable Neural Networks**
Sinitsin, A., Plokhotnyuk, V., Pyrkin, D., Popov, S. & Babenko, A. (2020).
[Link to Paper](http://arxiv.org/abs/2004.00345)
7. **Editing Factual Knowledge in Language Models**
De Cao, N., Aziz, W. & Titov, I. (2021).
[Link to Paper](http://arxiv.org/abs/2104.08164)
8. **MQuAKE: Assessing Knowledge Editing in Language Models via Multi-Hop Questions**
Zhong, Z., Wu, Z., Manning, C. D., Potts, C. & Chen, D. (2023).
[Link to Paper](http://arxiv.org/abs/2305.14795)
9. **Evaluating the ripple effects of knowledge editing in language models**
Cohen, R., Biran, E., Yoran, O., Globerson, A. & Geva, M. (2023).
[Link to Paper](https://transacl.org/ojs/index.php/tacl/article/view/3736)
10. **Language Models: Problems, Methods, and Opportunities**
(2023).
[Link to Paper](http://arxiv.org/abs/2305.13172)
11. **Fundamental problems with model editing: How should rational belief revision work in LLMs?**
Hase, P., Hofweber, T., Zhou, X., Stengel-Eskin, E. & Bansal, M. (2024).
[Link to Paper](https://scholar.google.com/citations?view_op=view_citation&hl=en&citation_for_view=FO90FgMAAAAAJ:M3ejUd6NZC8C)
12. **Can We Edit Multimodal Large Language Models?**
Cheng, S., Tian, B., Liu, Q., Chen, X., Wang, Y., Chen, H. & Zhang, N. (2023).
[Link to Paper](https://arxiv.org/abs/2305.14795)
Citations:
[1] https://proceedings.neurips.cc/paper_files/paper/2022/hash/6f1d43d5a82a37e89b
Recent work on representing “Feynman diagrams as computational graphs” has sparked an intriguing idea: Let’s map AI computation to Feynman diagrams to visualize and optimize AI architectures.
💡 By leveraging Meta’s LLM Compiler, we can create a powerful interpreter between quantum field theory techniques and AI model design.
𝐇𝐞𝐫𝐞'𝐬 𝐡𝐨𝐰 𝐢𝐭 𝐰𝐨𝐫𝐤𝐬:
1. Represent AI models as Feynman-like diagrams, with nodes as computation units (e.g., transformer blocks) and edges showing data flow.
2. Use the LLM Compiler to analyze these diagrams, suggesting optimizations based on both structure and underlying computations.
3. Instead of integrating traditional LLVMs we swap it out for Meta’s LLM compiler for a multi-level optimization approach:
- 𝐇𝐢𝐠𝐡-𝐥𝐞𝐯𝐞𝐥: LLM-driven architectural changes
- 𝐌𝐢𝐝-𝐥𝐞𝐯𝐞𝐥: Standard compiler optimizations
- 𝐋𝐨𝐰-𝐥𝐞𝐯𝐞𝐥: Hardware-specific tweaks
𝐓𝐡𝐢𝐬 𝐚𝐩𝐩𝐫𝐨𝐚𝐜𝐡 𝐨𝐟𝐟𝐞𝐫𝐬 𝐬𝐞𝐯𝐞𝐫𝐚𝐥 𝐤𝐞𝐲 𝐚𝐝𝐯𝐚𝐧𝐭𝐚𝐠𝐞𝐬:
1. 𝐄𝐧𝐡𝐚𝐧𝐜𝐞𝐝 𝐢𝐧𝐭𝐞𝐫𝐩𝐫𝐞𝐭𝐚𝐛𝐢𝐥𝐢𝐭𝐲: Feynman diagrams provide a visual language for complex AI systems, crucial for debugging and regulatory compliance.
2. 𝐂𝐫𝐨𝐬𝐬-𝐝𝐨𝐦𝐚𝐢𝐧 𝐢𝐧𝐬𝐢𝐠𝐡𝐭𝐬: The LLM's capabilities to compile and optimize models inspired by QFT principles.
3. 𝐇𝐚𝐫𝐝𝐰𝐚𝐫𝐞-𝐚𝐰𝐚𝐫𝐞 𝐝𝐞𝐬𝐢𝐠𝐧: Optimizations can be tailored to specific GPU or TPU architectures, improving efficiency.
4. 𝐈𝐭𝐞𝐫𝐚𝐭𝐢𝐯𝐞 𝐫𝐞𝐟𝐢𝐧𝐞𝐦𝐞𝐧𝐭: Continuous learning from optimization patterns leads to increasingly sophisticated improvements over time.
Of course, there are challenges. Representing very deep networks or handling the complexity of recurrent connections could be tricky. But I believe the potential benefits outweigh these hurdles.
💡 Now, here's where we can take it to the next level: Combine this Feynman diagram approach with LLM-based intelligent optimization, like Meta's LLM Compiler. We could create a powerful system where both human designers and AI systems work with the same visual language.
🪄 Imagine an LLM analyzing these AI Feynman diagrams, suggesting optimizations, and even generating or modifying code directly. This could bridge the gap between high-level model architecture and low-level implementation details, potentially leading to more efficient and interpretable AI systems.
This approach could be particularly powerful in domains like hashtag#explainableAI and hashtag#AIsafety, where understanding the decision-making process is crucial.
I'm incredibly excited about this direction. It could be a major leap towards more intuitive and powerful ways of developing AI, bringing together experts from physics, AI, and visual design.