solve the pdf concate error

Merge branch 'frontier' of github.com:binary-husky/chatgpt_academic into frontier
remove logging extra
2025-12-06 22:46:48 +00:00 · 2024-10-13 07:36:36 +00:00 · 2024-10-01 11:59:14 +00:00 · 2024-10-01 11:57:47 +00:00 · 2024-09-28 18:05:34 +08:00 · 2024-09-23 15:16:13 +00:00
--- a/.github/workflows/build-with-all-capacity-beta.yml
+++ b/.github/workflows/build-with-all-capacity-beta.yml
@@ -1,14 +1,14 @@
 # https://docs.github.com/en/actions/publishing-packages/publishing-docker-images#publishing-images-to-github-packages
-name: build-with-latex-arm
+name: build-with-all-capacity-beta

 on:
  push:
    branches:
-      - "master"
+      - 'master'

 env:
  REGISTRY: ghcr.io
-  IMAGE_NAME: ${{ github.repository }}_with_latex_arm
+  IMAGE_NAME: ${{ github.repository }}_with_all_capacity_beta

 jobs:
  build-and-push-image:
@@ -18,17 +18,11 @@ jobs:
      packages: write

    steps:
-      - name: Set up QEMU
-        uses: docker/setup-qemu-action@v3
-
-      - name: Set up Docker Buildx
-        uses: docker/setup-buildx-action@v3
-
      - name: Checkout repository
-        uses: actions/checkout@v4
+        uses: actions/checkout@v3

      - name: Log in to the Container registry
-        uses: docker/login-action@v3
+        uses: docker/login-action@v2
        with:
          registry: ${{ env.REGISTRY }}
          username: ${{ github.actor }}
@@ -41,11 +35,10 @@ jobs:
          images: ${{ env.REGISTRY }}/${{ env.IMAGE_NAME }}

      - name: Build and push Docker image
-        uses: docker/build-push-action@v6
+        uses: docker/build-push-action@v4
        with:
          context: .
          push: true
-          platforms: linux/arm64
-          file: docs/GithubAction+NoLocal+Latex+Arm
+          file: docs/GithubAction+AllCapacityBeta
          tags: ${{ steps.meta.outputs.tags }}
          labels: ${{ steps.meta.outputs.labels }}
--- a/.github/workflows/build-with-jittorllms.yml
+++ b/.github/workflows/build-with-jittorllms.yml
@@ -0,0 +1,44 @@
+# https://docs.github.com/en/actions/publishing-packages/publishing-docker-images#publishing-images-to-github-packages
+name: build-with-jittorllms
+
+on:
+  push:
+    branches:
+      - 'master'
+
+env:
+  REGISTRY: ghcr.io
+  IMAGE_NAME: ${{ github.repository }}_jittorllms
+
+jobs:
+  build-and-push-image:
+    runs-on: ubuntu-latest
+    permissions:
+      contents: read
+      packages: write
+
+    steps:
+      - name: Checkout repository
+        uses: actions/checkout@v3
+
+      - name: Log in to the Container registry
+        uses: docker/login-action@v2
+        with:
+          registry: ${{ env.REGISTRY }}
+          username: ${{ github.actor }}
+          password: ${{ secrets.GITHUB_TOKEN }}
+
+      - name: Extract metadata (tags, labels) for Docker
+        id: meta
+        uses: docker/metadata-action@v4
+        with:
+          images: ${{ env.REGISTRY }}/${{ env.IMAGE_NAME }}
+
+      - name: Build and push Docker image
+        uses: docker/build-push-action@v4
+        with:
+          context: .
+          push: true
+          file: docs/GithubAction+JittorLLMs
+          tags: ${{ steps.meta.outputs.tags }}
+          labels: ${{ steps.meta.outputs.labels }}
--- a/README.md
+++ b/README.md
@@ -1,6 +1,5 @@
 > [!IMPORTANT]
-> 2024.10.10: 突发停电，紧急恢复了提供[whl包](https://drive.google.com/file/d/19U_hsLoMrjOlQSzYS3pzWX9fTzyusArP/view?usp=sharing)的文件服务器  
-> 2024.10.8: 版本3.90加入对llama-index的初步支持，版本3.80加入插件二级菜单功能（详见wiki）  
+> 2024.6.1: 版本3.80加入插件二级菜单功能（详见wiki）  
 > 2024.5.1: 加入Doc2x翻译PDF论文的功能，[查看详情](https://github.com/binary-husky/gpt_academic/wiki/Doc2x)  
 > 2024.3.11: 全力支持Qwen、GLM、DeepseekCoder等中文大语言模型！ SoVits语音克隆模块，[查看详情](https://www.bilibili.com/video/BV1Rp421S7tF/) 
 > 2024.1.17: 安装依赖时，请选择`requirements.txt`中**指定的版本**。 安装命令：`pip install -r requirements.txt`。本项目完全开源免费，您可通过订阅[在线服务](https://github.com/binary-husky/gpt_academic/wiki/online)的方式鼓励本项目的发展。
--- a/crazy_functional.py
+++ b/crazy_functional.py
@@ -6,6 +6,7 @@ from loguru import logger
 def get_crazy_functions():
    from crazy_functions.读文章写摘要 import 读文章写摘要
    from crazy_functions.生成函数注释 import 批量生成函数注释
+    from crazy_functions.Rag_Interface import Rag问答
    from crazy_functions.SourceCode_Analyse import 解析项目本身
    from crazy_functions.SourceCode_Analyse import 解析一个Python项目
    from crazy_functions.SourceCode_Analyse import 解析一个Matlab项目
@@ -51,6 +52,13 @@ def get_crazy_functions():
    from crazy_functions.SourceCode_Comment import 注释Python项目

    function_plugins = {
+        "Rag智能召回": {
+            "Group": "对话",
+            "Color": "stop",
+            "AsButton": False,
+            "Info": "将问答数据记录到向量库中，作为长期参考。",
+            "Function": HotReload(Rag问答),
+        },
        "虚空终端": {
            "Group": "对话|编程|学术|智能体",
            "Color": "stop",
@@ -699,31 +707,6 @@ def get_crazy_functions():
        logger.error(trimmed_format_exc())
        logger.error("Load function plugin failed")

-    try:
-        from crazy_functions.Rag_Interface import Rag问答
-
-        function_plugins.update(
-            {
-                "Rag智能召回": {
-                    "Group": "对话",
-                    "Color": "stop",
-                    "AsButton": False,
-                    "Info": "将问答数据记录到向量库中，作为长期参考。",
-                    "Function": HotReload(Rag问答),
-                },
-            }
-        )
-    except:
-        logger.error(trimmed_format_exc())
-        logger.error("Load function plugin failed")
-
-
-    
-
-
-
-
-
    # try:
    #     from crazy_functions.高级功能函数模板 import 测试图表渲染
    #     function_plugins.update({
--- a/crazy_functions/Latex_Function.py
+++ b/crazy_functions/Latex_Function.py
@@ -138,43 +138,25 @@ def arxiv_download(chatbot, history, txt, allow_cache=True):
    cached_translation_pdf = check_cached_translation_pdf(arxiv_id)
    if cached_translation_pdf and allow_cache: return cached_translation_pdf, arxiv_id

-    extract_dst = pj(ARXIV_CACHE_DIR, arxiv_id, 'extract')
+    url_tar = url_.replace('/abs/', '/e-print/')
    translation_dir = pj(ARXIV_CACHE_DIR, arxiv_id, 'e-print')
-    dst = pj(translation_dir, arxiv_id + '.tar')
+    extract_dst = pj(ARXIV_CACHE_DIR, arxiv_id, 'extract')
    os.makedirs(translation_dir, exist_ok=True)
-    # <-------------- download arxiv source file ------------->

-    def fix_url_and_download():
-        # for url_tar in [url_.replace('/abs/', '/e-print/'), url_.replace('/abs/', '/src/')]:
-        for url_tar in [url_.replace('/abs/', '/src/'), url_.replace('/abs/', '/e-print/')]:
+    # <-------------- download arxiv source file ------------->
+    dst = pj(translation_dir, arxiv_id + '.tar')
+    if os.path.exists(dst):
+        yield from update_ui_lastest_msg("调用缓存", chatbot=chatbot, history=history)  # 刷新界面
+    else:
+        yield from update_ui_lastest_msg("开始下载", chatbot=chatbot, history=history)  # 刷新界面
        proxies = get_conf('proxies')
        r = requests.get(url_tar, proxies=proxies)
-            if r.status_code == 200:
        with open(dst, 'wb+') as f:
            f.write(r.content)
-                return True
-        return False
-
-    if os.path.exists(dst) and allow_cache:
-        yield from update_ui_lastest_msg(f"调用缓存 {arxiv_id}", chatbot=chatbot, history=history)  # 刷新界面
-        success = True
-    else:
-        yield from update_ui_lastest_msg(f"开始下载 {arxiv_id}", chatbot=chatbot, history=history)  # 刷新界面
-        success = fix_url_and_download()
-        yield from update_ui_lastest_msg(f"下载完成 {arxiv_id}", chatbot=chatbot, history=history)  # 刷新界面
-
-
-    if not success:
-        yield from update_ui_lastest_msg(f"下载失败 {arxiv_id}", chatbot=chatbot, history=history)
-        raise tarfile.ReadError(f"论文下载失败 {arxiv_id}")
-
    # <-------------- extract file ------------->
+    yield from update_ui_lastest_msg("下载完成", chatbot=chatbot, history=history)  # 刷新界面
    from toolbox import extract_archive
-    try:
    extract_archive(file_path=dst, dest_dir=extract_dst)
-    except tarfile.ReadError:
-        os.remove(dst)
-        raise tarfile.ReadError(f"论文下载失败")
    return extract_dst, arxiv_id


--- a/crazy_functions/Rag_Interface.py
+++ b/crazy_functions/Rag_Interface.py
@@ -2,7 +2,20 @@ from toolbox import CatchException, update_ui, get_conf, get_log_folder, update_
 from crazy_functions.crazy_utils import input_clipping
 from crazy_functions.crazy_utils import request_gpt_model_in_new_thread_with_ui_alive

+VECTOR_STORE_TYPE = "Milvus"
+
+if VECTOR_STORE_TYPE == "Milvus":
+    try:
+        from crazy_functions.rag_fns.milvus_worker import MilvusRagWorker as LlamaIndexRagWorker
+    except:
+        VECTOR_STORE_TYPE = "Simple"
+
+if VECTOR_STORE_TYPE == "Simple":
+    from crazy_functions.rag_fns.llama_index_worker import LlamaIndexRagWorker
+
+
 RAG_WORKER_REGISTER = {}
+
 MAX_HISTORY_ROUND = 5
 MAX_CONTEXT_TOKEN_LIMIT = 4096
 REMEMBER_PREVIEW = 1000
@@ -10,16 +23,6 @@ REMEMBER_PREVIEW = 1000
@CatchException
 def Rag问答(txt, llm_kwargs, plugin_kwargs, chatbot, history, system_prompt, user_request):

-    # import vector store lib
-    VECTOR_STORE_TYPE = "Milvus"
-    if VECTOR_STORE_TYPE == "Milvus":
-        try:
-            from crazy_functions.rag_fns.milvus_worker import MilvusRagWorker as LlamaIndexRagWorker
-        except:
-            VECTOR_STORE_TYPE = "Simple"
-    if VECTOR_STORE_TYPE == "Simple":
-        from crazy_functions.rag_fns.llama_index_worker import LlamaIndexRagWorker
-
    # 1. we retrieve rag worker from global context
    user_name = chatbot.get_user()
    checkpoint_dir = get_log_folder(user_name, plugin_name='experimental_rag')
--- a/crazy_functions/latex_fns/latex_toolbox.py
+++ b/crazy_functions/latex_fns/latex_toolbox.py
@@ -644,17 +644,8 @@ def run_in_subprocess(func):


 def _merge_pdfs(pdf1_path, pdf2_path, output_path):
-    try:
-        logger.info("Merging PDFs using _merge_pdfs_ng")
-        _merge_pdfs_ng(pdf1_path, pdf2_path, output_path)
-    except:
-        logger.info("Merging PDFs using _merge_pdfs_legacy")
-        _merge_pdfs_legacy(pdf1_path, pdf2_path, output_path)
-
-
-def _merge_pdfs_ng(pdf1_path, pdf2_path, output_path):
    import PyPDF2  # PyPDF2这个库有严重的内存泄露问题，把它放到子进程中运行，从而方便内存的释放
-    from PyPDF2.generic import NameObject, TextStringObject, ArrayObject, FloatObject, NumberObject
+    from PyPDF2.generic import NameObject, TextStringObject,ArrayObject,FloatObject,NumberObject

    Percent = 1
    # raise RuntimeError('PyPDF2 has a serious memory leak problem, please use other tools to merge PDF files.')
@@ -697,203 +688,65 @@ def _merge_pdfs_ng(pdf1_path, pdf2_path, output_path):
                    ),
                    0,
                )
-                if "/Annots" in new_page:
-                    annotations = new_page["/Annots"]
+                if '/Annots' in page1:
+                    page1_annot_id = [annot.idnum for annot in page1['/Annots']]
+                else:
+                    page1_annot_id = []
+
+                if '/Annots' in page2:
+                    page2_annot_id = [annot.idnum for annot in page2['/Annots']]
+                else:
+                    page2_annot_id = []
+                if '/Annots' in new_page:
+                    annotations = new_page['/Annots']
                    for i, annot in enumerate(annotations):
                        annot_obj = annot.get_object()

                        # 检查注释类型是否是链接（/Link）
-                        if annot_obj.get("/Subtype") == "/Link":
+                        if annot_obj.get('/Subtype') == '/Link':
                            # 检查是否为内部链接跳转（/GoTo）或外部URI链接（/URI）
-                            action = annot_obj.get("/A")
+                            action = annot_obj.get('/A')
                            if action:

-                                if "/S" in action and action["/S"] == "/GoTo":
+                                if '/S' in action and action['/S'] == '/GoTo':
                                    # 内部链接：跳转到文档中的某个页面
-                                    dest = action.get("/D")  # 目标页或目标位置
-                                    # if dest and annot.idnum in page2_annot_id:
-                                    if dest in pdf2_reader.named_destinations:
+                                    dest = action.get('/D')  # 目标页或目标位置
+                                    if dest and annot.idnum in page2_annot_id:
                                        # 获取原始文件中跳转信息，包括跳转页面
-                                        destination = pdf2_reader.named_destinations[
-                                            dest
-                                        ]
-                                        page_number = (
-                                            pdf2_reader.get_destination_page_number(
-                                                destination
-                                            )
-                                        )
-                                        # 更新跳转信息，跳转到对应的页面和，指定坐标 (100, 150)，缩放比例为 100%
-                                        # “/D”:[10,'/XYZ',100,100,0]
-                                        if destination.dest_array[1] == "/XYZ":
-                                            annot_obj["/A"].update(
-                                                {
-                                                    NameObject("/D"): ArrayObject(
-                                                        [
-                                                            NumberObject(page_number),
-                                                            destination.dest_array[1],
-                                                            FloatObject(
-                                                                destination.dest_array[
-                                                                    2
-                                                                ]
-                                                                + int(
-                                                                    page1.mediaBox.getWidth()
-                                                                )
-                                                            ),
-                                                            destination.dest_array[3],
-                                                            destination.dest_array[4],
-                                                        ]
-                                                    )  # 确保键和值是 PdfObject
-                                                }
-                                            )
-                                        else:
-                                            annot_obj["/A"].update(
-                                                {
-                                                    NameObject("/D"): ArrayObject(
-                                                        [
-                                                            NumberObject(page_number),
-                                                            destination.dest_array[1],
-                                                        ]
-                                                    )  # 确保键和值是 PdfObject
-                                                }
-                                            )
-
-                                        rect = annot_obj.get("/Rect")
+                                        destination = pdf2_reader.named_destinations[dest]
+                                        page_number = pdf2_reader.get_destination_page_number(destination)
+                                        #更新跳转信息，跳转到对应的页面和，指定坐标 (100, 150)，缩放比例为 100%
+                                        #“/D”:[10,'/XYZ',100,100,0]
+                                        annot_obj['/A'].update({
+                                            NameObject("/D"): ArrayObject([NumberObject(page_number),destination.dest_array[1], FloatObject(destination.dest_array[2] + int(page1.mediaBox.getWidth())) ,destination.dest_array[3],destination.dest_array[4]])  # 确保键和值是 PdfObject
+                                        })
+                                        rect = annot_obj.get('/Rect')
                                        # 更新点击坐标
-                                        rect = ArrayObject(
-                                            [
-                                                FloatObject(
-                                                    rect[0]
-                                                    + int(page1.mediaBox.getWidth())
-                                                ),
-                                                rect[1],
-                                                FloatObject(
-                                                    rect[2]
-                                                    + int(page1.mediaBox.getWidth())
-                                                ),
-                                                rect[3],
-                                            ]
-                                        )
-                                        annot_obj.update(
-                                            {
-                                                NameObject(
-                                                    "/Rect"
-                                                ): rect  # 确保键和值是 PdfObject
-                                            }
-                                        )
-                                    # if dest and annot.idnum in page1_annot_id:
-                                    if dest in pdf1_reader.named_destinations:
-
+                                        rect = ArrayObject([FloatObject(rect[0]+ int(page1.mediaBox.getWidth())),rect[1],
+                                                            FloatObject(rect[2]+int(page1.mediaBox.getWidth())),rect[3] ])
+                                        annot_obj.update({
+                                            NameObject("/Rect"): rect  # 确保键和值是 PdfObject
+                                        })
+                                    if dest and annot.idnum in page1_annot_id:
                                        # 获取原始文件中跳转信息，包括跳转页面
-                                        destination = pdf1_reader.named_destinations[
-                                            dest
-                                        ]
-                                        page_number = (
-                                            pdf1_reader.get_destination_page_number(
-                                                destination
-                                            )
-                                        )
-                                        # 更新跳转信息，跳转到对应的页面和，指定坐标 (100, 150)，缩放比例为 100%
-                                        # “/D”:[10,'/XYZ',100,100,0]
-                                        if destination.dest_array[1] == "/XYZ":
-                                            annot_obj["/A"].update(
-                                                {
-                                                    NameObject("/D"): ArrayObject(
-                                                        [
-                                                            NumberObject(page_number),
-                                                            destination.dest_array[1],
-                                                            FloatObject(
-                                                                destination.dest_array[
-                                                                    2
-                                                                ]
-                                                            ),
-                                                            destination.dest_array[3],
-                                                            destination.dest_array[4],
-                                                        ]
-                                                    )  # 确保键和值是 PdfObject
-                                                }
-                                            )
-                                        else:
-                                            annot_obj["/A"].update(
-                                                {
-                                                    NameObject("/D"): ArrayObject(
-                                                        [
-                                                            NumberObject(page_number),
-                                                            destination.dest_array[1],
-                                                        ]
-                                                    )  # 确保键和值是 PdfObject
-                                                }
-                                            )
+                                        destination = pdf1_reader.named_destinations[dest]
+                                        page_number = pdf1_reader.get_destination_page_number(destination)
+                                        #更新跳转信息，跳转到对应的页面和，指定坐标 (100, 150)，缩放比例为 100%
+                                        #“/D”:[10,'/XYZ',100,100,0]
+                                        annot_obj['/A'].update({
+                                            NameObject("/D"): ArrayObject([NumberObject(page_number),destination.dest_array[1], FloatObject(destination.dest_array[2]) ,destination.dest_array[3],destination.dest_array[4]])  # 确保键和值是 PdfObject
+                                        })
+                                        rect = annot_obj.get('/Rect')
+                                        rect = ArrayObject([FloatObject(rect[0]),rect[1],
+                                                            FloatObject(rect[2]),rect[3] ])
+                                        annot_obj.update({
+                                            NameObject("/Rect"): rect  # 确保键和值是 PdfObject
+                                        })

-                                        rect = annot_obj.get("/Rect")
-                                        rect = ArrayObject(
-                                            [
-                                                FloatObject(rect[0]),
-                                                rect[1],
-                                                FloatObject(rect[2]),
-                                                rect[3],
-                                            ]
-                                        )
-                                        annot_obj.update(
-                                            {
-                                                NameObject(
-                                                    "/Rect"
-                                                ): rect  # 确保键和值是 PdfObject
-                                            }
-                                        )
-
-                                elif "/S" in action and action["/S"] == "/URI":
+                                elif '/S' in action and action['/S'] == '/URI':
                                    # 外部链接：跳转到某个URI
-                                    uri = action.get("/URI")
+                                    uri = action.get('/URI')
                output_writer.addPage(new_page)
-            # Save the merged PDF file
-            with open(output_path, "wb") as output_file:
-                output_writer.write(output_file)
-
-
-def _merge_pdfs_legacy(pdf1_path, pdf2_path, output_path):
-    import PyPDF2  # PyPDF2这个库有严重的内存泄露问题，把它放到子进程中运行，从而方便内存的释放
-
-    Percent = 0.95
-    # raise RuntimeError('PyPDF2 has a serious memory leak problem, please use other tools to merge PDF files.')
-    # Open the first PDF file
-    with open(pdf1_path, "rb") as pdf1_file:
-        pdf1_reader = PyPDF2.PdfFileReader(pdf1_file)
-        # Open the second PDF file
-        with open(pdf2_path, "rb") as pdf2_file:
-            pdf2_reader = PyPDF2.PdfFileReader(pdf2_file)
-            # Create a new PDF file to store the merged pages
-            output_writer = PyPDF2.PdfFileWriter()
-            # Determine the number of pages in each PDF file
-            num_pages = max(pdf1_reader.numPages, pdf2_reader.numPages)
-            # Merge the pages from the two PDF files
-            for page_num in range(num_pages):
-                # Add the page from the first PDF file
-                if page_num < pdf1_reader.numPages:
-                    page1 = pdf1_reader.getPage(page_num)
-                else:
-                    page1 = PyPDF2.PageObject.createBlankPage(pdf1_reader)
-                # Add the page from the second PDF file
-                if page_num < pdf2_reader.numPages:
-                    page2 = pdf2_reader.getPage(page_num)
-                else:
-                    page2 = PyPDF2.PageObject.createBlankPage(pdf1_reader)
-                # Create a new empty page with double width
-                new_page = PyPDF2.PageObject.createBlankPage(
-                    width=int(
-                        int(page1.mediaBox.getWidth())
-                        + int(page2.mediaBox.getWidth()) * Percent
-                    ),
-                    height=max(page1.mediaBox.getHeight(), page2.mediaBox.getHeight()),
-                )
-                new_page.mergeTranslatedPage(page1, 0, 0)
-                new_page.mergeTranslatedPage(
-                    page2,
-                    int(
-                        int(page1.mediaBox.getWidth())
-                        - int(page2.mediaBox.getWidth()) * (1 - Percent)
-                    ),
-                    0,
-                )
                output_writer.addPage(new_page)
            # Save the merged PDF file
            with open(output_path, "wb") as output_file:
--- a/docs/Dockerfile+JittorLLM
+++ b/docs/Dockerfile+JittorLLM
@@ -0,0 +1 @@
+# 此Dockerfile不再维护，请前往docs/GithubAction+JittorLLMs
--- a/docs/GithubAction+AllCapacityBeta
+++ b/docs/GithubAction+AllCapacityBeta
@@ -0,0 +1,57 @@
+# docker build -t gpt-academic-all-capacity -f docs/GithubAction+AllCapacity  --network=host --build-arg http_proxy=http://localhost:10881 --build-arg https_proxy=http://localhost:10881 .
+# docker build -t gpt-academic-all-capacity -f docs/GithubAction+AllCapacityBeta  --network=host .
+# docker run -it --net=host gpt-academic-all-capacity  bash
+
+# 从NVIDIA源，从而支持显卡（检查宿主的nvidia-smi中的cuda版本必须>=11.3）
+FROM fuqingxu/11.3.1-runtime-ubuntu20.04-with-texlive:latest
+
+# edge-tts需要的依赖，某些pip包所需的依赖
+RUN apt update && apt install ffmpeg build-essential -y
+
+# use python3 as the system default python
+WORKDIR /gpt
+RUN curl -sS https://bootstrap.pypa.io/get-pip.py | python3.8
+
+# # 非必要步骤，更换pip源 （以下三行，可以删除）
+# RUN echo '[global]' > /etc/pip.conf && \
+#     echo 'index-url = https://mirrors.aliyun.com/pypi/simple/' >> /etc/pip.conf && \
+#     echo 'trusted-host = mirrors.aliyun.com' >> /etc/pip.conf
+
+# 下载pytorch
+RUN python3 -m pip install torch torchvision --extra-index-url https://download.pytorch.org/whl/cu113
+# 准备pip依赖
+RUN python3 -m pip install openai numpy arxiv rich
+RUN python3 -m pip install colorama Markdown pygments pymupdf
+RUN python3 -m pip install python-docx moviepy pdfminer
+RUN python3 -m pip install zh_langchain==0.2.1 pypinyin
+RUN python3 -m pip install rarfile py7zr
+RUN python3 -m pip install aliyun-python-sdk-core==2.13.3 pyOpenSSL webrtcvad scipy git+https://github.com/aliyun/alibabacloud-nls-python-sdk.git
+# 下载分支
+WORKDIR /gpt
+RUN git clone --depth=1 https://github.com/binary-husky/gpt_academic.git
+WORKDIR /gpt/gpt_academic
+RUN git clone --depth=1 https://github.com/OpenLMLab/MOSS.git request_llms/moss
+
+RUN python3 -m pip install -r requirements.txt
+RUN python3 -m pip install -r request_llms/requirements_moss.txt
+RUN python3 -m pip install -r request_llms/requirements_qwen.txt
+RUN python3 -m pip install -r request_llms/requirements_chatglm.txt
+RUN python3 -m pip install -r request_llms/requirements_newbing.txt
+RUN python3 -m pip install nougat-ocr
+
+
+# 预热Tiktoken模块
+RUN python3  -c 'from check_proxy import warm_up_modules; warm_up_modules()'
+
+# 安装知识库插件的额外依赖
+RUN apt-get update && apt-get install libgl1 -y
+RUN pip3 install transformers protobuf langchain sentence-transformers  faiss-cpu nltk beautifulsoup4 bitsandbytes tabulate icetk --upgrade
+RUN pip3 install unstructured[all-docs] --upgrade
+RUN python3  -c 'from check_proxy import warm_up_vectordb; warm_up_vectordb()'
+RUN rm -rf /usr/local/lib/python3.8/dist-packages/tests
+
+
+# COPY .cache /root/.cache
+# COPY config_private.py config_private.py
+# 启动
+CMD ["python3", "-u", "main.py"]
--- a/docs/GithubAction+NoLocal+Latex+Arm
+++ b/docs/GithubAction+NoLocal+Latex+Arm
@@ -1,25 +0,0 @@
-# 此Dockerfile适用于“无本地模型”的环境构建，如果需要使用chatglm等本地模型，请参考 docs/Dockerfile+ChatGLM
-# - 1 修改 `config.py`
-# - 2 构建 docker build -t gpt-academic-nolocal-latex -f docs/GithubAction+NoLocal+Latex .
-# - 3 运行 docker run -v /home/fuqingxu/arxiv_cache:/root/arxiv_cache --rm -it --net=host gpt-academic-nolocal-latex
-
-FROM menghuan1918/ubuntu_uv_ctex:latest
-ENV DEBIAN_FRONTEND=noninteractive
-SHELL ["/bin/bash", "-c"]
-WORKDIR /gpt
-COPY . .
-RUN /root/.cargo/bin/uv venv --seed \
-    && source .venv/bin/activate \
-    && /root/.cargo/bin/uv pip install openai numpy arxiv rich colorama Markdown pygments pymupdf python-docx pdfminer \
-    && /root/.cargo/bin/uv pip install -r requirements.txt \
-    && /root/.cargo/bin/uv clean
-
-# 对齐python3
-RUN rm -f /usr/bin/python3 && ln -s /gpt/.venv/bin/python /usr/bin/python3
-RUN rm -f /usr/bin/python && ln -s /gpt/.venv/bin/python /usr/bin/python
-
-# 可选步骤，用于预热模块
-RUN python3 -c 'from check_proxy import warm_up_modules; warm_up_modules()'
-
-# 启动
-CMD ["python3", "-u", "main.py"]
--- a/request_llms/bridge_all.py
+++ b/request_llms/bridge_all.py
@@ -256,8 +256,6 @@ model_info = {
        "max_token": 128000,
        "tokenizer": tokenizer_gpt4,
        "token_cnt": get_token_num_gpt4,
-        "openai_disable_system_prompt": True,
-        "openai_disable_stream": True,
    },
    "o1-mini": {
        "fn_with_ui": chatgpt_ui,
@@ -266,8 +264,6 @@ model_info = {
        "max_token": 128000,
        "tokenizer": tokenizer_gpt4,
        "token_cnt": get_token_num_gpt4,
-        "openai_disable_system_prompt": True,
-        "openai_disable_stream": True,
    },

    "gpt-4-turbo": {
@@ -1285,3 +1281,4 @@ def predict(inputs:str, llm_kwargs:dict, plugin_kwargs:dict, chatbot,

    # 更新一下llm_kwargs的参数，否则会出现参数不匹配的问题
    yield from method(inputs, llm_kwargs, plugin_kwargs, chatbot, history, system_prompt, stream, additional_fn)
+
--- a/request_llms/bridge_chatgpt.py
+++ b/request_llms/bridge_chatgpt.py
@@ -202,13 +202,10 @@ def predict_no_ui_long_connection(inputs:str, llm_kwargs:dict, history:list=[],
                    if (time.time()-observe_window[1]) > watch_dog_patience:
                        raise RuntimeError("用户取消了程序。")
        else: raise RuntimeError("意外Json结构："+delta)
-
-    finish_reason = json_data.get('finish_reason', None) if json_data else None
-    if finish_reason == 'content_filter':
-        raise RuntimeError("由于提问含不合规内容被过滤。")
-    if finish_reason == 'length':
+    if json_data and json_data['finish_reason'] == 'content_filter':
+        raise RuntimeError("由于提问含不合规内容被Azure过滤。")
+    if json_data and json_data['finish_reason'] == 'length':
        raise ConnectionAbortedError("正常结束，但显示Token不足，导致输出不完整，请削减单次输入的文本量。")
-
    return result


@@ -539,3 +536,4 @@ def generate_payload(inputs:str, llm_kwargs:dict, history:list, system_prompt:st

    return headers,payload

+
--- a/requirements.txt
+++ b/requirements.txt
@@ -2,15 +2,14 @@ https://public.agent-matrix.com/publish/gradio-3.32.10-py3-none-any.whl
 fastapi==0.110
 gradio-client==0.8
 pypdf2==2.12.1
-httpx<=0.25.2
 zhipuai==2.0.1
 tiktoken>=0.3.3
 requests[socks]
-pydantic==2.9.2
+pydantic==2.5.2
+llama-index~=0.10
 protobuf==3.20
 transformers>=4.27.1,<4.42
 scipdf_parser>=0.52
-spacy==3.7.4
 anthropic>=0.18.1
 python-markdown-math
 pymdown-extensions
@@ -33,14 +32,3 @@ loguru
 arxiv
 numpy
 rich
-
-
-llama-index-core==0.10.68
-llama-index-legacy==0.9.48
-llama-index-readers-file==0.1.33
-llama-index-readers-llama-parse==0.1.6
-llama-index-embeddings-azure-openai==0.1.10
-llama-index-embeddings-openai==0.1.10
-llama-parse==0.4.9
-mdit-py-plugins>=0.3.3
-linkify-it-py==2.0.3
--- a/tests/test_anim_gen.py
+++ b/tests/test_anim_gen.py
@@ -1,12 +0,0 @@
-"""
-对项目中的各个插件进行测试。运行方法：直接运行 python tests/test_plugins.py
-"""
-
-import init_test
-import os, sys
-
-
-if __name__ == "__main__":
-    from test_utils import plugin_test
-
-    plugin_test(plugin='crazy_functions.数学动画生成manim->动画生成', main_input="A point moving along function culve y=sin(x), starting from x=0 and stop at x=4*\pi.")
--- a/4
+++ b/4
@@ -1,5 +1,5 @@
 {
-  "version": 3.90,
+  "version": 3.83,
  "show_feature": true,
-  "new_feature": "增加RAG组件 <-> 升级多合一主提交键"
+  "new_feature": "增加欢迎页面 <-> 优化图像生成插件 <-> 添加紫东太初大模型支持 <-> 保留主题选择 <-> 支持更复杂的插件框架 <-> 上传文件时显示进度条"
 }
作者	SHA1	备注	提交日期
binary-husky	7415d532d1	solve the pdf concate error	2024-10-13 07:36:36 +00:00
binary-husky	97eef45ab7	Merge branch 'frontier' of github.com:binary-husky/chatgpt_academic into frontier	2024-10-01 11:59:14 +00:00
binary-husky	0c0e2acb9b	remove logging extra	2024-10-01 11:57:47 +00:00
Ren Lifei	9fba8e0142	Added some modules to support openrouter (#1975 ) * Added some modules for supporting openrouter model Added some modules for supporting openrouter model * Update config.py * Update .gitignore * Update bridge_openrouter.py * Not changed actually * Refactor logging in bridge_openrouter.py --------- Co-authored-by: binary-husky <qingxu.fu@outlook.com>	2024-09-28 18:05:34 +08:00
binary-husky	7d7867fb64	remove comment	2024-09-23 15:16:13 +00:00
binary-husky	f9dbaa39fb	Merge branch 'frontier' of github.com:binary-husky/chatgpt_academic into frontier	2024-09-21 15:40:24 +00:00
binary-husky	bbc2288c5b	relax llama index version	2024-09-21 15:40:10 +00:00
Steven Moder	64ab916838	fix: loguru argument error with proxy enabled (#1977 )	2024-09-21 23:32:00 +08:00
binary-husky	8fe559da9f	update translation matrix	2024-09-21 14:56:10 +00:00
binary-husky	09fd22091a	fix: console output	2024-09-21 14:41:36 +00:00
binary-husky	e296719b23	Merge branch 'purge_print' into frontier	2024-09-16 09:56:25 +00:00
binary-husky	2f343179a2	logging -> loguru: final stage	2024-09-15 15:51:51 +00:00
binary-husky	4d9604f2e9	update social helper	2024-09-15 15:16:36 +00:00
binary-husky	bbf9e9f868	logging -> loguru stage 4	2024-09-14 16:00:09 +00:00
binary-husky	aa1f967dd7	support o1-preview and o1-mini	2024-09-13 03:11:53 +00:00
binary-husky	0d082327c8	logging -> loguru: stage 3	2024-09-11 08:49:55 +00:00
binary-husky	80acd9c875	import loguru: stage 2	2024-09-11 08:18:01 +00:00
binary-husky	17cd4f8210	logging sys to loguru: stage 1 complete	2024-09-11 03:30:30 +00:00
				`@@ -0,0 +1 @@`
				`# 此Dockerfile不再维护，请前往docs/GithubAction+JittorLLMs`