AIの実践的探求

【概要】

人工知能を活用して,画像認識や画像合成,日本語文書の処理などのタスクを実現する.コードを公開しているため,必要な調整が可能であり,関連する文書形式の変換処理なども提供している.これらのプログラムを通じて,人工知能の実践的な活用方法を体験できる.

【目次】

AIタスク

AIエージェント

Webブラウザで動作するAIプログラム集(Google Colaboratory版)

日本語AIチャット

コード生成LLM Qwen3-Coder-30B-A3B-Instruct

Alibaba Cloudがコード生成特化型AI「Qwen3-Coder」をオープンソースライセンス(Apache 2.0)でリリースしており,商用利用も可能である.

Hugging Face: https://huggingface.co/Qwen/Qwen3-Coder-30B-A3B-Instruct

GitHub: https://github.com/QwenLM/Qwen3-Coder

install: pip install -U --no-user transformers torch accelerate

from transformers import AutoModelForCausalLM, AutoTokenizer

model_name = "Qwen/Qwen3-Coder-30B-A3B-Instruct"

# load the tokenizer and the model
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForCausalLM.from_pretrained(
    model_name,
    torch_dtype="auto",
    device_map="auto"
)

# prepare the model input
prompt = "Write a quick sort algorithm."
messages = [
    {"role": "user", "content": prompt}
]
text = tokenizer.apply_chat_template(
    messages,
    tokenize=False,
    add_generation_prompt=True,
)
model_inputs = tokenizer([text], return_tensors="pt").to(model.device)

# conduct text completion
generated_ids = model.generate(
    **model_inputs,
    max_new_tokens=65536
)
output_ids = generated_ids[0][len(model_inputs.input_ids[0]):].tolist()

content = tokenizer.decode(output_ids, skip_special_tokens=True)

print("content:", content)

Function Calling(関数呼び出し,モデルが必要に応じて外部ツールを呼び出す機能)の例

# Your tool implementation
def square_the_number(num: float) -> dict:
    return num ** 2

# Define Tools
tools = [
    {
        "type": "function",
        "function": {
            "name": "square_the_number",
            "description": "output the square of the number.",
            "parameters": {
                "type": "object",
                "required": ["input_num"],
                "properties": {
                    "input_num": {
                        "type": "number",
                        "description": "input_num is a number that will be squared"
                    }
                },
            }
        }
    }
]

from openai import OpenAI

# Define LLM
client = OpenAI(
    # Use a custom endpoint compatible with OpenAI API
    base_url="http://localhost:8000/v1",  # api_base
    api_key="EMPTY"
)

messages = [{"role": "user", "content": "square the number 1024"}]

completion = client.chat.completions.create(
    messages=messages,
    model="Qwen3-Coder-30B-A3B-Instruct",
    max_tokens=65536,
    tools=tools,
)

print(completion.choices[0])

vLLM(高速な推論サーバー)を使ったサーバー起動の例

pip install -U --no-user vllm
vllm serve "Qwen/Qwen3-Coder-30B-A3B-Instruct"

動画生成AI

関連タスク(データ処理など)

オンライン・コニュニケーション

夜間画像の画質改善

マークダウンへの変換

PDFへの変換

PDF をマークダウン(Markdown)に変換するツール

準備作業

  1. Python のインストール
  2. 必要なライブラリのインストール:

    Windows で,管理者権限でコマンドプロンプトを起動する (手順:Windowsキーまたはスタートメニュー → cmd と入力 → 右クリック → 「管理者として実行」)。

    次のコマンドを実行

    pip install -U --no-user pymupdf4llm pywin32
    
  3. プログラムの作成

    pdf2md.py という名前で,以下の内容のPythonファイルを作成する.

    from sys import argv
    import pymupdf4llm
    
    def convert_pdf_to_markdown(input_pdf, output_md):
        """
        Convert PDF to Markdown format.
    
        Args:
            input_pdf (str): Path to input PDF file
            output_md (str): Path to output Markdown file
        """
        page_num = 1
    
        try:
            # Open output file in write mode with UTF-8 encoding
            with open(output_md, 'w', encoding='utf-8') as f:
                while True:
                    try:
                        # Convert each slide to Markdown
                        md_text = pymupdf4llm.to_markdown(
                            input_pdf,
                            pages=[page_num - 1],
                            show_progress=False,
                            margins=0
                        )
    
                        # Write slide content and slide number
                        f.write(f"\n# 【スライド番号 {page_num:02d}】\n\n")  # Added extra newline for better spacing
                        f.write(md_text)
                        page_num += 1
    
                    except Exception as e:
                        if page_num == 1:
                            # If first page fails, there might be an issue with the file
                            raise Exception(f"PDFファイルの処理中にエラーが発生しました: {str(e)}")
                        # Exit when no more slides exist
                        break
    
        except Exception as e:
            print(f"エラーが発生しました: {str(e)}")
            exit(1)
    
    def main():
        """Main function to handle command line arguments and run conversion."""
        if len(argv) != 3:
            print("使用方法: python script.py 入力PDF.pdf 出力マークダウン.md")
            exit(1)
    
        input_pdf = argv[1]
        output_md = argv[2]
    
        # Ensure output file has .md extension
        if not output_md.endswith('.md'):
            output_md += '.md'
    
        # Ensure input file has .pdf extension
        if not input_pdf.endswith('.pdf'):
            input_pdf += '.pdf'
    
        try:
            convert_pdf_to_markdown(input_pdf, output_md)
            print(f"変換が完了しました。出力ファイル: {output_md}")
        except Exception as e:
            print(f"変換に失敗しました: {str(e)}")
            exit(1)
    
    if __name__ == "__main__":
        main()
    

PPTX, PPTファイルをMarkdownに変換するツール

ppt2md.py という名前で,以下の内容のPythonファイルを作成する.

フォルダ内のPPTX, PPTファイルを一括してMarkdownに変換するツール

pptindir2md.py という名前で,以下の内容のPythonファイルを作成する.

import pymupdf4llm
import win32com.client
import os
import time
import psutil  # PowerPointプロセスを確認するためにpsutilを使用

# カレントディレクトリを取得
input_dir = os.getcwd()

# ディレクトリ内のすべての.pptまたは.pptxファイルをリストアップ
ppt_files = [f for f in os.listdir(input_dir) if f.lower().endswith(('.ppt', '.pptx'))]

if not ppt_files:
    print("指定されたディレクトリに.pptまたは.pptxファイルが見つかりませんでした。")
    exit()

# PowerPointアプリケーションを処理する関数
def convert_ppt_to_pdf(input_file):
    ppt = win32com.client.Dispatch("PowerPoint.Application")
    ppt.Visible = 1  # PowerPointウィンドウを表示したい場合は1

    pdf_file = input_file.rsplit('.', 1)[0] + '.pdf'
    print(f"PDF出力先: {pdf_file}")

    # プレゼンテーションを開く
    presentation = ppt.Presentations.Open(os.path.abspath(input_file))

    # SaveAsメソッドでPDFとして保存 (32はPDF形式を指定)
    presentation.SaveAs(os.path.abspath(pdf_file), 32)  # 32はPDF形式

    presentation.Close()  # プレゼンテーションを閉じる
    ppt.Quit()  # PowerPointアプリケーションを終了

    # 完全にPowerPointプロセスが終了したことを確認
    time.sleep(1)  # 少し待機してから次の処理に進む

    # PowerPointプロセスが残っていないか確認
    for proc in psutil.process_iter(attrs=['pid', 'name']):
        if proc.info['name'] == 'POWERPNT.EXE':
            print("PowerPointプロセスが残っているので,終了します")
            proc.terminate()  # 必要に応じてPowerPointプロセスを強制終了する

    time.sleep(5)  # 5秒待機(必要ならば)
    return pdf_file

# 各PowerPointファイルを処理
for ppt_file in ppt_files:
    input_file = os.path.join(input_dir, ppt_file)
    file_ext = os.path.splitext(input_file)[1].lower()

    # PowerPointファイルの場合,PDFに変換
    if file_ext in ['.ppt', '.pptx']:
        try:
            # PowerPointをPDFに変換
            pdf_file = convert_ppt_to_pdf(input_file)
            input_file = pdf_file
        except Exception as e:
            print(f"PowerPointの変換でエラーが発生しました: {e}")
            continue  # エラーが発生しても次のファイルに進む

    # 出力ファイル名を作成(拡張子をmdに変更)
    output_file = os.path.splitext(input_file)[0] + '.md'

    # 出力ファイルを作成/オープン
    with open(output_file, 'w', encoding='utf-8') as f:
        page_num = 1
        while True:
            try:
                md_text = pymupdf4llm.to_markdown(input_file, pages=[page_num - 1])
                f.write(md_text)
                f.write(f"\n#【ページ番号 {page_num:02d}】\n")
                page_num += 1
            except Exception:
                page_num -= 1
                break

    print(f"Markdownファイルを保存しました: {page_num} ページ, {output_file}")

使用方法

  1. コマンドプロンプトで以下のように実行する.
    python pdf2markdown.py 変換したいPDFファイル名
    python ppt2markdown.py 変換したいppt, pptxファイル名
    python pptindir2md.py
    
    1. PDFファイル名は,拡張子(.pdf)を含めた完全なファイル名を指定する
    2. 変換されたMarkdownテキストは標準出力に表示される.必要に応じて,出力を別のファイルにリダイレクトできる

PDFファイルをHTMLファイルに一括変換

以下の手順を実行することで,指定したディレクトリ内にあるすべてのPDFファイルをHTML形式に一括変換できる. 以下のコマンドは,各PDFファイルを1つずつ処理し,同じ名前のHTMLファイルを作成する.

cd /home/www/ai/ae
for i in *.pdf; do
  echo "変換中: $i"
  pdf2htmlEX "$i" "$(basename "$i" .pdf).html"
done

Web ページの改善

HTML のタグの対応関係のチェック

tag_check.py

URL のリストについて,OKかリダイレクトかエラーを判定

ユースケース:サーチエンジンで404になっているファイルのリストを取得し,このプログラムで状況をチェックする.

url_checker.py

robots.txt では「content="noindex,nofollow,noarchive"」を含むHTML ファイルを Disallow に指定する

次のプログラムを実行し,結果を robots.txt に追加する.

サイトマップチェック

sitemap.xml の lastmod, changefreq, priority 属性を更新する.

次のプログラムを実行する.

Windowsでの環境構築

Python プログラムのexe化(Windows 上)

Pythonプログラムのexe化ガイド

プログラムのデプロイ,共有

Hello,World表示プログラムのデプロイ・共有