次に、テキストデータの次元削減を行います。ここでは、TF-IDFを用いてテキストをベクトル化した後、次元削減を行います。
# テキストデータの次元削減
from sklearn.feature_extraction.text import TfidfVectorizer
from sklearn.decomposition import PCA
import numpy as np
# サンプルテキストデータ
documents = [“This is the first sentence.”, “This is the second.”, “And this is the third one.”]
# TF-IDFでベクトル化
vectorizer = TfidfVectorizer()
X = vectorizer.fit_transform(documents)
# PCA
pca = PCA(n_components=2)
X_reduced = pca.fit_transform(X.toarray())
# 結果を出力
print(X_reduced)
コメント