]> code.communitydata.science - cdsc_reddit.git/blobdiff - similarities/Makefile
Some improvements to run affinity clustering on larger dataset and
[cdsc_reddit.git] / similarities / Makefile
index 89a908f968649334c0402bd5e88a754dfa8532a9..d5187c98d8a5b7944d328ee1246d87961715c88e 100644 (file)
@@ -1,2 +1,5 @@
 /gscratch/comdata/output/reddit_similarity/subreddit_comment_authors_10000.parquet: cosine_similarities.py /gscratch/comdata/output/reddit_similarity/tfidf/comment_authors.parquet
-       start_spark_and_run.sh 1 cosine_similarities.py author --outfile=/gscratch/comdata/output/reddit_similarity/subreddit_comment_authors_10000.parquet
+       start_spark_and_run.sh 1 cosine_similarities.py author --outfile=/gscratch/comdata/output/reddit_similarity/subreddit_comment_authors_10000.feather
+
+/gscratch/comdata/output/reddit_similarity/comment_terms_10000_weekly.parquet: cosine_similarities.py /gscratch/comdata/output/reddit_similarity/tfidf/comment_authors.parquet
+       start_spark_and_run.sh 1 weekly_cosine_similarities.py term --outfile=/gscratch/comdata/output/reddit_similarity/subreddit_comment_terms_10000_weely.parquet

Community Data Science Collective || Want to submit a patch?