WEKA: StringToWordVector results is different between CUI and my java code
I am using this setting in GUI and JAVA:
-R
last
-W
1000
-prune-rate
-1.0
-C
-I
-N
0
-S
-stemmer
weka.core.stemmers.NullStemmer
-M
20
-tokenizer
weka.core.tokenizers.WordTokenizer -delimiters " \r\n\t.,;:\'\"()?!"
JAVA:
@relation 'testing2-weka.filters.unsupervised.attribute.StringToWord开发者_如何学PythonVector-R2-W1000-prune-rate-1.0-C-I-N0-S-stemmerweka.core.stemmers.NullStemmer-M20-tokenizerweka.core.tokenizers.WordTokenizer -delimiters \" \\r\\n\\t.,;:\\\'\\\"()?!\"'
GUI:
weka.filters.unsupervised.attribute.StringToWordVector-R2-W1000-prune-rate-1.0-C-I-N0-S-stemmerweka.core.stemmers.NullStemmer-M20-tokenizerweka.core.tokenizers.WordTokenizer -delimiters " \r\n\t.,;:\'\"()?!"
And unfortunately am getting different results: the number of attributes I got in GUI is (86) and number of attributes in JAVA is around (300).
For same relation data-set.
This is the code i used:
BufferedReader reader = new BufferedReader(new FileReader("newTemp.arff"));
Instances dataRaw = new Instances(reader);
reader.close();
StringToWordVector filter = new StringToWordVector();
filter.setInputFormat(dataRaw);
filter.setAttributeIndices("last");
filter.setDoNotOperateOnPerClassBasis(false);
filter.setOutputWordCounts(true);
filter.setWordsToKeep(1000);
filter.setUseStoplist(true);
filter.setIDFTransform(true);
filter.setMinTermFreq(20);
filter.setDoNotOperateOnPerClassBasis(false);
filter.setPeriodicPruning(-1);
String[] options = filter.getOptions();
for(int i=0;i<options.length;i++) {
if (options[i].length() > 0)
System.out.println(options[i]);
}
Instances dataFiltered = Filter.useFilter(dataRaw, filter);
System.out.println("\n\n=====> Filtered data:<===\n\n" + dataFiltered.toString());
I don't know what is wrong. please help am coming sinking.
精彩评论